FLP Wiki

CourtListener is a vast archive of legal information, but because much of it is crowdsourced, it's difficult to know where the data is deep or shallow.

That said, CourtListener is an initiative of Free Law Project, a non-profit that sometimes consults with organizations to gather data. In these arrangements, we may make specialized data sets that we add to CourtListener.

Known Datasets

People and organizations regularly use our tools to develop their own datasets without telling us they've done so. If they do, we may never know that their dataset is in CourtListener.

Below are some datasets we know about that we can discuss publicly:

  1. Broad Data — In early 2023, we added short descriptions and basic metadata for about 150M documents in the PACER system. This data was drawn from over three million PACER RSS feeds gathered and contributed by Troller BK in support of our mission.

  2. Broad Data — In early 2021, we scraped basic metadata from every unsealed bankruptcy, civil, and criminal case in the PACER system. This data, gathered on behalf of a major media organization, spans about 60M cases, and includes the case name, date filed, date terminated, and judge.

  3. Broad Data — In May 2019, we downloaded nearly all civil district court dockets filed from January 1, 2016 to November 9, 2018. Dockets with natures of suit related to civil rights, prisoner petitions or patent law were excluded.

  4. Bankruptcy Filings — In 2020, we gathered $1.9M worth of content from the Southern District of Illinois Bankruptcy Court. For every case from 2007 to 2017, we downloaded the docket sheet, initial petition, and docket entries containing the word "final."

  5. Fair Labor Standards Act Data (FLSA) — We worked with a start-up to create an extensive collection of labor-related dockets and initial complaints. The date range for this collection is from 2009 to 2017.

  6. Export Control — We have a large, random sample of dockets related to export-controlled technology collected on behalf of a major Department of Defense policy organization.

  7. Invoices — We have a large collection of documents described by the word "invoice" that can be used for machine learning.

Planned Datasets

We frequently receive requests for the data sets listed below. If you are interested in working on one of these data sets, please get in touch. We're happy to work together on these and to discuss an appropriate level of publicity (including none at all):

  1. Patent Litigation — We do not currently parse patent numbers from case law or filings, but probably should. Once completed, it would be possible to query and alert by patent number. We'd also like to eventually add the PTAB, TTAB, and ITC decisions.

  2. Bankruptcy Data — We have millions of bankruptcy filings, but we do not parse them for relevant information. We'd like to do so using a combination of heuristic and AI approaches.

  3. SCOTUS Data — We have a growing collection of recent Supreme Court filings and cases, but we'd like to ingest and digitize the older content as well.

  4. Punishments & Damages — AI makes it feasible to extract punishment and damage information from cases to build a dataset for analysis.

  5. Court Forms & Rules — A dataset of all rules and court forms would be invaluable for expanding our access to justice work. We'd like to create these datasets.

  6. State Bar Court Decisions — Some state bar associations issue decisions with ethical guidance. We'd like to add these to CourtListener.

Finally, we have some data sources in our storehouse that we have not yet merged into CourtListener. These can be treasure troves of data and are detailed on our project board dedicated to the topic.

If any of these datasets would be useful to you, your research, or your practice area, we welcome partnership discussions.

2 views Last updated 7 hours, 20 minutes ago
Creator: mike