Artificial Intelligence Policies for Development of AI Systems
This policy applies to the development of AI-powered tools, systems, or features internally by or in association with FLP. It applies to all AI system development activities, including but not limited to the use of traditional machine learning techniques, the development of generative AI models, and the use of third-party AI systems to assist in development.
These terms are applicable to all FLP employees, contractors, and any other individual performing AI development work and/or services for FLP.
Foundational Rule: FLP’s overall mission is access to justice. Every AI system we deploy for external users must meet a higher standard than our internal tooling because failures with external systems affect real people, real legal outcomes, and public trust in FLP. If a system cannot be evaluated, explained, and monitored, it is not ready to deploy.
Principles
Accuracy over Speed: Legal information errors have real-world consequences. No model enters production without a Model Card documenting how it was evaluated, with rigor proportionate to the stakes of the deployment. Where a held-out test set representative of the deployment context is not used, the Model Card must state what was done instead and why that is sufficient for this use.
Clear Documentation and Reproducibility: Developers have the latitude to explore diverse design approaches, modeling strategies, and evaluation methods. That freedom comes with an obligation: key design choices, tradeoffs, and their rationale must be documented so they can be reviewed, reproduced, and learned from.
Transparency to Users: Users must be able to distinguish AI-generated or AI-inferred assertions from verified source data. This applies wherever a model's output makes a claim the user may rely on in place of reading the source: generated text, summaries, answers, and classifications, or scores about a document's content or legal effect.
Increased Oversight for High-Stakes Outputs: Outputs that inform legal decisions must have defined quality thresholds and human review or appeal pathways. Always consult the project-specific guidance and model cards.
Proportionate Safeguards: The level of pre-deployment review, monitoring, and documentation required scales with the stakes of the system.
Consultation and Governance Checkpoints
Before beginning development on any AI-driven feature or product, teams must consult with FLP’s internal AI team. This ensures alignment on approach, awareness of prior work, and early identification of risks before design decisions are locked in.
In addition, the following checkpoints are required for any AI system deployed to external users. These checkpoints ensure that FLP’s AI system development aligns with its AI governance principles and standards.
| Checkpoint | When Required | Applies To |
|---|---|---|
| AI team consult | Before development begins on any AI-driven feature or product | Initiating team + AI team |
| Consult and update project-specific guidance | Continuous | Initiation team, AI team, feature owner, and reviewer |
| Model Card | Before any model enters production, updated with each subsequent evaluation or release | Feature owner |
| Evaluation report | Before production + at each major update | Feature owner + reviewer |
| Periodic quality review | At minimum annually per deployed model | Feature owner |
System Development Standards
All AI systems that will eventually be deployed to external users must meet the following standards:
Documentation & Reproducibility
All AI-powered systems, pipelines, and workflows must have a completed Model Card to ensure reproducibility. All subsequent major changes and upgrades must be reflected in either Model Card Section 2 Base Model or Section 7 Version History of the Model Card.
Periodic & Ongoing Evaluation
Deployed models and pipelines must be monitored for performance degradation and formally re-evaluated at least annually, or sooner when a significant distribution shift is detected, a major update is released, or the system’s scope changes. Evaluation results must be recorded in Section 4.2 Metrics and Section 4.3 Failure Analysis of the Model Card.
Bias and Fairness
AI systems must be developed with bias and fairness in mind, consistent with FLP’s mission of equitable access to justice. Relevant considerations and any known limitations must be documented in Section 3.1 Training Dataset and Section 5 Known Limitations of the Model Card.
Training & Evaluation Data
Models must be trained, finetuned, and evaluated on data representative of actual production usage. All training and design decisions should be documented in Section 3 Data and Section 4 Training & Evaluation of the Model Card along with any accompanying rationale.
Evaluation data must be assessed against (1) the overall distribution of production usage, and (2) the distribution across individual features where performance variation would be meaningful (e.g., jurisdiction, court level, case type, filing source, document length). Where evaluation data underrepresents a known feature stratum, it must be acknowledged in Section 3.2 Validation and Test Dataset of the Model Card. Evaluation data must also be screened for leakage (where the evaluation data is also present in the training data) to preserve the integrity of performance measurements and prevent artificially inflated metrics. Training data must be properly balanced according to machine learning best practices and documented in Section 3.1 Training Dataset of the Model Card.
Data Standards
All datasets used to train, finetune, and evaluate models for deployment to external users must meet the following minimum standards:
Provenance Documented
Source, collection method, date range, and any applicable licenses or use restrictions must be recorded in Section 3 Data of the Model Card.
Nonpublic Personally Identifiable Information (PII) Screened
Training data must be screened for the following PII unless that PII appears in publicly accessible court records or similar public documents:
(a) information that identifies, relates to, describes, is reasonably capable of being associated with, or could reasonably be linked, directly or indirectly, with a particular consumer or household. This includes, but is not limited to an individual’s name, address, identification number, search history, location data;
(b) inferences drawn from information about an individual, and
(c) any data from sealed, expunged, or restricted court records.
Data failing screening must be remediated or excluded.
Copyright Reviewed
Avoid training systems on copyrightable data to the extent possible. When in doubt whether certain information may be used to train an AI system, consult with the internal AI team.
Use of User Data Disclosed
To the extent user queries and interactions from FLP’s production systems are used for model training and system improvements, clear disclosure and documentation must be in place. To the extent possible, FLP will seek to honor requests from users to opt out of the use of their information for training.
Questions?
Contact: Rachel Gao — rachel@free.law