Artificial Intelligence Development Policy
Bridge Policy — Interim Guidance
Effective: March 2026 | Expires: April 2026
Applies to: Engineering, Data, and Product Teams
Foundational Rule: FLP's mission is access to justice. Every AI system we deploy for external users must meet a higher standard than internal tooling — because failures affect real people, real legal outcomes, and public trust in our data. If a system cannot be evaluated, explained, and monitored, it is not ready to deploy.
Core Principles
| Principle | Description |
|---|---|
| Consult the AI Team First | Before beginning development on any AI-driven feature or product, teams must consult with FLP's internal AI team. This ensures alignment on approach, awareness of prior work, and early identification of risks before design decisions are locked in. |
| Accuracy Over Speed | Legal information errors have real-world consequences. No model enters production without documented evaluation against a held-out test set representative of the deployment context. |
| Clear Documentation and Reproducibility | Developers have latitude to explore design approaches, modeling strategies, and evaluation methods. That freedom comes with an obligation: key design choices, tradeoffs, and their rationale must be documented so they can be reviewed, reproduced, and learned from. |
| Transparency to Users | All AI-generated content served to external users must be clearly labeled as such. Users must be able to distinguish AI output from verified legal data. |
| Human Oversight at High-Stakes Outputs | Outputs that inform legal decisions must have defined quality thresholds and human review or appeal pathways. |
| Proportionate Safeguards | The level of pre-deployment review, monitoring, and documentation required scales with the stakes of the system. |
System Development Standards
| Standard | Requirement |
|---|---|
| Training & Evaluation Data | Models must be trained and finetuned following machine learning best practices and evaluated on data representative of actual production usage. |
| Documentation & Reproducibility | All AI-powered systems, pipelines, and workflows must have a completed Model Card to ensure reproducibility. All subsequent major changes and upgrades must be reflected in the Model Card. |
| Periodic & Ongoing Evaluation | Deployed models and pipelines must be monitored for performance degradation and formally re-evaluated at least annually, or sooner when a significant distribution shift is detected, a major update is released, or the system's scope changes. Evaluation results must be recorded in the Model Card. |
| Bias and Fairness | AI systems must be developed with bias and fairness in mind, consistent with FLP's mission of equitable access to justice. Relevant considerations and any known limitations must be documented in the Model Card. |
Data Standards
All datasets used to train, finetuned, and evaluate models for external-facing systems must meet the following minimum standards:
| Standard | Requirement |
|---|---|
| Provenance Documented | Source, collection method, date range, and any applicable licenses or use restrictions must be recorded in the Model Card. |
| PII Screened | Training data must be screened for private personally identifiable information, including names, contact information, and any data from sealed, expunged, or restricted court records. Data failing screening must be remediated or excluded. |
| Copyright Reviewed | The right to use potentially copyrighted data for model training must be confirmed before use. If in doubt, seek confirmation before proceeding. |
| Representativeness Assessed | Evaluation data must be assessed against (1) the overall distribution of production usage, and (2) the distribution across individual features where performance variation would be meaningful (e.g., jurisdiction, court level, case type, filing source, document length). Where evaluation data underrepresents a known feature stratum, it must be acknowledged in the Model Card. Evaluation data must also be screened for leakage to preserve the integrity of performance measurements and preventing artificially inflated metrics. Training data must be properly balanced according to machine learning best practices and documented in the Model Card. |
| Use of User Data Disclosed | To the extent user queries and interactions from FLP's production systems are used for model training and system improvements, clear disclosure and documentation must be in place, with opt-out options available upon request. |
Governance Checkpoints
The following checkpoints are required for all AI systems deployed to external users:
| Checkpoint | When Required | Owner |
|---|---|---|
| AI team consult | Before development begins on any AI-driven feature or product | Initiating team + AI team |
| Model Card | Before any model enters production, updated with each subsequent evaluation or release | Model owner |
| Evaluation report | Before production + at each major update | Model owner + reviewer |
| Periodic quality review | At minimum annually per deployed model | Model owner |
If you are using AI as an efficiency tool, be sure to read and adhere to the AI Use Policy.
Questions? Contact: Rachel Gao — rachel@free.law
Model Card Template · This bridge policy expires April 2026 and will be superseded by FLP's official AI Development Policy.