FLP Wiki

Artificial Intelligence Development Policy

Bridge Policy — Interim Guidance

Effective: March 2026 | Expires: April 2026

Applies to: Engineering, Data, and Product Teams


Foundational Rule: FLP's mission is access to justice. Every AI system we deploy for external users must meet a higher standard than internal tooling — because failures affect real people, real legal outcomes, and public trust in our data. If a system cannot be evaluated, explained, and monitored, it is not ready to deploy.


Core Principles

Principle Description
Consult the AI Team First Before beginning development on any AI-driven feature or product, teams must consult with FLP's internal AI team. This ensures alignment on approach, awareness of prior work, and early identification of risks before design decisions are locked in.
Accuracy Over Speed Legal information errors have real-world consequences. No model enters production without documented evaluation against a held-out test set representative of the deployment context.
Clear Documentation and Reproducibility Developers have latitude to explore design approaches, modeling strategies, and evaluation methods. That freedom comes with an obligation: key design choices, tradeoffs, and their rationale must be documented so they can be reviewed, reproduced, and learned from.
Transparency to Users All AI-generated content served to external users must be clearly labeled as such. Users must be able to distinguish AI output from verified legal data.
Human Oversight at High-Stakes Outputs Outputs that inform legal decisions must have defined quality thresholds and human review or appeal pathways.
Proportionate Safeguards The level of pre-deployment review, monitoring, and documentation required scales with the stakes of the system.

System Development Standards

Standard Requirement
Training & Evaluation Data Models must be trained and finetuned following machine learning best practices and evaluated on data representative of actual production usage.
Documentation & Reproducibility All AI-powered systems, pipelines, and workflows must have a completed Model Card to ensure reproducibility. All subsequent major changes and upgrades must be reflected in the Model Card.
Periodic & Ongoing Evaluation Deployed models and pipelines must be monitored for performance degradation and formally re-evaluated at least annually, or sooner when a significant distribution shift is detected, a major update is released, or the system's scope changes. Evaluation results must be recorded in the Model Card.
Bias and Fairness AI systems must be developed with bias and fairness in mind, consistent with FLP's mission of equitable access to justice. Relevant considerations and any known limitations must be documented in the Model Card.

Data Standards

All datasets used to train, finetuned, and evaluate models for external-facing systems must meet the following minimum standards:

Standard Requirement
Provenance Documented Source, collection method, date range, and any applicable licenses or use restrictions must be recorded in the Model Card.
PII Screened Training data must be screened for private personally identifiable information, including names, contact information, and any data from sealed, expunged, or restricted court records. Data failing screening must be remediated or excluded.
Copyright Reviewed The right to use potentially copyrighted data for model training must be confirmed before use. If in doubt, seek confirmation before proceeding.
Representativeness Assessed Evaluation data must be assessed against (1) the overall distribution of production usage, and (2) the distribution across individual features where performance variation would be meaningful (e.g., jurisdiction, court level, case type, filing source, document length). Where evaluation data underrepresents a known feature stratum, it must be acknowledged in the Model Card. Evaluation data must also be screened for leakage to preserve the integrity of performance measurements and preventing artificially inflated metrics. Training data must be properly balanced according to machine learning best practices and documented in the Model Card.
Use of User Data Disclosed To the extent user queries and interactions from FLP's production systems are used for model training and system improvements, clear disclosure and documentation must be in place, with opt-out options available upon request.

Governance Checkpoints

The following checkpoints are required for all AI systems deployed to external users:

Checkpoint When Required Owner
AI team consult Before development begins on any AI-driven feature or product Initiating team + AI team
Model Card Before any model enters production, updated with each subsequent evaluation or release Model owner
Evaluation report Before production + at each major update Model owner + reviewer
Periodic quality review At minimum annually per deployed model Model owner

If you are using AI as an efficiency tool, be sure to read and adhere to the AI Use Policy.


Questions? Contact: Rachel Gao — rachel@free.law

Model Card Template · This bridge policy expires April 2026 and will be superseded by FLP's official AI Development Policy.

53 views Last updated 3 months ago
Creator: rachel