About
How new.legal evaluates legal AI
We do not publish a numeric score. A two-digit ranking of Harvey against an eDiscovery tool, or a copilot against a multi-agent workflow, would look precise and be mostly theatre. Until we have a defensible benchmark, we show evidence coverage and classify autonomy ourselves.
Autonomy
This is an editorial classification. Vendors cannot set it, and calling a product “agentic” in marketing does not move it up the ladder. Copilot is not a slight; it is often the right tool.
- AI-enabled tool
- Software that uses models for a legal task without planning a multi-step job. Search, extract, classify.
- Copilot / assistant
- Assists a lawyer inside a document or research environment. The human drives each step.
- Agent
- Plans and executes a bounded legal task — research a question, review a set — with lawyer-in-the-loop gates.
- Multi-agent
- Specialised agents coordinate on a matter: retrieve, draft, check. A person still owns the output.
- Autonomous workflow
- Completes an end-to-end work product with limited intervention. The rarest class, and the one we will not award from a homepage claim.
Evidence coverage
Evidence coverage describes how well a listing is supported by attributable sources. It is not a product rating. Well documented, Documented or Thin file is not quality, recommendation, performance, market ranking or vendor verification — only how much we have on file.
Only items with a public source URL count. Unsourced claims are labelled “Source not publicly linkable” and weigh 0. A vendor announcing its own product is not independent evidence. Independence is taken from the source, then the weight is halved if the item is Historic (publication date more than 24 months old, or a source whose date cannot be read from the page).
Independence
- Independent
- Trade press, regulator, court or mainstream press. A current sourced independent item weighs 1.0.
- Customer
- The adopting organisation’s own site — a law firm or in-house team describing its own deployment. Current sourced weight 0.75.
- Vendor
- The company’s own newsroom, blog, homepage or PR wire. Current sourced weight 0.4.
- Aggregator
- A database, directory or marketplace listing — Crunchbase, CB Insights, Legaltech Hub — not independent reporting. Current sourced weight 0.5.
Weights
| Source | Current | Historic (>24 months) |
|---|---|---|
| Independent | 1.0 | 0.5 |
| Customer | 0.75 | 0.375 |
| Vendor | 0.4 | 0.2 |
| Aggregator | 0.5 | 0.25 |
| Unsourced | 0 | 0 |
The weighted score is the sum of those weights. A reader can reproduce any grade from the sources on the profile. The evidence header lists every non-zero class on the file, flags historic items (half weight), and shows the computed weight next to the band.
Vendor, customer and aggregator weight is capped at the weight of current independent sources. Excess is discarded, not banked. A product with no independent sources scores zero, however many vendor pages it publishes. A company cannot improve its own grade by publishing.
One event is one weighted item. Items that describe the same fact — same subject, event type and month — are clustered. The highest-independence, earliest-dated source carries the weight; the others sit under it as corroborating links at weight 0.
- Well documented
- Weighted score of 2.5 or more, and at least two current independent sources. One independent current item plus vendor pages can clear 2.5 on weight alone — that is not enough. The top grade needs a second independent source, not a homepage plus a single press mention.
- Documented
- Weighted score of at least 1.0, and not Well documented. One current independent source produces this grade (1.0). Four vendor pages on top of that independent source cannot push it over Well documented: the vendor cap holds the extra weight at the independent total. Two vendor homepages with no independent source score 0.
- Thin file
- Weighted score below 1.0. A vendor homepage alone is Thin file. A file with no independent sources is Thin file, regardless of vendor material. Unsourced claims never lift it.
Ratings are recorded in an evidence history table before they change: the reason is item added, item sourced, item expired, weighting revised or band criteria revised. Historic items remain on the profile so the file can be read; they simply count less.
Dates
A date is read when it appears on the source page (byline or dated URL). It is inferred when we assigned a date from context — a contemporaneous report, a last-checked observation — not from the page itself. Where no publication date can be read, the item is labelled “Date not stated” and counts as Historic. Where a page date and an in-body date disagree, we take the earlier in-body date and mark it inferred. A restated wire date in the body is not a new publication.
What we take, and what we skip
Independent trades and aggregators are polled every night. Vendor and customer newsrooms are polled weekly. Vendor posts that only record a brand tie-up, sponsorship, event attendance, hiring round or thought-leadership piece — nothing about a product capability, an adoption or a funding fact — are rejected as promotional. A partnership or integration announcement still counts.
Organisations
Organisation profiles are independent intelligence records of law firms, in-house legal teams, vendors, academia, professional bodies and government. They are not rankings, not paid, and not claimable. Suggest a correction if a fact is wrong. The organisation type is institutional identity (what it is), not a label for legal-AI activity. Evidence coverage uses the same weighting as products — Well documented, Documented or Thin file — and an empty file has no grade. Zero independent sources scores zero. Still not a rating.
Firm Zero Framework™
Firm Zero Framework™ classifications (legal capability → sub-capability → use case) are new.legal analysis, not source facts. Where the source uses different words, we keep the original terminology alongside the classification. A product is not a child of a use case; mappings are many-to-many and evidence-backed.
Reports and publications
A substantive report appears as an intelligence update with a direct source link. Publishing a report is evidence that the organisation published the report. It is not automatically evidence of implementation, product adoption, outcomes or internal AI deployment.
What we use
- Public evidence
- Company announcements, independent reporting, product documentation and regulatory filings, each labelled and linked when a public URL exists. We do not invent sources.
- Vendor-supplied information
- After a profile is claimed, the vendor may correct facts, set the product URL, add a demo request and state pricing. Claimed and Verified never move a listing up a list. Autonomy stays editorial.
- new.legal view
- The labelled paragraph on a profile is ours. It is opinion, can be wrong, and is separate from the factual description.
- Last checked
- Legal AI products move quickly. Each listing shows when we last reviewed the file.
- User signals
- Profile views, outbound product visits, compares and searches will later feed a genuine “Trending” list. They are not used to order Notable agents today. Nobody can pay for outbound traffic.
- Automated analysis
- When you ask a question, Grok extracts intent and ranks matches for that query. That order is match quality, not a market-wide score. Outputs can be wrong. It is not legal advice.
What we will not do
Independent. Nobody can pay to rank higher. Sponsored events are paid and labelled on the advertisement.
When a numbered methodology exists, it will live on this page, with the evidence count and last checked date on each profile. Until then we would rather look unfinished than scientific.