OpenAI Launches Astra for Law With a 230-Million-Document Legal Index, and Harvey Is Already Building on It
OpenAI launched Astra for Law on September 17, pairing GPT-6 Astra with a 230-million-URL legal research index, and rival legal AI startups Harvey and Legora are already signed on as API customers rather than competitors.
OpenAI launched Astra for Law on September 17, a version of its GPT-6 Astra model wrapped around a legal research index spanning more than 230 million URLs of US case law, statutes, regulations, court rules, and administrative decisions. The number that matters more than the index size is the benchmark result: on Vals AI's Legal Research Bench, Astra for Law passed 54 percent of 200 test questions, against 38.7 percent for the same underlying GPT-6 Astra model using nothing but plain web search. That's the actual product being sold here, not the model itself but the model paired with a search corpus built specifically so it doesn't have to guess where the accurate version of a statute lives.
Where the index comes from
A large share of the case-law data comes through a partnership with the Free Law Project, the nonprofit behind CourtListener, which OpenAI says covers more than 99.9 percent of published US precedential case law. That detail matters more than it might seem. Legal AI tools have a well-documented history of hallucinating citations to cases that don't exist, a problem serious enough that multiple US judges have sanctioned lawyers this decade for filing briefs with fabricated case law generated by an AI assistant. Grounding the model in a specific, near-complete, dated corpus rather than an open web search is OpenAI's answer to that failure mode, even if it doesn't eliminate the risk entirely.
Rollout, and who's already building on it
Astra for Law starts in limited release through OpenAI's Trusted Access program inside ChatGPT and Codex, showing up as "GPT-6 Astra Law" for the law firms selected in. An API version is coming later, with no pricing or date attached yet. On the data-handling side, OpenAI is offering zero data retention on the API and excluding ChatGPT Enterprise usage from human review by default, both baseline requirements for any law firm handling privileged client material.
The more interesting detail is who's already plugged into it. Harvey Technologies and Legora AB, two of the best-funded legal AI startups building their own products on top of foundation models, are listed as API customers building on Astra for Law rather than as competitors OpenAI is trying to displace. That's the same pattern OpenAI has run in other verticals: rather than only selling directly to end users, it's positioning itself as the underlying research layer that specialist legal AI companies build their own products on top of, collecting revenue either way. Thomson Reuters is integrating its own HighQ matter context into the system and has previewed a connector to CoCounsel, its existing legal AI product, which suggests even a company with a competing legal AI offering sees more upside in plugging into OpenAI's index than staying separate from it.
Twenty-six partner-built plugins connect ChatGPT to tools firms already use day to day, including Relativity, Clio, and iManage, plus nine additional plugins built by outside lawyers and engineers contributing 47 custom skills between them. Some individual firms built their own tools on top of the platform ahead of launch: Sullivan & Cromwell built an agreement analyzer that checks contracts against the firm's own playbooks and precedent library, Ropes & Gray built a deal diligence tool, and Cooley built a product called GO Public aimed at IPO preparation work. None of that is generic chatbot functionality. It's the kind of workflow-specific tooling that only gets built once a firm trusts the underlying data enough to wire real deal work into it.
The governance question nobody skips in legal AI
OpenAI is working with law firm Latham & Watkins specifically on governance design, covering information permissions, ethical walls, and client instructions, the unglamorous compliance infrastructure that determines whether a firm's conflict-of-interest and confidentiality obligations survive contact with an AI tool that can, in principle, see everything in a shared workspace. Ethical walls in particular are a hard technical problem for any AI system operating across a firm's full document set: the model needs enough context to be useful on a given matter while being reliably blocked from information tied to a client the firm represents adversely elsewhere. Getting a major law firm to co-design that governance layer, rather than having OpenAI build it in isolation and hope it satisfies bar association scrutiny later, is a meaningfully different approach than most enterprise AI rollouts bother with.
The part still worth being skeptical about
A 54 percent pass rate on a 200-question benchmark is a real improvement over plain web search, and it's also, read the other way, a tool that gets nearly half of a curated legal research benchmark wrong. Legal research is a field where a confidently wrong citation can end up in a filed court document, and OpenAI's own framing, that this is a research assistant sitting inside human-supervised workflows rather than an autonomous legal advisor, is doing real work to manage that risk. Whether firms keep up that level of human review once a tool with Sullivan & Cromwell and Ropes & Gray's names attached to custom features starts feeling reliable in daily use is the question that won't get answered by a benchmark score, only by how carefully the profession polices itself once the novelty wears off.

