
Anthropic says its Claude artificial intelligence models have successfully designed new protein binders against 14 of 15 tested targets, marking a significant expansion of generative AI into early-stage drug research.
The experiment involved Claude Opus 4.8 and an unreleased model called Mythos Preview. The models were asked to create small proteins capable of attaching to selected biological targets, a process widely used in pharmaceutical research, diagnostics and biotechnology.
The designs were physically produced and tested in laboratories operated by Adaptyv Bio and Twist Bioscience. Of the 1,320 protein designs evaluated, 354 successfully bound to their intended targets. That produced an overall laboratory-confirmed success rate of about 26.8%.
Anthropic said the models generated working binders for 14 targets. A fifteenth target produced no successful binder, while results for another target were excluded because the protein displayed aggregation and non-specific binding during testing.
Success Rates Exceed Common Protein Design Benchmarks
Claude’s performance varied depending on the model and how the design campaign was organised.
When multiple targets were handled together during a 48-hour session, Mythos Preview recorded a hit rate of 26.7%, while Opus 4.8 achieved 22.6%. Anthropic said conventional protein-design campaigns commonly produce hit rates of around 10% to 15%.
Mythos Preview performed better when it was assigned one target at a time. Across separate 24-hour sessions, the model reached an overall hit rate of 35.1%.
The AI systems were instructed to create 30 candidate binders for each target. Claude selected possible binding sites, generated protein sequences and structures, ran computational optimisation cycles and screened the resulting candidates before they were sent for laboratory testing.
The models did not perform the entire process using language-based reasoning alone. Claude coordinated several publicly available specialist tools for protein structure generation, sequence design, folding prediction and candidate ranking. It effectively acted as an autonomous research agent managing a complex computational workflow.
The campaign involved substantial computing resources. The multi-target tests were allowed up to 12,500 Nvidia H100 GPU hours over 48 hours. Separate single-target runs received up to 2,500 H100 hours for each target.
Human involvement after the initial prompt was limited mainly to granting system permissions, resolving infrastructure issues and arranging laboratory validation.
Claude Produces High-Affinity Binders
Several of the AI-designed proteins compared favourably with results from earlier protein-design competitions.
Against RBX1, a protein involved in the controlled destruction of regulatory proteins, Mythos Preview recorded a 40% hit rate. Previous competition participants had produced a hit rate of 3.7% for the same target. Claude’s top design also showed stronger binding than the winning entry from that competition.
The models produced high-affinity binders for at least six targets and matched or exceeded the best previously reported binding strength for at least four.
Laboratory results also showed that about 95% of Claude’s designs could be successfully expressed, meaning the digital sequences could be turned into physical proteins for testing.
For TREM2, the models achieved a hit rate of 80%, compared with 38.3% in an earlier design competition. On 15-PGDH, Claude produced a binder with an affinity of 33.4 nanomolar, improving sharply on a previous result of 1.7 micromolar.
Opus 4.8 also generated binders for TNF-alpha, an inflammation-related protein targeted by several established medicines. Some of those designs bound to human, mouse and cynomolgus monkey versions of the protein, a potentially useful property when candidates move into animal studies.
Mythos Preview, despite performing better overall, was unsuccessful on the same target. Anthropic said it had not established why the older Opus model performed better in that case.
Laboratory Testing Remains Essential
The results do not mean Claude has independently created finished medicines.
Protein binding is an early step in drug development. A successful candidate must still demonstrate stability, selectivity, manufacturability, safety and the intended biological effect. It must then pass cell studies, animal testing and several stages of human clinical trials.
Anthropic also noted that the small protein binders used in the experiment are not currently a standard drug format. The trial was designed primarily to test whether a general-purpose AI model could coordinate existing scientific tools and deliver candidates that worked in a physical laboratory.
The work was an open-loop experiment. Claude produced the designs, after which external laboratories tested them. The model did not receive the results and automatically launch a second round of improved candidates.
Closing that loop is expected to be the next step. Such a system would design proteins, review laboratory results, identify failures and refine the next batch without restarting the process manually.
Claude Also Analyses Chemistry Data
Anthropic separately tested Claude Opus 5 on analytical chemistry work involving nuclear magnetic resonance and liquid chromatography-mass spectrometry data.
Given raw files and short instructions, Claude completed two analyses in 23 minutes and 19 minutes. Its purity calculation of 96.4% closely matched the contract laboratory’s figure of 96.33%.
The model also generated a written report and proposed a follow-up experiment similar to one independently selected by the laboratory. Anthropic said these results show that AI systems can handle routine scientific analysis alongside computational design work.
Bottom Line
Claude’s protein-design trial produced 354 confirmed binders from 1,320 candidates and delivered working designs against 14 tested targets. The results indicate that general AI models can coordinate specialist scientific software and manage parts of a protein-engineering campaign with limited human direction.
The experiment remains far removed from producing an approved medicine. Its immediate significance lies in reducing the specialist effort required to generate and screen early-stage protein candidates, while leaving laboratory validation and the longer drug-development process firmly in human hands.