The protein-LLM hype problem: what two new studies actually found

Two independent teams — UT Austin, and a Microsoft / NVIDIA / Profluent / Caltech / Duke / Oxford collaboration (FLIP2) — reached the same conclusion: protein language models rarely beat simple baselines when real experimental data exists, and their rankings collapse to near-random when pushing toward genuinely new function.

Published
read time
category
Jinbei Li profile image.
Jinbei Li - Founder, CEO & CSO of ENZIDIA

Evidence continues to accumulate against the protein-LLM hype. Two rigorous studies from recent weeks caught the attention of our team. Separate groups, same conclusion.

One from UT Austin (Ellington and Wilke labs), one from Microsoft / NVIDIA / Profluent / Caltech / Duke University / University of Oxford (FLIP2).

What they found:

- 𝐃𝐚𝐭𝐚 𝐢𝐬 𝐤𝐢𝐧𝐠. 𝘞𝘩𝘦𝘯 𝘺𝘰𝘶 𝘩𝘢𝘷𝘦 𝘳𝘦𝘢𝘭 𝘦𝘹𝘱𝘦𝘳𝘪𝘮𝘦𝘯𝘵𝘢𝘭 𝘥𝘢𝘵𝘢, expensive models rarely earn their cost. FLIP2 showed plain linear regression often matched or beat fine-tuned protein language models. A fine-tuned pLM was the best option on fewer than half the tasks tested.

- 𝐙𝐞𝐫𝐨-𝐬𝐡𝐨𝐭 𝐢𝐬 𝐚 𝐟𝐢𝐥𝐭𝐞𝐫, 𝐧𝐨𝐭 𝐚 𝐫𝐚𝐧𝐤𝐞𝐫. 𝘞𝘩𝘦𝘯 𝘺𝘰𝘶 𝘥𝘰𝘯'𝘵 𝘩𝘢𝘷𝘦 𝘦𝘹𝘱𝘦𝘳𝘪𝘮𝘦𝘯𝘵𝘢𝘭 𝘥𝘢𝘵𝘢, the models can flag a broken protein, but inside the working set, when your goal is to find the best variants, their ranking drops to roughly random. (This is an obvious consequence when you understand the training data being unlabled.)

- 𝐏𝐮𝐬𝐡 𝐭𝐨𝐰𝐚𝐫𝐝 𝐠𝐞𝐧𝐮𝐢𝐧𝐞𝐥𝐲 𝐧𝐞𝐰 𝐟𝐮𝐧𝐜𝐭𝐢𝐨𝐧 𝐚𝐧𝐝 𝐭𝐡𝐞𝐲 𝐛𝐫𝐞𝐚𝐤. On mutations meant to redirect a protein to a new substrate or new chemistry, top models came out slightly negatively correlated with measured results. In other words, they point you at the wrong variants. These models have only ever seen what evolution already built.

- 𝐓𝐡𝐞𝐲 𝐦𝐨𝐬𝐭𝐥𝐲 𝐚𝐠𝐫𝐞𝐞 𝐰𝐢𝐭𝐡 𝐞𝐚𝐜𝐡 𝐨𝐭𝐡𝐞𝐫, 𝐧𝐨𝐭 𝐰𝐢𝐭𝐡 𝐫𝐞𝐚𝐥𝐢𝐭𝐲. Predictions correlate more across models than with ground truth, because they all read the same natural sequence data.

The conclusion: the bottleneck is not the algorithm. It is high-quality labeled data. Both papers more or less end there.

None of this is new. It gets shown again and again in rigorous works in the field. It just rarely seems to reach the ecosystem beyond the real scientists, where "foundation model for biology" keeps being hyped while real breakthroughs in data generation barely register a ripple.

The best case scenario I hope for: the ecosystem catches up to what rigorous science is finding, gets a hold of the mania, and directs capital toward real innovations rather than the next hype company.

The worst case I fear: things go on unchanged until five years from now, when the hyped claims don't materialize, and capital loses confidence, -- not just in the hype companies, but in the entire field of proteins, enzymes, and biotech. We've seen that in synthetic biology thanks to the disservice of a few irresponsible hype players.

Regardless of what happens, ENZIDIA will keep building the way we believe is right and responsible. I just hope there isn’t this much waste of resource and talent when there are pressing challenges to be solved for humanity.

Discussion credit to: Bruce Wittmann, Adam Meyer

View post on LinkedIn >< Back to all posts