FireProtDB 2.0: why more data do not make every prediction more accurate
FireProtDB 2.0 expands measurement datasets and variant types. Compatible conditions and record provenance matter more for practical interpretation than database size. Foldiff’s current lookup does not imply use of all version 2.0 data.
What changed in the database?
In the 2026 release paper the authors describe moving from a single-substitution database to a more flexible schema. FireProtDB 2.0 contains approximately 5.47 million experiments and supports multiple substitutions, insertions, deletions and WT values.
These numbers describe the authors’ database, not Foldiff’s training or lookup coverage. The current product uses its existing integration and saved evidence. Adopting the new schema requires separate validation of record matching.
Why records are not interchangeable
Databases combine absolute and relative endpoints, constructs and methods. The authors highlight large small-domain datasets and proteolysis-derived estimates. These cannot automatically be treated as direct measurements of the same thermodynamic property.
The practical question is whether a record matches your sample. Check sequence, numbering, tags, domains and original conditions. Matching substitution names alone do not guarantee matching experimental context.
Read the Foldiff evidence lookup
The Evidence tab shows reported measurements separately from predictors. Disagreement with a calculation calls for reviewing conditions and constructs. A missing record must not be treated as confirmation of stability.
The saved subtilisin example shows eight measured single substitutions; two match the checked list and disagree with the predicted direction. This describes one lookup, not an error percentage for the entire platform or FireProtDB.
Avoid leakage between training and evaluation
Evaluation can be too optimistic when similar proteins or the same measurements appear in both training and evaluation sets. Protein-family splits, explicit data provenance and independent new experiments are therefore important.
The current FireProtDB lookup is not claimed as independent scientific validation of Foldiff’s main model. Server numerical checks establish implementation reproducibility, not transfer to your enzyme. A laboratory pilot remains a separate source of evidence.
What to keep for future analysis
- Accession and actual construct sequence.
- Endpoint definition and sign convention.
- Conditions, units, batch identifier and replicate type.
- Raw WT and variant measurements, not just final means.
- Calculation version and links to the records used.
This set lets you reanalyze a series after database updates and distinguish data changes from algorithm changes. The beta exports measurements as CSV and JSON; automatic retraining is not enabled.
Sources and review
- FireProtDB 2.0, Nucleic Acids Research, 2026 release
- Official FireProtDB database
- Foldiff data and limits
Links checked on 2 October 2026. Prepared with AI assistance and checked against Foldiff code; no external scientific review has been performed.
Try the workflow on a public example
Explore candidates, create a series and analyze illustrative measurements.
Try Foldiff ↗