- A docking score is a protocol-specific surrogate, not an experimental binding free energy.
- Pose recovery, affinity prediction, congeneric ranking and screening enrichment require separate validation.
- Its defensible role is to generate falsifiable poses and experimental priorities—not convert a raw score into Kd.
1. The same unit does not imply the same physical quantity
For a simple equilibrium, the standard binding free energy can be related to dissociation through \Delta G^\circ = RT\ln(K_d/C^\circ). This quantity describes a free-energy difference between full unbound and bound ensembles. It includes solvent reorganisation, conformational changes in both partners, translational, rotational and conformational entropy, and sometimes coupled changes in protonation.[6] A docking score is instead a fast surrogate designed to guide search and ranking. AutoDock Vina, for example, combines a simple scoring function with stochastic global search and local optimisation. Its documentation reports “predicted binding affinity” in kcal/mol while also describing the engine as based on a simple scoring function.[1,2] Matching units are therefore not evidence of one-to-one thermodynamic calibration, and scores from different engines do not share a universal scale.
2. Pose recovery and affinity prediction are different tests
Docking first samples poses and then scores them. A scoring function cannot rescue a native-like pose that was never sampled; conversely, sampling can generate a good pose that scoring fails to rank first. CASF-2016 formalised this distinction with four endpoints: docking power for pose selection, scoring power for absolute affinity, ranking power within a target series, and screening power for separating binders from non-binders. Across 285 curated complexes and 25 scoring functions, pose-selection performance was generally more promising than scoring, ranking, or screening performance.[7] A foundational multi-target assessment likewise found no single program that performed consistently well for every target.[3] Successful redocking is useful, but it validates only a narrow pose-reproduction task in a known holo pocket; it does not establish quantitative affinity accuracy.
3. Molecular state and protocol are part of the prediction
The result can change when an apo structure is replaced with a holo conformation, a side chain is flipped, a missing loop is rebuilt, or a metal, cofactor, or structural water is removed. Ligand protonation, tautomerism, stereochemistry, and charge assignment are equally consequential. A systematic comparison showed that docking programs could struggle to select the correct protomer or stereoisomer from plausible alternatives.[4] Modern affinity-benchmark guidance therefore treats bond orders, stereochemistry, ionisation and tautomeric states, pocket protonation, cofactors, and crystallographic waters as required preparation decisions.[8] Scores generated from different molecular states may not answer the same chemical question even when every command-line option is identical. The risk is particularly acute for metal coordination, covalent mechanisms, large induced-fit changes, and water-mediated recognition.
4. Experimental affinity is conditional and noisy too
Kd, Ki, and activity readouts arise from specific constructs and assays; they are not context-free properties of a ligand. Combining public measurements across laboratories, temperatures, substrate concentrations, constructs, and analysis procedures adds heterogeneity. An analysis of public Ki data directly quantified this experimental uncertainty.[5] A model trained on mixed labels can consequently learn assay differences or dataset correlates in addition to molecular recognition. Target-level calibration is strongest when the reference measurements use a consistent assay format and conditions, preserve units and censoring limits, and report replication or uncertainty. Heterogeneous Kd, Ki, and functional measurements should not be merged as if they were interchangeable ground truth.[5,8]
5. Machine learning changes the approximation, not the need for an applicability domain
Machine-learning scoring functions can learn valuable patterns from large complex datasets, but random splits frequently place related ligands, proteins, or pockets on both sides of the train–test boundary. A 2022 evaluation of 12 representative models observed sequential performance loss from random cross-validation to sequence-based and then pocket-Pfam splits, exposing weak cross-target generalisation.[9] A 2023 study similarly warned that a model can exploit dataset shortcuts rather than interatomic recognition and used strict similarity filtering to interrogate what had actually been learned.[10] In a 2025 comparison spanning CASF, low-ligand-bias, congeneric-series, and out-of-distribution tests, Pearson correlations for two representative models fell from 0.83 to 0.57 and from 0.76 to 0.55 on the OOD test. The authors used these examples to show why a single benchmark can overstate operational performance.[12] This evidence does not make AI scoring useless; it makes training-set proximity, strict splits, external targets, and uncertainty part of every credible claim.
6. A top-ranked pose still needs chemical and experimental review
PoseBusters showed that ligand RMSD alone can miss abnormal bonds, internal clashes, severe protein–ligand overlap, or implausible strain in both classical and deep-learning docking outputs. Generalisation also weakened for targets unlike the training proteins.[11] Before selecting compounds, teams should therefore inspect internal geometry, clashes, strain, unsatisfied polar groups, key interactions, and the pocket environment, then ask whether conclusions survive changes in random seed, receptor conformation, and scoring model. Orthogonal experiments answer different questions: biochemical or biophysical assays test binding and affinity, structural experiments test the pose, and cellular assays test functional consequences. A docking score cannot substitute for all three.
Common misconceptions
- “A score of −10 kcal/mol can be converted directly to a much tighter Kd than −8.” A raw score is not a thermodynamic measurement. An empirical mapping is meaningful only after target-, protocol-, and assay-specific external calibration.
- “Scores are comparable across programs because they use kcal/mol.” Feature terms, weights, zero points, search procedures, and preparation rules differ. Any combination requires recalibration on the same validation set.
- “Successful redocking proves performance on new scaffolds.” Redocking in a known holo pocket is a narrow in-distribution test; new scaffolds, apo structures, and unseen pockets impose harder generalisation demands.
- “An AI confidence or affinity head has solved the physics.” Outputs remain sensitive to training-set similarity, label noise, and pose validity and require physical checks plus prospective experiments.[9–12]
Conclusion
The most productive interpretation of a docking score is as a low-cost, auditable hypothesis-ranking device—not an uncalibrated affinity instrument. A credible workflow places the number inside input quality control, an explicit applicability domain, controls, repeated calculations, pose inspection, and experimental validation. Docking can help decide what to test first; by itself it cannot prove that a molecule binds, quantify how tightly it binds, or establish cellular or clinical efficacy.
Decision checklist: trust a ranking only within its validated frame
Define the endpoint first
pose generation, virtual-screen enrichment, congeneric ranking, or absolute affinity prediction.
Check 2
Lock and record receptor state, pocket definition, ligand protomer/tautomer, cofactors, metals, structural waters, and covalent assumptions.
Check 3
Redock a reference ligand and test known actives plus weak/inactive controls or decoys; do not rely on one RMSD value.
Check 4
Repeat stochastic runs and perturb consequential settings; examine rank, interaction, and pose stability rather than the single best run.
Check 5
Calibrate only within the same target, related chemistry, fixed protocol, and homogeneous assay, with a genuinely held-out external set.
Check 6
Apply geometry, clash, strain, protonation, and medicinal-chemistry review; retain diverse backups rather than only the top N scores.
Check 7
Include positive, negative, and mechanistic controls in the experiment and describe the docking result as a priority or falsifiable pose hypothesis.
Interpretive boundaries to retain
- Claim: Initial prioritisation under one fixed protocol · Defensible strength: Useful · Minimum supporting evidence: Quality-controlled inputs, internal controls, replicates, pose review
- Claim: A possible binding mode or key residue · Defensible strength: Hypothesis · Minimum supporting evidence: Agreement across conformations/methods and a falsifiable structural or mutational test
- Claim: Ranking close analogues for one target · Defensible strength: Conditional · Minimum supporting evidence: Homogeneous assay calibration, strict holdout set, stated applicability domain
- Claim: Absolute Kd/Ki, cross-target comparison, or certain potency · Defensible strength: Not from docking alone · Minimum supporting evidence: A validated quantitative/free-energy workflow plus experimental measurement
- Claim: Cellular efficacy, selectivity, safety, or clinical benefit · Defensible strength: Outside scope · Minimum supporting evidence: Additional biochemical, cellular, PK/toxicology, and clinical evidence
Verified sources
- AutoDock Vina project. AutoDock Vina documentation: Molecular docking program. Official documentation, release 1.2.x. Accessed 2026-08-16.2026
- Eberhardt J, Santos-Martins D, Tillack AF, Forli S. AutoDock Vina 1.2.0: New Docking Methods, Expanded Force Field, and Python Bindings. *J Chem Inf Model.* 2021;61:3891–3898. DOI: [10.1021/acs.jcim.1c00203](2021 · DOI 10.1021/acs.jcim.1c00203
- Warren GL, Andrews CW, Capelli A-M, et al. A Critical Assessment of Docking Programs and Scoring Functions. *J Med Chem.* 2006;49:5912–5931. DOI: [10.1021/jm050362n](2006 · DOI 10.1021/jm050362n
- ten Brink T, Exner TE. Influence of Protonation, Tautomeric, and Stereoisomeric States on Protein–Ligand Docking Results. *J Chem Inf Model.* 2009;49:1535–1546. DOI: [10.1021/ci800420z](2009 · DOI 10.1021/ci800420z
- Kramer C, Kalliokoski T, Gedeck P, Vulpetti A. The Experimental Uncertainty of Heterogeneous Public Ki Data. *J Med Chem.* 2012;55:5165–5173. DOI: [10.1021/jm300131x](2012 · DOI 10.1021/jm300131x
- Pantsar T, Poso A. Binding Affinity via Docking: Fact and Fiction. *Molecules.* 2018;23:1899. DOI: [10.3390/molecules23081899](2018 · DOI 10.3390/molecules23081899
- Su M, Yang Q, Du Y, et al. Comparative Assessment of Scoring Functions: The CASF-2016 Update. *J Chem Inf Model.* 2019;59:895–913. DOI: [10.1021/acs.jcim.8b00545](2016 · DOI 10.1021/acs.jcim.8b00545
- Hahn DF, Bayly CI, Boby ML, et al. Best Practices for Constructing, Preparing, and Evaluating Protein-Ligand Binding Affinity Benchmarks [Article v1.0]. *Living J Comput Mol Sci.* 2022;4:1497. DOI: [10.33011/livecoms.4.1.1497](2022 · DOI 10.33011/livecoms.4.1.1497
- Zhu H, Yang J, Huang N. Assessment of the Generalization Abilities of Machine-Learning Scoring Functions for Structure-Based Virtual Screening. *J Chem Inf Model.* 2022;62:5485–5502. DOI: [10.1021/acs.jcim.2c01149](2022 · DOI 10.1021/acs.jcim.2c01149
- Scantlebury J, Vost L, Carbery A, et al. A Small Step Toward Generalizability: Training a Machine Learning Scoring Function for Structure-Based Virtual Screening. *J Chem Inf Model.* 2023;63:2960–2974. DOI: [10.1021/acs.jcim.3c00322](2023 · DOI 10.1021/acs.jcim.3c00322
- Buttenschoen M, Morris GM, Deane CM. PoseBusters: AI-based Docking Methods Fail to Generate Physically Valid Poses or Generalise to Novel Sequences. *Chem Sci.* 2024;15:3130–3139. DOI: [10.1039/D3SC04185A](2024 · DOI 10.1039/D3SC04185A
- Valsson Í, Warren MT, Deane CM, et al. Narrowing the Gap between Machine Learning Scoring Functions and Free Energy Perturbation Using Augmented Data. *Commun Chem.* 2025;8:41. DOI: [10.1038/s42004-025-01428-y](2025 · DOI 10.1038/s42004-025-01428-y
Search updated 2026-08-16. This is an evidence-led narrative methods review, not a registered systematic review or meta-analysis; citations prioritise primary papers, official documentation and standards.
