Machine Learning-Based Prediction of Drug–Target Binding Affinity Using Molecular Structures and Protein Sequences: A Comparative Analysis of Davis and KIBA Datasets

Authors

  • Sheo Kumar Department of Computer Science & Engineering, Dr. B. R. Ambedkar National Institute of Technology Jalandhar, Jalandhar, India Author

DOI:

https://doi.org/10.53555/76xc4919

Keywords:

Drug–target binding affinity, Machine learning, Molecular structures, Protein sequences, Davis and KIBA datasets

Abstract

Determination of the drug–target binding affinity (DTA) is a critical part of the drug discovery process, allowing for the fast selection of therapeutic candidates with high binding affinity to a target, thus minimizing the reliance on expensive experimental screening. This paper proposed and evaluated three machine learning algorithms (Random Forest, XGBoost, and LightGBM) for the prediction of binding affinity based on molecular structure and protein sequences from the Davis and KIBA benchmark datasets. Molecular fingerprints and physicochemical descriptors of drug molecules were used, while amino acid composition and sequence-derived features of protein targets were used. The data was preprocessed, features engineered, and models trained and tested using standard regression assessment measures such as Mean Absolute Error (MAE), Root Mean Squared Error (RMSE), coefficient of determination (R²), Pearson correlation coefficient, and Concordance Index. We showed that the method (LightGBM or Random Forest) that could predict best is dependent on the dataset, as LightGBM outperformed Random Forest on the Davis dataset and the opposite was true for the KIBA dataset, showing that algorithm performance is dependent on the dataset and interaction complexity. Further, feature importance analysis showed that molecular descriptors as well as protein sequence derived features were both important in the prediction accuracy. In general, the proposed framework is a reliable, repeatable and computationally inexpensive method for predicting a drug-target affinity that can be used for virtual screening and prioritizing leads in the early drug discovery phase. The results also provide insights into which models are suitable for which benchmark dataset based on its characteristics.

 

References

Downloads

Published

2026-07-28