In Molecular diversity
Cysteinyl leukotrienes 1 (CysLT1) receptor is a promising drug target for rhinitis or other allergic diseases. In our study, we built classification models to predict bioactivities of CysLT1 receptor antagonists. We built a dataset with 503 CysLT1 receptor antagonists which were divided into two groups: highly active molecules (IC50 < 1000 nM) and weakly active molecules (IC50 ≥ 1000 nM). The molecules were characterized by several descriptors including CORINA descriptors, MACCS fingerprints, Morgan fingerprint and molecular SMILES. For CORINA descriptors and two types of fingerprints, we used the random forests (RF) and deep neural networks (DNN) to build models. For molecular SMILES, we used recurrent neural networks (RNN) with the self-attention to build models. The accuracies of test sets for all models reached 85%, and the accuracy of the best model (Model 2C) was 93%. In addition, we made structure-activity relationship (SAR) analyses on CysLT1 receptor antagonists, which were based on the output from the random forest models and RNN model. It was found that highly active antagonists usually contained the common substructures such as tetrazoles, indoles and quinolines. These substructures may improve the bioactivity of the CysLT1 receptor antagonists.
Wang Hongzhao, Qin Zijian, Yan Aixia
Classification models, Cysteinyl leukotrienes 1 (CysLT1) receptor, Deep neural network (DNN), Morgan fingerprint, Random forest (RF), Recurrent neural network with self-attention (self-attention RNN), Structure–activity relationship (SAR)