A virtual screening method for small molecule inhibitors targeting YTHDF1 protein

By obtaining multiple conformations of the YTHDF1 protein, using the bait generated by the DeepCoy algorithm, combining the PLEC fingerprint features and the machine learning algorithm to construct a method, by obtaining the diversity of the YTHDF1 protein, using the bait molecules generated by the deep learning algorithm, combining the PLEC indicators, and implementing multiple constructed specificity scoring functions, a virtual screening method for the YTHDF1 protein targeting multiple conformations of the YTHDF1 protein was screened out, which solved the accuracy and efficiency of the virtual screening of the YTHDF1 protein targeting multiple conformations of the YTHDF1 protein in the existing technology, screened out small molecule inhibitors targeting the YTHDF1 protein, and achieved efficient drug screening.

CN119479782BActive Publication Date: 2025-09-30SHENZHEN UNIV
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202411409808.5
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2024-10-10
Publication Date
2025-09-30
Estimated Expiration
2044-10-10

AI Technical Summary

Technical Problem

Existing technologies have low accuracy and efficiency in screening small molecule inhibitors targeting YTHDF1 protein. Traditional methods cannot fully capture the complex interactions in protein-ligand binding, especially in the case of protein structural diversity and binding site flexibility.

Method used

By obtaining multiple conformations of the YTHDF1 protein, the DeepCoy algorithm was used to generate bait molecules, and the specificity scoring function was constructed by combining the PLEC fingerprint features and machine learning algorithm. High-throughput molecular docking and affinity prediction were performed to screen small molecule inhibitors targeting the YTHDF1 protein.

Benefits of technology

It significantly improved the screening accuracy and efficiency of small molecule inhibitors targeting YTHDF1 protein, provided candidate molecules with high affinity and stability, and laid the foundation for drug development for diseases such as cancer.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119479782B_ABST
    Figure CN119479782B_ABST
Patent Text Reader

Abstract

The present invention belongs to the field of drug design and specifically discloses a virtual screening method for small molecule inhibitors targeting the YTHDF1 protein, comprising: obtaining the YTHDF1 protein structure and generating multiple conformational proteins through homology modeling; obtaining a small molecule that binds to the YTHDF1 protein and generating a bait molecule to form a ligand molecule; docking the ligand molecule with the binding site on YTHDF1 and characterizing the complex through PLEC fingerprinting; using a machine learning algorithm to construct a specificity scoring function, scoring and screening the complexes of the docked compounds to be screened and the YTHDF1 molecule, and then performing affinity prediction to screen small molecule inhibitors targeting YTHDF1. The virtual screening method provided by the present invention can screen YTHDF1-targeting molecules with high affinity on a large scale and accurately, providing candidate molecules for drug development for YTHDF1-related diseases and laying the foundation for clinical trials and drug optimization.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of drug design, and in particular to a virtual screening method for small molecule inhibitors targeting YTHDF1 protein. Background Art

[0002] Tumors are one of the major diseases that threaten human health. Their occurrence and development involve not only gene mutations and deletions but are also closely related to imbalances in epigenetic regulation. Epigenetic modifications primarily regulate gene expression through pathways such as DNA methylation, histone modification, and RNA modification, and abnormalities in these modifications often trigger the occurrence and progression of tumors. N6-methyladenosine (m6A) is one of the most common RNA modifications, regulating gene expression by affecting mRNA splicing, translation, and degradation. Its recognition depends on m6A reader proteins, of which YTHDF1 is a key reader protein and is widely involved in the pathological processes of diseases such as inflammation and cancer.

[0003] The key role of YTHDF1 in cancer and its potential as a potential therapeutic target have prompted researchers to try to develop new anti-tumor drugs by screening its small molecule inhibitors. However, the number of existing YTHDF1 inhibitors is limited, and their screening and identification processes are relatively complicated. Ebselen is the first reported small molecule inhibitor that can bind to the YTHDF1 protein and interfere with its recognition function of m6A-modified RNA. However, the mode of action of Ebselen depends on environmental factors (such as a reducing environment) and may interact with YTHDF1 in a reversible manner under certain conditions. Therefore, it is of great significance to further explore and develop efficient and specific YTHDF1 inhibitors.

[0004] Traditional drug screening typically uses scoring functions based on force fields, knowledge bases, or empirical formulas. While these methods can identify certain active molecules, they cannot fully capture the complex interactions in protein-ligand binding, especially given the structural diversity and flexibility of proteins and binding sites. These limitations significantly reduce the efficiency and accuracy of virtual screening methods.

[0005] Therefore, the existing technology still needs to be improved and developed. Summary of the Invention

[0006] In view of the above-mentioned deficiencies in the prior art, the purpose of the present invention is to provide a virtual screening method for small molecule inhibitors targeting YTHDF1 protein, aiming to solve the problem of how to effectively improve the accuracy and efficiency of screening small molecule inhibitors targeting YTHDF1.

[0007] The technical solutions of the present invention are as follows:

[0008] The present invention provides a virtual screening method for small molecule inhibitors targeting YTHDF1 protein, comprising the steps of:

[0009] (a) YTHDF1 protein structure preparation: The crystal structure of YTHDF1 protein was obtained, and homology modeling of YTHDF1 protein was performed using the SWISS-MODEL platform to generate YTHDF1 protein in multiple conformations;

[0010] (b) Construction of ligand molecule dataset: Active small molecules and inactive small molecules that bind to the YTHDF1 protein are obtained, and inactive decoy molecules are generated based on the active small molecules to construct a ligand molecule dataset;

[0011] (c) Molecular docking and complex characterization: high-throughput molecular docking was performed between the molecules in the ligand molecule set and the binding sites on the YTHDF1 protein in various conformations using a molecular docking program, and the resulting protein-ligand complex was characterized by PLEC fingerprinting;

[0012] (d) Construction of specificity scoring function: Based on the obtained PLEC fingerprint dataset, a specificity scoring function for YTHDF1 protein was constructed using a machine learning algorithm;

[0013] (e) Virtual screening and affinity prediction: The compounds in the library of compounds to be screened are docked onto YTHDF1 proteins of various conformations to obtain the PLEC fingerprints of the complexes. After scoring and sorting using the specific scoring function, affinity calculation is performed to finally screen out the small molecule inhibitors targeting the YTHDF1 protein.

[0014] Preferably, in step (a), the crystal structure of the YTHDF1 protein is from the Protein Data Bank database, numbered 4RCJ, and up to ten different conformations of the YTHDF1 protein are generated by homology modeling.

[0015] Traditional virtual screening often relies solely on molecular docking of proteins in a single conformation, ignoring the protein's dynamics and conformational changes. This study, using homology modeling of the YTHDF1 protein, generated 10 different conformations. These conformations represent the protein's natural flexibility and dynamics, accounting for its dynamic changes. This allows molecular docking and virtual screening to fully consider the protein's behavior in different states, enabling a more comprehensive assessment of the binding ability of compounds to the YTHDF1 protein, significantly improving the accuracy and credibility of screening results.

[0016] Optionally, in step (b), the active small molecules and inactive small molecules that bind to the YTHDF1 protein are from PubChem, BindingDB and Chembl databases, and duplicate molecules are removed based on the SMILE string.

[0017] Preferably, in step (b), inactive bait molecules are generated based on the active small molecules, specifically comprising: using the DeepCoy algorithm to generate inactive bait molecules with similar physical and chemical properties to the active small molecules but different structures, using the SMILES of the active small molecules as input, first generating 100 bait molecules based on each active small molecule, and then retaining 50 bait molecules for each active small molecule based on the degree of embedding of the active small molecules in space.

[0018] Typically, directly obtaining active and inactive molecules may rely solely on simple structural similarity or collecting known molecules from the literature. However, the present invention uses the deep learning tool DeepCoy to generate inactive decoy molecules that are structurally dissimilar but have matching physicochemical properties. These decoy molecules are used to expand the dataset of inactive molecules to establish a large-scale equilibrium dataset containing active and inactive molecules, thereby ensuring that the model has sufficient robustness and discrimination when identifying active small molecules, further enhancing the ability to distinguish between active and inactive molecules. Therefore, the decoy molecules generated by DeepCoy in the present invention can effectively improve the machine learning model's ability to discriminate active small molecules, making virtual screening more challenging, thereby improving the model's ability to capture truly active molecules.

[0019] Preferably, in step (c), the molecular docking program includes the Smina molecular docking program, and in the Smina molecular docking program, the exhaustiveness parameter is set to 8 and the num_modes value is set to 10.

[0020] Preferably, in step (c), when the protein-ligand complex is characterized by using PLEC fingerprint, the parameter settings of the PLEC fingerprint include: a cutoff distance of The protein depth is 4, the ligand depth is 2, and the vector length is 1092.

[0021] Unlike traditional 2D or 3D structural features, the present invention utilizes extended connectivity fingerprints to characterize protein-ligand complexes in a high-dimensional manner. PLEC fingerprint features can capture the spatial information of protein-ligand interactions and represent the structural details of the complex as a 2048-dimensional high-dimensional feature vector. This approach is more expressive than traditional molecular descriptors and can capture more complex molecular interactions. Therefore, through the use of PLEC features, these high-dimensional features can more comprehensively characterize the spatial configuration and interaction patterns of protein-ligand complexes, providing high-quality input features for subsequent machine learning models, thereby achieving higher accuracy and reliability than traditional virtual screening methods.

[0022] Optionally, in step (d), the machine learning algorithm includes at least one of random forest, extreme gradient boosting, support vector machine and artificial neural network.

[0023] This paper proposes combining at least one of the following machine learning methods: random forest (RF), extreme gradient boosting (XGBoost), support vector machine (SVM), and artificial neural network (ANN) to construct a scoring function. Alternatively, multiple machine learning algorithms can be combined. After training different models through multiple rounds of cross-validation, an ensemble learning method is used to select the optimal model as the final scoring function. This multi-model integration strategy can improve the model's generalization ability and predictive accuracy.

[0024] Preferably, the machine learning algorithm is an artificial neural network, and the artificial neural network is trained and evaluated by cross-validation.

[0025] ANN has significant advantages in processing complex protein-ligand interactions, automatic feature extraction, processing large-scale data, and generalization capabilities. ANN automatically extracts complex features and patterns through multiple layers of neurons and nonlinear activation functions, which is very important in the YTHDF1 protein-ligand complex. This is because the three-dimensional structure of the complex and its interactions with molecules usually have highly complex nonlinear characteristics. ANN extracts meaningful information from these complex structural relationships, thereby improving the accuracy of predictions. In order to improve the generalization ability of the model, the present invention selects an artificial neural network model for training as a scoring function for predicting the activity of the compound. It can more accurately predict the binding activity of the compound to the YTHDF1 target, significantly improving the effectiveness of virtual screening.

[0026] Optionally, in step (e), the compound library to be screened is selected from the ZINC20 database, comprising 45,000 small molecule compounds.

[0027] Optionally, in step (e), the complexes are scored and ranked using the specificity scoring function to screen out the top 1% of candidate compounds, and the affinity of the candidate compounds is calculated using a logistic regression model to ultimately screen out small molecule inhibitors targeting YTHDF1 with high affinity.

[0028] According to the specificity scoring function constructed above, 45,000 compounds in the library of compounds to be screened were screened for activity. The specificity scoring function is based on the previously trained ANN model and is specifically optimized for the target protein, which can more accurately predict the activity of the compound. In order to further confirm the reliability of the screening results and the practical application potential of the compounds, the present invention also predicts the binding affinity of the top 1% of the screened molecules. The binding affinity prediction can help evaluate the binding strength between these compound molecules and the target protein, and then judge their drug development potential. Therefore, introducing affinity prediction into the later stage of virtual screening and coordinating it with the scoring function can ensure that the screened molecules not only have high affinity but also have good stability.

[0029] Beneficial effects:

[0030] The present invention provides a virtual screening method for small-molecule inhibitors targeting the YTHDF1 protein. This method utilizes homology modeling of multi-conformational proteins and the DeepCoy algorithm to generate decoy molecules, enhancing the specificity scoring function's ability to discriminate and capture active small molecules, thereby increasing the diversity and accuracy of virtual screening. Furthermore, by introducing the PLEC fingerprint feature and combining it with a machine learning algorithm, the generalization and predictive performance of the target-specific scoring function are significantly improved. Furthermore, by incorporating affinity prediction into the later stages of virtual screening and combining it with the specificity scoring function, the resulting molecules are guaranteed to possess both high affinity and good stability.

[0031] Therefore, the present invention combines the advanced technologies of bioinformatics, machine learning and cheminformatics to propose an effective virtual screening method for small molecule inhibitors targeting YTHDF1 protein, which can screen YTHDF1 targeting molecules with high affinity on a large scale and accurately, providing important candidate molecules for the drug development of YTHDF1-related diseases (such as cancer, etc.), and laying a solid foundation for subsequent clinical trials and drug optimization. BRIEF DESCRIPTION OF THE DRAWINGS

[0032] Figure 1 is the chemical structural formula of the small molecule inhibitor targeting YTHDF1 protein finally screened out in the examples of the present invention.

[0033] Figure 2 This is a docking diagram of the small molecule inhibitor targeting YTHDF1 protein and YTHDF1 protein that was finally screened out in the examples of the present invention. DETAILED DESCRIPTION

[0034] The present invention provides a virtual screening method for small molecule inhibitors targeting the YTHDF1 protein. To make the objectives, technical solutions, and effects of the present invention more clear and explicit, the present invention is further described in detail below. It should be understood that the specific embodiments described herein are only intended to illustrate the present invention and are not intended to limit the present invention.

[0035] Unless otherwise defined, all technical terms and scientific terms used herein have the same meaning as those commonly understood by those skilled in the art of the present invention. The terms used in the specification of the present invention herein are only for the purpose of describing specific embodiments and are not intended to limit the present invention.

[0036] The following describes it in detail through specific examples.

[0037] Example 1

[0038] (1) Preparation of protein structure

[0039] Protein structure preparation: The crystal structure of YTHDF1 (PDB ID: 4RCJ) was selected as the target, and the protein structure was further optimized through homology modeling. Using the SWISS-MODEL platform, up to 10 different conformations of the YTHDF1 protein were generated. These conformations represent the protein's natural flexibility and dynamics, capturing its diverse activity states and ensuring that the conformational diversity of YTHDF1 is accounted for during molecular docking. This step ensures that the model is not limited to a single conformation, enhancing the reliability and specificity of small molecule screening.

[0040] As shown in Table 1 , the YTHDF1 protein encoding 4RCJ was selected because it contains no missing key amino acids in the loop region that binds to mRNA.

[0041] Table 1

[0042]

[0043]

[0044] Based on the YTHDF1 protein of 4RCJ, 10 more conformations of YTHDF1 protein were constructed by homology modeling, for a total of 11 conformational protein structures, as shown in Table 2:

[0045] Table 2

[0046]

[0047] RMSD (root mean square deviation) is a key metric for assessing the structural similarity between proteins and ligands. We first gradually fill in the data set based on the RMSD values ​​of various protein conformations and 4RCJs. This prioritizes protein-ligand complexes that are closest to the 4RCJ structure, ensuring the representativeness and diversity of the complexes in the dataset. Furthermore, the use of RMSD helps select conformations that are closest to the experimental structure, enhancing the accuracy of subsequent models.

[0048] (2) Construction of ligand molecule set

[0049] We searched for YTHDF1 protein-binding molecules from Pubchem, BindingDB, and Chembl databases. These molecules were previously reported in the literature and were classified into active small molecules (30) and inactive small molecules (8) based on their activity. Duplicate molecules were removed using the SMILE string.

[0050] To supplement the number of inactive molecules, thereby better training the scoring function MLSFs and improving the model's ability to discriminate between active and inactive molecules, the DeepCoy algorithm, a graph-generating neural network, was used to generate decoy molecules with very similar physical and chemical properties to the active molecules, but with different structures. Using the SMILES of the active molecules as input, 100 decoy molecules were generated for each active small molecule. Finally, 50 decoy molecules were retained for each active small molecule based on the embedding level of the active small molecule in space. The constructed ligand molecule set includes both active and inactive molecule sets, with the inactive molecule set consisting of both inactive small molecules and decoy molecules.

[0051] (3) Molecular docking and complex characterization

[0052] To generate interaction patterns between all molecules and cGAS or kRAS, molecular docking programs can be used to generate all protein-ligand complexes. Autodock Vina is a well-known molecular docking software, and Smina is its optimized version with better speed and stronger scalability.

[0053] The Smina molecular docking program was used to dock the 11 YTHDF1 protein structures with molecules from the ligand set. The ligand search space was defined by the mRNA coordinates of the YTHDF1 protein binding site. This means that the molecules were induced to bind to the optimal binding site using the mRNA coordinates of the YTHDF1 protein itself. In the Smina molecular docking program, the exhaustiveness parameter was set to 8, which enabled a more detailed search for binding modes. The value of num_modes was set to 10 to generate multiple conformational complex structures for each molecule docked with the YTHDF1 protein for subsequent training of the enhanced model. Smina used gradient and Monte Carlo steps to search and rank poses, and the final docked poses were saved in SDF format.

[0054] Then, select specific PLEC fingerprint (Protein-Ligand Extended Connectivity Fingerprints) parameter settings, including: cutoff distance The protein depth is 4, the ligand depth is 2, and the vector length is 1092 to ensure accurate modeling of the spatial conformation of the complex. The PLEC fingerprint captures the spatial correlation between the protein and ligand, encoding the extended connectivity environment between atoms within a specific distance range. Based on the neighborhood information of the molecular graph, this fingerprint encodes the structural information of each atom and its surroundings in the protein and ligand into a fixed-length binary vector.

[0055] (4) Constructing a target-specific scoring function

[0056] Based on the obtained PLEC fingerprint dataset of protein-ligand complexes, an ANN (artificial neural network) was used to construct a specific scoring function for YTHDF1.

[0057] ANNs offer significant advantages in handling complex protein-ligand interactions, automatic feature extraction, processing large amounts of data, and generalization, making them ideal for constructing efficient scoring functions in virtual screening. Using multiple layers of neurons and nonlinear activation functions, ANNs automatically extract complex features and patterns. This is particularly important in the case of the YTHDF1 protein-ligand complex, as the complex's three-dimensional structure and its interactions with other molecules often exhibit highly complex, nonlinear characteristics. ANNs excel at extracting meaningful information from these complex structural relationships, thereby improving prediction accuracy. Compared to traditional machine learning models that rely on manually extracted features, ANNs, through their multi-layered architecture, can automatically learn deep-level features from input data. This automated feature extraction capability reduces reliance on domain knowledge and can capture underlying features of protein-ligand interactions, such as the spatial distribution of binding sites and atomic interactions. Furthermore, ANNs scale well to large datasets and can process thousands to millions of training samples. For the vast number of compounds in virtual screening, ANNs can efficiently train and infer across such a large sample space, significantly improving virtual screening efficiency and hit rates. In terms of generalization ability, after proper regularization and parameter adjustment, ANN can generalize well to unseen molecules. This means that an ANN model that performs well on the training set can also obtain good prediction results on unknown molecules in actual applications, thereby improving the practicality of the model.

[0058] Based on the performance of models trained with different amounts of data, the present invention chose to enhance the model's generalization by incorporating a collection of all 11 protein-ligand complexes with different conformations into the dataset. This data augmentation strategy significantly improved the model's prediction accuracy compared to using only a single protein structure. By incorporating data from multiple protein complexes, the model learns more about the differences between different complexes, thereby improving its predictive power for unseen data.

[0059] The dataset was divided into a training set and a test set based on the similarity of the complexes' PLEC fingerprints. The 4:1 ratio division ensured the diversity of the compounds in the training and test sets. The PLEC fingerprint division method avoided the problem of overly similar training and test sets caused by traditional random division, thereby enabling the model to generalize better. During the model training phase, classification and regression tasks were trained separately. The goal of the classification task is to distinguish between active and inactive molecules in the compound library. By learning the characteristics of the molecules, the model can quickly screen out potential active molecules. The regression task focuses more on predicting the binding affinity of active molecules and quantifying the strength of their interaction with the target protein. This combination of classification and regression allows for a more comprehensive assessment of molecular activity and further prediction of its actual level of biological activity.

[0060] (5) Virtual screening and affinity prediction

[0061] During the virtual screening phase, high-throughput molecular docking was performed using 45,000 compounds from the ZINC20 database against 11 YTHDF1 protein conformations. As a widely used public compound library, the ZINC20 database provides a rich resource of screenable molecules suitable for virtual screening in drug development.

[0062] Using a previously constructed target-specific scoring function, the resulting complexes were scored and ranked to initially screen for potential active compounds. After ranking the active molecules according to the probability (0–1) assessed by the scoring function, a total of 971 active molecules were screened from 45,000 compounds, achieving a hit rate of approximately 2.1%. This is a remarkable result for a virtual screening task.

[0063] To further confirm the reliability of the screening results and the practical application potential of the compounds, a logistic regression model was trained to predict the binding affinity of the top 1% of active candidate molecules. Binding affinity prediction can help assess the binding strength between these molecules and the target protein, further verify their stability and affinity for YTHDF1 binding, and thus determine their potential for drug development.

[0064] Finally, the 10 most promising small molecule inhibitors of YTHDF1 were screened out. The chemical structures of these 10 small molecule inhibitors targeting YTHDF1 protein are as follows: Figure 1 The corresponding predicted affinity values ​​are shown in Table 3:

[0065] Table 3

[0066]

[0067]

[0068] Figure 2 The following is a docking diagram of the selected small molecule inhibitors targeting YTHDF1 protein and YTHDF1 protein. Combined with the affinity prediction results in Table 3, it reflects the energy state and interaction mode of the binding of the final screened molecular compounds with YTHDF1 protein. The specific analysis is as follows:

[0069] The predicted affinity value of ZINC00024551 was 8.825, and this molecule formed hydrogen bonds and hydrophobic interactions with the amino acid residues around the site. Figure 2 It is shown in the results that the molecule forms π-π stacking interactions with aromatic amino acids (such as tryptophan, tyrosine, etc.), which help to improve the binding stability of the molecule.

[0070] ZINC00025054 has a predicted affinity of 9.555, the highest predicted affinity, indicating strong binding. This is likely due to the molecule forming hydrogen bonds with multiple polar residues and potential interactions with carbonyl groups on the protein backbone. Furthermore, it forms favorable van der Waals interactions with surrounding hydrophobic amino acids.

[0071] The predicted affinity value of ZINC00028212 is 9.058, and this molecule is likely to bind to the vicinity of the catalytic residue through multiple polar interactions. Figure 2 Its location is shown to be close to key catalytic residues, suggesting that the molecule may inhibit YTHDF1 function by hindering RNA binding.

[0072] The predicted affinity value of ZINC00039097 is 9.080, Figure 2 The molecule is shown to bind to a key lysine or arginine residue through hydrogen bonds, and these interactions play a key role in stabilizing the binding of the inhibitor in the active site.

[0073] The predicted affinity value of ZINC00057125 is 9.484, and this molecule interacts with both polar residues and hydrophobic regions. Figure 2 suggests that it may occupy an important region for RNA recognition, thereby effectively competing with RNA for binding.

[0074] The predicted affinity value of ZINC00057311 is 9.045. This molecule may be stabilized in the binding site through hydrogen bonds and hydrophobic interactions, especially the conserved residues around the target region, thereby blocking the RNA binding function of YTHDF1.

[0075] The predicted affinity value of ZINC00057394 is 9.107, Figure 2 The results show that the molecule interacts with both hydrophobic and polar residues in the binding site, indicating that it may affect the function of YTHDF1 by binding to the RNA binding region.

[0076] The predicted affinity value of ZINC00079466 is 9.176, and the molecule tightly binds to key residues of the binding site through hydrogen bonds and hydrophobic interactions, suggesting that it may inhibit YTHDF1 by competing for the binding of natural RNA.

[0077] The predicted affinity of ZINC00105325 is 9.175, Figure 2 It is shown that the molecule forms multiple hydrogen bonds and van der Waals interactions, tightly binding to the active site and exhibiting strong binding ability.

[0078] The predicted affinity of ZINC00112982 is 8.983, and this molecule forms a complex hydrogen bond network and non-polar interactions with residues in the RNA binding domain, which may alter the ability of YTHDF1 protein to bind its natural substrate.

[0079] According to the structural characteristics of the above 10 molecules, they can be divided into the following five categories:

[0080] (1) Nucleoside analogs: These compounds have a nucleoside-like structure, typically consisting of a sugar and a base. They readily interact with the RNA binding site of the YTHDF1 protein and competitively inhibit the binding of m6A-modified RNA. These molecules have a simple structure and often interact with their targets through hydrogen bonds and π-π stacking. For example, ZINC00112982 has a pyrimidine base and a sugar structure. ZINC00057311 contains a similar nucleoside structure. ZINC00039097 has a sugar and aromatic base structure that is highly similar to the RNA binding domain of the YTHDF1 protein.

[0081] (2) Aromatic rings and their derivatives: These compounds contain multiple aromatic or heterocyclic structures that bind to aromatic residues at the YTHDF1 binding site through π-π interactions and hydrophobic interactions. Their structures are relatively complex, and their inhibitory effects are enhanced through multi-point binding. For example, ZINC00057125 contains multiple aromatic rings and hydroxyl groups. ZINC00079466 has a complex structure, contains multiple heterocyclic rings, and can interact with multiple regions of the protein.

[0082] (3) Carbohydrate derivatives: These molecules contain sugar molecules with multiple hydroxyl groups that can form stable interactions with polar residues in the YTHDF1 binding domain through hydrogen bonds. For example, ZINC00057394: The sugar backbone is bound to multiple hydroxyl groups. ZINC00024551: The sugar and alkyne structures provide flexibility. ZINC00028212: Phenyl groups are attached to the sugar backbone and may form hydrophobic and hydrogen bonding interactions with YTHDF1.

[0083] (4) Carboxyl and amide-containing molecules: These compounds typically contain carboxyl and amide groups, bind to YTHDF1 through hydrogen bonds, and interfere with its recognition of m6A-modified RNA. For example, ZINC00025054: contains carboxyl and sugar groups, which can form hydrogen bonds with YTHDF1. ZINC00105325: contains a bicyclic and sugar structure, providing multiple interaction points.

[0084] (5) Molecules containing alkyne groups: The alkyne structure gives these compounds flexibility, allowing them to adapt to the shape of the protein binding pocket, thereby interfering with the function of the protein. For example, ZINC00024551: its alkyne group provides a specific interaction with the YTHDF1 binding pocket.

[0085] It can be seen that the present invention not only significantly improves the accuracy and efficiency of YTHDF1 small molecule inhibitor screening through the virtual screening method, but also provides a new technical path for drug development.

[0086] It should be understood that the application of the present invention is not limited to the above examples. For those skilled in the art, improvements or changes can be made based on the above description. All these improvements and changes should fall within the scope of protection of the claims attached to the present invention.

Claims

1. A virtual screening method for small molecule inhibitors targeting YTHDF1 protein, characterized in that: Including steps: (a) YTHDF1 protein structure preparation: The crystal structure of YTHDF1 protein was obtained, and homology modeling of YTHDF1 protein was performed using the SWISS-MODEL platform to generate YTHDF1 protein in multiple conformations; (b) Construction of ligand molecule dataset: Active small molecules and inactive small molecules that bind to the YTHDF1 protein are obtained, and inactive decoy molecules are generated based on the active small molecules to construct a ligand molecule dataset; (c) Molecular docking and complex characterization: high-throughput molecular docking is performed between the molecules in the ligand molecule set and the binding sites on the YTHDF1 protein in various conformations using a molecular docking program, and the resulting protein-ligand complex is characterized by PLEC fingerprinting; (d) Construction of specificity scoring function: Based on the obtained PLEC fingerprint dataset, a specificity scoring function for YTHDF1 protein was constructed using a machine learning algorithm; (e) Virtual screening and affinity prediction: Compounds in the library to be screened are docked onto YTHDF1 proteins in various conformations to obtain PLEC fingerprints of the complexes. After scoring and ranking the complexes using the specificity scoring function, affinity calculations are performed to ultimately identify small molecule inhibitors targeting the YTHDF1 protein. In step (b), inactive decoy molecules are generated based on the active small molecules, specifically comprising: using the DeepCoy algorithm to generate inactive decoy molecules with similar physical and chemical properties but different structures to the active small molecules, using the SMILES of the active small molecules as input, first generating 100 decoy molecules for each active small molecule, then retaining 50 decoy molecules for each active small molecule based on the degree of embedding of the active small molecules in space; In step (c), the molecular docking program includes the Smina molecular docking program, and in the Smina molecular docking program, the exhaustiveness parameter is set to 8 and the num_modes value is set to 10; In step (e), the complexes are scored and ranked using the specificity scoring function to screen out the top 1% of candidate compounds, and the affinity of the candidate compounds is then calculated using a logistic regression model to ultimately screen out small molecule inhibitors targeting YTHDF1 with high affinity.

2. The virtual screening method for small molecule inhibitors targeting YTHDF1 protein according to claim 1, characterized in that: In step (a), the crystal structure of the YTHDF1 protein is obtained from the Protein Data Bank database and is numbered 4RCJ. Ten YTHDF1 protein conformations are generated by homology modeling.

3. The virtual screening method for small molecule inhibitors targeting YTHDF1 protein according to claim 1, characterized in that: In step (b), active and inactive small molecules that bind to the YTHDF1 protein were obtained from the PubChem, BindingDB, and Chembl databases, and duplicate molecules were removed based on the SMILE string.

4. The virtual screening method for small molecule inhibitors targeting YTHDF1 protein according to claim 1, characterized in that: In step (c), when the protein-ligand complex is characterized using PLEC fingerprint, the parameter settings of the PLEC fingerprint include: a cutoff distance of 4.5 Å, a protein depth of 4, a ligand depth of 2, and a vector length of 1092.

5. The virtual screening method for small molecule inhibitors targeting YTHDF1 protein according to claim 1, characterized in that: In step (d), the machine learning algorithm includes at least one of random forest, extreme gradient boosting, support vector machine and artificial neural network.

6. The virtual screening method for small molecule inhibitors targeting YTHDF1 protein according to claim 1, characterized in that: In step (d), the machine learning algorithm is an artificial neural network, and the artificial neural network is trained and evaluated through cross-validation.

7. The virtual screening method for small molecule inhibitors targeting YTHDF1 protein according to claim 1, characterized in that: In step (e), the compound library to be screened is selected from the ZINC20 database, which contains 45,000 small molecule compounds.

Citation Information

Patent Citations

  • Molecular docking result screening method based on positive compound residue contribution similarity

    CN113380320A

  • Virtual screening method and device of compound

    JP2008174503A