Screening method and device of targeted RNA (Ribonucleic Acid) small molecules
By combining feature mapping models and molecular dynamics simulations, the problems of long experimental cycles and high false negative rates in the screening of targeted RNA small molecules have been solved, achieving efficient and accurate screening of RNA small molecules.
Patent Information
- Application Number
- CN202510930801.6
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-07-07
- Publication Date
- 2025-10-28
AI Technical Summary
Existing technologies for screening small RNA molecules suffer from problems such as long experimental cycles, high costs, and high false negative rates. In particular, the molecular dynamics simulation calculation cycle is too long, making it difficult to handle batch screening tasks.
A method combining feature mapping model and molecular dynamics simulation was adopted. The feature mapping model was constructed by Uni-Mol model and E(3)-equivariant graph neural network, and the physical feature encoder was trained to simulate RNA conformation diversity and screen small molecules, thereby reducing computational cost and improving screening efficiency.
It improves the richness of RNA conformations and the accuracy of screening, reduces the false negative rate of small molecules, improves the efficiency of batch screening, and reduces computational time costs.
Smart Images

Figure CN120853683A_ABST
Abstract
Description
Technical Field
[0001] This invention relates to the field of data processing technology, and in particular to a method and apparatus for screening targeted RNA small molecules. Background Technology
[0002] Ribonucleic acid (RNA) plays a crucial role in various life activities, including gene expression regulation and catalytic biochemical reactions. In recent years, with a deeper understanding of the biological functions of RNA, its potential as a drug target has been rediscovered and explored. Developing small-molecule drugs that can specifically recognize and regulate the function of specific RNA molecules has become a cutting-edge and hot topic in innovative drug development.
[0003] Small molecules that bind to RNA are called target RNA small molecules. Batch screening of target RNA small molecules is an important task in the field of small molecule drug design. However, using traditional experimental methods for batch screening suffers from problems such as long experimental cycles and high experimental costs.
[0004] To improve screening efficiency, Structure-Based Virtual Screening (SBVS) technology can be used. Molecular docking is the core technology of SBVS, which analyzes the binding ability of RNA-small molecule complexes (such as affinity analysis) by simulating the binding conformation of various small molecules in the binding pocket of RNA targets and then performs screening based on the analysis results. However, molecular docking technology often treats the RNA target site as a rigid structure during the simulation process. This approach ignores the conformational diversity of the flexible structure of RNA and results in a limited number of complex conformations, which may lead to the missed detection of some effective small molecules.
[0005] Molecular dynamics simulations can fully capture the diverse conformations that occur during RNA-small molecule binding. In principle, using molecular dynamics simulations for targeted RNA-small molecule screening can indeed reduce the false negative rate. However, simulating the entire RNA-small molecule binding process also suffers from excessively long computation cycles; simulating even a single RNA-small molecule complex can require hours or even days of computation. Therefore, relying solely on molecular dynamics simulations to handle batch screening tasks is impractical.
[0006] Based on this situation, we designed a screening scheme that can improve both RNA conformation diversity and processing efficiency: First, a feature mapping model is trained to learn the mapping relationship between any RNA conformation and the physicochemical features of its preferred small molecules; then, accelerated molecular dynamics simulations are used to obtain multiple RNA conformations forming a conformation set; next, the feature mapping model is used to map the physicochemical features of the preferred small molecules in the conformation set one by one to obtain a feature set; finally, batch screening of small molecules is performed based on the feature set. This new scheme expands the range of physicochemical features (feature set) of suitable small molecules by increasing the richness of RNA conformations (conformation set), thereby increasing the range of suitable small molecules and reducing the false negative rate. Furthermore, although this scheme uses molecular dynamics simulation technology, it is limited to short-time simulations of RNA conformation diversity (single simulation cycle in microseconds or nanoseconds), thus improving both the richness and accuracy of simulated conformations while significantly reducing computational cycle costs and improving overall processing efficiency. The specific implementation method of this new scheme is the technical problem that this invention aims to solve. Summary of the Invention
[0007] The purpose of this invention is to address the shortcomings of existing technologies by providing a method, apparatus, electronic device, and computer-readable storage medium for screening targeted RNA small molecules. This invention first constructs a property prediction model based on the Uni-Mol model to predict the physicochemical properties of small molecules, and then constructs a feature mapping model based on an E(3)-equivariant graph neural network to learn the physicochemical features of RNA-adaptive small molecules. Next, a first dataset is constructed through large-scale data collection of the physicochemical properties of drug small molecules, and the property prediction model is trained based on this dataset. At the end of training, a physicochemical feature encoder is extracted from the model. Then, a second dataset is initialized through large-scale data collection of RNA-small molecule complexes, and the physicochemical feature encoder is used to label the second dataset using a tensor. Finally, the feature mapping is trained based on the second dataset. The model is designed to perform a series of steps. After both models are trained, a rapid simulation method based on molecular dynamics simulation is used to expand the user-specified RNA conformations into homologous conformations, resulting in a sampled conformation set. This set is then subjected to conformation clustering and typical conformation extraction to obtain a typical conformation set. Next, a feature mapping model and a physical feature encoder are used to construct feature sets from the typical conformation set and the user-specified small molecule database, resulting in a mapped feature set and a physical feature set. Finally, based on the feature similarity comparison results of the mapped and physical feature sets, small molecule conformations in the small molecule database that match the user-specified RNA conformations are screened and ranked to obtain the preferred conformation sequences, which are then fed back to the user. This invention improves the richness and accuracy of RNA conformations, reduces the small molecule false negative rate, reduces simulation time costs, and increases batch screening efficiency.
[0008] To achieve the above objectives, a first aspect of the present invention provides a method for screening targeted RNA small molecules, the method comprising:
[0009] A property prediction model for predicting the physicochemical properties of small molecules is constructed based on the Uni-Mol model; and a feature mapping model for learning the physicochemical features of RNA-adapted small molecules is constructed based on a class of E(3)-equivariant graph neural networks; the property prediction model is used to predict a set of specified physicochemical properties of the current molecule based on the small molecule conformation input to the model to obtain the corresponding physicochemical property vector; the feature mapping model is used to perform feature mapping on the physicochemical features of the current RNA-adapted small molecule based on the target RNA conformation input to the model to obtain the corresponding physicochemical feature mapping tensor.
[0010] The first dataset was constructed by collecting large amounts of data on the physicochemical properties of small drug molecules;
[0011] The feature prediction model is trained based on the first dataset; and the corresponding materialized feature encoder is extracted from the feature prediction model at the end of the training.
[0012] The second dataset is initialized by collecting large amounts of data on RNA-small molecule complexes and then labeled with a label tensor using the physical feature encoder described above.
[0013] The feature mapping model is trained based on the second dataset;
[0014] After both models have been trained, the RNA conformation input by the user will be used as the corresponding current RNA conformation; and the small molecule database specified by the user will be used as the corresponding current molecular library.
[0015] Molecular dynamics simulations are performed on the current RNA conformation to obtain a corresponding sampled conformation set; and conformation clustering and typical conformation extraction are performed on the sampled conformation set to obtain a corresponding typical conformation set;
[0016] Based on the feature mapping model and the materialized feature encoder, feature sets are constructed for the typical conformation set and the current molecular library to obtain the corresponding mapping feature set and materialized feature set;
[0017] Based on the mapping feature set and the physical feature set, the molecular conformations in the current molecular library that are compatible with the current RNA conformation are screened to obtain the preferred conformation sequence, which is then fed back to the current user.
[0018] Preferably, the small molecule conformation and the target RNA conformation are each an atomic-level three-dimensional conformation;
[0019] The atomic-level three-dimensional conformation includes a first set of atoms; the first set of atoms includes multiple first atoms; the first atom includes an atom identifier, an atom element type, and atom three-dimensional coordinates.
[0020] The length of the materialized characteristic vector is equal to the total number N of the characteristic types in the preset materialized characteristic set. F Equal; N F The value is a preset positive integer; the materialized characteristic vector includes N. F The set of physical and chemical properties includes N predicted values. F Each physical and chemical property type corresponds to a class of physical or chemical properties; the predicted value of each physical and chemical property corresponds one-to-one with the physical and chemical property type.
[0021] The materialized feature mapping tensor has a shape of D1×L1, where D1 is a preset first feature dimension, L1 is a preset total number of first sub-vectors, and L1=N. F The materialized feature mapping tensor consists of L1 first mapping vectors of length D1; the first mapping vectors correspond one-to-one with the materialized feature types of the materialized feature set.
[0022] The first dataset includes multiple first data records; each first data record corresponds to a class of small molecules; the first data record includes a first molecule sequence, a first training conformation, and a first label vector; the first molecule sequence is the one-dimensional molecular sequence of the current small molecule; the first training conformation is the atomic-level three-dimensional conformation of the current small molecule; the first label vector is the physicochemical property label vector of the current small molecule, consisting of N... F It consists of a set of physical property tag values; each physical property tag value corresponds one-to-one with the physical property type of the physical property set;
[0023] The second dataset includes multiple second data records; each second data record corresponds to a class of RNA-small molecule complexes; the second data record includes a first complex conformation, a first RNA conformation sampling set, a first small molecule conformation, and a first tag tensor; the first complex conformation is the atomic-level three-dimensional conformation of the current RNA-small molecule complex; the first small molecule conformation is the atomic-level three-dimensional conformation of the small molecule in the current RNA-small molecule complex, and the first small molecule conformation is a local conformation of the first complex conformation; the remaining local conformations in the first complex conformation other than the first small molecule conformation are the atomic-level three-dimensional conformations of the RNA in the current RNA-small molecule complex, denoted as the first RNA conformation. The first RNA conformation sampling set includes multiple first RNA sampling conformations; each first RNA sampling conformation is a local conformation of the first RNA conformation; the sampling rule is as follows: in the first complex conformation, with the geometric center point of the first small molecule conformation as the center of a sphere, corresponding sampling spherical regions are set according to the radii of each sphere in a preset set of sphere radii, and the local conformations of the first RNA conformation in each of the sampling spherical regions are extracted as a corresponding first RNA sampling conformation; the set of sphere radii includes multiple sphere radii; the first label tensor is the physical feature tensor corresponding to the first small molecule conformation, which is obtained by the physical feature encoder performing feature encoding processing on the first small molecule conformation;
[0024] The current RNA conformation is an atomic-level three-dimensional conformation;
[0025] The current molecular library includes multiple first molecular conformations; each first molecular conformation is the atomic-level three-dimensional conformation of a class of small molecules;
[0026] The sampled conformation set includes multiple first sampled conformations; the typical conformation set includes one or more first typical conformations.
[0027] The mapping feature set includes one or more first mapping features; each of the first mapping features corresponds one-to-one with the first typical conformation.
[0028] The set of physical features includes multiple first physical features; each of the first physical features corresponds one-to-one with the first molecular conformation.
[0029] The preferred conformation sequence is formed by sequentially sorting multiple first preferred conformations.
[0030] Preferably, the model input terminal of the property prediction model is used to receive the small molecule conformation, and the model output terminal is used to output the corresponding physicochemical property vector;
[0031] The feature prediction model includes the materialized feature encoder and the materialized feature prediction network; the input of the materialized feature encoder is connected to the input of the model, and the output is connected to the materialized feature prediction network; the output of the materialized feature prediction network is connected to the output of the model.
[0032] The materialized feature encoder is composed of a first embedding encoding module, a first encoder, and a first feature mapping network connected sequentially; the first encoder is implemented based on the Uni-Mol model; the Uni-Mol model has been pre-trained; the first feature mapping network is composed of a first pooling layer and a first MLP model connected sequentially.
[0033] The materialization feature encoder is used to encode the materialization features of the small molecule conformation to obtain the corresponding materialization feature tensor and send it to the materialization property prediction network. Specifically, the first embedding encoding module performs embedding encoding processing on the small molecule conformation according to the embedding encoding rules of the Uni-Mol model input data to obtain the corresponding first embedding encoding tensor and sends it to the first encoder; the first encoder performs atomic-level feature encoding processing on the first embedding encoding tensor to obtain the corresponding first feature encoding tensor and sends it to the first feature mapping network; the first feature mapping network performs global pooling processing on each feature dimension of the first feature encoding tensor through the first pooling layer to obtain the corresponding first pooling feature vector, and performs vector mapping of the first pooling feature vector on the materialization feature space through the first MLP model to obtain the corresponding materialization feature tensor and sends it to the materialization property prediction network; the shape of the materialization feature tensor is D1×L1; the materialization feature tensor is composed of L1 vectors of length D1 as first feature vectors; the first feature vectors correspond one-to-one with the materialization property types of the materialization property set.
[0034] The physical property prediction network has built-in N F Each first characteristic prediction model corresponds one-to-one with the physical characteristic type of the physical characteristic set; each first characteristic prediction model is implemented based on an MLP model; each first characteristic prediction model corresponds one-to-one with the first feature vector.
[0035] The physical property prediction network is used to input each of the first feature vectors in the physical feature tensor into the corresponding first property prediction model to predict the corresponding physical property, obtain the corresponding predicted value of the physical property, and then use the obtained N... F The predicted values of the physical and chemical properties are combined to form the corresponding physical and chemical property vector and output.
[0036] Preferably, the input end of the feature mapping model is used to receive the target RNA conformation, and the output end is used to output the corresponding materialized feature mapping tensor.
[0037] The feature mapping model includes a second embedding encoding module, a second encoder, and a second feature mapping network; the second encoder is implemented based on a class of E(3)-equivariant graph neural networks; the E(3)-equivariant graph neural networks include EGNN model, SE(3)-Transformer model, EGAT model, and Cormorant model; the second feature mapping network is formed by sequentially connecting a second pooling layer and a second MLP model;
[0038] The input of the second embedding encoding module is connected to the input of the model, and the output is connected to the input of the second encoder; the output of the second encoder is connected to the input of the second feature mapping network; the output of the second feature mapping network is connected to the output of the model.
[0039] The second embedding coding module performs embedding coding processing on the target RNA conformation according to the embedding coding rules of the model input data of the second encoder to obtain the corresponding second embedding coding tensor, which is then sent to the second encoder.
[0040] The second encoder performs atomic-level feature encoding processing on the second embedding encoding tensor to obtain the corresponding second feature encoding tensor, which is then sent to the second feature mapping network.
[0041] The second feature mapping network performs global pooling on each feature dimension of the second feature encoding tensor through the second pooling layer to obtain the corresponding second pooling feature vector, and then performs vector mapping of the second pooling feature vector to the materialized feature space through the second MLP model to obtain the corresponding materialized feature mapping tensor and output it.
[0042] Preferably, the step of constructing the first dataset by collecting large amounts of data on the physicochemical properties of small drug molecules specifically includes:
[0043] Step 51: Collect large amounts of data on the molecular sequence and / or molecular conformation of drug small molecules and the physicochemical properties that satisfy the set of physicochemical properties through the publicly available small molecule data collection method to obtain the corresponding first collection dataset;
[0044] The disclosed small molecule data acquisition methods include multiple publicly available small molecule databases, multiple publicly available small molecule experimental databases, and multiple publicly available databases of scientific and technological literature / technical papers in the fields of biochemistry and medicine.
[0045] The first acquisition dataset includes multiple first acquisition records; each first acquisition record corresponds to a class of drug small molecules; the first acquisition record includes a small molecule sequence, a small molecule conformation, and a small molecule property set; the small molecule sequence is the one-dimensional molecular sequence of the current drug small molecule; the small molecule conformation is the atomic-level three-dimensional conformation of the current drug small molecule; the small molecule property set consists of N... F The set consists of small molecule characteristic values, which correspond one-to-one with the physicochemical characteristic types of the physicochemical characteristic set; each of the small molecule sequences and small molecule characteristic sets in the first collection record is a required option and cannot be empty, while the small molecule conformation is an optional option and can be empty;
[0046] Step 52: Identify whether the small molecule conformation of each of the first acquisition records in the first acquisition dataset is empty. If so, perform atomic-level three-dimensional conformation simulation based on the small molecule sequence of the current first acquisition record using a preset molecular dynamics simulation tool, and use the simulation result as the corresponding current simulation conformation. Then, set the small molecule conformation of the current first acquisition record based on the current simulation conformation.
[0047] The molecular dynamics simulation tools include GROMACS, AMBER, NAMD, CHARMM, and LAMMPS.
[0048] Step 53: Take each of the first acquisition records in the first acquisition dataset as the corresponding current acquisition record; and take the small molecule sequence, the small molecule conformation and the small molecule characteristic set of the current acquisition record as a corresponding first molecule sequence, the first training conformation and the first label vector to form a corresponding first data record; and take all the obtained first data records to form the corresponding first dataset.
[0049] Preferably, training the feature prediction model based on the first dataset specifically includes:
[0050] Step 61: Based on a preset first segmentation ratio, the first dataset is randomly divided into two sub-datasets, denoted as the first training set and the first evaluation set.
[0051] Wherein, both the first training set and the first evaluation set are composed of multiple first data records; the ratio of the total number of records in the first training set and the first evaluation set satisfies the first segmentation ratio;
[0052] Step 62: Take each of the first data records in the first training set as the corresponding current training record; and take the first training conformation of the current training record as the current small molecule conformation and input it into the property prediction model for processing to obtain the corresponding physical property vector; and form a corresponding first prediction-label pair by the current physical property vector and the first label vector of the current training record.
[0053] Step 63: Input all the obtained first prediction-label pairs into the preset first model loss function to calculate the corresponding first loss value;
[0054] The first model loss function is implemented based on the L1 loss function or the L2 loss function;
[0055] Step 64: Identify whether the first loss value meets the preset first loss value range; if it does, proceed to step 65; if it does not, modulate the model parameters of the feature prediction model in one round based on the preset first model optimizer in the direction of minimizing the first model loss function, and return to step 62 to continue training when the first round of modulation ends.
[0056] The first model optimizer includes the Adam optimizer and the SGD optimizer.
[0057] Step 65: Take each of the first data records in the first evaluation set as the corresponding current evaluation record; and take the first training conformation of the current evaluation record as the current small molecule conformation and input it into the property prediction model for processing to obtain the corresponding physical and chemical property vector; and form a corresponding second prediction-label pair by the current physical and chemical property vector and the first label vector of the current evaluation record; and input all the obtained second prediction-label pairs into the preset first model evaluation function to calculate the corresponding first evaluation value;
[0058] The first model evaluation function is implemented based on the MAE function, MSE function, or RMSE function.
[0059] Step 66: Identify whether the first evaluation value meets the preset first evaluation value range; if not, return to step 61 to continue training; if it does, confirm that the training of the feature prediction model has ended.
[0060] Preferably, the step of initializing the second dataset through large-scale acquisition of RNA-small molecule complexes and using the physical feature encoder to perform label tensor annotation on the second dataset specifically includes:
[0061] Step 71: Collect large data on the three-dimensional conformation of the RNA-small molecule complex using the publicly available RNA-small molecule complex data acquisition method to obtain the corresponding first collection conformation set.
[0062] The disclosed RNA-small molecule complex data acquisition methods include multiple publicly available RNA-small molecule complex databases, multiple publicly available RNA-small molecule complex experimental databases, and multiple publicly available scientific and technological literature / technical paper databases in the fields of biochemistry and medicine.
[0063] The first set of acquisition conformations includes multiple first acquisition conformations; each first acquisition conformation is an atomic-level three-dimensional conformation of a class of RNA-small molecule complexes; each first acquisition conformation consists of a pair of RNA subconformations and small molecule subconformations;
[0064] Step 72: Each of the first acquired conformations in the first acquired conformation set is taken as the corresponding current complex conformation; the current complex conformation, along with the corresponding RNA sub-conformation and the small molecule sub-conformation, is taken as a set of corresponding first complex conformations, first RNA conformations, and first small molecule conformations; within the current first complex conformation, with the geometric center point of the current first small molecule conformation as the sphere center, corresponding sampling spherical regions are set according to the sphere radii of the set of sphere radii, and local conformations of the current first RNA conformation within each of the sampling spherical regions are extracted as a corresponding first RNA sampling conformation; all the first RNA sampling conformations obtained this time form a corresponding first RNA conformation sampling set; a corresponding first label tensor is set to empty; the first complex conformation corresponding to the current complex conformation, the first RNA conformation sampling set, the first small molecule conformation, and the first label tensor form a corresponding second data record; and all the obtained second data records form the initialized second dataset.
[0065] Step 73: Take each of the second data records in the second dataset as the corresponding current record; input the first small molecule conformation of the current record as the current small molecule conformation into the materialization feature encoder for processing to obtain the corresponding materialization feature tensor; and set the first label tensor of the current record based on the current materialization feature tensor.
[0066] Preferably, training the feature mapping model based on the second dataset specifically includes:
[0067] Step 81: Based on the preset second segmentation ratio, the second dataset is randomly divided into two subsets, which are denoted as the corresponding second training set and second evaluation set.
[0068] The second training set and the second evaluation set are each composed of multiple second data records; the ratio of the total number of records in the second training set and the second evaluation set satisfies the second segmentation ratio.
[0069] Step 82: Take each of the second data records in the second training set as the corresponding current training record; and take each of the first RNA sampling conformations in the first RNA conformation sampling set of the current training record as the current target RNA conformation and input them into the feature mapping model for processing to obtain the corresponding materialized feature mapping tensor; and form a corresponding third prediction-label pair with each of the materialized feature mapping tensors corresponding to the current training record and the first label tensor of the current training record.
[0070] Step 83: Input all the obtained third prediction-label pairs into the preset second model loss function to calculate the corresponding second loss value;
[0071] The second model loss function is based on the L2 loss function.
[0072] Step 84: Identify whether the second loss value meets the preset range of the second loss value; if it does, proceed to step 85; if it does not, adjust the model parameters of the feature mapping model in one round based on the preset second model optimizer in the direction of minimizing the second model loss function, and return to step 82 to continue training when the adjustment is completed.
[0073] The second model optimizer includes the Adam optimizer and the SGD optimizer;
[0074] Step 85: Take each of the second data records in the second evaluation set as the corresponding current evaluation record; and take each of the first RNA sampling conformations in the first RNA conformation sampling set of the current evaluation record as the current target RNA conformation and input them into the feature mapping model for processing to obtain the corresponding materialized feature mapping tensor; and form a corresponding fourth prediction-label pair with each of the materialized feature mapping tensors corresponding to the current evaluation record and the first label tensor of the current evaluation record; and input all the obtained fourth prediction-label pairs into the preset second model evaluation function to calculate the corresponding second evaluation value.
[0075] The second model evaluation function is implemented based on the RMSE function;
[0076] Step 86: Identify whether the second evaluation value meets the preset second evaluation value range; if not, proceed to step 81 to continue training; if it meets the range, confirm that the training of the feature mapping model has ended.
[0077] Preferably, the step of performing molecular dynamics simulations on the current RNA conformation to obtain the corresponding sampled conformation set specifically includes:
[0078] Using a preset molecular dynamics simulation tool, a simulation of the current RNA conformation is performed within a preset fixed simulation time using a preset rapid simulation method. During the simulation, the dynamically changing simulated conformation is collected in real time at preset sampling time intervals. Each real-time simulated conformation collected is taken as a corresponding first sampled conformation. At the end of this simulation round, the current RNA conformation is also taken as a corresponding first sampled conformation, and all the obtained first sampled conformations form the corresponding sampled conformation set.
[0079] The molecular dynamics simulation tools include GROMACS, AMBER, NAMD, CHARMM, and LAMMPS; the fast simulation methods include Gaussian accelerated molecular dynamics simulation, meta-dynamics simulation, high-temperature molecular dynamics simulation, and temperature replication-exchange molecular dynamics simulation.
[0080] Preferably, the step of performing conformation clustering and typical conformation extraction on the sampled conformation set to obtain the corresponding typical conformation set specifically includes:
[0081] The total number of the first sampled conformations in the sampled conformation set is counted to obtain the corresponding total number of conformations N. M ; and based on the total number N of the conformations M Construct an error matrix D; the error matrix D includes N M ×N M d matrix elements i,j ; 1 ≤ row index i ≤ N M 1 ≤ column index j ≤ N M Each row or column of the error matrix D corresponds one-to-one with the first sampling configuration; the matrix elements d on the diagonal of the error matrix D i,j=i All are 0; the symmetric elements of each pair of diagonals in the error matrix D are equal: d i,j≠i =d j≠i,i ;
[0082] The first sampled conformations of the sampled conformation set are combined in pairs to obtain multiple conformation pairs; the two first sampled conformations of each conformation pair are aligned, and the root mean square error of the atomic coordinates of the two aligned conformations is calculated to obtain the corresponding conformation pair error; each conformation pair error is recorded as the corresponding current error, and the two diagonal symmetric elements in the error matrix D corresponding to the current error are recorded as the corresponding current element pair, and the two matrix elements of the current element pair are set based on the current error;
[0083] And for all matrix elements d of the error matrix D i,j Perform mean calculation and use the result as the corresponding error threshold d. hold ;
[0084] And based on the fully linked hierarchical clustering algorithm, according to the error threshold d hold Clustering all the first sampled conformations in the sampled conformation set yields multiple corresponding first conformation clusters; wherein each first conformation cluster includes one or more first sampled conformations; in each first conformation cluster containing two or more first sampled conformations, the error between any two first sampled conformations does not exceed the error threshold d. hold Between every two first conformation clusters, there exist at least two conformation pairs of the first sampled conformations with errors greater than the error threshold d. hold ;
[0085] Each of the first conformation clusters is taken as the corresponding current cluster; the total number of the first sampled conformations in the current cluster is counted to obtain the total number of conformations in the current cluster; the total number of conformations in the current cluster is identified; if the total number of conformations in the current cluster is 1, the unique first sampled conformation in the current cluster is taken as the first typical conformation of the current cluster; if the total number of conformations in the current cluster is 2, any first sampled conformation in the current cluster is taken as the first typical conformation of the current cluster; if the total number of conformations in the current cluster is greater than 2, the sum of the conformation pair errors of each first sampled conformation in the current cluster and the other first sampled conformations in the cluster is calculated to obtain the corresponding total conformation error, and the first sampled conformation corresponding to the smallest total conformation error is taken as the first typical conformation of the current cluster; and all the obtained first typical conformations form the corresponding typical conformation set.
[0086] Preferably, the step of constructing corresponding mapping feature sets and materialized feature sets for the typical conformation set and the current molecular library based on the feature mapping model and the materialized feature encoder specifically involves:
[0087] Each of the first typical conformations in the typical conformation set is used as the current target RNA conformation and input into the feature mapping model for processing to obtain the corresponding materialized feature mapping tensor. Each of the materialized feature mapping tensors obtained this time is used as a corresponding first mapping feature, and all the obtained first mapping features form the corresponding mapping feature set.
[0088] The first molecular conformations of the current molecular library are input into the physical feature encoder as the current small molecular conformations to obtain the corresponding physical feature tensors. Each physical feature tensor obtained this time is used as a corresponding first physical feature, and all the obtained first physical features form the corresponding physical feature set.
[0089] Preferably, the step of screening molecular conformations in the current molecular library that match the current RNA conformation based on the mapping feature set and the physical feature set to obtain a preferred conformation sequence and then feeding it back to the current user specifically includes:
[0090] Each of the first materialized features in the materialized feature set is taken as the corresponding current feature; and based on a preset similarity algorithm, the similarity between the current feature and each of the first mapping features in the mapping feature set is calculated to obtain the corresponding first similarity; and the largest of the first similarities is taken as the corresponding small molecule fit; the similarity algorithm includes a cosine vector similarity algorithm.
[0091] The first molecular conformations in the current molecular library are sorted in descending order of small molecule fitness; the first S-position first molecular conformations are taken as the corresponding first preferred conformations; and the S first preferred conformations are sorted in descending order of small molecule fitness to obtain the corresponding preferred conformation sequence and fed back to the current user; S is a pre-set positive integer.
[0092] A second aspect of the present invention provides an apparatus for implementing the screening method for targeted RNA small molecules described in the first aspect above. The apparatus includes: a model building module, a first data acquisition module, a first model training module, a second data acquisition module, a second model training module, a data receiving module, an RNA conformation expansion module, a conformation feature encoding module, and a screening feedback module.
[0093] The model building module constructs a property prediction model based on the Uni-Mol model to predict the physicochemical properties of small molecules; and constructs a feature mapping model based on a class of E(3)-equivariant graph neural networks to learn the physicochemical features of RNA-adapted small molecules; the property prediction model is used to predict a set of specified physicochemical properties of the current molecule based on the small molecule conformation input to the model to obtain the corresponding physicochemical property vector; the feature mapping model is used to perform feature mapping on the physicochemical features of the current RNA-adapted small molecule based on the target RNA conformation input to the model to obtain the corresponding physicochemical feature mapping tensor.
[0094] The first data acquisition module constructs a first dataset by collecting large amounts of data on the physicochemical properties of small drug molecules;
[0095] The first model training module trains the feature prediction model based on the first dataset; and extracts the corresponding materialized feature encoder from the feature prediction model at the end of training;
[0096] The second data acquisition module initializes the second dataset by acquiring large amounts of RNA-small molecule complex data and uses the physical and chemical feature encoder to perform label tensor annotation on the second dataset;
[0097] The second model training module trains the feature mapping model based on the second dataset;
[0098] The data receiving module is used to take the RNA conformation input by the user as the corresponding current RNA conformation after both models have been trained; and to take the small molecule database specified by the user as the corresponding current molecule library.
[0099] The RNA conformation expansion module performs molecular dynamics simulation on the current RNA conformation to obtain a corresponding sampled conformation set; and performs conformation clustering and typical conformation extraction on the sampled conformation set to obtain a corresponding typical conformation set.
[0100] The conformation feature encoding module constructs corresponding mapping feature sets and materialization feature sets based on the feature mapping model and the materialization feature encoder for the typical conformation set and the current molecular library;
[0101] The screening and feedback module filters molecular conformations in the current molecular library that match the current RNA conformation based on the mapping feature set and the physical feature set, and obtains the preferred conformation sequence to feed back to the current user.
[0102] A third aspect of the present invention provides an electronic device, including: a memory, a processor, and a transceiver;
[0103] The processor is used to couple with the memory, read and execute instructions in the memory to implement the steps of the method described in the first aspect above;
[0104] The transceiver is coupled to the processor, and the processor controls the transceiver to send and receive messages.
[0105] A fourth aspect of the present invention provides a computer-readable storage medium storing computer instructions that, when executed by a computer, cause the computer to perform the instructions described in the first aspect.
[0106] This invention provides a method, apparatus, electronic device, and computer-readable storage medium for screening targeted RNA small molecules. As described above, this invention first constructs a property prediction model based on the Uni-Mol model to predict the physicochemical properties of small molecules, and then constructs a feature mapping model based on the E(3)-equivariant graph neural network to learn the physicochemical features of RNA-adaptive small molecules. Next, a first dataset is constructed through large-scale data collection of the physicochemical properties of drug small molecules, and the property prediction model is trained based on this dataset. At the end of training, a physicochemical feature encoder is extracted from the model. Then, a second dataset is initialized through large-scale data collection of RNA-small molecule complexes, and the physicochemical feature encoder is used to label the second dataset with tensors. Finally, the model is trained based on the second dataset. The process involves training a feature mapping model. After both models are trained, a rapid simulation method based on molecular dynamics simulation is used to expand the user-specified RNA conformation into a homologous conformation set. This sampled conformation set is then subjected to conformation clustering and typical conformation extraction to obtain a typical conformation set. Next, the feature mapping model and physical feature encoder are used to construct feature sets from the typical conformation set and the user-specified small molecule database, resulting in corresponding mapped feature sets and physical feature sets. Finally, based on the feature similarity comparison results of the mapped feature sets and physical feature sets, small molecule conformations in the small molecule database that match the user-specified RNA conformation are screened and ranked to obtain the preferred conformation sequences, which are then fed back to the user. This invention improves the richness and accuracy of RNA conformations, reduces the small molecule false negative rate, lowers simulation time costs, and improves the overall processing efficiency of batch screening. Attached Figure Description
[0107] Figure 1 This is a schematic diagram of a screening method for targeted RNA small molecules provided in Embodiment 1 of the present invention;
[0108] Figure 2 This is a schematic diagram of the feature prediction model provided in Embodiment 1 of the present invention.
[0109] Figure 3This is a schematic diagram of the feature mapping model provided in Embodiment 1 of the present invention;
[0110] Figure 4 This is a module structure diagram of a screening device for targeting small RNA molecules provided in Embodiment 2 of the present invention;
[0111] Figure 5 This is a schematic diagram of the structure of an electronic device provided in Embodiment 3 of the present invention. Detailed Implementation
[0112] To make the objectives, technical solutions, and advantages of this invention clearer, the invention will be further described in detail below with reference to the accompanying drawings. Obviously, the described embodiments are merely some embodiments of this invention, and not all embodiments. Based on the embodiments of this invention, all other embodiments obtained by those skilled in the art without creative effort are within the scope of protection of this invention.
[0113] Embodiment 1 of the present invention provides a method for screening targeted RNA small molecules, such as... Figure 1 The schematic diagram shows a method for screening targeted RNA small molecules provided in Embodiment 1 of the present invention. The method mainly includes the following steps:
[0114] Step 1: Construct a property prediction model based on the Uni-Mol model to predict the physicochemical properties of small molecules; and construct a feature mapping model based on a class of E(3)-equivariant graph neural networks to learn the physicochemical characteristics of RNA-adapted small molecules.
[0115] Here, the property prediction model of this invention is used to predict the corresponding physicochemical property vector of a specified set of physicochemical properties of the current molecule based on the small molecule conformation input to the model.
[0116] Among them, the small molecule conformation is an atomic-level three-dimensional conformation.
[0117] The atomic-level three-dimensional conformation in this embodiment of the invention is a type of custom three-dimensional conformation data type, whose data structure includes a first atom set; the first atom set includes multiple first atoms; the first atom includes an atom identifier, an atom element type, and atom three-dimensional coordinates.
[0118] The vector length of the materialized property vector is equal to the total number N of property types in the preset materialized property set. F Equal, N F It is a preset positive integer; the materialized property vector includes N. F One predicted value for physical and chemical properties; the set of physical and chemical properties includes N. FEach physicochemical property type corresponds to a class of physical or chemical properties; the predicted value of the physicochemical property corresponds one-to-one with the physicochemical property type. The physical or chemical properties mentioned here include solubility, melting point, boiling point, partition coefficient, polar surface area, acid-base dissociation constant, stability, chirality, reactivity, etc.
[0119] like Figure 2 As shown in the schematic diagram of the characteristic prediction model provided in Embodiment 1 of the present invention, the model input end of the characteristic prediction model is used to receive the small molecule conformation, and the model output end is used to output the corresponding physicochemical characteristic vector.
[0120] like Figure 2 As shown, the model components of the feature prediction model include: a materialized feature encoder and a materialized feature prediction network.
[0121] The connection relationships of the model components in the feature prediction model are as follows: the input end of the materialized feature encoder is connected to the model input end, and the output end is connected to the materialized feature prediction network; the output end of the materialized feature prediction network is connected to the model output end.
[0122] The features of the model components in the feature prediction model are shown below.
[0123] 1) Materialized feature encoder:
[0124] like Figure 2 As shown, the materialized feature encoder consists of a first embedding encoding module, a first encoder, and a first feature mapping network sequentially connected. The first encoder is implemented based on the Uni-Mol model. The first feature mapping network consists of a first pooling layer and a first MLP model sequentially connected.
[0125] It should be noted that the Uni-Mol model is a general encoding model used for feature encoding of the 3D conformation of the model input. Its model input embedding encoding rules, model inference principles, and model pre-training scheme can be understood through technical document B, "A Universal 3D Molecular Representation Learning Framework". The Uni-Mol model in this embodiment of the invention has been pre-trained.
[0126] The materialization feature encoder is used to encode the materialization features of small molecule conformations to obtain the corresponding materialization feature tensor, which is then sent to the materialization property prediction network. Specifically, the first embedding encoding module performs embedding encoding processing on the small molecule conformation according to the embedding encoding rules of the Uni-Mol model input data to obtain the corresponding first embedding encoding tensor, which is then sent to the first encoder. The first encoder performs atomic-level feature encoding processing on the first embedding encoding tensor to obtain the corresponding first feature encoding tensor, which is then sent to the first feature mapping network. The first feature mapping network performs global pooling processing on each feature dimension of the first feature encoding tensor through the first pooling layer to obtain the corresponding first pooling feature vector, and then performs vector mapping of the first pooling feature vector to the materialization feature space through the first MLP model to obtain the corresponding materialization feature tensor, which is then sent to the materialization property prediction network. The shape of the materialization feature tensor is D1×L1; the materialization feature tensor consists of L1 first feature vectors of length D1; and the first feature vectors correspond one-to-one with the materialization property types in the materialization property set.
[0127] 2) Physicochemical property prediction network:
[0128] like Figure 2 As shown, the physical property prediction network has N built-in. F Each first characteristic prediction model corresponds one-to-one with the physical characteristic type of the physical characteristic set; each first characteristic prediction model is implemented based on an MLP model; each first characteristic prediction model corresponds one-to-one with the first feature vector.
[0129] The material properties prediction network is used to input each first eigenvector in the material features tensor into the corresponding first property prediction model to predict the corresponding material properties and obtain the corresponding predicted material properties value. The resulting N... F The predicted values of each physical and chemical property are combined to form the corresponding physical and chemical property vector and output.
[0130] The feature mapping model in this embodiment of the invention is used to perform feature mapping on the physical and chemical features of the small molecules adapted to the current RNA based on the target RNA conformation input by the model, so as to obtain the corresponding physical and chemical feature mapping tensor.
[0131] The target RNA conformation is an atomic-level three-dimensional conformation.
[0132] The materialized feature mapping tensor has a shape of D1×L1, where D1 is the preset first feature dimension, L1 is the preset total number of first sub-vectors, and L1=N. F The materialized feature mapping tensor consists of L1 vectors of length D1, each representing a first mapping vector. Each first mapping vector corresponds one-to-one with a materialized feature type in the materialized feature set.
[0133] like Figure 3As shown in the schematic diagram of the feature mapping model provided in Embodiment 1 of the present invention, the model input end of the feature mapping model is used to receive the target RNA conformation, and the model output end is used to output the corresponding materialized feature mapping tensor.
[0134] like Figure 3 As shown, the model components of the feature mapping model include: a second embedding encoding module, a second encoder, and a second feature mapping network.
[0135] The connection relationships of the model components in the feature mapping model are as follows: the input end of the second embedding encoding module is connected to the input end of the model, and the output end is connected to the input end of the second encoder; the output end of the second encoder is connected to the input end of the second feature mapping network; and the output end of the second feature mapping network is connected to the output end of the model.
[0136] The features of the model components in the feature mapping model are shown below.
[0137] 1) Second Embedded Encoding Module:
[0138] The second embedding coding module performs embedding coding processing on the target RNA conformation according to the embedding coding rules of the model input data of the second encoder, and sends the corresponding second embedding coding tensor to the second encoder.
[0139] 2) Second encoder:
[0140] The second encoder in this embodiment of the invention is implemented based on a class of E(3)-equivariant graph neural networks; the E(3)-equivariant graph neural networks mentioned herein include EGNN model, SE(3)-Transformer model, EGAT model, Cormorant model, etc.
[0141] The second encoder performs atomic-level feature encoding processing on the second embedding encoding tensor to obtain the corresponding second feature encoding tensor, which is then sent to the second feature mapping network.
[0142] 3) Second feature mapping network:
[0143] The second feature mapping network in this embodiment of the invention is formed by sequentially connecting a second pooling layer and a second MLP model.
[0144] The second feature mapping network performs global pooling on each feature dimension of the second feature encoding tensor through the second pooling layer to obtain the corresponding second pooling feature vector. Then, it performs vector mapping on the second pooling feature vector through the second MLP model to obtain the corresponding materialized feature mapping tensor and outputs it.
[0145] Step 2: Construct the first dataset by collecting large amounts of data on the physicochemical properties of small drug molecules;
[0146] The first dataset includes multiple first data records; each first data record corresponds to a class of small molecules; each first data record includes a first molecule sequence, a first training conformation, and a first label vector; the first molecule sequence is the one-dimensional molecular sequence of the current small molecule; the first training conformation is the atomic-level three-dimensional conformation of the current small molecule; the first label vector is the physicochemical property label vector of the current small molecule, consisting of N... F It consists of a set of physical property tag values; each physical property tag value corresponds one-to-one with the physical property type of the physical property set;
[0147] Specifically, it includes: Step 21, collecting the molecular sequence and / or molecular conformation of drug small molecules and the physicochemical properties that satisfy the physicochemical property set through the publicly available small molecule data collection method to obtain the corresponding first collection dataset;
[0148] Among them, the publicly available small molecule data acquisition channels include multiple publicly available small molecule databases, multiple publicly available small molecule experimental databases, and multiple publicly available scientific and technological literature / technical paper databases in the fields of biochemistry and medicine.
[0149] The first acquisition dataset includes multiple first acquisition records; each first acquisition record corresponds to a class of drug small molecules; the first acquisition record includes the small molecule sequence, small molecule conformation, and small molecule property set; the small molecule sequence is the one-dimensional molecular sequence of the current drug small molecule; the small molecule conformation is the atomic-level three-dimensional conformation of the current drug small molecule; the small molecule property set consists of N... F It consists of small molecule characteristic values, and each small molecule characteristic value corresponds one-to-one with the physicochemical characteristic type of the physicochemical characteristic set; the small molecule sequence and small molecule characteristic set of each first acquisition record are mandatory and cannot be empty, while the small molecule conformation is optional and can be empty;
[0150] Step 22: Identify whether the small molecule conformation of each first acquisition record in the first acquisition dataset is empty. If so, perform atomic-level three-dimensional conformation simulation based on the small molecule sequence of the current first acquisition record using a preset molecular dynamics simulation tool, and use the simulation result as the corresponding current simulation conformation. Then, set the small molecule conformation of the current first acquisition record based on the current simulation conformation.
[0151] Here, the molecular dynamics simulation tools in this embodiment of the invention include GROMACS, AMBER, NAMD, CHARMM, and LAMMPS.
[0152] Step 23: Take each of the first acquisition records in the first acquisition dataset as the corresponding current acquisition record; and take the small molecule sequence, small molecule conformation and small molecule characteristic set of the current acquisition record as a set of corresponding first molecule sequence, first training conformation and first label vector to form a corresponding first data record; and take all the obtained first data records to form the corresponding first dataset.
[0153] Step 3: Train the feature prediction model based on the first dataset; and extract the corresponding materialized feature encoder from the feature prediction model at the end of training.
[0154] Specifically, this includes: Step 31, training a feature prediction model based on the first dataset;
[0155] Specifically, it includes: Step 311, randomly dividing the first dataset into two subsets based on a preset first segmentation ratio, denoted as the first training set and the first evaluation set;
[0156] Wherein, the first segmentation ratio is a pre-set ratio parameter, such as 8:2; both the first training set and the first evaluation set consist of multiple first data records; the ratio of the total number of records in the first training set and the first evaluation set satisfies the first segmentation ratio;
[0157] Step 312: Take each first data record of the first training set as the corresponding current training record; take the first training conformation of the current training record as the current small molecule conformation input property prediction model for processing to obtain the corresponding physical property vector; and form a corresponding first prediction-label pair by the current physical property vector and the first label vector of the current training record.
[0158] Step 313: Input all the obtained first prediction-label pairs into the preset first model loss function to calculate the corresponding first loss value;
[0159] The first model loss function is implemented based on either the L1 loss function or the L2 loss function.
[0160] Step 314: Identify whether the first loss value meets the preset first loss value range; if it does, proceed to step 315; if it does not, based on the preset first model optimizer, perform a round of modulation on the model parameters of the feature prediction model in the direction of minimizing the first model loss function, and return to step 312 to continue training when the first round of modulation ends.
[0161] The first loss value range is a pre-set numerical range; the first model optimizer includes the Adam optimizer and the SGD optimizer.
[0162] Step 315: Take each first data record of the first evaluation set as the corresponding current evaluation record; take the first training conformation of the current evaluation record as the current small molecule conformation input property prediction model for processing to obtain the corresponding physicochemical property vector; and form a corresponding second prediction-label pair by the current physicochemical property vector and the first label vector of the current evaluation record; and input all the obtained second prediction-label pairs into the preset first model evaluation function to calculate the corresponding first evaluation value.
[0163] The first model evaluation function is implemented based on the MAE function, MSE function, or RMSE function.
[0164] Step 316: Identify whether the first evaluation value meets the preset first evaluation value range; if not, return to step 311 to continue training; if it meets the range, confirm that the training of the feature prediction model has ended.
[0165] Here, the first evaluation value range is a pre-set numerical range;
[0166] Step 32, and at the end of training, extract the corresponding materialized feature encoder from the feature prediction model.
[0167] Step 4: Initialize the second dataset by collecting large amounts of data on RNA-small molecule complexes and use a physical and chemical feature encoder to label the second dataset with tensors;
[0168] The second dataset includes multiple second data records; each second data record corresponds to a class of RNA-small molecule complexes; each second data record includes a first complex conformation, a first RNA conformation sampling set, a first small molecule conformation, and a first tag tensor; the first complex conformation is the atomic-level three-dimensional conformation of the current RNA-small molecule complex; the first small molecule conformation is the atomic-level three-dimensional conformation of the small molecules in the current RNA-small molecule complex, and the first small molecule conformation is a local conformation of the first complex conformation; the remaining local conformations in the first complex conformation, excluding the first small molecule conformation, represent the current RNA-small molecule complex. - The atomic-level three-dimensional conformation of RNA in a small molecule complex is denoted as the first RNA conformation; the first RNA conformation sampling set includes multiple first RNA sampling conformations; each first RNA sampling conformation is a local conformation of the first RNA conformation; the sampling rule is as follows: in the first complex conformation, with the geometric center point of the first small molecule conformation as the center of a sphere, corresponding sampling spherical regions are set according to the radii of each sphere in a preset set of sphere radii, and the local conformations of the first RNA conformation in each sampling spherical region are extracted as a corresponding first RNA sampling conformation; the set of sphere radii includes multiple sphere radii, such as... The first label tensor is the physical feature tensor corresponding to the first small molecule conformation, which is obtained by the physical feature encoder performing feature encoding on the first small molecule conformation.
[0169] The current step 4 specifically includes:
[0170] Step 41: The three-dimensional conformation of the RNA-small molecule complex is collected using publicly available RNA-small molecule complex data acquisition methods to obtain the corresponding first collection conformation set.
[0171] Among them, the publicly available RNA-small molecule complex data acquisition methods include multiple publicly available RNA-small molecule complex databases, multiple publicly available RNA-small molecule complex experimental databases, and multiple publicly available scientific literature / technical paper databases in the fields of biochemistry and medicine.
[0172] The first acquisition conformation set includes multiple first acquisition conformations; each first acquisition conformation is an atomic-level three-dimensional conformation of a class of RNA-small molecule complexes; each first acquisition conformation consists of a pair of RNA subconformations and small molecule subconformations;
[0173] Step 42: Each first-collected conformation in the first-collected conformation set is taken as the corresponding current complex conformation; the current complex conformation, along with the corresponding RNA sub-conformation and small molecule sub-conformation, are taken as a set of corresponding first complex conformation, first RNA conformation, and first small molecule conformation; within the current first complex conformation, with the geometric center point of the current first small molecule conformation as the sphere center, corresponding sampling spherical regions are set according to the sphere radii of each sphere in the sphere radius set, and the local conformations of the current first RNA conformation within each sampling spherical region are extracted as a corresponding first RNA sampling conformation; all the first RNA sampling conformations obtained this time form a corresponding first RNA conformation sampling set; a corresponding first label tensor is set to empty; the first complex conformation, the first RNA conformation sampling set, the first small molecule conformation, and the first label tensor corresponding to the current complex conformation form a corresponding second data record; and all the obtained second data records form the initialized second dataset.
[0174] Step 43: Take each second data record of the second dataset as the corresponding current record; input the first small molecule conformation of the current record as the current small molecule conformation into the materialization feature encoder for processing to obtain the corresponding materialization feature tensor; and set the first label tensor of the current record based on the current materialization feature tensor.
[0175] Step 5: Train the feature mapping model based on the second dataset;
[0176] Specifically, it includes: Step 51, randomly dividing the second dataset into two subsets based on a preset second segmentation ratio, denoted as the corresponding second training set and second evaluation set;
[0177] The second segmentation ratio is a pre-set ratio parameter, such as 8:2; both the second training set and the second evaluation set consist of multiple second data records; the ratio of the total number of records in the second training set and the second evaluation set satisfies the second segmentation ratio.
[0178] Step 52: Take each second data record of the second training set as the corresponding current training record; and take each first RNA sample conformation of the first RNA conformation sampling set of the current training record as the current target RNA conformation and input it into the feature mapping model to obtain the corresponding materialized feature mapping tensor; and form a corresponding third prediction-label pair with each materialized feature mapping tensor corresponding to the current training record and the first label tensor of the current training record.
[0179] Step 53: Input all the obtained third prediction-label pairs into the preset second model loss function to calculate the corresponding second loss value;
[0180] The second model loss function is based on the L2 loss function.
[0181] Step 54: Identify whether the second loss value meets the preset range of the second loss value; if it does, proceed to step 55; if it does not, adjust the model parameters of the feature mapping model in one round based on the preset second model optimizer in the direction of minimizing the second model loss function, and return to step 52 to continue training when the adjustment ends.
[0182] The second loss value range is a pre-set numerical range; the second model optimizer includes the Adam optimizer and the SGD optimizer.
[0183] Step 55: Take each second data record of the second evaluation set as the corresponding current evaluation record; take each first RNA sample conformation of the first RNA conformation sampling set of the current evaluation record as the current target RNA conformation and input it into the feature mapping model to obtain the corresponding materialized feature mapping tensor; and form a corresponding fourth prediction-label pair with each materialized feature mapping tensor corresponding to the current evaluation record and the first label tensor of the current evaluation record; and input all the obtained fourth prediction-label pairs into the preset second model evaluation function to calculate the corresponding second evaluation value.
[0184] The second model evaluation function is based on the RMSE function.
[0185] Step 56: Identify whether the second evaluation value meets the preset range of the second evaluation value; if not, proceed to step 51 to continue training; if it meets the range, confirm that the training of the feature mapping model has ended.
[0186] Here, the second evaluation value range is a pre-set numerical range.
[0187] Step 6: After both models have been trained, the RNA conformation input by the user is used as the corresponding current RNA conformation; and the small molecule database specified by the user is used as the corresponding current molecular library.
[0188] The current RNA conformation is an atomic-level three-dimensional conformation. The current molecular library includes multiple first-molecule conformations; each first-molecule conformation is an atomic-level three-dimensional conformation of a class of small molecules.
[0189] Step 7: Perform molecular dynamics simulation on the current RNA conformation to obtain the corresponding sampled conformation set; and perform conformation clustering and typical conformation extraction on the sampled conformation set to obtain the corresponding typical conformation set.
[0190] Specifically, this includes: Step 71, performing molecular dynamics simulations on the current RNA conformation to obtain the corresponding sampled conformation set;
[0191] The sampled conformation set includes multiple first sampled conformations;
[0192] Specifically, this includes: using molecular dynamics simulation tools, performing a simulation of the current RNA conformation within a preset fixed simulation time using a preset rapid simulation method; and during the simulation, collecting real-time conformation data of the dynamically changing simulated conformation at preset sampling time intervals; and using each collected real-time simulated conformation as a corresponding first sample conformation; and at the end of this simulation round, also using the current RNA conformation as a corresponding first sample conformation, and combining all the obtained first sample conformations to form a corresponding sample conformation set;
[0193] The fast simulation methods include Gaussian Accelerated Molecular Dynamics (GaMD), Metadynamics (MetaD), High-Temperature Molecular Dynamics (HT-MD), and Replica Exchange Molecular Dynamics with Temperature (REMD or T-REMD); the simulation duration and sampling time interval are fixed to two pre-set duration parameters.
[0194] Step 72, and perform conformation clustering and typical conformation extraction on the sampled conformation set to obtain the corresponding typical conformation set;
[0195] The typical conformation set includes one or more first typical conformations;
[0196] Specifically, this includes: step 721, which involves statistically analyzing the total number of the first sampled conformations in the sampled conformation set to obtain the corresponding total number of conformations N. M ; and based on the total number of conformations N M Construct the error matrix D;
[0197] The error matrix D includes N M ×N M d matrix elements i,j ; 1 ≤ row index i ≤ N M 1 ≤ column index j ≤ N M Each row or column of the error matrix D corresponds one-to-one with the first sampling configuration; the matrix elements d on the diagonal of the error matrix D i,j=i All are 0; the symmetric elements of each pair of diagonals in the error matrix D are equal: d i,j≠i =d j≠i,i ;
[0198] Step 722: Combine the first sampled conformations of the sampled conformation set in pairs to obtain multiple conformation pairs; perform conformation alignment on the two first sampled conformations of each conformation pair, and calculate the root mean square error of the atomic coordinates of the two aligned conformations to obtain the corresponding conformation pair error; record each conformation pair error as the corresponding current error, and record the two diagonal symmetric elements in the error matrix D corresponding to the current error as the corresponding current element pair, and set the two matrix elements of the current element pair based on the current error;
[0199] Step 723, and perform a check on all matrix elements d of the error matrix D. i,j Perform mean calculation and use the result as the corresponding error threshold d. hold ;
[0200] Step 724, and based on the fully linked hierarchical clustering algorithm, according to the error threshold d hold Clustering all first sampled conformations in the sampled conformation set yields multiple corresponding first conformation clusters;
[0201] Each first conformation cluster includes one or more first sampled conformations; in each first conformation cluster containing two or more first sampled conformations, the conformation pair error between any two first sampled conformations does not exceed the error threshold d. hold There exists at least two conformation pairs between every two first conformation clusters where the error of the first sampled conformation is greater than the error threshold d. hold ;
[0202] Step 725: Each first conformation cluster is taken as the corresponding current cluster; the total number of first sampled conformations in the current cluster is counted to obtain the total number of conformations in the current cluster; the total number of conformations in the current cluster is identified; if the total number of conformations in the current cluster is 1, the unique first sampled conformation in the current cluster is taken as the first typical conformation of the current cluster; if the total number of conformations in the current cluster is 2, any first sampled conformation in the current cluster is taken as the first typical conformation of the current cluster; if the total number of conformations in the current cluster is greater than 2, the sum of conformation pair errors between each first sampled conformation in the current cluster and other first sampled conformations in the cluster is calculated to obtain the corresponding total conformation error, and the first sampled conformation corresponding to the smallest total conformation error is taken as the first typical conformation of the current cluster.
[0203] Step 726, and the corresponding set of typical conformations is formed by all the first typical conformations obtained.
[0204] Step 8: Based on the feature mapping model and the materialized feature encoder, construct the feature sets of the typical conformation set and the current molecular library to obtain the corresponding mapped feature set and materialized feature set;
[0205] The mapping feature set includes one or more first mapping features; each first mapping feature corresponds one-to-one with the first typical conformation of the typical conformation set; the physical feature set includes multiple first physical features; each first physical feature corresponds one-to-one with the first molecular conformation of the current molecular library.
[0206] Specifically, step 81 involves taking each of the first typical conformations of the typical conformation set as the current target RNA conformation input feature mapping model for processing to obtain the corresponding materialized feature mapping tensor, and taking each of the materialized feature mapping tensors obtained this time as a corresponding first mapping feature, and forming the corresponding mapping feature set by all the obtained first mapping features.
[0207] Step 82: Input each first molecular conformation of the current molecular library as the current small molecule conformation into the physical feature encoder for processing to obtain the corresponding physical feature tensor. Then, take each physical feature tensor obtained this time as a corresponding first physical feature, and form the corresponding physical feature set by all the obtained first physical features.
[0208] Step 9: Based on the mapping feature set and the physical feature set, screen the molecular conformations in the current molecular library that are compatible with the current RNA conformation to obtain the preferred conformation sequence and feed it back to the current user.
[0209] The preferred conformation sequence is formed by sequentially sorting multiple first preferred conformations;
[0210] Specifically, it includes: step 91, taking each first materialized feature of the materialized feature set as the corresponding current feature; and based on the preset similarity algorithm, calculating the similarity between the current feature and each first mapping feature of the mapping feature set to obtain the corresponding first similarity; and taking the largest first similarity as the corresponding small molecule fit.
[0211] Among them, the similarity algorithm includes the cosine vector similarity algorithm;
[0212] Step 92: Sort all first molecular conformations in the current molecular library in descending order of small molecule fitness; take the first S first molecular conformations as the corresponding first preferred conformations; sort the S first preferred conformations in descending order of small molecule fitness to obtain the corresponding preferred conformation sequence and feed it back to the current user.
[0213] Here, S is a pre-set positive integer.
[0214] Figure 4 This is a module structure diagram of a screening device for targeting small RNA molecules provided in Embodiment 2 of the present invention. This device can be a terminal device or server implementing the aforementioned method embodiments, or it can be a device that enables the aforementioned terminal device or server to implement the aforementioned method embodiments. For example, the device can be a device or chip system of the aforementioned terminal device or server. Figure 4 As shown, the device includes: a model building module 201, a first data acquisition module 202, a first model training module 203, a second data acquisition module 204, a second model training module 205, a data receiving module 206, an RNA conformation expansion module 207, a conformation feature encoding module 208, and a screening feedback module 209.
[0215] The model building module 201 constructs a property prediction model based on the Uni-Mol model to predict the physicochemical properties of small molecules; and constructs a feature mapping model based on a class of E(3)-equivariant graph neural networks to learn the physicochemical features of RNA-adapted small molecules; the property prediction model is used to predict a set of specified physicochemical properties of the current molecule based on the small molecule conformation input by the model to obtain the corresponding physicochemical property vector; the feature mapping model is used to perform feature mapping on the physicochemical features of the current RNA-adapted small molecule based on the target RNA conformation input by the model to obtain the corresponding physicochemical feature mapping tensor.
[0216] The first data acquisition module 202 constructs the first dataset by collecting big data on the physicochemical properties of small drug molecules.
[0217] The first model training module 203 trains a feature prediction model based on the first dataset; and extracts the corresponding materialized feature encoder from the feature prediction model at the end of training.
[0218] The second data acquisition module 204 initializes the second dataset by acquiring large amounts of RNA-small molecule complex data and uses a physical and chemical feature encoder to label the second dataset with tensors.
[0219] The second model training module 205 trains the feature mapping model based on the second dataset.
[0220] The data receiving module 206 is used to take the RNA conformation input by the user as the corresponding current RNA conformation after both models have been trained, and to take the small molecule database specified by the user as the corresponding current molecule library.
[0221] The RNA conformation expansion module 207 performs molecular dynamics simulations on the current RNA conformation to obtain the corresponding sampled conformation set; and performs conformation clustering and typical conformation extraction on the sampled conformation set to obtain the corresponding typical conformation set.
[0222] The conformation feature encoding module 208 constructs the corresponding mapping feature set and physical feature set based on the feature mapping model and the physical feature encoder for the typical conformation set and the current molecular library.
[0223] The screening feedback module 209 filters molecular conformations in the current molecular library that are compatible with the current RNA conformation based on the mapping feature set and the physical feature set, and then feeds back the preferred conformation sequence to the current user.
[0224] The present invention provides a screening device for targeting small RNA molecules, which can perform the method steps in the above method embodiments. Its implementation principle and technical effect are similar, and will not be repeated here.
[0225] It should be noted that the division of the various modules in the above device is merely a logical functional division. In actual implementation, they can be fully or partially integrated into a single physical entity, or they can be physically separated. Furthermore, these modules can be implemented entirely in software via processing elements; they can be fully implemented in hardware; or some modules can be implemented by processing elements calling software, while others are implemented in hardware. For example, the model building module can be a separate processing element, or it can be integrated into a chip in the above device. Alternatively, it can be stored as program code in the memory of the above device, and its functions can be called and executed by a processing element. The implementation of other modules is similar. Moreover, these modules can be fully or partially integrated together, or they can be implemented independently. The processing element described here can be an integrated circuit with signal processing capabilities. In the implementation process, each step of the above method or each of the above modules can be completed through integrated logic circuits in the hardware of the processor element or through software instructions.
[0226] For example, these modules can be one or more integrated circuits configured to implement the above methods, such as one or more application-specific integrated circuits (ASICs), one or more digital signal processors (DSPs), or one or more field-programmable gate arrays (FPGAs). As another example, when a module is implemented using processing element scheduler code, the processing element can be a general-purpose processor, such as a central processing unit (CPU) or other processor capable of calling program code. Furthermore, these modules can be integrated together as a system-on-a-chip (SOC).
[0227] In the above embodiments, implementation can be achieved, in whole or in part, through software, hardware, firmware, or any combination thereof. When implemented in software, it can be implemented, in whole or in part, as a computer program product. This computer program product includes one or more computer instructions. When these computer program instructions are loaded and executed on a computer, all or part of the processes or functions described in the foregoing method embodiments are generated. The computer described above can be a general-purpose computer, a special-purpose computer, a computer network, or other programmable device. The aforementioned computer instructions can be stored in a computer-readable storage medium or transmitted from one computer-readable storage medium to another. For example, the aforementioned computer instructions can be transmitted from one website, computer, server, or data center to another via wired (e.g., coaxial cable, fiber optic, Digital Subscriber Line (DSL)) or wireless (e.g., infrared, wireless, Bluetooth, microwave, etc.) means. The aforementioned computer-readable storage medium can be any available medium that a computer can access, or a data storage device such as a server or data center that integrates one or more available media. The aforementioned available media can be magnetic media (e.g., floppy disks, hard disks, magnetic tapes), optical media (e.g., DVDs), or semiconductor media (e.g., solid-state disks (SSDs)).
[0228] Figure 5This is a schematic diagram of an electronic device provided in Embodiment 3 of the present invention. This electronic device can be a terminal device or server implementing the methods of the aforementioned embodiments, or it can be a terminal device or server connected to the aforementioned terminal device or server implementing the methods of the aforementioned embodiments. Figure 5 As shown, the electronic device may include: a processor 301 (e.g., CPU), a memory 302, and a transceiver 303; the transceiver 303 is coupled to the processor 301, and the processor 301 controls the transmission and reception operations of the transceiver 303. The memory 302 may store various instructions for performing various processing functions and implementing the processing steps described in the foregoing embodiments. Preferably, the electronic device involved in the embodiments of the present invention further includes: a power supply 304, a system bus 305, and a communication port 306. The system bus 305 is used to realize communication connections between components. The communication port 306 is used for communication between the electronic device and other peripherals.
[0229] exist Figure 5 The system bus 305 mentioned can be a Peripheral Component Interconnect (PCI) bus or an Extended Industry Standard Architecture (EISA) bus, etc. This system bus can be divided into address bus, data bus, control bus, etc. For ease of representation, it is represented by only one thick line in the figure, but this does not indicate that there is only one bus or one type of bus. The communication interface is used to enable communication between the database access device and other devices (e.g., clients, read-write libraries, and read-only libraries). Memory may include Random Access Memory (RAM) and may also include Non-Volatile Memory, such as at least one disk storage device.
[0230] The processors mentioned above can be general-purpose processors, including central processing units (CPUs), network processors (NPs), graphics processing units (GPUs), etc.; they can also be digital signal processors (DSPs), application-specific integrated circuits (ASICs), field-programmable gate arrays (FPGAs), or other programmable logic devices, discrete gate or transistor logic devices, or discrete hardware components.
[0231] It should be noted that the embodiments of the present invention also provide a computer-readable storage medium storing instructions that, when run on a computer, cause the computer to perform the methods and processes provided in the above embodiments.
[0232] This invention provides a method, apparatus, electronic device, and computer-readable storage medium for screening targeted RNA small molecules. As described above, this invention first constructs a property prediction model based on the Uni-Mol model to predict the physicochemical properties of small molecules, and then constructs a feature mapping model based on the E(3)-equivariant graph neural network to learn the physicochemical features of RNA-adaptive small molecules. Next, a first dataset is constructed through large-scale data collection of the physicochemical properties of drug small molecules, and the property prediction model is trained based on this dataset. At the end of training, a physicochemical feature encoder is extracted from the model. Then, a second dataset is initialized through large-scale data collection of RNA-small molecule complexes, and the physicochemical feature encoder is used to label the second dataset with tensors. Finally, the model is trained based on the second dataset. The process involves training a feature mapping model. After both models are trained, a rapid simulation method based on molecular dynamics simulation is used to expand the user-specified RNA conformation into a homologous conformation set. This sampled conformation set is then subjected to conformation clustering and typical conformation extraction to obtain a typical conformation set. Next, the feature mapping model and physical feature encoder are used to construct feature sets from the typical conformation set and the user-specified small molecule database, resulting in corresponding mapped feature sets and physical feature sets. Finally, based on the feature similarity comparison results of the mapped feature sets and physical feature sets, small molecule conformations in the small molecule database that match the user-specified RNA conformation are screened and ranked to obtain the preferred conformation sequences, which are then fed back to the user. This invention improves the richness and accuracy of RNA conformations, reduces the small molecule false negative rate, lowers simulation time costs, and improves the overall processing efficiency of batch screening.
[0233] The steps of the methods or algorithms described in conjunction with the embodiments disclosed herein can be implemented in hardware, a software module executed by a processor, or a combination of both. The software module can be located in random access memory (RAM), main memory, read-only memory (ROM), electrically programmable ROM, electrically erasable programmable ROM, registers, hard disk, removable disk, CD-ROM, or any other form of storage medium known in the art.
[0234] The specific implementation methods described above further illustrate the objectives, technical solutions and beneficial effects of the present invention in detail. It should be understood that the above description is only a specific implementation method of the present invention and is not intended to limit the scope of protection of the present invention. Any modifications, equivalent substitutions, improvements, etc. made within the spirit and principles of the present invention should be included in the scope of protection of the present invention.
Claims
1. A method for screening targeted RNA small molecules, characterized in that, The method includes: A property prediction model for predicting the physicochemical properties of small molecules is constructed based on the Uni-Mol model; and a feature mapping model for learning the physicochemical features of RNA-adapted small molecules is constructed based on a class of E(3)-equivariant graph neural networks; the property prediction model is used to predict a set of specified physicochemical properties of the current molecule based on the small molecule conformation input to the model to obtain the corresponding physicochemical property vector; the feature mapping model is used to perform feature mapping on the physicochemical features of the small molecule adapted to the current RNA based on the target RNA conformation input to the model to obtain the corresponding physicochemical feature mapping tensor. The first dataset was constructed by collecting large amounts of data on the physicochemical properties of small drug molecules; The feature prediction model is trained based on the first dataset; and the corresponding materialized feature encoder is extracted from the feature prediction model at the end of the training. The second dataset is initialized by collecting large amounts of data on RNA-small molecule complexes and then labeled with a label tensor using the physical feature encoder described above. The feature mapping model is trained based on the second dataset; After both models have been trained, the RNA conformation input by the user will be used as the corresponding current RNA conformation; and the small molecule database specified by the user will be used as the corresponding current molecular library. Molecular dynamics simulations are performed on the current RNA conformation to obtain a corresponding sampled conformation set; and conformation clustering and typical conformation extraction are performed on the sampled conformation set to obtain a corresponding typical conformation set; Based on the feature mapping model and the materialized feature encoder, feature sets are constructed for the typical conformation set and the current molecular library to obtain the corresponding mapping feature set and materialized feature set; Based on the mapping feature set and the physical feature set, the molecular conformations in the current molecular library that are compatible with the current RNA conformation are screened to obtain the preferred conformation sequence, which is then fed back to the current user.
2. The method for screening targeted RNA small molecules according to claim 1, characterized in that, The small molecule conformation and the target RNA conformation are each an atomic-level three-dimensional conformation; The atomic-level three-dimensional conformation includes a first set of atoms; the first set of atoms includes multiple first atoms; the first atom includes an atom identifier, an atom element type, and atom three-dimensional coordinates. The length of the materialized characteristic vector is equal to the total number N of the characteristic types in the preset materialized characteristic set. F Equal; N F The value is a preset positive integer; the materialized characteristic vector includes N. F The set of physical and chemical properties includes N predicted values. F Each physical and chemical property type corresponds to a class of physical or chemical properties; the predicted value of each physical and chemical property corresponds one-to-one with the physical and chemical property type. The materialized feature mapping tensor has a shape of D1×L1, where D1 is a preset first feature dimension, L1 is a preset total number of first sub-vectors, and L1=N. F The materialized feature mapping tensor consists of L1 first mapping vectors of length D1; the first mapping vectors correspond one-to-one with the materialized feature types of the materialized feature set. The first dataset includes multiple first data records; each first data record corresponds to a class of small molecules; the first data record includes a first molecule sequence, a first training conformation, and a first label vector; the first molecule sequence is the one-dimensional molecular sequence of the current small molecule; the first training conformation is the atomic-level three-dimensional conformation of the current small molecule; the first label vector is the physicochemical property label vector of the current small molecule, consisting of N... F It consists of a set of physical property tag values; each physical property tag value corresponds one-to-one with the physical property type of the physical property set; The second dataset includes multiple second data records; each second data record corresponds to a class of RNA-small molecule complexes; the second data record includes a first complex conformation, a first RNA conformation sampling set, a first small molecule conformation, and a first tag tensor; the first complex conformation is the atomic-level three-dimensional conformation of the current RNA-small molecule complex; The first small molecule conformation is the atomic-level three-dimensional conformation of the small molecule in the current RNA-small molecule complex, and the first small molecule conformation is a local conformation of the first complex conformation; the remaining local conformations in the first complex conformation other than the first small molecule conformation are the atomic-level three-dimensional conformations of the RNA in the current RNA-small molecule complex, denoted as the first RNA conformation; the first RNA conformation sampling set includes multiple first RNA sampling conformations; each first RNA sampling conformation is a local conformation of the first RNA conformation; the sampling rule is: in the first complex conformation, with the geometric center point of the first small molecule conformation as the center of a sphere, corresponding sampling spherical regions are set according to the radii of each sphere in a preset sphere radius set, and the local conformations of the first RNA conformation in each of the sampling spherical regions are extracted as a corresponding first RNA sampling conformation; the sphere radius set includes multiple sphere radii; the first label tensor is the physical feature tensor corresponding to the first small molecule conformation, obtained by the physical feature encoder performing feature encoding processing on the first small molecule conformation. The current RNA conformation is an atomic-level three-dimensional conformation; The current molecular library includes multiple first molecular conformations; each first molecular conformation is the atomic-level three-dimensional conformation of a class of small molecules; The sampled conformation set includes multiple first sampled conformations; the typical conformation set includes one or more first typical conformations. The mapping feature set includes one or more first mapping features; each of the first mapping features corresponds one-to-one with the first typical conformation. The set of physical features includes multiple first physical features; each of the first physical features corresponds one-to-one with the first molecular conformation. The preferred conformation sequence is formed by sequentially sorting multiple first preferred conformations.
3. The method for screening targeted RNA small molecules according to claim 2, characterized in that, The input terminal of the property prediction model is used to receive the small molecule conformation, and the output terminal is used to output the corresponding physicochemical property vector. The feature prediction model includes the materialized feature encoder and the materialized feature prediction network; the input of the materialized feature encoder is connected to the input of the model, and the output is connected to the materialized feature prediction network; the output of the materialized feature prediction network is connected to the output of the model. The materialized feature encoder is composed of a first embedding encoding module, a first encoder, and a first feature mapping network connected sequentially; the first encoder is implemented based on the Uni-Mol model; the Uni-Mol model has been pre-trained; the first feature mapping network is composed of a first pooling layer and a first MLP model connected sequentially. The materialization feature encoder is used to encode the materialization features of the small molecule conformation to obtain the corresponding materialization feature tensor and send it to the materialization property prediction network. Specifically, the first embedding encoding module performs embedding encoding processing on the small molecule conformation according to the embedding encoding rules of the Uni-Mol model input data to obtain the corresponding first embedding encoding tensor and sends it to the first encoder; the first encoder performs atomic-level feature encoding processing on the first embedding encoding tensor to obtain the corresponding first feature encoding tensor and sends it to the first feature mapping network; the first feature mapping network performs global pooling processing on each feature dimension of the first feature encoding tensor through the first pooling layer to obtain the corresponding first pooling feature vector, and performs vector mapping of the first pooling feature vector on the materialization feature space through the first MLP model to obtain the corresponding materialization feature tensor and sends it to the materialization property prediction network; the shape of the materialization feature tensor is D1×L1; the materialization feature tensor is composed of L1 first feature vectors of length D1; the first feature vectors correspond one-to-one with the materialization property types of the materialization property set. The physical property prediction network has built-in N F Each first characteristic prediction model corresponds one-to-one with the physical characteristic type of the physical characteristic set; each first characteristic prediction model is implemented based on an MLP model; each first characteristic prediction model corresponds one-to-one with the first feature vector. The physical property prediction network is used to input each of the first feature vectors in the physical feature tensor into the corresponding first property prediction model to predict the corresponding physical property, obtain the corresponding predicted value of the physical property, and then use the obtained N... F The predicted values of the physical and chemical properties are combined to form the corresponding physical and chemical property vector and output.
4. The method for screening targeted RNA small molecules according to claim 2, characterized in that, The input end of the feature mapping model is used to receive the target RNA conformation, and the output end is used to output the corresponding materialized feature mapping tensor. The feature mapping model includes a second embedding encoding module, a second encoder, and a second feature mapping network; the second encoder is implemented based on a class of E(3)-equivariant graph neural networks; the E(3)-equivariant graph neural networks include EGNN model, SE(3)-Transformer model, EGAT model, and Cormorant model; the second feature mapping network is formed by sequentially connecting a second pooling layer and a second MLP model; The input of the second embedding encoding module is connected to the input of the model, and the output is connected to the input of the second encoder; the output of the second encoder is connected to the input of the second feature mapping network; the output of the second feature mapping network is connected to the output of the model. The second embedding coding module performs embedding coding processing on the target RNA conformation according to the embedding coding rules of the model input data of the second encoder to obtain the corresponding second embedding coding tensor, which is then sent to the second encoder. The second encoder performs atomic-level feature encoding processing on the second embedding encoding tensor to obtain the corresponding second feature encoding tensor, which is then sent to the second feature mapping network. The second feature mapping network performs global pooling on each feature dimension of the second feature encoding tensor through the second pooling layer to obtain the corresponding second pooling feature vector, and then performs vector mapping of the second pooling feature vector to the materialized feature space through the second MLP model to obtain the corresponding materialized feature mapping tensor and output it.
5. The method for screening targeted RNA small molecules according to claim 2, characterized in that, The construction of the first dataset through big data collection of the physicochemical properties of small drug molecules specifically includes: Step 51: Collect large amounts of data on the molecular sequence and / or molecular conformation of drug small molecules and the physicochemical properties that satisfy the set of physicochemical properties through the publicly available small molecule data collection method to obtain the corresponding first collection dataset; The disclosed small molecule data acquisition methods include multiple publicly available small molecule databases, multiple publicly available small molecule experimental databases, and multiple publicly available databases of scientific and technological literature / technical papers in the fields of biochemistry and medicine. The first acquisition dataset includes multiple first acquisition records; each first acquisition record corresponds to a class of drug small molecules; the first acquisition record includes a small molecule sequence, a small molecule conformation, and a small molecule property set; the small molecule sequence is the one-dimensional molecular sequence of the current drug small molecule; the small molecule conformation is the atomic-level three-dimensional conformation of the current drug small molecule; the small molecule property set consists of N... F The set consists of small molecule characteristic values, which correspond one-to-one with the physicochemical characteristic types of the physicochemical characteristic set; each of the small molecule sequences and small molecule characteristic sets in the first collection record is a required option and cannot be empty, while the small molecule conformation is an optional option and can be empty; Step 52: Identify whether the small molecule conformation of each of the first acquisition records in the first acquisition dataset is empty. If so, perform atomic-level three-dimensional conformation simulation based on the small molecule sequence of the current first acquisition record using a preset molecular dynamics simulation tool, and use the simulation result as the corresponding current simulation conformation. Then, set the small molecule conformation of the current first acquisition record based on the current simulation conformation. The molecular dynamics simulation tools include GROMACS, AMBER, NAMD, CHARMM, and LAMMPS. Step 53: Take each of the first acquisition records in the first acquisition dataset as the corresponding current acquisition record; and take the small molecule sequence, the small molecule conformation and the small molecule characteristic set of the current acquisition record as a corresponding first molecule sequence, the first training conformation and the first label vector to form a corresponding first data record; and take all the obtained first data records to form the corresponding first dataset.
6. The method for screening targeted RNA small molecules according to claim 2, characterized in that, Training the feature prediction model based on the first dataset specifically includes: Step 61: Based on a preset first segmentation ratio, the first dataset is randomly divided into two sub-datasets, denoted as the first training set and the first evaluation set. Wherein, both the first training set and the first evaluation set are composed of multiple first data records; the ratio of the total number of records in the first training set and the first evaluation set satisfies the first segmentation ratio; Step 62: Take each of the first data records in the first training set as the corresponding current training record; and take the first training conformation of the current training record as the current small molecule conformation and input it into the property prediction model for processing to obtain the corresponding physical property vector; and form a corresponding first prediction-label pair by the current physical property vector and the first label vector of the current training record. Step 63: Input all the obtained first prediction-label pairs into the preset first model loss function to calculate the corresponding first loss value; The first model loss function is implemented based on the L1 loss function or the L2 loss function; Step 64: Identify whether the first loss value meets the preset first loss value range; if it does, proceed to step 65; if it does not, modulate the model parameters of the feature prediction model in one round based on the preset first model optimizer in the direction of minimizing the first model loss function, and return to step 62 to continue training when the first round of modulation ends. The first model optimizer includes the Adam optimizer and the SGD optimizer. Step 65: Take each of the first data records in the first evaluation set as the corresponding current evaluation record; and take the first training conformation of the current evaluation record as the current small molecule conformation and input it into the property prediction model for processing to obtain the corresponding physical and chemical property vector; and form a corresponding second prediction-label pair by the current physical and chemical property vector and the first label vector of the current evaluation record; and input all the obtained second prediction-label pairs into the preset first model evaluation function to calculate the corresponding first evaluation value; The first model evaluation function is implemented based on the MAE function, MSE function, or RMSE function. Step 66: Identify whether the first evaluation value meets the preset first evaluation value range; if not, return to step 61 to continue training; if it does, confirm that the training of the feature prediction model has ended.
7. The method for screening targeted RNA small molecules according to claim 2, characterized in that, The process of initializing the second dataset through large-scale acquisition of RNA-small molecule complexes and then using the physical feature encoder to perform label tensor annotation on the second dataset specifically includes: Step 71: Collect large data on the three-dimensional conformation of the RNA-small molecule complex using the publicly available RNA-small molecule complex data acquisition method to obtain the corresponding first collection conformation set. The disclosed RNA-small molecule complex data acquisition methods include multiple publicly available RNA-small molecule complex databases, multiple publicly available RNA-small molecule complex experimental databases, and multiple publicly available scientific and technological literature / technical paper databases in the fields of biochemistry and medicine. The first set of acquisition conformations includes multiple first acquisition conformations; each first acquisition conformation is an atomic-level three-dimensional conformation of a class of RNA-small molecule complexes; each first acquisition conformation consists of a pair of RNA subconformations and small molecule subconformations; Step 72: Each of the first acquired conformations in the first acquired conformation set is taken as the corresponding current complex conformation; the current complex conformation, along with the corresponding RNA sub-conformation and the small molecule sub-conformation, is taken as a set of corresponding first complex conformations, first RNA conformations, and first small molecule conformations; within the current first complex conformation, with the geometric center point of the current first small molecule conformation as the sphere center, corresponding sampling spherical regions are set according to the sphere radii of the set of sphere radii, and local conformations of the current first RNA conformation within each of the sampling spherical regions are extracted as a corresponding first RNA sampling conformation; all the first RNA sampling conformations obtained this time form a corresponding first RNA conformation sampling set; a corresponding first label tensor is set to empty; the first complex conformation corresponding to the current complex conformation, the first RNA conformation sampling set, the first small molecule conformation, and the first label tensor form a corresponding second data record; and all the obtained second data records form the initialized second dataset. Step 73: Take each of the second data records in the second dataset as the corresponding current record; input the first small molecule conformation of the current record as the current small molecule conformation into the materialization feature encoder for processing to obtain the corresponding materialization feature tensor; and set the first label tensor of the current record based on the current materialization feature tensor.
8. The method for screening targeted RNA small molecules according to claim 2, characterized in that, The step of training the feature mapping model based on the second dataset specifically includes: Step 81: Based on the preset second segmentation ratio, the second dataset is randomly divided into two subsets, which are denoted as the corresponding second training set and second evaluation set. The second training set and the second evaluation set are each composed of multiple second data records; the ratio of the total number of records in the second training set and the second evaluation set satisfies the second segmentation ratio. Step 82: Take each of the second data records in the second training set as the corresponding current training record; and take each of the first RNA sampling conformations in the first RNA conformation sampling set of the current training record as the current target RNA conformation and input them into the feature mapping model for processing to obtain the corresponding materialized feature mapping tensor; and form a corresponding third prediction-label pair with each of the materialized feature mapping tensors corresponding to the current training record and the first label tensor of the current training record. Step 83: Input all the obtained third prediction-label pairs into the preset second model loss function to calculate the corresponding second loss value; The second model loss function is based on the L2 loss function. Step 84: Identify whether the second loss value meets the preset range of the second loss value; if it does, proceed to step 85; if it does not, adjust the model parameters of the feature mapping model in one round based on the preset second model optimizer in the direction of minimizing the second model loss function, and return to step 82 to continue training when the adjustment is completed. The second model optimizer includes the Adam optimizer and the SGD optimizer; Step 85: Take each of the second data records in the second evaluation set as the corresponding current evaluation record; and take each of the first RNA sampling conformations in the first RNA conformation sampling set of the current evaluation record as the current target RNA conformation and input them into the feature mapping model for processing to obtain the corresponding materialized feature mapping tensor; and form a corresponding fourth prediction-label pair with each of the materialized feature mapping tensors corresponding to the current evaluation record and the first label tensor of the current evaluation record; and input all the obtained fourth prediction-label pairs into the preset second model evaluation function to calculate the corresponding second evaluation value. The second model evaluation function is implemented based on the RMSE function; Step 86: Identify whether the second evaluation value meets the preset second evaluation value range; if not, proceed to step 81 to continue training; if it meets the range, confirm that the training of the feature mapping model has ended.
9. The method for screening targeted RNA small molecules according to claim 2, characterized in that, The process of performing molecular dynamics simulations on the current RNA conformation to obtain the corresponding sampled conformation set specifically includes: Using a preset molecular dynamics simulation tool, a simulation of the current RNA conformation is performed within a preset fixed simulation time using a preset rapid simulation method. During the simulation, the dynamically changing simulated conformation is collected in real time at preset sampling time intervals. Each real-time simulated conformation collected is taken as a corresponding first sampled conformation. At the end of this simulation round, the current RNA conformation is also taken as a corresponding first sampled conformation, and all the obtained first sampled conformations form the corresponding sampled conformation set. The molecular dynamics simulation tools include GROMACS, AMBER, NAMD, CHARMM, and LAMMPS; the fast simulation methods include Gaussian accelerated molecular dynamics simulation, meta-dynamics simulation, high-temperature molecular dynamics simulation, and temperature replication-exchange molecular dynamics simulation.
10. The method for screening targeted RNA small molecules according to claim 2, characterized in that, The process of performing conformation clustering and typical conformation extraction on the sampled conformation set to obtain the corresponding typical conformation set specifically includes: The total number of the first sampled conformations in the sampled conformation set is counted to obtain the corresponding total number of conformations N. M ; and based on the total number N of the conformations M Construct an error matrix D; the error matrix D includes N M ×N M d matrix elements i,j ; 1 ≤ row index i ≤ N M 1 ≤ column index j ≤ N M Each row or column of the error matrix D corresponds one-to-one with the first sampling configuration; the matrix elements d on the diagonal of the error matrix D i,j=i All are 0; the symmetric elements of each pair of diagonals in the error matrix D are equal: d i,j≠i =d j≠i,i ; The first sampled conformations of the sampled conformation set are combined in pairs to obtain multiple conformation pairs; the two first sampled conformations of each conformation pair are aligned, and the root mean square error of the atomic coordinates of the two aligned conformations is calculated to obtain the corresponding conformation pair error; each conformation pair error is recorded as the corresponding current error, and the two diagonal symmetric elements in the error matrix D corresponding to the current error are recorded as the corresponding current element pair, and the two matrix elements of the current element pair are set based on the current error; And for all matrix elements d of the error matrix D i,j Perform mean calculation and use the result as the corresponding error threshold d. hold ; And based on the fully linked hierarchical clustering algorithm, according to the error threshold d hold Clustering all the first sampled conformations in the sampled conformation set yields multiple corresponding first conformation clusters; wherein each first conformation cluster includes one or more first sampled conformations; in each first conformation cluster containing two or more first sampled conformations, the error between any two first sampled conformations does not exceed the error threshold d. hold Between every two first conformation clusters, there exist at least two conformation pairs of the first sampled conformations with errors greater than the error threshold d. hold ; Each of the first conformation clusters is taken as the corresponding current cluster; the total number of the first sampled conformations in the current cluster is counted to obtain the total number of conformations in the current cluster; the total number of conformations in the current cluster is identified; if the total number of conformations in the current cluster is 1, the unique first sampled conformation in the current cluster is taken as the first typical conformation of the current cluster; if the total number of conformations in the current cluster is 2, any first sampled conformation in the current cluster is taken as the first typical conformation of the current cluster; if the total number of conformations in the current cluster is greater than 2, the sum of the conformation pair errors of each first sampled conformation in the current cluster and the other first sampled conformations in the cluster is calculated to obtain the corresponding total conformation error, and the first sampled conformation corresponding to the smallest total conformation error is taken as the first typical conformation of the current cluster; and all the obtained first typical conformations form the corresponding typical conformation set.
11. The method for screening targeted RNA small molecules according to claim 2, characterized in that, The process of constructing corresponding mapping feature sets and materialized feature sets based on the feature mapping model and the materialized feature encoder for the typical conformation set and the current molecular library is as follows: Each of the first typical conformations in the typical conformation set is used as the current target RNA conformation and input into the feature mapping model for processing to obtain the corresponding materialized feature mapping tensor. Each of the materialized feature mapping tensors obtained this time is used as a corresponding first mapping feature, and all the obtained first mapping features form the corresponding mapping feature set. The first molecular conformations of the current molecular library are input into the physical feature encoder as the current small molecular conformations to obtain the corresponding physical feature tensors. Each physical feature tensor obtained this time is used as a corresponding first physical feature, and all the obtained first physical features form the corresponding physical feature set.
12. The method for screening targeted RNA small molecules according to claim 2, characterized in that, The process of selecting preferred conformation sequences from the current molecular library that match the current RNA conformation based on the mapping feature set and the physical feature set, and then feeding these sequences back to the current user, specifically includes: Each of the first materialized features in the materialized feature set is taken as the corresponding current feature; and based on a preset similarity algorithm, the similarity between the current feature and each of the first mapping features in the mapping feature set is calculated to obtain the corresponding first similarity; and the largest of the first similarities is taken as the corresponding small molecule fit; the similarity algorithm includes a cosine vector similarity algorithm. The first molecular conformations in the current molecular library are sorted in descending order of small molecule fitness; the first S-position first molecular conformations are taken as the corresponding first preferred conformations; and the S first preferred conformations are sorted in descending order of small molecule fitness to obtain the corresponding preferred conformation sequence and fed back to the current user; S is a pre-set positive integer.
13. An apparatus for performing the screening method for targeted RNA small molecules according to any one of claims 1-12, characterized in that, The device includes: a model building module, a first data acquisition module, a first model training module, a second data acquisition module, a second model training module, a data receiving module, an RNA conformation expansion module, a conformation feature encoding module, and a screening feedback module; The model building module constructs a property prediction model based on the Uni-Mol model to predict the physicochemical properties of small molecules; and constructs a feature mapping model based on a class of E(3)-equivariant graph neural networks to learn the physicochemical features of RNA-adapted small molecules; the property prediction model is used to predict a set of specified physicochemical properties of the current molecule based on the small molecule conformation input to the model to obtain the corresponding physicochemical property vector; the feature mapping model is used to perform feature mapping on the physicochemical features of the current RNA-adapted small molecule based on the target RNA conformation input to the model to obtain the corresponding physicochemical feature mapping tensor. The first data acquisition module constructs a first dataset by collecting large amounts of data on the physicochemical properties of small drug molecules; The first model training module trains the feature prediction model based on the first dataset; and extracts the corresponding materialized feature encoder from the feature prediction model at the end of training; The second data acquisition module initializes the second dataset by acquiring large amounts of RNA-small molecule complex data and uses the physical and chemical feature encoder to perform label tensor annotation on the second dataset; The second model training module trains the feature mapping model based on the second dataset; The data receiving module is used to take the RNA conformation input by the user as the corresponding current RNA conformation after both models have been trained; and to take the small molecule database specified by the user as the corresponding current molecule library. The RNA conformation expansion module performs molecular dynamics simulation on the current RNA conformation to obtain a corresponding sampled conformation set; and performs conformation clustering and typical conformation extraction on the sampled conformation set to obtain a corresponding typical conformation set. The conformation feature encoding module constructs corresponding mapping feature sets and materialization feature sets based on the feature mapping model and the materialization feature encoder for the typical conformation set and the current molecular library; The screening and feedback module filters molecular conformations in the current molecular library that match the current RNA conformation based on the mapping feature set and the physical feature set, and obtains the preferred conformation sequence to feed back to the current user.
14. An electronic device, characterized in that, include: Memory, processor, and transceiver; The processor is configured to be coupled to the memory, read and execute instructions in the memory to implement the method according to any one of claims 1-12; The transceiver is coupled to the processor, and the processor controls the transceiver to send and receive messages.
15. A computer-readable storage medium, characterized in that, The computer-readable storage medium stores computer instructions that, when executed by a computer, cause the computer to perform the method according to any one of claims 1-12.