Prediction method, system, storage medium and equipment for interaction relationship between small molecule and RNA (Ribonucleic Acid)

Through the deep learning model SMRTnet, an encoder combining RNA and small molecules and a multimodal data fusion module, the problem of predicting the interaction relationship between small and RNA in the existing technology is solved, and efficient drug screening and targeted RNA drug development is achieved.

CN120199320APending Publication Date: 2025-06-24TSINGHUA UNIVERSITY
View PDF 0 Cites 1 Cited by

Patent Information

Application Number
CN202510263239.6
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-03-06
Publication Date
2025-06-24

AI Technical Summary

Technical Problem

The prior art is difficult to effectively predict the interaction relationship between small molecules and RNA, especially in the absence of defined RNA tertiary structure and binding site information, which limits the application of drug screening.

Method used

The SMRTnet model based on deep learning is used to obtain the sequence and structural characteristics of RNA and small molecules through RNA encoder and small molecule encoder, and the multimodal data fusion module is used to learn the binding mode of small molecules and RNA to predict their interaction relationship.

Benefits of technology

Accurate prediction of the interaction relationship between small molecules and RNA is achieved, the hit rate of drug screening is improved, and the development process of small molecule drugs targeting RNA is promoted.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120199320A_ABST
    Figure CN120199320A_ABST
Patent Text Reader

Abstract

The invention provides a prediction method and system for an interaction relationship between a small molecule and RNA, a storage medium and equipment. The method comprises the following steps: acquiring input data, and inputting the input data into an SMRTnet model; on the basis of the input data, by utilizing the SMRTnet model, obtaining sequences and structural characteristics of RNA and small molecules, and integrating paired combination information of the RNA and the small molecules to obtain combination scores of the small molecules and the RNA; and obtaining a final binding score by using an integrated scoring strategy according to the binding score of the small molecule and the RNA, so as to predict the interaction relationship between the small molecule and the RNA and the binding site of the small molecule and the RNA. According to the prediction method for the interaction relationship between the small molecules and the RNA, the hit rate of drug screening is increased, and the research and development process of small molecule drugs of targeted RNA is promoted.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention belongs to the technical fields of artificial intelligence and life sciences, and particularly relates to a method, system, storage medium and device for predicting the interaction relationship between small molecules and RNA. Background Art

[0002] Small molecules can bind to RNA and regulate its function, providing broad prospects for the treatment of human diseases and infectious diseases. Currently, molecular docking methods are often widely used to predict the optimal binding pose of small molecules in known protein active pockets. Recent algorithm improvements have enabled AutoDock Vina, RLDOCK, NLDock, and rDock to have the ability to perform molecular docking on DNA / RNA. Recently, some deep learning has also been applied to predict small molecule drugs targeting RNA. Among them, RNAmigos predicts the molecular fingerprint of small molecules that match a given RNA binding site through a graph neural network, and then searches the library for the small molecule that best matches this molecular fingerprint. RNAmigos2 ranks the binding probabilities of multiple small molecules at the RNA binding site by using a variational autoencoder and a graph neural network. RLaffinity uses a three-dimensional convolutional neural network to predict the binding affinity of small molecule and RNA complexes.

[0003] However, current molecular docking methods require selecting appropriate stance parameters and scoring functions according to ligands and receptors, and deep learning methods cannot directly achieve the classification task of predicting whether small molecules and RNA bind. In addition, all current tools require prior obtained RNA tertiary structure and binding site information. However, since many disease-related RNAs lack a defined tertiary structure and only a few RNAs have known active sites, this greatly limits their practical applications in drug screening. Summary of the Invention

[0004] Aiming at the deficiencies of the prior art, the present invention provides a method, system, storage medium and device for predicting the interaction relationship between small molecules and RNA. By using the SMRTnet model based on deep learning, RNA and small molecule sequence and structure features are obtained through an RNA encoder and a small molecule encoder, and a multimodal data fusion module is used to learn the binding mode of small molecules and RNA for predicting the interaction relationship between small molecules and RNA, which can accurately predict the interaction relationship between small molecules and RNA, improve the hit rate of drug screening, and promote the research and development process of small molecule drugs targeting RNA.

[0005] The present invention is achieved through the following technical solutions:

[0006] Obtain input data and input it into the SMRTnet model;

[0007] Based on the input data, using the SMRTnet model, obtain the RNA and small molecule sequence and structure features, integrate the pairwise binding information of RNA and small molecules, and obtain the binding score of small molecules and RNA;

[0008] According to the binding score of small molecules and RNA, use an integrated scoring strategy to obtain the final binding score and predict the interaction relationship between small molecules and RNA.

[0009] Optionally,

[0010] The input data includes:

[0011] RNA sequence, RNA secondary structure, and the SMILES of small molecules.

[0012] Optionally,

[0013] The step of based on the input data, using the SMRTnet model, obtaining the RNA and small molecule sequence and structure features, integrating the pairwise binding information of RNA and small molecules, and obtaining the binding score of small molecules and RNA includes:

[0014] Based on the input data, through the RNA encoder and small molecule encoder, encode the sequences and structures of RNA and small molecules to obtain the RNA and small molecule sequence and structure features;

[0015] Based on the RNA and small molecule sequence and structure features, integrate the pairwise binding information of RNA and small molecules through a multimodal data fusion module;

[0016] Based on the pairwise binding information of RNA and small molecules, obtain the binding score of small molecules and RNA through a decoder.

[0017] Optionally,

[0018] The step of based on the input data, through the RNA encoder and small molecule encoder, encoding the sequences and structures of RNA and small molecules includes:

[0019] Encode the sequence and structure features of RNA through an RNA encoder;

[0020] Encode the SMILES and structure features of small molecules through a small molecule encoder.

[0021] Optionally,

[0022] The RNA encoder includes: an RNA sequence feature encoder and an RNA structure feature encoder;

[0023] The small molecule encoder includes: a small molecule SMILES feature encoder and a small molecule structure feature encoder.

[0024] Optionally,

[0025] The multimodal data fusion module is a three-layer feature fusion module based on the attention mechanism, including: one layer of cross-attention and two layers of self-attention.

[0026] Optionally,

[0027] The method further includes:

[0028] By integrating a scoring strategy, calculating the median of the binding scores of the small molecule and RNA using five models constructed based on five-fold cross-validation as the final binding score.

[0029] Optionally,

[0030] The method further includes:

[0031] Adopting a gradient backpropagation mechanism to calculate the contribution of each nucleotide binding and predicting the binding sites on RNA at the single-nucleotide resolution.

[0032] Optionally,

[0033] The method further includes:

[0034] Training the SMRTnet model using the training dataset; and / or,

[0035] Evaluating the SMRTnet model using the benchmark dataset through known interaction relationships.

[0036] The present invention also provides a prediction system for the interaction relationship between a small molecule and RNA for implementing the foregoing method, and the system includes:

[0037] A data acquisition module for acquiring input data and inputting it into the SMRTnet model;

[0038] The SMRTnet module for obtaining the sequence and structural features of RNA and small molecules based on the input data using the SMRTnet model, and integrating the pairwise binding information of RNA and small molecules to obtain the binding score of the small molecule and RNA;

[0039] A binding score module for obtaining the final binding score using an integrated scoring strategy according to the binding score of the small molecule and RNA to predict the interaction relationship between the small molecule and RNA.

[0040] Optionally,

[0041] The input data includes:

[0042] RNA sequence, RNA secondary structure, and the SMILES of the small molecule.

[0043] Optionally,

[0044] The SMRTnet module is further configured to:

[0045] Based on the input data, encode the sequences and structures of RNA and small molecules through an RNA encoder and a small molecule encoder to obtain RNA and small molecule sequence and structure features;

[0046] Based on the RNA and small molecule sequence and structure features, integrate the pairwise binding information of RNA and small molecules through a multi-modal data fusion module;

[0047] Based on the pairwise binding information of the RNA and small molecules, obtain the binding scores of small molecules and RNA through a decoder.

[0048] Optionally,

[0049] The SMRTnet module is further configured to:

[0050] Encode the sequence and structure features of RNA through an RNA encoder;

[0051] Encode the SMILES and structure features of small molecules through a small molecule encoder.

[0052] Optionally,

[0053] The binding score module is further configured to:

[0054] Through an integrated scoring strategy, calculate the median of the binding scores of the small molecules and RNA using five models constructed based on five-fold cross-validation as the final binding score.

[0055] The present invention also provides a computer-readable storage medium storing one or more programs, which when executed, can implement the aforementioned method for predicting the interaction relationship between small molecules and RNA.

[0056] The present invention also provides a device including a processor, a communication interface, a computer-readable storage medium, and a communication bus; wherein, the processor, the communication interface, and the computer-readable storage medium communicate with each other through the communication bus;

[0057] The processor is used to execute the programs stored in the computer-readable storage medium.

[0058] Compared with the prior art, the present invention has the following advantages:

[0059] 1. The prediction method for the interaction relationship between small molecules and RNA proposed by the present invention uses the SMRTnet model based on deep learning to obtain the sequence and structural features of RNA and small molecules through an RNA encoder and a small molecule encoder, and uses a multimodal data fusion module to learn the binding mode between small molecules and RNA, so as to predict the interaction relationship between small molecules and RNA, which can accurately predict the interaction relationship between small molecules and RNA, improve the hit rate of drug screening, and promote the R & D process of small molecule drugs targeting RNA.

[0060] 2. By using an RNA encoder and a small molecule encoder, the sequences and structures of RNA and small molecules are encoded to obtain the sequence and structural features of RNA and small molecules, and the binding mode is learned through a multimodal data fusion module, thereby improving the accuracy of predicting the interaction relationship between small molecules and RNA.

[0061] 3. Using a training dataset and a benchmark dataset to train the SMRTnet model and verify the results output by the SMRTnet model, thereby improving the stability and accuracy of model calculation.

[0062] 4. Through an integrated scoring strategy, the median of the binding scores of the small molecules and RNA calculated by five models constructed based on five-fold cross-validation is used as the final binding score, thereby solving the model instability caused by limited data and inconsistent data distribution.

[0063] Other features and advantages of the present invention will be described in the subsequent specification, and part of them will become obvious from the specification or be understood by implementing the present invention. The objectives and other advantages of the present invention can be realized and obtained through the structures pointed out in the specification, claims and drawings. BRIEF DESCRIPTION OF THE DRAWINGS

[0064] In order to more clearly illustrate the technical solutions in the embodiments of the present application or the prior art, the following will briefly introduce the drawings required for use in the description of the embodiments or the prior art. Obviously, the following drawings are some embodiments of the present application. For those of ordinary skill in the art, other drawings can be obtained based on these drawings without creative efforts.

[0065] Figure 1 It shows a schematic flow chart of the prediction method for the interaction relationship between small molecules and RNA;

[0066] Figure 2 It shows a schematic block diagram of the structure of the prediction system for the interaction relationship between small molecules and RNA;

[0067] Figure 3 It is a schematic flow chart of the prediction method for the interaction relationship between small molecules and RNA according to the embodiment of the present invention;

[0068] Figure 4 Schematic diagram of the SMRTnet model framework according to an embodiment of the present invention;

[0069] Figure 5 Schematic diagram of the RNA sequence encoder according to an embodiment of the present invention;

[0070] Figure 6 Schematic diagram of the RNA structure encoder according to an embodiment of the present invention;

[0071] Figure 7 Schematic diagram of the small molecule structure encoder according to an embodiment of the present invention;

[0072] Figure 8 Schematic diagram of the multi-modal data fusion module according to an embodiment of the present invention;

[0073] Figure 9 Schematic diagram of the training data set and the data set processing method according to an embodiment of the present invention;

[0074] Figure 10 Schematic diagram of the integrated scoring strategy according to an embodiment of the present invention;

[0075] Figure 11 Schematic diagram of the performance of the model according to an embodiment of the present invention on the training data set;

[0076] Figure 12 Schematic diagram of the performance of the model according to an embodiment of the present invention on the benchmark data set;

[0077] Figure 13 Schematic diagram of the performance of the SMRTnet model according to an embodiment of the present invention on different types of RNA;

[0078] Figure 14 Schematic diagram of the performance comparison between the SMRTnet model according to an embodiment of the present invention and other different computational tools;

[0079] Figure 15 Schematic diagram of the comparison between the SMRTnet binding site prediction and the known interaction relationship according to an embodiment of the present invention;

[0080] Figure 16 Schematic diagram of the prediction result and experimental verification hit rate of the SMRTnet according to an embodiment of the present invention on the screening data set;

[0081] Figure 17 Schematic diagram of the binding site prediction of the SMRTnet for Compound 1 on the MYC IRES according to an embodiment of the present invention;

[0082] Figure 18Schematic diagram of the prediction results and experimental verification performance of SMRTnet in the MYC IRES mutation dataset in the embodiments of the present invention;

[0083] Figure 19 Schematic diagram of the structure of a device in the embodiments of the present invention. Specific embodiments

[0084] To make the objectives, technical solutions and advantages of the embodiments of the present invention clearer, the technical solutions in the embodiments of the present invention will be clearly and completely described below with reference to the accompanying drawings in the embodiments of the present invention. Obviously, the described embodiments are some, but not all, of the embodiments of the present invention. All other embodiments obtained by those of ordinary skill in the art based on the embodiments of the present invention without creative efforts shall fall within the protection scope of the present invention.

[0085] See attached Figure 1 , the method of the present invention includes:

[0086] S1. Obtain input data and input it into the SMRTnet model;

[0087] Among them, the input data includes:

[0088] RNA sequence, RNA secondary structure, and SMILES of small molecules.

[0089] S2. Based on the input data, use the SMRTnet model to obtain RNA and small molecule sequence and structure features, and integrate the paired binding information of RNA and small molecules to obtain the binding score of small molecules and RNA;

[0090] Among them, based on the input data, the RNA encoder and the small molecule encoder are used to encode the sequences and structures of RNA and small molecules to obtain RNA and small molecule sequence and structure features;

[0091] Based on the RNA and small molecule sequence and structure features, the paired binding information of RNA and small molecules is integrated through a multimodal data fusion module;

[0092] Based on the paired binding information of RNA and small molecules, the binding score of small molecules and RNA is obtained through a decoder;

[0093] Among them, the encoding of the sequences and structures of RNA and small molecules through the RNA encoder and the small molecule encoder based on the input data includes:

[0094] Encoding the sequence and structure features of RNA through an RNA encoder;

[0095] Encode the SMILES and structural features of small molecules through a small molecule encoder;

[0096] Among them, the RNA encoder includes: an RNA sequence feature encoder and an RNA structural feature encoder;

[0097] The small molecule encoder includes: a small molecule SMILES feature encoder and a small molecule structural feature encoder,

[0098] Among them, the multimodal data fusion module is a three-layer feature fusion module based on the attention mechanism, including: one layer of cross-attention and two layers of self-attention.

[0099] S3. According to the binding scores of the small molecule and RNA, use an integrated scoring strategy to obtain the final binding score and predict the interaction relationship between the small molecule and RNA.

[0100] Among them, through the integrated scoring strategy, the median of the binding scores of the small molecule and RNA calculated by five models constructed based on five-fold cross-validation is used as the final binding score;

[0101] Among them, the gradient backpropagation mechanism is used to calculate the contribution of each nucleotide binding, and the binding sites on RNA are predicted at the single nucleotide resolution;

[0102] Among them, the SMRTnet model is trained using the training dataset; and / or,

[0103] The SMRTnet model is evaluated using the benchmark dataset through the known interaction relationship.

[0104] Specifically,

[0105] 1. Obtain input data and input it into the SMRTnet model.

[0106] The input data includes: RNA sequence, RNA secondary structure, and the SMILES (Simplified Molecular Input Line Entry System) of small molecules.

[0107] In this embodiment, SMRTnet is a deep learning-based model that can learn the binding pattern between small molecules and RNA from structures containing at least one RNA and one small molecule determined by techniques such as NMR, X-ray, and CryoEM in the Protein Data Bank, and is used to predict the interaction relationship between small molecules and RNA. SMRTnet learns the binding pattern between small molecules and RNA in a supervised learning manner by minimizing the error between the prediction result and the actual result. SMRTnet takes the RNA sequence and its secondary structure, and the Simplified Molecular Input Line Entry System (SMILES) of the small molecule as inputs; the output is the binding probability between the small molecule and the RNA target, and the potential binding site information of the small molecule on the RNA is obtained through the gradient backpropagation mechanism.

[0108] In this embodiment, the SMRTnet model is a binary classification model with three inputs (RNA sequence, RNA secondary structure, and small molecule SMILES encoding) and one output (the binding score between the small molecule and the RNA). In addition, SMRTnet can also quantify the input contribution through gradient backpropagation to identify the potential binding sites of small molecules on the RNA.

[0109] Second, based on the input data, using the SMRTnet model, obtain the sequence and structural features of RNA and small molecules, and integrate the pairwise binding information of RNA and small molecules to obtain the binding score between small molecules and RNA.

[0110] 1. SMRTnet model.

[0111] In this embodiment, based on the input data, using the SMRTnet model, obtain the sequence and structural features of RNA and small molecules, and integrate the pairwise binding information of RNA and small molecules to obtain the binding score between small molecules and RNA, including:

[0112] Based on the input data, encode the sequences and structures of RNA and small molecules through an RNA encoder and a small molecule encoder to obtain the sequence and structural features of RNA and small molecules;

[0113] Based on the sequence and structural features of RNA and small molecules, integrate the pairwise binding information of RNA and small molecules through a multimodal data fusion module;

[0114] Based on the pairwise binding information of RNA and small molecules, obtain the binding score between small molecules and RNA through a decoder.

[0115] The framework of the SMRTnet model includes: an RNA encoder, a small molecule encoder, a multimodal data fusion module, and a decoder.

[0116] The features of small molecules and RNA sequences and structures are obtained through a small molecule encoder and an RNA encoder respectively.

[0117] In this embodiment, the RNA encoder encodes the sequence and structural features of RNA;

[0118] The small molecule encoder encodes the SMILES and structural features of small molecules.

[0119] The RNA encoder integrates an RNA language model (RNASwan-seq) and a two-layer convolutional neural network (CNN) with a residual neural network (ResNet) to extract nucleotide and base pairing information as the representation of the input RNA.

[0120] The small molecule encoder combines a published chemical language model (MoLFormer) and a three-layer graph attention network (GAT) to capture the atomic composition and chemical structure as the representation of the input small molecule.

[0121] In this embodiment, the multi-modal data fusion module is a three-layer feature fusion module based on the attention mechanism, including: one layer of cross-attention and two layers of self-attention.

[0122] The multi-modal data fusion module gradually integrates the pairing binding information by using cross-attention neural networks and self-attention neural networks, captures the complex interactions between the RNA representation and the small molecule representation when defining the interaction relationship between small molecules and RNA, and outputs an interaction representation, which is passed to the decoder of the fully connected neural network to predict the final binding score.

[0123] The binding score of the SMRTnet model is calculated as follows:

[0124] SmrtNet(x, y, z) = σ(f fc (f fusion (f cnn-res (x, y), f rnaswan-seq (x), f gat (z‘), f molformer (z))))

[0125] Among them, x is the RNA sequence, y is the RNA secondary structure in dot-bracket notation, z is the standard SMILES encoding of the small molecule, and z‘ is the two-dimensional structure diagram of the small molecule generated by RDkit. σ is a Sigmoid function used to normalize the output to the range of 0-1, and f fc is a fully connected layer, f fusion is an attention-based feature fusion module, f cnn-res is an RNA structure feature encoder based on a convolutional neural network (CNN) and a residual network (ResNet).rnaswan-seq is an RNA sequence feature encoder based on an RNA language model, f molformer is a small molecule SMILES feature encoder based on a chemical language model, f gat is a small molecule structure feature encoder based on a graph attention network (GAT). The model output is transformed from the output value f through an activation layer fc is obtained.

[0126] 1.1. RNA Encoder Encoding

[0127] The fragmented RNA sequence and its secondary structure are input into the RNA encoder, which uses a self-developed RNA language model (RNASwan-seq) and a two-layer convolutional neural network (CNN) with a residual neural network (ResNet) to encode the sequence and structural features of RNA respectively.

[0128] The RNA encoder includes: an RNA sequence feature encoder and an RNA structure feature encoder.

[0129] (1) RNA Sequence Feature Encoder

[0130] To enhance the representation of RNA sequences, we developed an RNA language model (RNASwan-seq).

[0131] In this embodiment, the custom dataset used for RNASwan-seq pre-training contains a total of 470 million RNA sequences (including ncRNA and mRNA), representing a comprehensive collection of RNA types from a wide range of organisms; by using MMSeqs2, duplicate RNA sequences are removed according to 100% sequence similarity, and all RNA sequences are preprocessed, finally obtaining 214 million unique RNA sequences, which are composed of 5 symbols including A, U, C, G, and N; through a random splitting strategy, retaining 30% sequence similarity, the obtained data is divided into a training set and a test set for self-supervised training of RNASwan-seq.

[0132] RNASwan-seq consists of 30 transformer encoder blocks with rotary position embeddings (RoPE). Each encoder block includes a layer with a hidden size of 640 and 20 attention heads. Layer normalization is applied before each block, while residual connections are applied after each block.

[0133] In the pre-training stage, a random cropping mechanism is introduced, and 1024 nucleotide fragments are cut out from the full-length RNA sequence in each iteration; 15% of the nucleotide tokens are randomly selected for potential replacement: 80% of the selected tokens are replaced with [MASK], 10% remain unchanged, and 10% are replaced with randomly selected tokens. The model is trained using masked language modeling (MLM) to predict the originally masked tokens with cross-entropy loss. In addition, adopting a fast attention mechanism can significantly improve the training speed of the model. The training process can be represented by the following objective function:

[0134]

[0135] where M represents the indices of the tokens randomly sampled from each input sequence x, which are masked. For each masked token, given the masked sequence as context, the function minimizes the negative log-likelihood of the true nucleotide x i .

[0136] In this embodiment, RNASwan-seq is trained on 16 40GB A100 graphics processing units (GPUs). The pre-trained model is integrated into SMRTnet as an RNA sequence encoder, and a fine-tuning strategy is adopted. For an RNA sequence of a given length L, the output tensor of RNASwan-seq is represented as an L×640 matrix.

[0137] (2) RNA structure feature encoder:

[0138] To obtain more comprehensive RNA sequence and structure information, we designed an RNA structure feature encoder. In the RNA structure feature encoder, the RNA sequence (x) is encoded using one-hot encoding and combined with the dot-bracket notation of the RNA secondary structure (y), using token encoding as the fifth dimension. Thus, the 5D vector consists of 4D x and 1D y. The embedding calculation of the RNA structure feature encoder is as follows:

[0139] f cnn-res (x, y) = ((f R (f SE (f C (x')))))

[0140] where f C is a convolutional block, f SE is a squeeze-and-excitation block, and f R is a residual block. The three blocks are defined as follows:

[0141] f C (x') = ReLU(BN(Conv K,P (x')))

[0142] In f C , ReLU is the rectified linear unit activation function, BN is the batch normalization layer, and Conv K,P is a 2D convolutional layer with learnable kernels and padding. These configurations ensure that the output shape of the convolutional layer matches the input shape.

[0143]

[0144] In f SE , the squeeze-and-excitation block, as a channel-wise self-attention mechanism, promotes the detection of binding patterns through weight recalibration. is the channel multiplication of the SE block between the input and the learned C-dimensional vector. Initially, the SE block first uses the global average pooling function f sq to compress the global sequence context information into channel statistics. Subsequently, it excites this statistic into a set of channel weights scaled between 0 and 1 by applying a non-linear transformation f ex that includes two fully-connected layers and a ReLU activation function.

[0145] f R (x) = f R1 (AvgPool(f R2 (x)))

[0146] In f R , f R1 and f R2 are residual blocks with 1D and 2D convolutional kernels. f R1 is designed to learn the combined patterns of sequences and structures, while f R2 is designed to capture the spatial background features for precisely locating binding sites. The AvgPool function is an average pooling layer that converts a 2D feature map into a 1D vector.

[0147] 1.2. Small molecule encoder encoding.

[0148] The SMILES encoding of small molecules is input into the small molecule encoder, and the small molecule SMILES encoder uses the published chemical language model (MoLFormer) and a three-layer graph attention network (GAT) to encode the SMILES and structural features of small molecules.

[0149] The small molecule encoder includes: a small molecule SMILES feature encoder and a small molecule structural feature encoder.

[0150] (1) Small molecule SMILES feature encoder.

[0151] In this embodiment, an advanced large language model MoLFormer is introduced as a small molecule SMILES feature encoder to learn the compressed representation of chemical molecules. This model adopts a linear attention mechanism and embeds rotational positions on the SMILES sequences of 1.1 billion unlabeled molecules in the PubChem and ZINC datasets. Specifically, MoLFormer utilizes an encoder based on a linear attention transformer, including 12 layers and 12 attention heads per layer, with a hidden state size of 768. Therefore, each small molecule SMILES encoding is converted into an L * 768 vector matrix based on MoLFormer. Since MoLFormer sets a checkpoint every 5000 iterations, the final model with a fine-tuning strategy can be used for model training.

[0152] (2) Small molecule structure feature encoder:

[0153] In this embodiment, a three-layer GAT block is designed as a small molecule structure feature encoder, which can adaptively learn the weights of each edge and represent each node through message passing of small molecules. Specifically, each canonical SMILES string is converted into its 2D molecular graph G by using the RDKit package. Then, the drug graph is represented as G=(V, E), where V is the node represented by drug atoms and E represents the set of edges between nodes. Each node is represented by a 74-dimensional vector based on the DGL LifeSci package, including atomic type, number of atomic connections, number of implicit hydrogens, formal charge, number of radical electrons, atomic hybridization, total number of hydrogens, and whether the atom is aromatic. Molecules with fewer nodes will contain zero-filled virtual nodes. The embedding calculation of the small molecule structure feature encoder is as follows:

[0154]

[0155] To integrate the graph structure information of small molecules, a masked attention mechanism is applied based on the adjacency matrix to ensure that the model only focuses on adjacent nodes; the attention weights of all potential neighbors are normalized by using the SoftMax function to make them comparable among different nodes; the atomic feature vectors are updated by aggregating the features of atoms connected by chemical bonds, and the feature representation of small molecules is obtained.

[0156] 1.3. Multimodal data fusion module.

[0157] The multimodal data fusion module is a three-layer feature fusion module based on attention to enhance the learning of output feature information between different encoders. The gradient backpropagation mechanism is adopted to calculate the contribution of each nucleotide binding, so that the model can calculate and predict the binding sites on RNA at the single nucleotide level.

[0158] The multi-modal data fusion module includes a cross-attention layer and two self-attention layers with different parameters. Each layer contains a residual network and a layer normalization block, and the entire connection block integrates the features of the input of the next layer. This module effectively integrates the embeddings of the small molecule and RNA encoders through progressive layer-by-layer information exchange, enhancing the learning of binding patterns. The multi-modal data fusion module is defined as follows:

[0159] f fusion (E r ,E s ,E m ,E t )=f f3 (f f2 (f f1 (E s ,E r ,E r ),f f1 (E r ,E s ,E s )),f f2 ((f f1 (E t ,E m ,E m ),f f1 (E m ,E t ,E t )))

[0160] Among them, E r is the output of the RNA sequence feature encoder, E s is the output of the RNA structure feature encoder, E m is the output of the small molecule SMILES feature encoder, and E t is the output of the small molecule structure feature encoder.

[0161] The first feature fusion layer f f1 includes a cross-attention module, while the second (f f2 ) and the third (f f3 ) feature fusion layers both include self-attention modules. The structures of the three feature fusion layers are as follows:

[0162] For f f1 , a cross-attention mechanism based on ViLBERT is implemented, which extends the BERT architecture to a multi-modal, two-stream model. Here, four independent input streams interact with each other through the cross-attention block to produce four outputs. The output of each input is calculated as follows:

[0163]

[0164] Q=f ln(f mlp (y)); K = V = f ln (f mlp (x))

[0165] Among them, Q represents the output of a small molecule SMILES feature encoder or an RNA sequence feature encoder, a small molecule structure feature encoder or an RNA structure feature encoder, while K and V represent the paired encoder outputs of each corresponding Q. f ln is a layer normalization block, and f mlp is a multi-layer perceptron. The output embedding representation dimension c is 128, and the number of attention heads h is 2.

[0166] For f f2 , a self-attention mechanism is used to further enhance the integration of the features of small molecules and RNAs, thereby generating two outputs. The output is calculated as follows:

[0167]

[0168] Q = K = V = [f ln (embedding1), (f ln (embedding2)]

[0169] Among them, Q, K, and V come from the same source, representing the output of f f1 , f ln is a layer normalization block, and [] is a concatenation block. The output embedding representation dimension c is 128, and the number of attention heads h is 2.

[0170] For f f3 , the same framework as f f2 is adopted, but different hyperparameters are used to further enhance the exchange of binding information, generating a single embedding representation for model prediction for the decoder to predict the binding score. The output is calculated as follows:

[0171]

[0172] Q = K = V = [f ln (sequence), (f ln (structure)]

[0173] Among them, Q, K, and V come from the same source, and f f2 represents the output from. f ln is a layer normalization block, [] is a concatenation block, and f b is a batch normalization block. The output embedding representation dimension c is 128, and the number of attention heads h is 2.

[0174] 2. Training and evaluation of the SMRTnet model.

[0175] (1) Input of the SMRTnet model.

[0176] During model training, the input data is defined as B, and the number of samples is defined as N. Each sample consists of an RNA sequence, an RNA secondary structure, a small molecule encoded by SMILES, and a label. Each RNA sequence is a 31nt sequence composed of four bases (A, U, C, and G). The RNA structure in each sample is represented by a 31nt dot-bracket notation with three symbols ("(", ".", and ")"). The small molecule is encoded in canonical SMILES and calibrated by RDKit. The label in each sample is binary, with two symbols ("1" for positive samples and "0" for negative samples). In this example, a sliding window strategy (size = 31, stride = 1) is also introduced in the inference stage to process input RNAs longer than 31nt.

[0177] (2) Training strategy of SMRTnet

[0178] The SMRTnet model is trained to optimize its parameters by minimizing a loss function that includes L2 regularization of all parameters and binary cross-entropy (BCE) loss between the target label T and the prediction Y in the training set:

[0179]

[0180] where, t i is the target label, y i is the predicted binding score, represents all the parameters of SMRTnet, and N is the batch size during model training.

[0181] In this example, the model parameters can be optimized using the Adam optimizer, which is an extension of the stochastic gradient descent algorithm that can adaptively adjust the step size and requires minimal hyperparameter tuning. Additionally, the learning rate can be adjusted through a warm-up scheme with a linear scaling rule to address optimization challenges in the early stages of training.

[0182] To prevent overfitting, a batch normalization layer follows each convolutional layer, and a dropout layer follows each residual block. The L2 norm of all parameters is used as a weight decay term to further reduce overfitting. Early stopping is used to automatically stop the SMRTnet training when the validation performance stops improving after M epochs. Specifically, the AUC score of the model on the validation set is monitored at each time point. After 100 epochs of iteration, if the validation loss fails to decrease for 20 consecutive epochs, early stopping is activated. Then, the best model with the lowest validation loss is selected to evaluate the test set. Finally, 5-fold cross-validation (CV) is adopted in the model training, following the same strategy.

[0183] (3) Training dataset and benchmark dataset.

[0184] In this embodiment, the SMRTnet model can also be trained using the training dataset.

[0185] Training dataset:

[0186] The training dataset is derived from the structural data of the interaction between small molecules and RNA in 1,061 PDBs. Each interaction site between a small molecule and RNA may involve multiple RNA fragments. The secondary structures of these fragments are obtained, and a total of 8,672 interaction pairs of RNA fragments and small molecules are generated. The RNA fragments and small molecules are randomly paired, and non-interaction pairs are created as negative samples. In this embodiment, all positive samples and negative samples extracted at different ratios (1:1, 1:2, 1:3, 1:4, 1:5, and 1:10) are used for model training.

[0187] In this embodiment, to train the SMRTnet model, 2,477 structures containing at least one RNA and a small molecule are screened out from 195,340 structures in the PDB. After removing metal ions, ligands unrelated to treatment, structures with fewer than 31 RNA residues, and binding sites where more than 50% of the residues are proteins within finally, 1,061 high-quality SRI structures are obtained.

[0188] The DSSR tool is used to convert the three-dimensional RNA structures of these 1,061 SRIs into secondary structures, and the atomium tool is used to obtain the binding sites of RNA residues around the small molecule within The binding sites are extended by 15 nucleotides towards the 5' and 3' ends respectively to obtain 31-nucleotide RNA fragments; the RDKit tool is used to convert the chemical structure of the small molecule into a SMILES code, and the RNA fragments are paired with the corresponding SMILES as positive samples. The RNA fragments and small molecules are randomly paired, and non-interaction pairs are created as negative samples.

[0189] After screening out the paired known interactions, a specific positive-negative sample ratio (1:1, 1:2, 1:3, 1:4, 1:5, and 1:10) is maintained. The benchmark dataset is divided into a training set (80%), a validation set (10%), and a test set (10%) according to the small molecule-based segmentation strategy, ensuring that the SMILES in the test set do not appear in the training set and the validation set.

[0190] In this embodiment, 5-fold cross-validation can also be used to evaluate the stability of the model, ensuring that the test set and the validation set in each fold do not contain the same small molecules.

[0191] Benchmark dataset:

[0192] In this embodiment, the SMRTnet model can also be evaluated by using the benchmark dataset through known interaction relationships.

[0193] The benchmark dataset contains 5 published and experimentally verified datasets, consisting of a total of 2,011 interaction and non - interaction relationships, including 524 data from the NALDB database, 459 data from SMMRNA, 423 data from R - SIM, 389 data from R - BINDv2.0, and 216 data from newly published articles. The benchmark dataset is used to evaluate the generalization ability of the model and does not participate in model training.

[0194] In this embodiment, in order to evaluate the generalization ability of SMRTnet, the benchmark dataset collected experimental verification data of 2,011 small molecule - RNA interaction / non - interaction pairs, including 1,609 interaction pairs and 402 non - interaction small molecule - RNA pairs. These small molecule - RNA pairs do not contain data from the SMRTnet dataset (duplicate removal by 100% similarity).

[0195] In this embodiment, 1,795 interaction / non - interaction pairs were obtained from 178 papers in 4 published databases, including R - BINDv2.0, SMMRNA, NALDB, and R - SIM. And 216 interaction / non - interaction pairs not included in the published databases were obtained from 22 published papers, called the "New collection" subset.

[0196] In addition to obtaining the RNA sequence and its secondary structure from the corresponding papers, the RNAsturcutre tool can also be used to predict the secondary structure according to the RNA sequence.

[0197] Meanwhile, draw the chemical structure of the small molecule according to the relevant papers and convert it into SMILES encoding using OPENBABEL.

[0198] III. According to the binding scores of the small molecule and RNA, use an integrated scoring strategy to obtain the final binding score and predict the interaction relationship between the small molecule and RNA.

[0199] Through the integrated scoring strategy, calculate the median of the binding scores of the small molecule and RNA using five models constructed based on five - fold cross - validation as the final binding score.

[0200] In this embodiment, an integrated scoring strategy can be adopted to address the model instability caused by limited data and inconsistent data distribution, and highly reliable results can be ranked according to five models generated by 5-fold cross-validation.

[0201]

[0202] Among them, F represents the integrated scoring strategy, and f median (r i , s i ) represents the median score of the five models. By inputting RNA r and small molecule s, the binding score of each r i and s i is calculated using a sliding window strategy (window size = 31 nt, step size = 1 nt). For a 31-nt RNA sequence, the median predicted value of the five models is calculated as the binding score corresponding to the SRI. For RNA sequences longer than 31 nt, a sliding window strategy (window size = 31 nt, step size = 1 nt) can be applied to calculate the binding score of each window. For RNA sequences with lengths between 31 nt and 41 nt (31 < L ≤ 40), the maximum binding score (S) among all windows can be used as the final binding score corresponding to the SRI. For input RNA with a length of 41 nt or longer (L > 40), the maximum binding score of the potential binding region can be used. The potential binding region is defined as: if the binding scores of at least four consecutive windows are greater than 0.5, it is considered that the model has found a potential binding region, and the maximum value of this region is used as the final binding score corresponding to the SRI. On the contrary, if there is no potential binding region (i.e., fewer than four consecutive windows have binding scores greater than 0.5), the minimum binding score (S) among all windows is used as the final score, and the RNA and small molecule will be classified as non-binding.

[0203] In this embodiment, the prediction of the binding site is also carried out through the gradient backpropagation mechanism.

[0204] Gradient-weighted class activation mapping (Grad-CAM) is a method widely used in image classification. It uses the gradient information in the backpropagation mechanism to generate an interpretable heat map, also known as a saliency map. In addition, the SmoothGrad algorithm calculates the average gradient by perturbing the input 20 times, thereby reducing the noise interference in the input and smoothing the gradient. Here, these methods are applied to quantify the contribution of each nucleotide in the SRI binding. A higher signal value indicates a greater importance of each base in the binding.

[0205] In this embodiment, the SmoothGrad algorithm is applied to obtain a 1×31 gradient based on the RNA sequence encoder and a 5×31 gradient based on the RNA structure encoder. Then, the gradients from the RNA sequence encoder and the RNA structure encoder are combined to obtain a 2×31 gradient, where the 1×31 part represents the gradient of the RNA sequence and the other 1×31 part represents the gradient of the RNA structure.

[0206] In this embodiment, when an input is fed into SMRTnet, the gradient g(x) of each encoder is calculated as follows:

[0207]

[0208] where g represents the Grad-CAM algorithm, represents the SmoothGrad algorithm, n equals 20, indicating 20 rounds of perturbation with small Gaussian noise ((N(0, σ 2 ))), and ⊙ represents the multiplication operation. To identify the highly concerned regions (HAR) on the RNA, the Savitzky-Golay filter can be applied to convert the discrete encoded signal (the gradient of the RNA encoder) into a continuous form, thereby generating potential binding sites. Finally, the gradient values are normalized by Min-Max normalization to the range of 0 to 1.

[0209] In this embodiment, the drug screening ability of the model and the validation of the predicted binding sites on the RNA can also be evaluated by screening the dataset, including:

[0210] The SMRTnet drug screening dataset contains 10 disease-related RNA targets and 7,350 unique small molecule natural product compounds, which can evaluate the drug screening ability of the model for subsequent experimental verification.

[0211] The MYC IRES mutation dataset consists of 20 mutants of MYC IRES, which can evaluate the interpretability of the model for the validation of the predicted binding sites.

[0212] The present application will be further elaborated in detail below in conjunction with the accompanying drawings and specific embodiments.

[0213] Refer to the attached Figure 2 figure, which shows the structure of a prediction system for the interaction relationship between small molecules and RNA for implementing the above method, including a data acquisition module, an SMRTnet module, and a binding score module.

[0214] The data acquisition module is used to acquire input data and input it into the SMRTnet model;

[0215] The SMRTnet module is used to obtain RNA and small molecule sequence and structural features based on the input data by using the SMRTnet model, integrate the paired binding information of RNA and small molecules, and obtain the binding score between small molecules and RNA.

[0216] The binding score module is used to obtain the final binding score using an integrated scoring strategy based on the binding score between small molecules and RNA, and predict the interaction relationship between small molecules and RNA.

[0217] Figure 3 This is a schematic flowchart of the method for predicting the interaction relationship between small molecules and RNA in an embodiment of the present invention. In this embodiment, the RNA sequence and its secondary structure and the SMILES of small molecules are used as inputs, and the SMRTnet model based on a deep neural network is used to output the binding score between small molecules and RNA and predict potential binding sites. The prediction results can be verified through model evaluation and application.

[0218] Figure 4 This is a schematic diagram of the SMRTnet model framework in an embodiment of the present invention. In this embodiment, the SMRTnet model mainly includes: an RNA encoder, a small molecule encoder, a multi-modal data fusion module, and a decoder. The RNA encoder uses an RNA language model (RNASwan-seq) to process the RNA sequence and uses a convolutional neural network and a residual neural network to process the RNA secondary structure; the small molecule encoder uses a chemical language model (MoLFormer) to process the small molecule SMILES encoding and uses a three-layer graph neural network to process the chemical structure of small molecules; the multi-modal data fusion (MDF) module gradually integrates paired binding information using an attention-based neural network; a fully connected neural network is used as the decoder to predict the binding score of the input RNA and small molecule pair, and an integrated scoring strategy is introduced to obtain the final binding score.

[0219] Figure 5 This is a schematic diagram of the RNA sequence encoder in an embodiment of the present invention. In this embodiment, RNASwan-seq is trained on 16 40GB A100 graphics processing units (GPUs). The pre-trained model is integrated into SMRTnet as the RNA sequence encoder, and a fine-tuning strategy is adopted. For an RNA sequence with a given length of L, the output tensor of RNASwan-seq is represented as a matrix of L×640.

[0220] Figure 6 This is a schematic diagram of the RNA structure encoder in an embodiment of the present invention. In this embodiment, the RNA sequence (x) is encoded using one-hot encoding and combined with the dot-bracket notation of the RNA secondary structure (y), and marker encoding is used as the fifth dimension.

[0221] Figure 7 Schematic diagram of the small molecule structure encoder of an embodiment of the present invention. In this embodiment, the small molecule structure feature encoder is a three-layer GAT block, which can adaptively learn the weight of each edge and represent each node through small molecule message passing. Each standardized SMILES string is converted into a small molecule 2D chemical structure formula by using the RDKit package; and then the small molecule 2D chemical structure formula is converted into a graph representation. Then, through the graph attention module, the drug graph is represented as G = (V, E), where V is the node represented by the drug atom and E represents the edge set between the nodes. Each node is represented by a 74-dimensional vector based on the DGL LifeSci software package, including the atom type, the number of atomic connections, the number of implicit hydrogens, the formal charge, the number of radical electrons, the atomic hybridization, the total number of hydrogens, and whether the atom is aromatic. Molecules with fewer nodes will contain zero-filled virtual nodes.

[0222] Figure 8 The figure is a schematic diagram of a multimodal data fusion module of an embodiment of the present invention. In this embodiment, the multimodal data fusion module includes a mutual attention layer and two self-attention layers with different parameters. Each layer contains a residual network and a layer normalization block, and the entire connection block integrates the features of the next layer input. The module effectively integrates the embedding of small molecules and RNA encoders through progressive layer-by-layer information exchange, enhancing the learning of binding patterns.

[0223] Figure 9 Schematic diagram of the training data set and data set processing method of the embodiment of the present invention. In this embodiment, all tertiary structures including RNA and small molecules were collected from the PDB database, and some small molecules that did not meet the specifications and small molecules were removed. After more than 50% of the residues are amino acids, 1,061 high-quality RNA-small molecule interaction structures remain. RNA bases are obtained and extended to 31 nt in length. Finally, 31 nt fragments and the secondary structures of these fragments are obtained, and the fragments are paired with the corresponding small molecules.

[0224] Figure 10 Schematic diagram of an integrated scoring strategy of an embodiment of the present invention. In this embodiment, an integrated scoring strategy is used to cope with model instability caused by limited data and inconsistent data distribution, and highly reliable results can be ranked according to five models generated by 5-fold cross validation.

[0225] Figure 11Schematic diagram of the performance of the model according to the embodiments of the present invention on the training dataset. In this embodiment, the receiver operating characteristic curve (ROC curve) of the SMRTnet model at a positive-negative sample ratio of 1:2 is evaluated through five-fold cross-validation and a compound-based partitioning strategy.

[0226] Figure 12 Schematic diagram of the performance of the model according to the embodiments of the present invention on the benchmark dataset. In this embodiment, 12-a represents the interaction and non-interaction of small molecule and RNA pairing in five experimental benchmarks, which is called SMRTnet-benchmark. 12-b represents the receiver operating characteristic curve (ROC curve) of SMRTnet on five SMRTnet-benchmark datasets.

[0227] Figure 13 Schematic diagram of the performance of the SMRTnet model according to the embodiments of the present invention on different types of RNA. It can be seen from 13-a the percentage distribution of eight RNA types in the SMRTnet-benchmark dataset. 13-b shows the performance of SMRTnet on eight RNA types, respectively showing the interacting and non-interacting pairs, as well as the overall model performance of each RNA type.

[0228] Figure 14 Schematic diagram of the performance comparison between the SMRTnet model according to the embodiments of the present invention and other different computational tools. The figure shows the overall performance of the SMRTnet model and the leading tools in the current field in the decoy evaluation task using the training dataset.

[0229] Figure 15 Schematic diagram of the comparison between the SMRTnet binding site prediction according to the embodiments of the present invention and the known interaction relationship. The figure shows the binding site predictions of MYC IRES, HIV-1 TAR element, CUG expansion in the HTT gene, and pre-miR18a. Two heatmap tracks show the response of the model at each nucleotide position and the high-concern regions (the upper track is the sequence response and the lower track is the structure response). In the RNA structure diagram below, the binding sites of known small molecules are shown, and the labels around the binding sites indicate the names of the small molecules.

[0230] Figure 16Schematic diagram of the prediction results of the SMRTnet model of the embodiments of the present invention on the screening dataset and the hit rate of experimental verification. As shown in the figure, 16-a represents the secondary structures of 10 disease-related RNA targets and the binding score distribution of 10 disease-related RNA targets predicted by SMRTnet. Each point represents the predicted binding score of a compound to a target (N = 7,350). The dashed line represents the classification threshold (0.704) of SMRTnet. 16-b represents the hit rate of each target determined based on the microscale thermophoresis experiment.

[0231] In this embodiment, based on the drug screening dataset, for 10 disease-related RNA targets, the binding scores of 7,350 natural compounds were predicted using SMRTnet, and the top 20 small molecules for each target were selected according to the binding scores (while requiring the scores to be higher than the classification threshold of 0.704), and a total of 190 predicted SRIs were screened for MST experimental verification. Finally, the interaction relationships between 40 predicted small molecules and RNA were verified through the binding inspection mode of MST, and the average verification rate was 21.1%. The hit rate of the SMRTnet model of the present invention far exceeds that of existing high-throughput experimental screening methods: the hit rate of ALIS is 0.0002% (1 / 50,000), the hit rate of SMM is 0.22% - 1.19% (24,572 compounds for 36 nucleic acids), the hit rate of SHAPE-MaP is 2.7% (41 / 1,500), etc. Therefore, SMRTnet can be used to significantly improve the hit rate of the binding of potential small molecule drugs to RNA targets in the virtual screening process, saving the labor, time, and money costs of drug screening.

[0232] For MYC RNA IRES, in addition to the above verification of the top-ranked prediction results, in this embodiment, a larger-scale experimental verification was also carried out. By randomly extracting 376 small molecule samples from the same natural compound library, covering different predicted binding scores. By using t-distributed stochastic neighbor embedding (t-SNE) to analyze the distribution of small molecules in the chemical structure space, it was found that there was no bias towards any specific chemical structure. A significant positive correlation was observed between the verification rate and the predicted binding score through MST experimental verification.

[0233] Figure 17 Schematic diagram of the binding site prediction of Compound 1 by SMRTnet of the embodiments of the present invention for MYC IRES. As shown in the figure, the potential binding site of Compound 1 and MYC IRES predicted by SMRTnet. In the RNA structure diagram, the marked high-concern regions can be seen, and the two heatmap tracks show the responses of the model at each nucleotide position (the upper track is the sequence response, and the lower track is the structure response).

[0234] Figure 18 Schematic diagram of the prediction results and experimental verification performance of SMRTnet in the MYC IRES mutation dataset of the embodiments of the present invention.

[0235] As shown in the figure, 18-a represents the secondary structures of 20 mutations of MYC IRES, including: fully paired structures, upper 1×1 internal loop structures, lower 1×1 internal loop structures, 2×2 internal loop structures, and 3×3 internal loop structures. The nucleotides marked in the structures indicate the positions and types of mutations. The bar charts in 18-b respectively represent the predicted binding score distributions of different mutated MYC IRES to Compound 1, and the bound and unbound samples verified by microscale thermophoresis experiments. It can be seen from this the average binding scores of the same binding site type and the trend of binding scores with the change of binding site type. 18-c represents the comprehensive performance of compound 1 for the computational prediction and experimental verification of 20 MYC IRES mutations.

[0236] In this embodiment, based on the MYC IRES mutation dataset, further studies were conducted on how sequence and structural changes affect small molecule-RNA interactions (SRIs). We designed 20 mutated MYC IRES RNAs by simultaneously changing the bases near the binding sites and classified them into five different binding site types: i) 3×3 internal loop structure; ii-iv) 2×2 or 1×1 internal loop structures; v) completely removing the internal loop to form a fully paired structure. The results show that when changing from the 2×2 internal loop structure to the 1×1 internal loop structure, the predicted binding scores gradually decrease, and reach the lowest when changing to the fully paired structure. On the contrary, when changing from the 2×2 internal loop structure to the 3×3 internal loop structure, the predicted binding scores increase. It is worth noting that this predicted binding score is highly consistent with the verification rate obtained from experimental verification. This not only demonstrates the accuracy of SMRTnet in predicting binding scores, but also highlights the reliability of binding site recognition.

[0237] In this embodiment, the binding sites predicted by SMRTnet were further verified by molecular docking. rDock was used to dock these small molecules with the tertiary structure of MYC IRES. The results show that all small molecules are accurately located at the binding sites predicted by SMRTnet.

[0238] In addition, the embodiments of the present application also provide a prediction device for the interaction relationship between small molecules and RNAs, including:

[0239] A data acquisition module that acquires input data and inputs it into the SMRTnet model; wherein, the input data includes: RNA sequences, RNA secondary structures, and SMILES of small molecules;

[0240] The SMRTnet module, based on the input data, utilizes the SMRTnet model to obtain RNA and small molecule sequence and structural features, and integrates the paired binding information of RNA and small molecules to obtain the binding score between small molecules and RNA. Among them, based on the input data, through the RNA encoder and the small molecule encoder, the sequences and structures of RNA and small molecules are encoded to obtain RNA and small molecule sequence and structural features. Based on the RNA and small molecule sequence and structural features, the paired binding information of RNA and small molecules is integrated through a multimodal data fusion module. Based on the paired binding information of RNA and small molecules, the binding score between small molecules and RNA is obtained through a decoder.

[0241] The binding score module, according to the binding score between small molecules and RNA, uses an integrated scoring strategy to obtain the final binding score and predict the interaction relationship between small molecules and RNA. Among them, through the integrated scoring strategy, the median of the binding scores of the small molecules and RNA calculated by five models constructed based on five-fold cross-validation is used as the final binding score.

[0242] Based on the same inventive concept, the present invention also provides a computer-readable storage medium storing one or more programs, which when executed, can implement the aforementioned method for predicting the interaction relationship between small molecules and RNA.

[0243] As Figure 19 shown, an embodiment of the present invention also provides a device, including a processor, a communication interface, a memory, and a communication bus. Among them, the processor, the communication interface, and the memory communicate with each other through the communication bus.

[0244] The memory is a computer-readable storage medium for storing one or more programs.

[0245] The processor is configured to execute the programs stored in the computer-readable storage medium.

[0246] This computer-readable storage medium may be included in the device / apparatus described in the above embodiments; or it may exist alone without being assembled into the device / apparatus.

[0247] Although the present application has been described in detail with reference to the foregoing embodiments, those of ordinary skill in the art should understand that: they can still modify the technical solutions recorded in the foregoing embodiments, or perform equivalent replacements on some of the technical features; and these modifications or replacements do not make the essence of the corresponding technical solutions deviate from the spirit and scope of the technical solutions of the embodiments of the present application.

Claims

1. A method for predicting the interaction between small molecules and RNA, characterized in that: include: Get input data and input it into the SMRTnet model; Based on the input data, using the SMRTnet model, the sequence and structural features of RNA and small molecules are obtained, and the pairwise binding information of RNA and small molecules is integrated to obtain the binding score of small molecules and RNA; According to the binding scores of the small molecule and RNA, an integrated scoring strategy is used to obtain the final binding score to predict the interaction relationship between the small molecule and RNA.

2. The method according to claim 1, characterized in that The input data includes: RNA sequence, RNA secondary structure, and SMILES of small molecules.

3. The method according to claim 2, characterized in that Based on the input data, the SMRTnet model is used to obtain RNA and small molecule sequence and structural features, and the paired binding information of RNA and small molecules is integrated to obtain the binding score of small molecules and RNA, including: Based on the input data, the sequences and structures of RNA and small molecules are encoded by RNA encoders and small molecule encoders to obtain the sequence and structural features of RNA and small molecules; Based on the sequence and structural characteristics of the RNA and small molecules, the paired binding information of the RNA and small molecules is integrated through a multimodal data fusion module; Based on the pairwise binding information of the RNA and the small molecule, the binding score of the small molecule and the RNA is obtained by the decoder.

4. The method according to claim 3, characterized in that Based on the input data, the RNA and small molecule sequences and structures are encoded by RNA encoder and small molecule encoder, including: The sequence and structural features of RNA are encoded by RNA encoders; Encoding of SMILES and structural features of small molecules is done by Small Molecule Encoder.

5. The method according to claim 4, characterized in that The RNA encoder includes: an RNA sequence feature encoder and an RNA structure feature encoder; The small molecule encoder includes: a small molecule SMILES feature encoder and a small molecule structural feature encoder.

6. The method according to claim 3, characterized in that The multimodal data fusion module is a three-layer feature fusion module based on the attention mechanism, including: one layer of mutual attention, and two layers of self-attention.

7. The method according to claim 1, characterized in that The method further comprises: The median binding score of the small molecule and RNA was calculated using an integrated scoring strategy using five models constructed based on five-fold cross validation as the final binding score.

8. The method according to any one of claims 1 to 7, characterized in that: The method further comprises: The gradient back-propagation mechanism is used to calculate the contribution of each nucleotide binding and predict the binding sites on RNA at single-nucleotide resolution.

9. The method according to any one of claims 1 to 7, characterized in that: The method further comprises: Training the SMRTnet model using a training data set; and / or, The SMRTnet model is evaluated using a benchmark dataset with known interactions.

10. A prediction system for the interaction between small molecules and RNA, characterized in that: The system comprises: The data acquisition module is used to obtain input data and input it into the SMRTnet model; A SMRTnet module, for obtaining RNA and small molecule sequence and structural features based on the input data and using the SMRTnet model, and integrating the pairwise binding information of RNA and small molecules to obtain the binding score of small molecules and RNA; The binding score module is used to obtain the final binding score based on the binding score of the small molecule and RNA using an integrated scoring strategy to predict the interaction relationship between the small molecule and RNA.

11. The system according to claim 10, characterized in that The input data includes: RNA sequence, RNA secondary structure, and SMILES of small molecules.

12. The system according to claim 11, characterized in that The SMRTnet module is also configured to: Based on the input data, the sequences and structures of RNA and small molecules are encoded by RNA encoders and small molecule encoders to obtain the sequence and structural features of RNA and small molecules; Based on the sequence and structural characteristics of the RNA and small molecules, the paired binding information of the RNA and small molecules is integrated through a multimodal data fusion module; Based on the pairwise binding information of the RNA and the small molecule, the binding score of the small molecule and the RNA is obtained by the decoder.

13. The system according to claim 12, characterized in that The SMRTnet module is also configured to: The sequence and structural features of RNA are encoded by RNA encoders; Encoding of SMILES and structural features of small molecules is done by Small Molecule Encoder.

14. The system according to claim 10, characterized in that The combined scoring module is further configured to: The median binding score of the small molecule and RNA was calculated using an integrated scoring strategy using five models constructed based on five-fold cross validation as the final binding score.

15. A computer-readable storage medium storing one or more programs, characterized in that: When the one or more programs are executed, the method for predicting the interaction between small molecules and RNA according to any one of claims 1 to 9 can be implemented.

16. An electronic device comprising a processor, a communication interface, the computer-readable storage medium of claim 15, and a communication bus; wherein: The processor, the communication interface, and the computer-readable storage medium communicate electronically with each other via a communication bus; characterized in that: The processor is configured to execute a program stored in a computer-readable storage medium.

Citation Information

Cited By

  • RNA biosensor, screening method of RNA biosensor, Ectoin-producing strain and screening method of Ectoin-producing strain

    CN121006359A