Analytical adversarial training-based circular RNA and RBP interaction site prediction method and system
By employing multi-scale feature encoding and adversarial training, this study addresses the issues of insufficient feature co-learning and poor robustness in predicting circular RNA-RBP interaction sites. It achieves high-precision and interpretable prediction results, identifies key regions of circular RNA-RBP interaction sites, and provides biological explanations.
Patent Information
- Application Number
- CN202511032701.8
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-07-25
- Publication Date
- 2025-12-05
AI Technical Summary
Existing methods for predicting the interaction sites between circular RNA and RBP suffer from a lack of feature information and a lack of collaborative learning mechanisms, insufficient model robustness against adversarial attacks, and a lack of biological interpretability, resulting in unstable prediction results and difficulty in providing biologically meaningful explanations.
Employing multi-scale feature encoding, context-related feature extraction, and adversarial training mechanisms, features are extracted through natural language processing, chemical feature extraction, and pre-trained language models. Combined with motif analysis, a fusion feature matrix is obtained and adversarial training is performed to extract attention weight data to identify conservative sequence patterns.
It improves prediction accuracy, enhances the model's ability to resist interference, provides biological interpretability, and can identify key regions of the interaction sites between circular RNA and RBP and provide functional explanations.
Smart Images

Figure CN121075413A_ABST
Abstract
Description
TECHNICAL FIELD
[0001] The application relates to the technical field of biological information, in particular to a circular RNA and RBP interaction site prediction method and system based on adversarial training with explainability. BACKGROUND
[0002] With the in-depth study of the regulatory function of circular RNA, it is found that the functional regulation of circRNA is highly dependent on the specific recognition of RNA binding proteins. CircRNA can interact with a variety of proteins, including transcription factors, RNA processing proteins, etc., and these interactions are almost involved in the whole process of the circRNA life cycle. Although high-throughput experimental techniques such as CLIP-seq have been widely used to detect RNA-protein interactions, these wet experimental methods have the significant shortcomings of high cost and long time consumption. In contrast, computational methods have obvious advantages in identifying RNA-protein interactions.
[0003] At present, machine learning methods have made some progress in predicting RBP binding sites, but the prediction of circRNA and RBP interaction sites is still in its infancy. The existing computational methods mainly have the following limitations: first, different sources of feature information are simply combined, and there is a lack of effective feature collaborative learning mechanism; second, the extraction ability of the model for circRNA sequence features is insufficient, and it is difficult to capture multi-scale feature information; third, the existing prediction model generally lacks explainability and cannot provide explanation information with biological significance. In addition, the robustness of the model in the face of adversarial samples needs to be improved.
[0004] The above content is only used to assist in understanding the technical solutions of the application and does not represent the acknowledgement of the above content as prior art. SUMMARY
[0005] The main purpose of the present application is to provide a circular RNA and RBP interaction site prediction method and system based on adversarial training with explainability, which aims to effectively integrate multi-scale feature information, improve the adversarial robustness of the model, provide biological explainability, and enhance the feature collaborative learning ability.
[0006] To achieve the above purpose, the application provides a circular RNA and RBP interaction site prediction method based on adversarial training with explainability, which comprises:
[0007] obtaining the interaction site sequence of circular RNA and RBP;
[0008] The interaction site sequence is subjected to multi-scale feature coding processing to obtain first feature data and second feature data, and context correlation features of the interaction site sequence are extracted through a pre-trained language model to obtain third feature data;
[0009] Feature fusion processing is performed based on the first feature data, the second feature data, and the third feature data to obtain a fusion feature matrix;
[0010] Adversarial training is performed based on the fusion feature matrix to obtain a trained prediction model;
[0011] The circular RNA sequence to be predicted is input into the prediction model to obtain a predicted interaction site result; at the same time, attention weight data in the prediction model is extracted, a conserved sequence pattern is identified through motif analysis, and biological interpretation information is obtained.
[0012] In an embodiment, the step of subjecting the interaction site sequence to multi-scale feature coding processing to obtain first feature data and second feature data comprises:
[0013] The interaction site sequence is subjected to natural language processing, the interaction site sequence is segmented into a plurality of nucleotide units, and word frequency is calculated to obtain first feature data;
[0014] The interaction site sequence is subjected to chemical feature extraction, and features are calculated based on the number of hydrogen bonds and the rotation angle parameters of nucleotides to obtain second feature data.
[0015] In an embodiment, the step of subjecting the interaction site sequence to natural language processing, segmenting the interaction site sequence into a plurality of nucleotide units, and calculating word frequency to obtain first feature data comprises:
[0016] The interaction site sequence is cut into a plurality of fixed-length nucleotide units through a sequence segmentation operation;
[0017] Word frequency statistical processing is performed on a plurality of the nucleotide units to calculate the number of times each nucleotide unit appears in the interaction site sequence;
[0018] Normalization is performed based on the word frequency statistical result to convert the word frequency into a vector representation to output first feature data.
[0019] In an embodiment, the step of subjecting the interaction site sequence to chemical feature extraction, and calculating features based on the number of hydrogen bonds and the rotation angle parameters of nucleotides to obtain second feature data comprises:
[0020] query the chemical parameters of the nucleotides corresponding to the interaction site sequence, and obtain basic chemical parameter data based on the number of hydrogen bonds and the rotation angle of the nucleotides;
[0021] Perform mathematical transformation operations based on the basic chemical parameter data, and calculate a normalization value for each nucleotide position;
[0022] Combine the mathematical transformation results into a feature matrix through vectorization operations to output second feature data.
[0023] In an embodiment, the step of extracting the context-related features of the interaction site sequence through the pre-trained language model to obtain third feature data includes:
[0024] Encode the interaction site sequence through a bidirectional Transformer architecture operation, and calculate self-attention weights according to a multi-head self-attention mechanism;
[0025] Perform residual connection operations according to the self-attention weights to obtain corresponding residual data;
[0026] Perform linear transformation on the residual data, concatenate and project the multi-head attention output into deep features to output third feature data.
[0027] In an embodiment, the step of performing feature fusion processing based on the first feature data, the second feature data, and the third feature data to obtain a fusion feature matrix includes:
[0028] Apply a convolution kernel to the first feature data to extract local features, and calculate the average value of spatially adjacent windows through average pooling to obtain first processed feature data;
[0029] Apply a convolution kernel to the second feature data to extract local features, and calculate the average value of spatially adjacent windows through average pooling to obtain second processed feature data;
[0030] Based on the first processed feature data, the second processed feature data, and the third feature data, calculate the weights using a sigmoid activation function, and perform feature fusion processing through linear combination operations to obtain a fusion feature matrix.
[0031] In an embodiment, the step of performing feature fusion processing based on the first processed feature data, the second processed feature data, and the third feature data using a sigmoid activation function to calculate the weights, and through linear combination operations to obtain a fusion feature matrix includes:
[0032] Input the first processed feature data and the second processed feature data into a sigmoid function through a parameter matrix multiplication operation;
[0033] Based on the output of the sigmoid function obtained, the fusion weight is obtained through a weight calculation operation;
[0034] Based on the fusion weight and the third feature data, a fusion feature matrix is output through a weighted summation operation.
[0035] In an embodiment, the step of performing adversarial training based on the fusion feature matrix to obtain a trained prediction model comprises:
[0036] Through a fast gradient method operation, the gradient of the cross-entropy loss function is calculated in the model training process to generate adversarial perturbation data;
[0037] Based on the adversarial perturbation data, a perturbation is added to the fusion feature matrix through a perturbation application operation;
[0038] Based on the fusion feature matrix after adding the perturbation, a trained prediction model is output through a loss function optimization operation using the cross-entropy loss to train the model parameters.
[0039] In an embodiment, the step of extracting the attention weight data in the prediction model to obtain biological interpretation information by identifying a conserved sequence pattern through motif analysis comprises:
[0040] The importance score data of each nucleotide position in the prediction model is obtained through an attention weight extraction operation;
[0041] Based on the importance score data, high weight fragment data higher than a preset importance threshold is extracted through a threshold screening operation;
[0042] Based on the high weight fragment data, a conserved sequence pattern is identified through an STREME tool operation;
[0043] Based on the conserved sequence pattern, a database comparison operation is performed to match a known functional motif database to obtain biological interpretation information.
[0044] In addition, to achieve the above-mentioned purpose, the application also proposes a circular RNA and RBP interaction site prediction system with explainability based on adversarial training, which comprises a memory, a processor, and a circular RNA and RBP interaction site prediction program with explainability based on adversarial training stored on the memory and executable on the processor, wherein the circular RNA and RBP interaction site prediction program with explainability based on adversarial training is configured to implement the steps of the circular RNA and RBP interaction site prediction method with explainability based on adversarial training.
[0045] This application proposes an interpretable method and system for predicting circular RNA-RBP interaction sites based on adversarial training. Through multi-scale feature encoding, context-related feature extraction, adversarial training mechanism, and motif analysis technology, it solves the problems of insufficient feature integration, poor model robustness, and lack of biological interpretability in existing methods. It has the advantages of effectively improving prediction accuracy, enhancing model anti-interference ability, and providing interpretable results. Attached Figure Description
[0046] The accompanying drawings, which are incorporated in and form part of this specification, illustrate embodiments consistent with this application and, together with the description, serve to explain the principles of this application.
[0047] To more clearly illustrate the technical solutions in the embodiments of this application or the prior art, the drawings used in the description of the embodiments or the prior art will be briefly introduced below. Obviously, for those skilled in the art, other drawings can be obtained based on these drawings without creative effort.
[0048] Figure 1 This is a flowchart illustrating an embodiment of the adversarial training-based method for predicting circular RNA-RBP interaction sites with interpretability, as described in this application.
[0049] Figure 2 For this application Figure 1 A detailed flowchart of step S200 is provided in one embodiment;
[0050] Figure 3 For this application Figure 2 Detailed flowchart of step S200A1;
[0051] Figure 4 For this application Figure 2 Detailed flowchart of step S200A2;
[0052] Figure 5 For this application Figure 1 A detailed flowchart of another embodiment of step S200 is provided;
[0053] Figure 6 For this application Figure 1 Detailed flowchart of step S300;
[0054] Figure 7 For this application Figure 6 A detailed flowchart of step S330;
[0055] Figure 8 For this application Figure 1 Detailed flowchart of step S400;
[0056] Figure 9 For the purpose of the present application Figure 1 The detailed flowchart of step S500 in the present application is shown in the figure;
[0057] Figure 10 The structural schematic diagram provided by an embodiment of the present application has an explainable mechanism based on the circular RNA and RBP interaction site prediction system based on adversarial training.
[0058] Explanation of reference numerals:
[0059] 10, memory; 20, processor.
[0060] The purpose of the present application, functional characteristics and advantages will be further described with reference to the embodiments and the accompanying drawings. DETAILED DESCRIPTION
[0061] The technical solutions in the present application will be described clearly and completely in the present application by combining the accompanying drawings. Obviously, the described embodiments are only part of the embodiments of the present application, not all. The components of the present application described and shown in the accompanying drawings can be arranged and designed in various configurations. Therefore, the following detailed description of the embodiments of the present application provided in the accompanying drawings is not intended to limit the scope of the claimed application, but only represents selected embodiments of the present application. Based on the embodiments of the present application, all other embodiments obtained by those skilled in the art without creative labor are within the scope of protection of the present application.
[0062] It should be understood that similar reference numerals and letters represent similar items in the following drawings, so once an item is defined in one drawing, it does not need to be further defined and explained in subsequent drawings. At the same time, in the description of the present application, the terms "first", "second" and the like are only used to distinguish the description, and cannot be understood as indicating or implying relative importance.
[0063] In the prior art, the prediction of the interaction site of circular RNA and RNA binding protein mainly relies on high-throughput experimental techniques, which has the defects of high cost and long cycle. Although the machine learning based computational model is gradually applied to this field, the existing method has limitations in feature fusion, and different sources of features often use simple splicing method, which fails to achieve effective collaborative learning. The traditional model lacks the capture of the context association of the sequence, and lacks the robustness training of the adversarial sample, resulting in poor stability of the prediction result. In addition, most prediction models lack explainable mechanism, which cannot reveal the biological association between conserved sequence pattern and binding site.
[0064] To solve the above problems, researchers found that the effective fusion of multi-scale features is the key to improve the prediction accuracy, but the existing methods are difficult to coordinate the complementary relationship between different features. Through analysis, it is found that natural language processing technology can capture the local statistical characteristics of the sequence, while chemical feature encoding can reflect the physical and chemical properties of nucleotides, and pre-trained language models are good at extracting long-distance dependencies. Further considering the robustness of the model, the introduction of the adversarial training mechanism can enhance the adaptability to noise data. Finally, the explainability analysis based on attention weight provides a new idea for revealing sequence motifs.
[0065] Based on this, the embodiment of the application provides a circular RNA and RBP interaction site prediction method based on adversarial training with explainability, referring to Figure 1 , the method comprises steps S100-S500, wherein:
[0066] Step S100, obtaining the interaction site sequence of circular RNA and RBP;
[0067] Step S200, performing multi-scale feature encoding processing on the interaction site sequence to obtain first feature data and second feature data, and extracting context correlation features of the interaction site sequence through a pre-trained language model to obtain third feature data;
[0068] Step S300, performing feature fusion processing based on the first feature data, the second feature data and the third feature data to obtain a fusion feature matrix;
[0069] Step S400, performing adversarial training based on the fusion feature matrix to obtain a trained prediction model;
[0070] Step S500, inputting the circular RNA sequence to be predicted into the prediction model to obtain the predicted interaction site result; at the same time, extracting attention weight data in the prediction model, identifying conserved sequence patterns through motif analysis, and obtaining biological explanation information.
[0071] In this embodiment, multi-scale feature encoding processing refers to extracting features from two dimensions of sequence statistical characteristics and chemical properties, for example, generating first feature data by using a k-mer word frequency statistical method, and generating second feature data by calculating the number of hydrogen bonds and the rotation angle parameter. Context-related feature extraction uses a pre-trained model based on a Transformer architecture, for example, using a BERT model to capture long-distance dependencies between nucleotides. Feature fusion processing uses a dynamic weight allocation mechanism, for example, calculating the fusion weight coefficients of different features by using a sigmoid function. Adversarial training enhances the robustness of the model by gradient perturbation, for example, adding adversarial noise to the feature matrix during the training process. Motif analysis uses a sliding window to screen high-weight fragments, for example, using the STREME tool to identify conserved sequence patterns.
[0072] Specifically, the interaction site sequence is first divided into k-mer units of a fixed length, and the frequency of each unit is counted and then converted into a normalized vector. At the same time, the number of hydrogen bonds and the dihedral angle parameter of each nucleotide are calculated, and after standardization, a chemical feature matrix is formed. A pre-trained language model encodes the original sequence and outputs an embedding vector containing context information. The three features are extracted by a convolutional layer to extract local patterns, and then dynamically fused by an attention mechanism to generate complementary fusion features. In the adversarial training phase, a perturbation signal is generated based on the loss function gradient, and the perturbed features are input into the classifier for iterative optimization. In the prediction phase, not only the binding site probability is output, but also the key nucleotide region is located by visualizing the attention weight, and the functional explanation is obtained by matching with the motif database.
[0073] Compared with the prior art, the traditional method uses single feature encoding, which leads to information loss, while the present scheme improves the feature representation ability through multi-scale feature collaborative learning. The conventional model lacks an adversarial training mechanism and is easily disturbed by noise, while the present scheme enhances the generalization performance of the model through gradient perturbation. The existing prediction system mostly uses a black box model, and the present scheme innovatively combines the attention mechanism and motif analysis to realize visual explanation of the prediction results.
[0074] Through the above technical solutions, the present application effectively integrates sequence statistical characteristics, chemical properties, and context semantic information, significantly improving the accuracy of binding site prediction. The dynamic feature fusion mechanism overcomes the feature conflict problem caused by simple splicing in traditional methods, and the adversarial training strategy enhances the robustness of the model to data noise. The combination of attention weight and motif analysis not only provides a prediction basis, but also identifies potential binding functional domains, providing guidance for subsequent biological experiments.
[0075] In a feasible implementation manner, with reference to Figure 2 , step S200 includes steps S200A1-S200A2, wherein:
[0076] In step S200A1, the interaction site sequence is subjected to natural language processing, the interaction site sequence is segmented into a plurality of nucleotide units, and word frequency is calculated to obtain first feature data.
[0077] In step S200A2, the interaction site sequence is subjected to chemical feature extraction, and features are calculated based on the number of hydrogen bonds and the rotation angle of nucleotides to obtain second feature data.
[0078] In this embodiment, natural language processing refers to converting biological sequences into computable language units, which can be specifically implemented by sequence segmentation and word frequency statistics. By cutting the sequence into fixed-length nucleotide units and counting their frequency, the repeated local patterns in the sequence can be captured, providing statistical features for subsequent models. Chemical feature extraction refers to extracting features from the physical and chemical properties of nucleotides, which can be specifically implemented by calculating the number of hydrogen bonds and the rotation angle. By querying the chemical parameters of nucleotides and calculating the normalized values, the molecular structure characteristics of the bases in the sequence can be reflected, providing chemical-level feature expression for the model.
[0079] Specifically, in the natural language processing process, the interaction site sequence is cut into fixed-length nucleotide units, for example, fragments of length 3. The number of occurrences of each fragment is calculated by word frequency statistics method, and is converted into vector form after normalization. In the chemical feature extraction process, the number of hydrogen bonds and the rotation angle parameters of each nucleotide are queried and recorded, for example, the number of hydrogen bonds of adenine is 2 and the rotation angle is 36 degrees. These parameters are transformed into normalized values after mathematical transformation, and then combined into a feature matrix through vectorization operation. The two kinds of features respectively encode the data from two dimensions of sequence statistical law and molecular structure characteristics.
[0080] Compared with the prior art, the traditional method usually only uses a single type of feature encoding, for example, only relying on word frequency statistics or chemical parameter calculation. While the present scheme simultaneously captures the statistical distribution law and chemical structure information of the sequence through multi-scale feature fusion, so that the model can understand the sequence characteristics from different dimensions, making up for the defects of insufficient expression ability of single feature.
[0081] Through the above technical scheme, the present application can effectively combine the statistical features and chemical attribute features of the sequence, enhance the comprehensiveness and discrimination of feature expression. Through multi-scale coding, the model can simultaneously identify high-frequency fragment patterns and molecular structure rules of bases in the sequence, providing a richer data basis for subsequent feature fusion and adversarial training, thereby improving the accuracy of interaction site prediction.
[0082] In a feasible implementation manner, with reference to Figure 3 Step S200A1 includes steps S200A11-S200A13.
[0083] Step S200A11, cutting the interaction site sequence into a plurality of fixed-length nucleotide units through a sequence segmentation operation;
[0084] Step S200A12, performing a word frequency statistical processing on the plurality of nucleotide units to calculate the number of occurrences of each nucleotide unit in the interaction site sequence;
[0085] Step S200A13, performing a normalization operation based on the word frequency statistical result to convert the word frequency into a vector representation to output the first feature data.
[0086] In this embodiment, the sequence segmentation operation refers to cutting the continuous RNA sequence into fixed-length short fragments, which can be realized by using a sliding window method, for example, setting the window length to 3 nucleotides and cutting with a step length of 1. This operation can convert long sequences into processable local units, facilitating subsequent statistical modeling. The word frequency statistical processing refers to counting the frequency of each cut nucleotide unit in the entire sequence, which can be realized by using a hash table or dictionary data structure. This processing can quantify the distribution characteristics of different local patterns in the sequence, reflecting the composition rules of the sequence. The normalization operation refers to converting the word frequency value into a standardized vector, which can be realized by using the maximum and minimum value normalization or logarithmic transformation method. This operation can eliminate the dimensional difference, making the word frequencies of different units comparable, facilitating machine learning model processing.
[0087] Specifically, in the sequence processing process, the original RNA sequence is first cut into a plurality of fixed-length nucleotide units by using a sliding window. For example, for the sequence "AGCUAG", when the window length is 3, four units "AGC", "GCU", "CUA" and "UAG" are obtained. Then the number of occurrences of each unit is counted, for example, "AGC" appears twice and "GCU" appears once. Finally, the statistical result is processed by logarithmic normalization, for example, the word frequency is converted into log(1+x) form and concatenated into a vector form for output. This process can convert the original sequence into numerical features with statistical significance, retaining the local composition information of the sequence.
[0088] Compared with the prior art, the traditional method usually directly uses single nucleotide or di-nucleotide frequency as a feature, which cannot effectively capture long-distance sequence patterns. However, the present scheme can extract local patterns containing more rich context information, such as the preference of tri-nucleotide combination, by fixed-length multi-nucleotide unit segmentation. At the same time, the word frequency statistics combined with the normalization processing method makes the feature expression not only retain the statistical characteristics of the original sequence, but also meet the input requirements of the machine learning model.
[0089] By the technical solution, local composition features of the RNA sequence can be effectively extracted, and learning ability of the model on sequence patterns is enhanced. By fixed-length unit segmentation, sequence rules of different scales can be captured; by word frequency statistics and normalization processing, numerical features with biological significance and suitable for model training can be generated, thereby providing high-quality input data for subsequent multi-feature fusion and adversarial training.
[0090] In an implementable embodiment, the reference Figure 4 , the step S200A2 comprises steps S200A21-S200A23, in which:
[0091] In the step S200A21, a chemical parameter of a nucleotide corresponding to the interaction site sequence is queried, and based on a number of hydrogen bonds and a rotation angle of the nucleotide, basic chemical parameter data is obtained;
[0092] In the step S200A22, a mathematical transformation operation is performed based on the basic chemical parameter data, and a normalized value is calculated for each nucleotide position;
[0093] In the step S200A23, the mathematical transformation result is combined into a feature matrix through a vectorization operation, and second feature data is output.
[0094] In the embodiment, the chemical parameter refers to a number of hydrogen bonds and a rotation angle of a glycosidic bond in a nucleotide molecule, and standardized numerical values recorded in a biochemistry database can be used for querying, for example, the number of hydrogen bonds of adenine is 2, and the rotation angle is 120 degrees. The mathematical transformation operation refers to a nonlinear conversion of the original chemical parameter, and logarithmic transformation or standardization processing can be used, for example, the number of hydrogen bonds is divided by the maximum value to eliminate the dimensional difference. The normalized value refers to mapping the chemical parameters of different dimensions to a unified numerical range, and the minimum-maximum scaling method can be used, for example, the rotation angle parameter is linearly converted to the interval of 0 to 1. The vectorization operation refers to arranging the features of each nucleotide position in sequence into a multidimensional array, and the matrix splicing method can be used, for example, the normalized values of each position are stacked into a row vector in sequence order.
[0095] Specifically, the chemical feature extraction process comprises three steps: firstly, the number of hydrogen bonds and the rotation angle parameter of each nucleotide are obtained from a preset database, for example, the number of hydrogen bonds of cytosine is 3, and the rotation angle is 90 degrees; then, the two parameters are respectively subjected to mathematical transformation, for example, the number of hydrogen bonds is subjected to logarithmic conversion to reduce the numerical difference, and the rotation angle is subjected to angle radian conversion; then, the transformed parameters are subjected to normalization processing, for example, the z-score standardization method is used to eliminate the dimensional influence of different parameters; finally, the processing results of each nucleotide position are combined into a two-dimensional matrix in sequence order, for example, the number of hydrogen bonds and the rotation angle of the i th nucleotide in the sequence are respectively taken as two column elements of the i th row of the matrix.
[0096] Compared with the prior art, the conventional method usually only uses a single parameter or does not perform standardization when extracting chemical features, for example, directly using the number of hydrogen bonds as a feature without considering its relevance to the rotation angle, resulting in dimensional differences and distribution deviations between features. The present scheme can eliminate the dimensional differences of different chemical properties while retaining the spatial conformation information of nucleotides, such as reflecting the influence of nucleotide spatial arrangement on the binding site through the rotation angle parameter, by joint extraction of multiple parameters and mathematical transformation operations.
[0097] Through the above technical scheme, the present application can effectively extract multi-dimensional features of nucleotide chemical properties, solve the technical problems of single chemical features and non-standardization in traditional methods, so that the subsequent model can be trained using chemical features with comparability and discrimination, for example, by using normalized hydrogen bond number and rotation angle parameters to enhance the model's recognition ability of sensitive regions of spatial conformation, while avoiding model training bias caused by parameter dimensional differences.
[0098] In a feasible implementation manner, with reference to Figure 5 , step S200 further includes steps S200B1-S200B3, wherein:
[0099] Step S200B1, encoding the interaction site sequence by a bidirectional Transformer architecture operation, and calculating self-attention weights according to a multi-head self-attention mechanism;
[0100] Step S200B2, residual connection operation according to the self-attention weights to obtain corresponding residual data;
[0101] Step S200B3, linear transformation of the residual data, concatenation and projection of multi-head attention output into deep features to output third feature data.
[0102] In this embodiment, the bidirectional Transformer architecture refers to an encoder-based deep learning model structure, which can be specifically implemented by a neural network module containing forward and backward information transmission layers, and functions to capture the bidirectional context dependency of each nucleotide position in the sequence. The multi-head self-attention mechanism refers to an operation of splitting attention calculation into multiple parallel subspaces, which can be specifically implemented by assigning different dimensions of query, key, and value vectors, and functions to extract the association patterns between nucleotides in the sequence from different semantic perspectives. The residual connection operation refers to an operation of establishing a cross-layer jump connection between neural network layers, which can be specifically implemented by element-wise addition of input features and transformed output features, and functions to alleviate the gradient vanishing problem in the training process of deep networks. The linear transformation refers to an operation of matrix multiplication and nonlinear activation processing of feature data, which can be specifically implemented by a combination of fully connected layers and activation functions, and functions to integrate the dispersed features output by the multi-head attention mechanism into a unified deep representation.
[0103] Specifically, in the process of obtaining the third feature data, the circular RNA sequence is first input into the pre-trained bidirectional Transformer model, and the sequence is abstracted layer by layer by the multi-layer encoder. In each layer of the encoder, the multi-head self-attention mechanism is used to calculate the association weight of each nucleotide position with other positions, for example, for a sequence of length N, an N x N attention matrix is generated to represent the global dependency relationship. Subsequently, the attention weight is multiplied by the value vector to obtain the weighted feature, which is fused with the original input through the residual connection to avoid information loss. Further, the outputs of multiple attention heads are spliced in the feature dimension, and are mapped to the target dimension through a linear projection layer, and finally the third feature data containing context semantic information is output.
[0104] Compared with the prior art, the traditional method usually uses a unidirectional recurrent neural network or a convolutional network to extract sequence features, which is difficult to capture long-distance dependency relationships. The bidirectional Transformer architecture introduced in the present scheme can utilize the context information before and after the sequence, for example, when processing the closed structure of circular RNA, it can effectively identify the interaction between the beginning and end of the sequence. In addition, compared with a single attention model, the multi-head attention mechanism can extract association features of different granularities from multiple subspaces, for example, it can simultaneously focus on the interaction patterns between local base pairing and global domains, thereby improving the feature expression capability.
[0105] By the technical solution, the application solves the problems of insufficient long-range dependence capture and single feature dimension in the sequence context modeling of the existing method. The deep context features extracted by the pre-trained language model can more accurately represent the base interaction rules implied in the circRNA sequence, providing high-discrimination input data for subsequent feature fusion and adversarial training, and thus improving the accuracy of interaction site prediction. At the same time, the visual output of the attention weight provides an explainability basis for motif analysis. For example, the weight distribution can be used to identify conserved sequence fragments with biological significance.
[0106] It can be understood that a single encoding scheme cannot completely simulate the real translation process of a protein, and therefore cannot fully capture all information of circRNA-RBP interaction, which greatly affects the recognition performance, especially when dealing with small-scale data sets. When processing the sequence text of circRNA-RBP interaction sites, characters or words at different positions carry different semantic information. Therefore, compared with single-scale encoding, multi-scale feature encoding can include more rich information, so that different features can be observed at different scales.
[0107] In a feasible implementation, with reference to Figure 6 , the step S300 includes steps S310-S330, in which:
[0108] In step S310, the first feature data is applied to a convolution kernel to extract local features, and the average values of spatial adjacent windows are calculated by average pooling to obtain first processed feature data.
[0109] In step S320, the second feature data is applied to a convolution kernel to extract local features, and the average values of spatial adjacent windows are calculated by average pooling to obtain second processed feature data.
[0110] In step S330, based on the first processed feature data, the second processed feature data and the third feature data, a sigmoid activation function is used to calculate the weight, and a linear combination operation is performed for feature fusion processing to obtain a fusion feature matrix.
[0111] In this embodiment, the convolution kernel extracts local features, which means that the local pattern between adjacent nucleotides in the sequence is captured by using convolution operation. Specifically, a one-dimensional convolution kernel can be used to calculate the feature map by sliding along the sequence direction. For example, the width of the convolution kernel can be set to 3-5 nucleotide units. The average pooling calculates the average value of the spatially adjacent window, which means that the dimensionality reduction processing is performed on the feature map after convolution. Specifically, a sliding average operation with a window length of 2-4 can be used to retain the main features of the local area and reduce the data dimension. The sigmoid activation function calculates the weight, which means that the feature is mapped to the 0-1 interval as a fusion weight through a nonlinear transformation. Specifically, a fully connected layer combined with a sigmoid function can be used to dynamically adjust the contribution of different feature sources. The linear combination operation performs feature fusion processing, which means that the weighted features are added element by element. Specifically, a combination of matrix point multiplication and addition operation can be used to realize the collaborative integration of multi-source features.
[0112] Specifically, the first feature data is decomposed into multiple local feature maps after convolution kernel processing. The feature values in each window are averaged by the average pooling operation to obtain the compressed first processed feature data. The second feature data is processed by convolution and pooling in the same way to generate the second processed feature data. The third feature data directly retains its original deep semantic features. Subsequently, the first processed feature data and the second processed feature data are input into a fully connected layer to generate an initial weight vector, which is normalized by a sigmoid function to obtain a fusion weight in the 0-1 interval. Finally, the third feature data is added element by element with the weighted first processed feature data and the second processed feature data to form a fusion feature matrix containing multi-scale features and context semantics.
[0113] Compared with the prior art, the traditional method usually uses feature splicing or simple weighted average for fusion, which easily ignores the nonlinear correlation between different feature sources, resulting in information redundancy or loss of key features. The present scheme enhances the local pattern capture ability through convolution operation, and combines with a dynamic weight calculation mechanism to adaptively balance the contribution ratio between sequence statistical features, chemical features and semantic features, effectively improving the discrimination of feature expression.
[0114] Through the above technical scheme, the present application solves the problem of low information integration efficiency in the multi-source feature fusion process. Local key information is extracted by convolution and pooling operations, and the dynamic weight adjustment mechanism is used to realize the collaborative complementation between features, so that the fused feature matrix can more comprehensively reflect the physical and chemical properties, statistical rules and context correlation characteristics of the sequence, providing high-discrimination input data for subsequent adversarial training, thereby improving the accuracy of interaction site prediction.
[0115] In a feasible implementation manner, reference is made to Figure 7, step S330 includes steps S331-S333, wherein:
[0116] In step S331, the first processing feature data and the second processing feature data are input into a sigmoid function through a parameter matrix multiplication operation.
[0117] In step S332, based on the output of the acquired sigmoid function, a fusion weight is acquired through a weight calculation operation.
[0118] In step S333, based on the fusion weight and the third feature data, a fusion feature matrix is output through a weighted summation operation.
[0119] In this embodiment, the parameter matrix multiplication operation refers to a process of linear transformation of feature data and a learnable parameter matrix, which can be specifically implemented by a matrix multiplication operation. This operation can map features of different scales to a unified space for subsequent fusion. The sigmoid function refers to a nonlinear activation function that compresses input values to the interval of 0-1, and the output thereof can be used as a dynamic weight to adjust the contribution of different features. The fusion weight refers to a dynamic distribution coefficient output by the activation function, which can be specifically implemented by an element-by-element multiplication and summation operation. This weight can adaptively adjust the importance of each modality according to the input feature. The weighted summation operation refers to a process of multiplying the weight and the feature by position and then accumulating, which can be specifically implemented by a vector dot product operation. This operation can achieve smooth integration of multi-source features.
[0120] Specifically, in the feature fusion stage, the first processing feature data and the second processing feature data processed by convolution and pooling are first multiplied by a trainable parameter matrix, and the product is input into a sigmoid function for nonlinear transformation. At this time, the output value of each feature position is constrained between 0 and 1, forming a preliminary weight distribution. Then, the two weight distributions are multiplied element by element to obtain a final fusion weight matrix. The weight matrix is multiplied by the third feature data (i.e., the context feature extracted by the language model) in the corresponding positions, and finally all weighted feature vectors are summed along the channel dimension to generate a fusion feature matrix with multi-scale features.
[0121] Compared with the prior art, the conventional feature fusion method usually adopts a fixed weight splicing or simple average strategy, which is difficult to adaptively adjust the contribution of different feature modalities. The present scheme introduces a learnable parameter matrix and a sigmoid function to realize a data-driven dynamic weight distribution mechanism, which can automatically adjust the fusion proportion of multi-source features according to the local features and global context relationship of a specific sequence, overcoming the limitations of static fusion strategies in complex biological sequence analysis.
[0122] By the technical solution, the application effectively solves the problem of rigid weight allocation in the multi-modal feature fusion process, so that the local chemical features, word frequency statistical features and global context features can be differentially fused at different sequence positions. The dynamic fusion mechanism significantly improves the recognition ability of the model for key sites in the circular RNA sequence, and provides a more discriminative feature representation basis for subsequent adversarial training.
[0123] It can be understood that, in order to more effectively utilize the features of different information, the application adopts a deep multi-scale feature fusion technology, which fuses deep features by jointly modeling different features. In multi-scale features, there may be semantic differences between features of different scales. If these features of different scales are fused, noise and other problems may be introduced, thereby producing negative effects. Therefore, after feature extraction, the application retains the mapping of features of different scales by a hierarchical feature fusion structure, so as to fully utilize these hierarchical features.
[0124] In a feasible implementation manner, referring to Figure 8 , the step S400 includes steps S410-S430, in which:
[0125] In step S410, by a fast gradient method operation, in the model training process, the gradient of the cross-entropy loss function is calculated to generate adversarial perturbation data;
[0126] In step S420, based on the adversarial perturbation data, by a perturbation application operation, perturbation is added to the fusion feature matrix;
[0127] In step S430, based on the fusion feature matrix after adding the perturbation, by a loss function optimization operation, the model parameters are trained using the cross-entropy loss to output a trained prediction model.
[0128] In this embodiment, the fast gradient method operation refers to an algorithm for generating adversarial samples based on the gradient direction of the loss function in the model training process, which can be implemented by using the FGSM algorithm. By calculating the gradient of the loss function with respect to the input feature and adding perturbation along the gradient direction, the robustness of the model to input perturbation is enhanced. The adversarial perturbation data refers to a small noise vector generated by gradient calculation, and the amplitude thereof can be controlled by a perturbation coefficient, for example, the perturbation coefficient is set to a value between 0.01 and 0.1, which is used to simulate data noise and improve the generalization ability of the model. The cross-entropy loss function is an optimization objective function for classification tasks, which guides parameter update by calculating the difference between the predicted probability distribution and the true label, and can be implemented in the form of multi-class cross-entropy, which is used to measure the prediction error of the model on the adversarial samples.
[0129] Specifically, in the model training phase, firstly, the cross-entropy loss value corresponding to the fusion feature matrix is calculated through forward propagation, and then the gradient matrix of the loss function to the input feature is calculated through back propagation. The sign direction of the gradient matrix is used to generate an adversarial perturbation vector, and the perturbation is superimposed on the original feature matrix in a weighted manner. Through multiple iterations to update the model parameters, the model can maintain stable classification performance on both adversarial samples and original samples. For example, at each parameter update, the loss gradient corresponding to the original input can be calculated first, the perturbed feature is generated, and the adversarial loss is calculated again. Finally, the two losses are weighted and summed for back propagation.
[0130] Compared with the prior art, the traditional model training method only optimizes the parameters for the original data, and is prone to performance degradation on noisy data or out-of-distribution samples. The present scheme introduces an adversarial training mechanism to construct challenging perturbed samples in the feature space, forcing the model to learn more general feature representations. This active defense mechanism can effectively improve the tolerance of the model to data noise and overcome the overfitting defect of traditional methods on complex biological sequence data.
[0131] Through the above technical solutions, the robustness and generalization ability of the prediction model in real scenarios can be significantly improved, so that the model can still maintain stable interaction site recognition accuracy when facing experimental noise or unknown sequence variations. At the same time, the perturbed features generated during the adversarial training process can be used as a data augmentation means to effectively alleviate the overfitting problem caused by small sample training data.
[0132] In a feasible implementation, with reference to Figure 9 , step S500 includes steps S510-S540, wherein:
[0133] Step S510, obtaining importance score data of each nucleotide position in the prediction model through attention weight extraction operation;
[0134] Step S520, based on the importance score data, extracting high-weight fragment data higher than a preset importance threshold through threshold screening operation;
[0135] Step S530, based on the high-weight fragment data, identifying a conserved sequence pattern through STREME tool operation;
[0136] Step S540, based on the conserved sequence pattern, matching with a known functional motif database through database comparison operation to obtain biological interpretation information.
[0137] In this embodiment, the attention weight extraction operation refers to extracting the contribution of each nucleotide position to the prediction result from the self-attention layer of the model. Specifically, the weight value of the corresponding position in the output matrix of the attention layer can be used as the importance score, thereby quantifying the influence of different sequence regions on the prediction result. The threshold screening operation refers to filtering out fragments with significantly higher importance than background noise according to a preset numerical value. Specifically, statistical methods can be used to calculate the distribution of global scores, for example, regions with a score higher than twice the standard deviation of the average value are defined as high weight fragments, thereby focusing on regions with potential biological significance. The STREME tool operation refers to using bioinformatics tools to mine sequence patterns in high weight fragments. Specifically, the probability model in the tool can be used to analyze high-frequency nucleotide combination patterns, thereby identifying conserved motif structures. The database comparison operation refers to matching the identified motifs with known functional motifs in public databases. Specifically, sequence similarity comparison algorithms or regular expression matching methods can be used, for example, the motifs are compared with entries in the JASPAR or MEME database, thereby associating with known biological functions or regulatory mechanisms.
[0138] Specifically, after the model prediction is completed, the attention weight value corresponding to each nucleotide position is first extracted from the output of the self-attention layer to form importance score data. Then, by setting a dynamic threshold, continuous sequence fragments with scores significantly higher than the average level are screened out, for example, high weight regions with a length of 5-15 nucleotides. These high weight fragments are input into the STREME tool, and the built-in expectation maximization algorithm is used to mine repeated sequence patterns to generate candidate motifs. Finally, the candidate motifs are compared with the functional motif database, and if a known motif is matched, the corresponding biological annotation information, such as transcription factor binding sites or splicing regulatory signals, is output.
[0139] In some specific embodiments, the preset importance threshold can be dynamically adjusted according to the distribution of the training data, for example, the mean and variance of the local region score are calculated using a sliding window, and then an adaptive threshold is determined. The parameter settings of the STREME tool can include the minimum motif length and the maximum number of motifs, for example, the minimum length is set to 4 nucleotides and the maximum number is set to 10. Regular expression matching can be used in the database comparison process, allowing a certain degree of mismatch to be compatible with sequence variations.
[0140] It can be understood that the accuracy of prediction has been the core and even the only standard for evaluating the performance of a model for a long time, and in order to obtain better classification accuracy, deeper and more complex deep neural network models have been designed in previous studies. Although the accuracy of these models can meet the application requirements, their robustness under adversarial attacks has not been fully studied. Adversarial training is a new regularization method that can be used to improve the robustness of the classifier under the worst-case perturbation. The present application introduces adversarial training and transfer learning, and combines attention mechanism to construct a deep learning model with high accuracy and robustness, and improves the prediction performance of circRNA-RBP interaction sites.
[0141] Compared with the prior art, the existing method usually only relies on the model prediction result and lacks explanation of the internal decision mechanism, or needs manual screening of key sequence fragments for motif analysis. The present scheme realizes a closed-loop process from model prediction to biological explanation by automatically extracting attention weights and combining bioinformatics tools, avoids subjective bias caused by manual intervention, and improves the efficiency of motif identification.
[0142] Through the above technical solutions, the present application can automatically identify the key sequence region affecting the decision in the prediction model, and associate it with the known functional motif, thereby providing a verifiable biological basis for the prediction result. This helps to reveal the potential molecular mechanism of circRNA and RBP interaction, such as discovering new binding sites or verifying existing regulatory patterns, and improves the credibility and practicality of the prediction model.
[0143] It should be noted that the above examples are only used to understand the present application and do not constitute a limitation on the present application, which is an interpretable circRNA and RBP interaction site prediction method based on adversarial training. More forms of simple transformation based on this technical concept are within the protection scope of the present application.
[0144] The present application also provides an interpretable circRNA and RBP interaction site prediction system based on adversarial training, which refers to Figure 10 The system comprises a memory 10, a processor 20, and an interpretable circRNA and RBP interaction site prediction program based on adversarial training stored on the memory 10 and executable on the processor 20, wherein the interpretable circRNA and RBP interaction site prediction program based on adversarial training is configured to implement the steps of the interpretable circRNA and RBP interaction site prediction method based on adversarial training.
[0145] The application provides an explainable circular RNA and RBP interaction site prediction system based on adversarial training. The explainable circular RNA and RBP interaction site prediction method based on adversarial training in the above embodiment can effectively integrate multi-scale feature information, improve the adversarial robustness of the model, provide biological explainability, and enhance the feature collaborative learning capability. Compared with the prior art, the application provides an explainable circular RNA and RBP interaction site prediction system based on adversarial training. The beneficial effects of the explainable circular RNA and RBP interaction site prediction method based on adversarial training provided in the above embodiment are the same, and other technical features of the explainable circular RNA and RBP interaction site prediction system are the same as the features disclosed in the above embodiment method. Here, no further description is given.
[0146] The above only describes some embodiments of the application, and does not limit the patent scope of the application. Any equivalent structural transformation made by using the content of the application specification and drawings, or direct / indirect application in other related technical fields is included in the patent protection scope of the application.
Claims
1. An anti-training-based circular RNA and RBP interaction site prediction method with explainability, characterized in that, The method comprises: obtaining the interaction site sequence of circular RNA and RBP; multi-scale feature coding processing is carried out on the interaction site sequence to obtain first feature data and second feature data, and the context correlation feature of the interaction site sequence is extracted through a pre-trained language model to obtain third feature data; based on the first feature data, the second feature data and the third feature data, the feature fusion processing is carried out to obtain the fusion feature matrix; based on the fusion feature matrix, the training is carried out to obtain the trained prediction model; the predicted interaction site result is obtained by inputting the circular RNA sequence to be predicted into the prediction model, and the attention weight data in the prediction model is extracted, the conserved sequence pattern is identified through motif analysis, and the biological interpretation information is obtained. 2.The method of claim 1, wherein the method is characterized by, The step of multi-scale feature coding processing on the interaction site sequence to obtain first feature data and second feature data comprises: the interaction site sequence is segmented into multiple nucleotide units and the word frequency is calculated through natural language processing to obtain the first feature data; the interaction site sequence is extracted for chemical feature, and the feature is calculated based on the number of hydrogen bonds and the rotation angle parameters of nucleotides to obtain the second feature data. 3.The method of claim 2, wherein the method is characterized by, The step of natural language processing of the interaction site sequence, segmenting the interaction site sequence into multiple nucleotide units and calculating the word frequency to obtain the first feature data comprises: the interaction site sequence is cut into multiple fixed length nucleotide units through sequence segmentation operation; word frequency statistical processing is carried out on multiple nucleotide units to calculate the number of each nucleotide unit appearing in the interaction site sequence; normalization operation is carried out based on the word frequency statistical result to convert the word frequency into vector representation to output the first feature data. 4.The method of claim 2, wherein the method is characterized by, The step of extracting chemical features of the interaction site sequence, calculating features based on the number of hydrogen bonds and the rotation angle parameters of nucleotides to obtain the second feature data comprises: query the chemical parameters of the nucleotides corresponding to the interaction site sequence, and obtain the basic chemical parameter data based on the number of hydrogen bonds and the rotation angle of the nucleotides; mathematical transformation operation is carried out based on the basic chemical parameter data to calculate the normalized value of each nucleotide position; the mathematical transformation result is combined into a feature matrix through vectorization operation to output the second feature data. 5.The method of claim 1, wherein the method is characterized by, The step of extracting the context correlation feature of the interaction site sequence through the pre-trained language model to obtain the third feature data comprises: the interaction site sequence is encoded through bidirectional Transformer architecture operation, and self-attention weight is calculated according to multi-head self-attention mechanism; residual connection operation is carried out according to the self-attention weight to obtain the corresponding residual data; the residual data is linearly transformed, the multi-head attention output is spliced and projected into deep features to output the third feature data. 6.The method of claim 1, wherein the method is characterized by, The step of performing feature fusion processing based on the first feature data, the second feature data, and the third feature data to obtain a fusion feature matrix comprises: applying a convolution kernel to the first feature data to extract local features, and calculating the average value of spatially adjacent windows through average pooling to obtain first processed feature data; applying a convolution kernel to the second feature data to extract local features, and calculating the average value of spatially adjacent windows through average pooling to obtain second processed feature data; based on the first processed feature data, the second processed feature data, and the third feature data, using a sigmoid activation function to calculate weights, and performing feature fusion processing through linear combination operation to obtain a fusion feature matrix.
7. The method for predicting circRNA-RBP interaction sites based on adversarial training with explainability according to claim 6, wherein, The step of performing feature fusion processing based on the first processed feature data, the second processed feature data, and the third feature data, using a sigmoid activation function to calculate weights, and performing feature fusion processing through linear combination operation to obtain a fusion feature matrix comprises: inputting the first processed feature data and the second processed feature data into a sigmoid function through a parameter matrix multiplication operation; based on the output of the obtained sigmoid function, obtaining fusion weights through a weight calculation operation; based on the fusion weights and the third feature data, outputting a fusion feature matrix through a weighted summation operation. 8.The method of claim 1, wherein the method is characterized by, The step of performing adversarial training based on the fusion feature matrix to obtain a trained prediction model comprises: through a fast gradient method operation, calculating the gradient of the cross-entropy loss function in the model training process to generate adversarial perturbation data; based on the adversarial perturbation data, adding perturbation to the fusion feature matrix through a perturbation application operation; based on the fusion feature matrix after adding the perturbation, training the model parameters using the cross-entropy loss through a loss function optimization operation to output a trained prediction model. 9.The method of claim 1, wherein the method is characterized by, The step of extracting attention weight data in the prediction model, identifying a conserved sequence pattern through motif analysis to obtain biological interpretation information comprises: obtaining importance score data of each nucleotide position in the prediction model through an attention weight extraction operation; based on the importance score data, extracting high weight fragment data higher than a preset importance threshold through a threshold screening operation; based on the high weight fragment data, identifying a conserved sequence pattern through an STREME tool operation; based on the conserved sequence pattern, matching with a known functional motif database through a database comparison operation to obtain biological interpretation information.
10. An explainable system for predicting circular RNA and RBP interaction sites based on adversarial training, comprising: The system comprises a memory, a processor, and an interpretable adversarial training based circRNA and RBP interaction site prediction program stored on the memory and executable on the processor, and the interpretable adversarial training based circRNA and RBP interaction site prediction program is configured to implement the steps of the interpretable adversarial training based circRNA and RBP interaction site prediction method according to any one of claims 1 to 9.