RNA structure prediction method, electronic device, storage medium and program product

CN119943155APending Publication Date: 2025-05-06FUDAN UNIVERSITY
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510020185.0
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Priority Date
2024-11-27
Filing Date
2025-01-06
Publication Date
2025-05-06

Smart Images

  • Figure CN119943155A_ABST
    Figure CN119943155A_ABST
Patent Text Reader

Abstract

The embodiment of the invention provides an RNA (Ribonucleic Acid) structure prediction method, electronic equipment, a storage medium and a program product. The method comprises the following steps: acquiring a to-be-detected base sequence; the to-be-detected base sequence is input into the RNA language model, an RNA secondary structure is obtained, the RNA language model is used for extracting feature information (such as the base arrangement sequence and the base type of each base) of the to-be-detected base sequence, and the RNA secondary structure corresponding to the feature information is determined according to the incidence relation between the feature information and the RNA secondary structure; therefore, the secondary structure of the RNA is accurately determined, and prediction errors are reduced.
Need to check novelty before this filing date? Find Prior Art

Description

[0001] This application claims the priority of the Chinese patent application filed with the China Patent Office on November 27, 2024, with application number 202411728765.7 and application name “RNA structure prediction method, electronic device, storage medium and program product”, all contents of which are incorporated by reference in this application. Technical Field

[0002] The present application relates to the field of computer technology, and in particular to an RNA structure prediction method, electronic equipment, storage medium and program product. Background Art

[0003] In recent years, high-throughput experimental methods have developed rapidly, resulting in the generation of a large number of new ribonucleic acids (RNAs). The prediction of RNA function has become a core issue in the field of molecular biology research.

[0004] The function of RNA is strongly related to its structure. Therefore, predicting the higher-level structure of RNA can help study the function and regulatory mechanism of RNA molecules. However, the traditional RNA secondary structure prediction method is time-consuming and expensive. Summary of the invention

[0005] The embodiments of the present application provide an RNA structure prediction method, electronic device, storage medium and program product, which solve the problems of low efficiency and low accuracy in predicting RNA secondary structure.

[0006] In the first aspect, the present application proposes an RNA structure prediction method, which is applied to electronic devices, and the method includes: obtaining a base sequence to be tested; inputting the base sequence to be tested into an RNA language model to obtain an RNA secondary structure, wherein the RNA language model is used to extract feature information of the base sequence to be tested, and determine the RNA secondary structure corresponding to the feature information based on the correlation between the feature information and the RNA secondary structure.

[0007] It can be understood that the characteristic information of the base sequence to be tested can be, for example, the base arrangement order and the base type of each base. The electronic device (such as the electronic device 1500 proposed below) can be based on the RNA language model with a large language model architecture, and deeply learn the characteristic information of the base sequence, such as the arrangement order of each base in the base sequence and the base type of each base, and the relationship between the characteristic information and the RNA secondary structure, such as base pairing rules, sequence conservation, etc. The electronic device can learn how to establish a mapping relationship between the characteristics of the base sequence of RNA and the corresponding RNA secondary structure based on the RNA language model. Thus, the electronic device can use the RNA language model to efficiently and accurately predict the RNA secondary structure corresponding to the base sequence to be tested, without querying the RNA sequence database (such as the RNA sequence database 01 proposed below). Compared with traditional methods, there are significant improvements in prediction accuracy, computational efficiency and adaptability, which provides strong support for RNA structural biology research, gene therapy, drug design and other fields.

[0008] In a possible implementation of the first aspect above, the association between the characteristic information and the RNA secondary structure includes the base pairing rule and sequence conservation of the RNA base sequence.

[0009] It can be understood that the basic pairing relationship (i.e., pairing rule) of RNA is as follows: adenine (A) is paired with uracil (U), and they are connected by two hydrogen bonds. Cytosine (C) is paired with guanine (G), and they are connected by three hydrogen bonds. The above pairing rule is the basis for predicting RNA secondary structure. The sequence conservation of RNA base sequences refers to the evolutionary relationship of RNA base sequences. Based on the RNA language model, by training on a large-scale sequence data set (such as the sequence data set provided by the RNAcentral database), the evolutionary relationship of RNA base sequences can be captured, that is, the sequence conservation of RNA base sequences can be determined. Furthermore, the trained RNA language model includes the characteristic information of the RNA base sequence (such as the base arrangement order and the type information of the base) and the correlation relationship with the RNA secondary structure (such as base pairing rules, sequence conservation). Based on the trained RNA language model, the RNA secondary structure of the RNA base sequence to be tested can be accurately predicted.

[0010] In a possible implementation of the first aspect above, the RNA secondary structure corresponding to the characteristic information is determined according to the association between the characteristic information and the RNA secondary structure, including: determining the pairing probability of each base in the base sequence to be tested according to the association between the characteristic information and the RNA secondary structure; and determining the RNA secondary structure corresponding to the characteristic information according to a probability sequence constituted by the pairing probability of each base.

[0011] Exemplarily, an electronic device (e.g., electronic device 1500 proposed below) can use the base pairing mode that satisfies the pairing probability condition as the predicted RNA secondary structure. It can be understood that RNA secondary structure prediction is to determine the pairing mode of RNA single strands, that is, to predict the hydrogen bond pairing mode between adjacent bases of the RNA internal single strand. Therefore, the base pairing mode with a higher pairing probability can be used as the predicted RNA secondary structure.

[0012] In some embodiments of the present application, the pairing probability of each of the above bases may be between 0 and 1, and is not limited here.

[0013] In some embodiments, the electronic device 1500 may use a base pairing pattern greater than a pairing threshold as a predicted RNA secondary structure. For example, a base pairing pattern with a pairing probability greater than 0.7 may be used as a predicted RNA secondary structure.

[0014] In other embodiments, the electronic device 1500 may use the base pairing pattern with the highest pairing probability as the predicted RNA secondary structure.

[0015] In a possible implementation of the first aspect, inputting the base sequence to be tested into the RNA language model includes: during the input of the base sequence to be tested, randomly performing a masking operation on a preset proportion of bases.

[0016] In some embodiments of the present application, the preset ratio may be, for example, 15%.

[0017] In a possible implementation of the first aspect above, the training method of the RNA language model includes: obtaining training data; preprocessing the obtained training data to obtain preprocessed training data; and training the RNA language model to be trained using the preprocessed training data to obtain a trained RNA language model.

[0018] In a possible implementation of the first aspect above, the preprocessing includes removing irrelevant characters and / or completing missing bases.

[0019] It can be understood that preprocessing the RNA base sequence is a key step before training the RNA language model to ensure the accuracy and consistency of the input data. RNA base sequences usually contain only four bases: adenine (A), uracil (U), cytosine (C), and guanine (G). During preprocessing, all non-RNA base characters need to be removed, such as numbers, letters (non-A, U, C, G), spaces, special symbols, etc. During the sequencing process, some lower-quality bases (such as N) may be generated, which may introduce noise in subsequent analysis. Therefore, these low-quality bases need to be removed or replaced with high-quality bases. If the proportion of N in a sequence is too high, you may need to consider deleting the sequence.

[0020] In a possible implementation of the first aspect above, preprocessing the acquired training data further includes: performing clustering processing on the training data to remove redundant base sequences in the training data.

[0021] It can be understood that the training data after removing redundant base sequences is used to train the RNA language model to be trained, which can improve the generalization ability of the trained RNA language model.

[0022] In a possible implementation of the first aspect above, the training method of the RNA language model also includes: applying a preset loss function to implement the training process of the RNA language model, wherein the preset loss function includes a cross entropy loss function.

[0023] It can be understood that the cross entropy loss function can effectively improve the prediction accuracy of the RNA language model.

[0024] In a possible implementation of the first aspect above, the RNA language model includes a feature extraction network and a secondary structure prediction network, wherein the feature extraction network is used to extract feature information of the base sequence to be tested; and the secondary structure prediction network is used to determine the RNA secondary structure corresponding to the feature information based on the correlation between the feature information and the RNA secondary structure.

[0025] Exemplarily, the neural network architecture of the feature extraction network can adopt a large language model as the core architecture. For example, the feature extraction network can be a neural network based on the BERT (bidirectional encoder representations from transformers) architecture, which includes 33 bidirectional Transformer layers. The secondary structure prediction network can include a multi-layer deep convolutional neural network (CNN) and a recurrent neural network with an attention mechanism (BiLSTM).

[0026] In a possible implementation of the first aspect above, extracting feature information of the base sequence to be tested includes: encoding the base sequence to be tested into a one-hot vector based on relative position encoding through a feature extraction network, and extracting feature information according to the one-hot vector.

[0027] It can be understood that the one-hot vector facilitates subsequent feature extraction and improves processing efficiency.

[0028] In a possible implementation of the first aspect above, the preprocessed training data is input into an RNA language model to obtain a trained RNA language model, including: inputting the preprocessed training data into a feature extraction network to extract feature data of each base; inputting the feature data of each base into a secondary structure prediction network to obtain a pairing probability of each base; determining a predicted RNA secondary structure based on the pairing probability; and updating the RNA language model based on the difference between the predicted RNA secondary structure and the RNA secondary structure in the training data.

[0029] It can be understood that the electronic device can obtain unified training data by preprocessing the training data, and then obtain a more accurate trained RNA language model.

[0030] In a possible implementation of the first aspect above, the pairing probability of each base is between 0-1.

[0031] In a possible implementation of the first aspect above, the secondary structure prediction network includes a feature extractor based on a Transformer architecture, wherein the feature extractor is capable of fusing 1-dimensional sequence information and 2-dimensional pairing information.

[0032] It can be understood that the electronic device utilizes the feature extractor in the secondary structure prediction network to fuse the 1-dimensional sequence information feature vector and the 2-dimensional pairing information feature vector, so as to better extract the features of the RNA base sequence information.

[0033] In a possible implementation of the first aspect above, the feature information of the base sequence to be tested includes a 1-dimensional sequence information attention map, and the 1-dimensional sequence information and the 2-dimensional pairing information are fused, including: reducing the dimension of the 1-dimensional sequence information attention map to obtain a 1-dimensional sequence information feature vector; performing an outer product operation on the 1-dimensional sequence information feature vector to obtain a 2-dimensional pairing information feature vector; adding the 1-dimensional sequence information feature vector to the 2-dimensional pairing information feature vector through an outer product operation to introduce the 2-dimensional pairing information; and adding the 2-dimensional pairing information feature vector as a bias item to the 1-dimensional sequence information attention map to introduce the 1-dimensional sequence information.

[0034] It can be understood that by adding the new 2D pairing information feature vector obtained by the outer product of the 1D sequence information feature vector, the 2D pairing information can be introduced, and by using the 2D pairing information feature vector as a bias term and adding it to the 1D sequence information attention map, the 1D sequence information can be introduced. Thus, more accurate feature information can be extracted.

[0035] In a second aspect, the present application also provides an electronic device, comprising: one or more processors; one or more memories; one or more memories storing one or more programs, and when the one or more programs are executed by the one or more processors, the electronic device executes the RNA structure prediction method provided by the first aspect and various possible implementations of the first aspect.

[0036] In a third aspect, the present application also provides a computer-readable storage medium having instructions stored thereon, which, when executed on a computer, causes the computer to execute the RNA structure prediction method provided by the first aspect and various possible implementations of the first aspect.

[0037] In a fourth aspect, an embodiment of the present application discloses a computer program product, including a computer program / instruction, which, when executed by a processor, implements the RNA structure prediction method provided by the first aspect and various possible implementations of the first aspect.

[0038] The beneficial effects of the second to fourth aspects mentioned above can be referred to the relevant descriptions in the first aspect mentioned above and various possible implementations of the first aspect, and will not be elaborated here. BRIEF DESCRIPTION OF THE DRAWINGS

[0039] Figure 1A A schematic diagram showing a scenario for predicting RNA structure;

[0040] Figure 1B A schematic diagram showing another scenario of predicting RNA structure;

[0041] Figure 2 A schematic diagram showing a specific implementation process of an RNA structure prediction method provided according to some embodiments of the present application is shown;

[0042] Figure 3 A schematic diagram of a training process of an RNA language model proposed according to an embodiment of the present application is shown;

[0043] Figure 4A A schematic diagram of the structure of an RNA language model proposed according to an embodiment of the present application is shown;

[0044] Figure 4B A schematic diagram of a feature extractor design for an RNA language model proposed in some embodiments of the present application is shown;

[0045] Figure 5 A schematic diagram of a specific implementation process of training an RNA language model according to some embodiments of the present application is shown;

[0046] Figure 6 A schematic structural diagram of an electronic device provided according to some embodiments of the present application is shown. DETAILED DESCRIPTION

[0047] In order to facilitate understanding of the technical solutions provided by the embodiments of the present application, the meanings of some related field terms involved in the embodiments of the present application are explained below.

[0048] (1) RNA secondary structure refers to the local spatial structure within a single RNA molecule, which is formed by hydrogen bonding between bases. This structure mainly involves base pairing within the RNA molecule, where some bases are connected to each other by hydrogen bonds to form double-stranded regions, while unpaired bases form single-stranded loops or protrusions. RNA secondary structure is a key factor in the function and stability of RNA molecules.

[0049] (2) AlphaFold2 is an AI-based tool for predicting the three-dimensional structure of RNA.

[0050] (3) MMSeqs2 is a software suite for searching and clustering large RNA and nucleic acid sequence sets. It is implemented in C++ and is open source software under the GPL-3.0 license.

[0051] (4) Relative position encoding mechanism: This is an encoding method used in the Transformer model and its derivative models to capture the relationship between different positions in a sequence. It allows the model to understand the relative distance between words or tokens, which is very important for understanding sentence structure and semantics.

[0052] (5) One-hot encoding: Also known as one-hot encoding, it can convert categorical or nominal variables into a form that can be better processed by machine learning algorithms. In this encoding, each category value is represented as a binary vector, with all positions except one that is 1, indicating the category, and all other positions are 0.

[0053] (6) AdamW optimizer: It is a variant of the Adam optimizer (adaptive moment estimation). It adds weight decay to Adam to solve the problem of model overfitting. AdamW optimizer has faster convergence speed and robustness when training deep learning models.

[0054] In order to make the purpose, technical solutions and advantages of the embodiments of the present application clearer, the technical solutions provided by the embodiments of the present application will be described in detail below in conjunction with the accompanying drawings and specific implementation methods.

[0055] It can be understood that the electronic device 1500 in the embodiment of the present application can be a terminal device, which can also be called a terminal, a user equipment (UE), a mobile station (MS), a mobile terminal (MT), etc. The terminal device can be a mobile phone, a smart TV, a wearable device, a tablet computer (pad), a computer with a wireless transceiver function, a virtual reality (VR) terminal device, an augmented reality (AR) terminal device, a wireless terminal in industrial control, a wireless terminal in self driving, a wireless terminal in remote medical surgery, a wireless terminal in smart grid, a wireless terminal in transportation safety, a wireless terminal in smart city, a wireless terminal in smart home, a vehicle-mounted terminal, etc.

[0056] The following uses the electronic device 1500 as the terminal 100 as an example, and describes in further detail the RNA structure prediction method provided in the embodiment of the present application in combination with relevant drawings.

[0057] Figure 1A A schematic diagram showing a scenario for predicting RNA structure. Figure 1B Schematic diagram showing another scenario for predicting RNA structure.

[0058] refer to Figure 1A The algorithm set in the terminal 100 reads the RNA primary structure (hereinafter referred to as the base sequence to be tested) input by the user and can return the predicted RNA secondary structure to the user.

[0059] As mentioned above, the reliability and efficiency of Terminal 100 in predicting RNA secondary structure are relatively low.

[0060] For example, refer to Figure 1B, the terminal 100 reads the base sequence to be tested and can send a query request to the RNA sequence database 01. The RNA sequence database 01 can be stored in the storage environment 201 provided by the server 200, and a large amount of RNA primary structure (i.e., the base sequence of the RNA) and the corresponding RNA secondary structure data are pre-stored in the RNA sequence database 01. The terminal 100 can find out the RNA secondary structure corresponding to the base sequence to be tested by comparing the base sequence to be tested with the base sequence in the RNA sequence database 01, and return the RNA secondary structure obtained by the query to the user.

[0061] However, RNA species are constantly increasing, and many base sequences in the RNA sequence database 01 have not yet determined their secondary structures. Therefore, it is not necessarily possible to find the secondary structure corresponding to the base sequence to be tested in the RNA sequence database 01. Figure 1B The method shown has low robustness in predicting RNA secondary structure. In addition, corresponding to a large number of RNA secondary structure prediction scenarios, a large number of base sequences to be tested can generate high-concurrency query requests. Limited by the hardware performance of the terminal 100, the design and architecture of the RNA sequence database 01, and the number of data connections between the terminal 100 and the server 200, the query efficiency of this method is relatively low.

[0062] In order to solve the problems of low efficiency and low accuracy in predicting RNA secondary structure, an embodiment of the present application provides an RNA prediction method, which is applied to electronic equipment, including: obtaining a base sequence to be tested; inputting the base sequence to be tested into an RNA language model to obtain the RNA secondary structure, wherein the RNA language model is used to extract characteristic information of the base sequence to be tested (such as the base arrangement order and the base type of each base), and determine the RNA secondary structure corresponding to the characteristic information based on the correlation between the characteristic information and the RNA secondary structure, thereby accurately determining the secondary structure of the RNA and reducing prediction errors.

[0063] The embodiment of the present application utilizes the powerful ability of the large language model in the field of natural language processing and applies it to the prediction of RNA secondary structure. It can be understood that the primary structure of RNA is a base sequence, and the base sequence can also be regarded as a "language", in which the base pairing relationship corresponds to the grammatical rules of the "language". Furthermore, the electronic device 1500 provided in the embodiment of the present application can be based on the RNA language model with a large language model architecture, and deeply learn the characteristic information of the base sequence, such as the arrangement order of each base in the base sequence and the base type of each base, and the relationship between the characteristic information and the RNA secondary structure, such as base pairing rules, sequence conservation, etc. The electronic device 1500 can learn how to establish a mapping relationship between the characteristics of the base sequence of RNA and the corresponding RNA secondary structure based on the RNA language model. Thus, the electronic device 1500 can use the RNA language model to efficiently and accurately predict the RNA secondary structure corresponding to the base sequence to be tested, without querying the RNA sequence database 01. Compared with traditional methods, there are significant improvements in prediction accuracy, computational efficiency and adaptability, which provides strong support for RNA structural biology research, gene therapy, drug design and other fields.

[0064] In the embodiment of the present application, the electronic device 1500 can input the RNA base sequence into the RNA language model, and predict the secondary structure of the RNA through the model's "understanding" and "analysis" of the sequence.

[0065] Combine the following Figure 2 The specific implementation process and principle of predicting RNA structure provided in the examples of the present application are described in detail.

[0066] Figure 2 A schematic diagram of a specific implementation process of an RNA structure prediction method provided according to some embodiments of the present application is shown.

[0067] Here, Figure 2 The execution subject of the specific implementation process shown may be the electronic device 1500, which will not be elaborated here.

[0068] refer to Figure 2 , the specific implementation process includes the following steps:

[0069] S201, obtaining a base sequence to be tested.

[0070] It can be understood that the base sequence to be tested is the base sequence of the RNA to be tested, which is the primary structure of the RNA to be tested.

[0071] In some embodiments, the electronic device 1500 may obtain the base sequence to be tested input by the user through an input control, and the input control may be, for example, an input box control.

[0072] S202, inputting the base sequence to be tested into an RNA language model, wherein the RNA language model is used to extract feature information of the base sequence to be tested, and determine the RNA secondary structure corresponding to the feature information according to the correlation between the feature information and the RNA secondary structure.

[0073] It can be understood that the electronic device 1500 can extract the corresponding features in the base sequence to be tested based on the above RNA language model, and match the features with the association between the pre-trained base sequence and the RNA secondary structure in the model, for example, according to the base pairing rules and sequence conservation, determine the most likely base pairing mode of each base in the base sequence to be tested, and then determine the pairing probability sequence corresponding to the base pairing mode, and then output the RNA secondary structure corresponding to the paired base. Among them, the association between the base sequence and the RNA secondary structure may include base pairing rules and sequence conservation.

[0074] It can be understood that since the electronic device 1500 has learned the correlation between the base sequence and the RNA secondary structure based on the RNA language model, the electronic device 1500 can predict the accurate RNA secondary structure without searching the RNA sequence database 01 mentioned above, which is conducive to large-scale RNA secondary structure prediction.

[0075] S203, obtain RNA secondary structure.

[0076] Exemplarily, the electronic device 1500 may determine the RNA secondary structure based on the pairing probability output by the RNA language model.

[0077] It is understood that the electronic device 1500 can use the base pairing mode that meets the pairing probability condition as the predicted RNA secondary structure. It is understood that RNA secondary structure prediction is to determine the pairing mode of RNA single strands, that is, to predict the hydrogen bond pairing mode between adjacent bases in the RNA internal single strand. Therefore, the base pairing mode with a higher pairing probability can be used as the predicted RNA secondary structure.

[0078] In some embodiments of the present application, the electronic device 1500 can obtain the pairing probabilities between the bases output by the RNA language model, and can determine the RNA secondary structure according to the probability sequence composed of the pairing probabilities corresponding to each base.

[0079] In some embodiments of the present application, the pairing probability of each of the above bases may be between 0 and 1, and is not limited here.

[0080] It can be understood that the bases in the RNA base sequence mainly include the following four types: adenine (A), uracil (U), cytosine (C) and guanine (G). These four bases can form a stable structure through a specific pairing relationship in the RNA base sequence. For example, adenine (A) can be paired with uracil (U), and they are connected by two hydrogen bonds. Cytosine (C) can be paired with guanine (G), and they are connected by three hydrogen bonds. These pairing relationships enable RNA molecules to form specific single-stranded or double-stranded structures, which is crucial for the role played by RNA in the transmission and expression of genetic information. The RNA secondary structure is the spatial structure formed after the above-mentioned base pairing, so the base pairing of the base sequence to be tested can be determined by predicting the paired base with a larger probability of pairing for each base to be tested in the RNA base sequence to be tested, and the constructed spatial structure is predicted based on the base pairing situation, that is, the corresponding RNA secondary structure is determined. Thus, the electronic device 1500 can accurately detect the RNA secondary structure based on the RNA language model.

[0081] According to the RNA structure prediction method shown in the above steps S201 to S203, the embodiment of the present application provides an RNA language model based on a large language model architecture, which can extract the characteristic information of the base sequence to be tested, and determine the RNA secondary structure corresponding to the characteristic information based on the correlation between the characteristic information and the RNA secondary structure, so that the electronic device 1500 can efficiently and accurately predict the RNA secondary structure based on the RNA language model.

[0082] In some embodiments of the present application, the RNA language model can be a deep learning algorithm including a multi-layer neural network, and the neural network of the deep learning algorithm can include: feedforward neural networks (FNN), where information flows in only one direction, from the input layer to the hidden layer, and then to the output layer; convolutional neural networks (CNN), which can extract local features through the convolution layer and reduce the spatial dimension of the features through the pooling layer; recurrent neural networks (RNN), which are used to process sequence data, such as time series or natural language, and can process the time dependency between input data; long short-term memory network (LSTM), a special RNN that can learn long-term dependency information and solve the gradient vanishing problem of traditional RNN; gated recurrent unit (GRU), which is similar to LSTM, but has a simpler structure and fewer parameters, and can sometimes be used as a substitute for LSTM; bidirectional long short-term memory network (bidirectional long short-term memory network, LSTM); BiLSTM is a special recurrent neural network (RNN) that can process sequence data and maintain long-term memory. BiLSTM runs two LSTMs on the time series simultaneously, one from front to back and the other from back to front, so that it can consider both past and future information at the same time. This structure enables BiLSTM to better capture the contextual relationship in the sequence data. BiLSTM is widely used in natural language processing tasks, such as sentiment analysis, machine translation, text generation and question-answering systems. It can capture the temporal dependency between the source language and the target language, as well as the contextual information of the dialogue or text in the dialogue system or news summary. Generative adversarial networks (GAN), composed of generators and discriminators, are used to generate new data samples, such as images, music, etc. Variational autoencoders (VAE), used to generate new data samples and learn the potential representation of the input data at the same time; Transformer, a model based on the self-attention mechanism, is widely used in natural language processing tasks; deep residual networks (ResNet), solve the problem of difficult training of deep networks by introducing residual learning.

[0083] In some embodiments of the present application, the electronic device 1500 can use the AdamW optimizer in combination with a large-scale RNA sequence database (such as the RNA sequence database 01 mentioned above) to perform model training on the RNA language model. In addition, the loss function used in the training process can be a cross entropy loss function, so that the trained RNA language model can accurately capture the key information in the base sequence, and then accurately predict the secondary structure of the RNA.

[0084] The specific implementation process of training the above-mentioned RNA language model in the embodiment of the present application is described in detail below in conjunction with the relevant drawings.

[0085] Figure 3 A schematic diagram of the training process of an RNA language model proposed according to an embodiment of the present application is shown.

[0086] Here, Figure 3 The execution subject of the specific implementation process shown may be the electronic device 1500, which will not be elaborated here.

[0087] refer to Figure 3 , the specific implementation process includes the following steps:

[0088] S301, obtain training data.

[0089] It can be understood that the training data can be Figure 1B The RNA data in the RNA sequence database 01 mentioned in the above may include a base sequence and a corresponding RNA secondary structure.

[0090] S302, preprocessing the acquired training data to obtain preprocessed training data.

[0091] In some embodiments of the present application, the electronic device 1500 may pre-process the input RNA base sequence. The pre-processing may include removing irrelevant characters, completing missing bases, etc., to ensure the accuracy and consistency of the input data.

[0092] It can be understood that preprocessing the RNA base sequence is a key step before training the RNA language model to ensure the accuracy and consistency of the input data. RNA base sequences usually contain only four bases: adenine (A), uracil (U), cytosine (C), and guanine (G). During preprocessing, all non-RNA base characters need to be removed, such as numbers, letters (non-A, U, C, G), spaces, special symbols, etc. During the sequencing process, some lower-quality bases (such as N) may be generated, which may introduce noise in subsequent analysis. Therefore, these low-quality bases need to be removed or replaced with high-quality bases. If the proportion of N in a sequence is too high, you may need to consider deleting the sequence.

[0093] S303, using the preprocessed training data to train the RNA language model to be trained, to obtain a trained RNA language model.

[0094] Exemplarily, the electronic device 1500 can input the preprocessed RNA base sequence into the RNA language model to be trained, use the RNA language model to be trained to extract features of the RNA base sequence, and then learn the association between the extracted features and the secondary structure. The association includes key information such as base pairing rules and sequence conservation in the RNA base sequence. It can be understood that these features will serve as the basis for subsequent predictions.

[0095] Exemplarily, the training data may be 6 million non-coding RNA base sequences from the RNAcentral database. In some embodiments, the preprocessing may also include clustering processing, for example, the electronic device 1500 may perform clustering processing on the acquired training data through MMSeqs2 to remove redundant base sequences in the training data. It is understood that the electronic device 1500 uses the clustered training data to train the RNA language model, which can further improve the generalization performance of the trained RNA language model.

[0096] It can be understood that the basic pairing relationship (i.e., pairing rule) of RNA is as follows: adenine (A) is paired with uracil (U), and they are connected by two hydrogen bonds. Cytosine (C) is paired with guanine (G), and they are connected by three hydrogen bonds. The above pairing rule is the basis for predicting RNA secondary structure. The sequence conservation of RNA base sequences refers to the evolutionary relationship of RNA base sequences. Based on the RNA language model, by training on a large-scale sequence data set (such as the sequence data set provided by the RNAcentral database), the evolutionary relationship of RNA base sequences can be captured, that is, the sequence conservation of RNA base sequences can be determined. Furthermore, the trained RNA language model includes the characteristic information of the RNA base sequence (such as the base arrangement order and the type information of the base) and the correlation relationship with the RNA secondary structure (such as base pairing rules, sequence conservation). Based on the trained RNA language model, the RNA secondary structure of the RNA base sequence to be tested can be accurately predicted.

[0097] It can be understood that the pairing rules corresponding to the above pairing relationship are not completely fixed, so it is necessary to further judge the pairing possibility of each base in the base sequence based on the pairing relationship. Therefore, the embodiment of the present application proposes an electronic device 1500 provided with an RNA language model. The electronic device 1500 can predict the pairing probability of each base sequence based on the RNA language model, so as to facilitate the prediction of the secondary structure of RNA.

[0098] In some embodiments of the present application, the above-mentioned training implementation process can apply a preset loss function to implement the training process of the RNA language model. For example, the electronic device 1500 can use the cross-entropy loss function to implement iterative optimization of the RNA language model.

[0099] It can be understood that according to the above specific implementation steps S301 to S303, the electronic device 1500 can obtain unified training data by pre-processing the training data, and then obtain a more accurate trained RNA language model.

[0100] Figure 4A A schematic diagram of the framework structure of an RNA language model proposed according to some embodiments of the present application is shown. Figure 4B A schematic diagram of a feature extractor design for an RNA language model proposed in some embodiments of the present application is shown.

[0101] In some embodiments of the present application, reference Figure 4A The RNA language model may include a feature extraction network and a secondary structure prediction network. The feature extraction network may be used to extract the features of each base of the base sequence to be tested, and the secondary structure prediction network may be used to predict the pairing of each base according to the features of each base extracted by the feature extraction network, and determine the corresponding RNA secondary structure.

[0102] Exemplarily, the neural network architecture of the feature extraction network can use a large language model as the core architecture. For example, the feature extraction network can be a neural network based on the BERT (bidirectional encoder representations from transformers) architecture, which includes 33 bidirectional Transformer layers.

[0103] In some embodiments of the present application, a relative position encoding mechanism can be introduced to improve the input layer of the feature extraction network. Relative position encoding is a way to encode RNA base sequences according to the relative relationship between positions. Relative position encoding takes into account the relative distance and relationship between different positions in the RNA base sequence, and uses learnable parameters to model these relationships. Relative position encoding can capture the relative information between positions by calculating the offset or relative position difference between different positions. Compared with absolute position encoding, relative position encoding pays more attention to the relative order and distance between positions in the sequence, and it can better handle position information in long sequences. Furthermore, after the base sequence to be tested is input into the input layer of the feature extraction network, the base sequence to be tested can be encoded into a one-hot vector through relative position encoding, which is convenient for subsequent feature extraction processing. And the one-hot vector can be extracted through 33 layers of Transformer layers to obtain a feature vector of 1280 dimensions.

[0104] In some embodiments of the present application, during the input of the base sequence to be tested, the electronic device 1500 can randomly perform a mask operation on a preset proportion of bases based on the RNA language model. For example, during the input of the base sequence to be tested, the feature extraction network can randomly perform a mask operation on a preset proportion (e.g., 15%) of bases, and the output layer of the feature extraction network can predict the base type at the masked position.

[0105] Exemplarily, the secondary structure prediction network may include a feature extractor. For example, based on the AlphaFold2 architecture advantage, a feature extractor based on the Transformer architecture may be designed. Furthermore, the electronic device 1500 may extract the features (e.g. Figure 4B The 1-dimensional sequence information attention map LM Features (L, 1280) shown in FIG. 1 is input into the secondary structure prediction network including the above-mentioned feature extractor for model training. During the training process, the secondary structure prediction network will learn how to map the features of the RNA base sequence to its corresponding secondary structure. After the training is completed, the RNA base sequence to be predicted is input into the RNA language model, and the RNA language model will predict the secondary structure of the RNA based on the learned features and mapping relationship.

[0106] For example, refer to Figure 4B The electronic device 1500 uses the secondary structure prediction network to embed the 1-dimensional sequence information attention map LM Features (L, 1280) extracted by the feature extraction network into a 1-dimensional sequence information feature vector (e.g. Figure 4BThe Seq features (L, 256) shown) and the 2-dimensional pairing information feature vector (e.g. Figure 4B Pair features(L,L,128) shown). Among them, the 1280-dimensional features LM Features(L, 1280) extracted by the feature extraction network can be reduced in dimension to obtain a 1-dimensional sequence information feature vector Seq features(L, 256), and the 1-dimensional sequence information feature vector Seq features(L, 256) can be obtained after the outer product of the 1-dimensional sequence information feature vector Seq features(L, 256) to obtain a 2-dimensional pairing information feature vector Pair features(L,L, 128). Furthermore, the electronic device 1500 uses the feature extractor in the secondary structure prediction network to fuse the 1-dimensional sequence information feature vector and the 2-dimensional pairing information feature vector, so as to better extract the features of the RNA base sequence information. For example, the 1-dimensional sequence information feature vector is added to the 2-dimensional pairing information feature vector through the outer product operation to introduce the 2-dimensional pairing information; for another example, the 2-dimensional pairing information feature vector is used as a bias term and is added to the 1-dimensional sequence information attention map LM Features(L, 1280) to introduce the 1-dimensional sequence information, thereby extracting a more accurate pairing feature (for example) that fuses the 1-dimensional sequence information and the 2-dimensional pairing information. Figure 4B After that, the fused pairing features can be input into a multi-layer deep convolutional neural network (CNN) to obtain the pairing probability of each base in the RNA base sequence, and thus predict the RNA secondary structure of the RNA base sequence to be tested.

[0107] In some embodiments of the present application, the above-mentioned secondary structure prediction network may include a multi-layer deep convolutional neural network (CNN) and a recurrent neural network (BiLSTM) with an attention mechanism. Exemplarily, the feature vector input to the secondary structure prediction network can be feature extracted by a multi-layer deep convolutional neural network (CNN). CNN can automatically learn local features in RNA sequences and convert them into higher-level feature representations through convolution operations. The extracted features are sequence modeled by a recurrent neural network (BiLSTM) with an attention mechanism. BiLSTM can consider the dependencies between adjacent elements in the feature sequence extracted by CNN, thereby improving the accuracy of the prediction, and is used for further sequence modeling and classification prediction of the extracted features.

[0108] Figure 5 A schematic diagram of a specific implementation process of training an RNA language model proposed according to some embodiments of the present application is shown.

[0109] Understandably, Figure 5 This is an example of a specific implementation of step S303 in the above text, and its execution subject may be the electronic device 1500, which will not be described in detail here.

[0110] refer to Figure 5 , the specific implementation steps include:

[0111] S303a, inputting the preprocessed training data into a feature extraction network to extract feature data of each base.

[0112] It can be understood that the electronic device 1500 can input the preprocessed training data into the feature extraction network, for example, input the preprocessed base sequence into the feature extraction network, and extract the feature data of each base in the base sequence.

[0113] In some embodiments, the training data of the feature extraction network of the RNA language model (e.g., a language network based on a large language model) can be 6 million non-coding RNA base sequences from the RNAcentral database. Here, the above training data can be clustered by MMSeqs2, which can remove redundant base sequences and improve the generalization performance of the RNA language model.

[0114] In some embodiments, the electronic device 1500 can train the feature extraction network for a total of about 25 rounds, with an initial learning rate of 1e-3. After each round, verification is performed on a randomly divided verification machine. Corresponding to the case where the first accuracy condition is met, for example, the accuracy rate decreases, the current learning rate is divided by 2 as the learning rate for the next round of training. Corresponding to the case where the second accuracy condition is met, for example, the total number of times the accuracy rate decreases reaches a preset threshold, for example, a total of 4 decreases, the training is stopped. In this way, the feature extraction network can learn the accurate conservation of RNA sequences, thereby replacing cumbersome database query operations.

[0115] S303b, inputting the characteristic data of each base into the secondary structure prediction network to obtain the pairing probability of each base.

[0116] Exemplarily, the input of the secondary structure prediction network is the features of each base in the RNA base sequence (L, 1280), and its output is the pairing probability of each base (L, L, probability). Thus, the electronic device 1500 can determine the pairing probability of each base in the training data based on the secondary structure prediction network.

[0117] In some embodiments of the present application, the pairing probability of each of the above bases may be between 0 and 1, and is not limited here.

[0118] In some embodiments of the present application, for the task of predicting RNA secondary structure, the electronic device 1500 can train the secondary structure prediction network on the TR0 training set of the bpRNA data set. This data set contains more than 10,000 sequences and is a benchmark data set in the task of RNA secondary structure prediction. It can provide more accurate and comprehensive data samples and has a better training effect.

[0119] In other embodiments, the electronic device 1500 can also use the secondary structure prediction network to predict the RNA secondary structure in the following manner: use different random initialization parameters for the secondary structure prediction network, perform training three times in total, obtain three different trained secondary structure prediction networks, and take the average of the prediction results of the three trained secondary structure prediction networks as the final result to improve the prediction accuracy.

[0120] S303c, determining the predicted RNA secondary structure based on the pairing probability.

[0121] Exemplarily, the electronic device 1500 can use the base pairing mode that satisfies the pairing probability condition as the predicted RNA secondary structure. It can be understood that RNA secondary structure prediction is to determine the pairing mode of RNA single strands, that is, to predict the hydrogen bond pairing mode between adjacent bases of the RNA internal single strand. Therefore, the base pairing mode with a higher pairing probability can be used as the predicted RNA secondary structure.

[0122] In some embodiments, the electronic device 1500 may use a base pairing pattern greater than a pairing threshold as a predicted RNA secondary structure. For example, a base pairing pattern with a pairing probability greater than 0.7 may be used as a predicted RNA secondary structure.

[0123] In other embodiments, the electronic device 1500 may use the base pairing pattern with the highest pairing probability as the predicted RNA secondary structure.

[0124] S303d, updating the RNA language model according to the difference between the predicted RNA secondary structure and the RNA secondary structure in the training data.

[0125] Exemplarily, the electronic device 1500 can compare the predicted RNA secondary structure with the known RNA secondary structure in the training data to determine the difference between the two, and can use a preset loss function to optimize the feature extraction network and the secondary structure prediction network to obtain an updated RNA language model.

[0126] In some embodiments, the preset loss function may include a cross entropy loss function.

[0127] It can be understood that through the above steps S303a to S303d, the electronic device 1500 can train the feature extraction network and the secondary structure prediction network in the RNA language model, so that the association between the base sequence and the RNA secondary structure is constructed in the RNA language model, thereby improving the model prediction robustness.

[0128] According to the RNA structure prediction method provided in the embodiments of the present application, the present application also provides a computer program product, which includes: computer program code, when the computer program code is executed on a computer, the computer implements the steps performed by the electronic device 1500 in any one of the above embodiments.

[0129] According to the RNA structure prediction method provided in the embodiments of the present application, the present application also provides a computer-readable medium, which stores a program code. When the program code is executed on a computer, the computer implements the steps performed by the electronic device 1500 in any one of the above embodiments.

[0130] Figure 6 A schematic structural diagram of an electronic device 1500 provided according to some embodiments of the present application is shown.

[0131] It can be understood that the electronic device 1500 can be the electronic device 1500 mentioned above.

[0132] like Figure 6 As shown, the electronic device 1500 includes one or more processors 1501, a system memory 1502, a non-volatile memory (NVM) 1503, a communication interface 1504, an input / output (I / O) device 1505, and a system control logic 1506 for coupling the processor 1501, the system memory 1502, the non-volatile memory 1503, the communication interface 1504 and the input / output (I / O) device 1505. Among them:

[0133] The processor 1501 may include one or more processing units, for example, a data processing unit or processing circuit that may include a central processing unit (CPU), a graphics processing unit (GPU), a digital signal processor (DSP), a microprocessor (MCU), an artificial intelligence (AI) processor or a programmable logic device (field programmable gate array, FPGA), a neural network processor (neural-network processing unit, NPU), etc. may include one or more single-core or multi-core processors. In some embodiments, the processor 1501 can be used to execute instructions to implement the above-mentioned RNA structure prediction method.

[0134] The system memory 1502 is a volatile memory, such as a random-access memory (RAM), a double data rate synchronous dynamic random access memory (DDR SDRAM), etc. The system memory is used to temporarily store data and / or instructions. For example, in some embodiments, the system memory 1502 can be used to store instructions, and can also be used to store original data objects and changed data objects.

[0135] The non-volatile memory 1503 may include one or more tangible, non-transitory computer-readable media for storing data and / or instructions. In some embodiments, the non-volatile memory 1503 may include any suitable non-volatile memory such as a flash memory and / or any suitable non-volatile storage device, such as a hard disk drive (HDD), a compact disc (CD), a digital versatile disc (DVD), a solid-state drive (SSD), etc. In some embodiments, the non-volatile memory 1503 may also be a removable storage medium, such as a secure digital (SD) memory card, etc. In other embodiments, the non-volatile memory 1503 may be used to store instructions, and may also be used to store original data objects and changed data objects.

[0136] In some embodiments, the system memory 1502 and the non-volatile memory 1503 may respectively include: a temporary copy and a permanent copy of the instruction 1507. The instruction 1507 may include: when executed by at least one of the processors 1501, the electronic device 1500 implements the RNA structure prediction method provided in various embodiments of the present application.

[0137] The communication interface 1504 may include a transceiver for providing a wired or wireless communication interface for the electronic device 1500, so as to communicate with any other suitable device through one or more networks. In some embodiments, the communication interface 1504 may be integrated into other components of the electronic device 1500, for example, the communication interface 1504 may be integrated into the processor 1501. In some embodiments, the electronic device 1500 may communicate with other devices through the communication interface 1504, for example, the electronic device 1500 may establish a communication connection with other devices through the communication interface 1504, so as to send data change requests to other devices, obtain original data objects, and send changed data objects through the communication connection.

[0138] Input / output (I / O) device 1505 may include input devices such as a keyboard, a mouse, etc., and output devices such as a display, etc. A user may interact with electronic device 1500 via input / output (I / O) device 1505. For example, business personnel may input / select content for data change via input / output (I / O) device 1505.

[0139] The system control logic 1506 may include any suitable interface controller to provide any suitable interface with other modules of the electronic device 1500. For example, in some embodiments, the system control logic 1506 may include one or more memory controllers to provide interfaces to the system memory 1502 and the non-volatile memory 1503.

[0140] In some embodiments, at least one of the processors 1501 may be packaged together with the logic of one or more controllers for the system control logic 1506 to form a system in package (SiP). In other embodiments, at least one of the processors 1501 may also be integrated with the logic of one or more controllers for the system control logic 1506 on the same chip to form a system-on-chip (SoC).

[0141] Understandably, Figure 6The structure of the electronic device 1500 shown is only an example. In other embodiments, the electronic device 1500 may include more or fewer components than shown, or combine some components, or separate some components, or arrange the components differently. The components shown in the figure may be implemented in hardware, software, or a combination of software and hardware.

[0142] The various embodiments of the mechanism disclosed in the present application can be implemented in hardware, software, firmware or a combination of these implementation methods. The embodiments of the present application can be implemented as a computer module or module code executed on a programmable system, and the programmable system includes at least one processor, a storage system (including volatile and non-volatile memory and / or storage elements), at least one input device and at least one output device.

[0143] The module code may be applied to input instructions to perform the functions described herein and generate output information. The output information may be applied to one or more output devices in a known manner. For purposes of this application, a processing system includes any system having a processor such as, for example, a digital signal processor (DSP), a microcontroller, an application specific integrated circuit (ASIC), or a microprocessor.

[0144] Module code can be implemented with high-level modular language or object-oriented programming language to communicate with the processing system. When necessary, module code can also be implemented with assembly language or machine language. In fact, the mechanism described in this application is not limited to the scope of any specific programming language. In either case, the language can be a compiled language or an interpreted language.

[0145] In some cases, the disclosed embodiments may be implemented in hardware, firmware, software, or any combination thereof. The disclosed embodiments may also be implemented as instructions carried or stored on one or more temporary or non-temporary machine-readable (e.g., computer-readable) storage media, which may be read and executed by one or more processors. For example, instructions may be distributed over a network or through other computer-readable media. Therefore, machine-readable media may include any mechanism for storing or transmitting information in a machine (e.g., computer) readable form, including, but not limited to, floppy disks, optical disks, optical disks, read-only memories (CD-ROMs), magneto-optical disks, read-only memories (ROMs), random access memories (RAMs), erasable programmable read-only memories (EPROMs), electrically erasable programmable read-only memories (EEPROMs), magnetic or optical cards, flash memory, or tangible machine-readable memories for transmitting information (e.g., carrier waves, infrared signals, digital signals, etc.) using the Internet in electrical, optical, acoustic, or other forms of propagation signals. Accordingly, machine-readable media include any type of machine-readable media suitable for storing or transmitting electronic instructions or information in a form readable by a machine (eg, a computer).

[0146] In the accompanying drawings, some structural or method features may be shown in a specific arrangement and / or order. However, it should be understood that such a specific arrangement and / or order may not be required. Instead, in some embodiments, these features may be arranged in a manner and / or order different from that shown in the illustrative drawings. In addition, the inclusion of structural or method features in a particular figure does not mean that such features are required in all embodiments, and in some embodiments, these features may not be included or may be combined with other features.

[0147] It should be noted that the units / modules mentioned in the various device embodiments of the present application are all logical units / modules. Physically, a logical unit / module can be a physical unit / module, or a part of a physical unit / module, or can be implemented as a combination of multiple physical units / modules. The physical implementation method of these logical units / modules themselves is not the most important. The combination of functions implemented by these logical units / modules is the key to solving the technical problems proposed by the present application. In addition, in order to highlight the innovative part of the present application, the above-mentioned device embodiments of the present application do not introduce units / modules that are not closely related to solving the technical problems proposed by the present application, which does not mean that there are no other units / modules in the above-mentioned device embodiments.

[0148] The various embodiments of the mechanism disclosed in the present application can be implemented in hardware, software, firmware or a combination of these implementation methods. The embodiments of the present application can be implemented as a computer program or program code executed on a programmable system, which includes at least one processor, a storage system (including volatile and non-volatile memory and / or storage elements), at least one input device and at least one output device.

[0149] Program code can be applied to input instructions to perform the functions described in this application and generate output information. The output information can be applied to one or more output devices in a known manner. For the purposes of this application, a processing system includes any system having a processor such as, for example, a digital signal processor (DSP), a microcontroller, an application specific integrated circuit (ASIC), or a microprocessor.

[0150] Program code can be implemented with high-level programming language or object-oriented programming language to communicate with the processing system. When necessary, program code can also be implemented with assembly language or machine language. In fact, the mechanism described in this application is not limited to the scope of any specific programming language. In either case, the language can be a compiled language or an interpreted language.

[0151] It should be noted that the units / modules mentioned in the various device embodiments of the present application are all logical units / modules. Physically, a logical unit / module can be a physical unit / module, or a part of a physical unit / module, or can be implemented as a combination of multiple physical units / modules. The physical implementation method of these logical units / modules themselves is not the most important. The combination of functions implemented by these logical units / modules is the key to solving the technical problems proposed by the present application. In addition, in order to highlight the innovative part of the present application, the above-mentioned device embodiments of the present application do not introduce units / modules that are not closely related to solving the technical problems proposed by the present application, which does not mean that there are no other units / modules in the above-mentioned device embodiments.

[0152] It should be noted that, in the examples and description of the present application, relational terms such as first and second, etc. are only used to distinguish one entity or operation from another entity or operation, and do not necessarily require or imply any such actual relationship or order between these entities or operations. Moreover, the terms "include", "comprise" or any other variants thereof are intended to cover non-exclusive inclusion, so that the process, method, article or equipment including a series of elements includes not only those elements, but also includes other elements not explicitly listed, or also includes elements inherent to such process, method, article or equipment. In the absence of further restrictions, the elements defined by the statement "comprise one" do not exclude the presence of other identical elements in the process, method, article or equipment including the elements.

[0153] References to "some embodiments" or "embodiments" in the specification mean that the specific features, structures, or characteristics described in conjunction with the embodiments are included in at least one exemplary implementation or technology disclosed according to the embodiments of the present application. The appearance of the phrase "in some embodiments" in various places in the specification does not necessarily all refer to the same embodiment.

[0154] In addition, the language used in this specification has been primarily selected for readability and instructional purposes and may not be selected to describe or limit the disclosed subject matter. Therefore, the present application embodiment disclosure is intended to illustrate rather than limit the scope of the concepts discussed herein.

Claims

1. A method for predicting RNA structure, characterized in that: Applied to electronic equipment, the method comprises: Obtaining a base sequence to be tested; The base sequence to be tested is input into the RNA language model to obtain the RNA secondary structure, The RNA language model is used to extract the characteristic information of the base sequence to be tested, and to determine the RNA secondary structure corresponding to the characteristic information according to the correlation between the characteristic information and the RNA secondary structure.

2. The method according to claim 1, characterized in that The correlation between the characteristic information and the RNA secondary structure includes the base pairing rule and sequence conservation of the RNA base sequence.

3. The method according to claim 1, characterized in that Determining the RNA secondary structure corresponding to the feature information according to the association relationship between the feature information and the RNA secondary structure includes: Determining the pairing probability of each base in the base sequence to be tested according to the correlation between the characteristic information and the RNA secondary structure; The RNA secondary structure corresponding to the characteristic information is determined according to a probability sequence formed by the pairing probability of each base.

4. The method according to claim 1, characterized in that: The step of inputting the base sequence to be tested into the RNA language model comprises: During the input of the base sequence to be tested, a masking operation is randomly performed on a preset proportion of bases.

5. The method according to claim 1, characterized in that The training method of the RNA language model includes: Get training data; Preprocessing the acquired training data to obtain preprocessed training data; The preprocessed training data is used to train the RNA language model to be trained to obtain a trained RNA language model.

6. The method according to claim 5, characterized in that The preprocessing includes removing irrelevant characters and / or filling in missing bases.

7. The method according to claim 5, characterized in that The preprocessing of the acquired training data further includes: The training data is clustered to remove redundant base sequences in the training data.

8. The method according to claim 5, characterized in that The training method of the RNA language model also includes: A preset loss function is applied to implement the training process of the RNA language model, wherein the preset loss function includes a cross entropy loss function.

9. The method according to claim 5, characterized in that The RNA language model includes a feature extraction network and a secondary structure prediction network, wherein: The feature extraction network is used to extract feature information of the base sequence to be tested; The secondary structure prediction network is used to determine the RNA secondary structure corresponding to the feature information according to the association relationship between the feature information and the RNA secondary structure.

10. The method according to claim 9, characterized in that The step of extracting characteristic information of the base sequence to be tested includes: The feature extraction network encodes the base sequence to be tested into a one-hot vector based on relative position encoding, and extracts the feature information according to the one-hot vector.

11. The method according to claim 9, characterized in that The step of inputting the preprocessed training data into the RNA language model to obtain a trained RNA language model comprises: Inputting the preprocessed training data into the feature extraction network to extract feature data of each base; Inputting the characteristic data of each base into the secondary structure prediction network to obtain the pairing probability of each base; Determining the predicted RNA secondary structure according to the pairing probability; The RNA language model is updated according to the difference between the predicted RNA secondary structure and the RNA secondary structure in the training data.

12. The method according to claim 11, characterized in that The pairing probability of each base is between 0 and 1.

13. The method according to claim 9, characterized in that The secondary structure prediction network includes a feature extractor based on a Transformer architecture, wherein the feature extractor is capable of fusing 1-dimensional sequence information and 2-dimensional pairing information.

14. The method according to claim 13, characterized in that The characteristic information of the base sequence to be tested includes a 1-dimensional sequence information attention map, and The fusion of 1-dimensional sequence information and 2-dimensional pairing information includes: Performing dimensionality reduction on the 1-dimensional sequence information attention map to obtain a 1-dimensional sequence information feature vector; Performing an outer product on the 1-dimensional sequence information feature vector to obtain a 2-dimensional pairing information feature vector; The 1-dimensional sequence information feature vector is added to the 2-dimensional pairing information feature vector through an outer product operation to introduce the 2-dimensional pairing information; The 2D pairing information feature vector is used as a bias term and added to the 1D sequence information attention map to introduce the 1D sequence information.

15. An electronic device, characterized in that: include: one or more processors; One or more memories; the one or more memories store one or more programs, and when the one or more programs are executed by the one or more processors, the electronic device executes the RNA structure prediction method according to any one of claims 1 to 14.

16. A computer readable medium, characterized in that The computer-readable medium stores instructions, which, when executed on a machine, enable the machine to perform the RNA structure prediction method according to any one of claims 1 to 14.

17. A computer program product, characterized in that The method comprises a computer program / instruction, which, when executed by a processor, implements the RNA structure prediction method according to any one of claims 1 to 14.