Method and system for predicting secondary structure of RNA

The method and system for predicting RNA secondary structure address the challenge of identifying complex RNA folding, enabling accurate target sequence identification for RNA therapeutics and improving disease treatment options.

WO2025116216A1PCT designated stage expired Publication Date: 2025-06-05SPIDERCORE INC
View PDF 3 Cites 0 Cited by

Patent Information

Application Number
PCT/KR2024/012960
Authority / Receiving Office
WO · WO
Patent Type
Applications
Current Assignee / Owner
Priority Date
2023-11-28
Filing Date
2024-08-29
Publication Date
2025-06-05

AI Technical Summary

Technical Problem

Current methods for developing RNA therapeutics for diseases such as genetic, cardiovascular, and cancer require accurate identification of the secondary structure of specific mRNAs, which is challenging due to the complex folding of RNA molecules.

Method used

A method and system utilizing a learning data generation unit, a secondary structure prediction unit, and a post-processing unit to predict the secondary structure of RNA. This involves generating learning data from a secondary structure database, training an artificial neural network to predict hydrogen bonding relationships, and refining the predictions through post-processing to generate a secondary structure prediction matrix.

Benefits of technology

The system effectively predicts the secondary structure of RNA, enabling more accurate identification of target mRNA sequences for RNA therapeutics, which is crucial for developing effective treatments for various diseases.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure KR2024012960_05062025_PF_FP_ABST
    Figure KR2024012960_05062025_PF_FP_ABST
Patent Text Reader

Abstract

In the RNA secondary structure prediction method, the base sequences of multiple RNAs are each used as input data; a symmetric matrix representing the intrinsic hydrogen bonding relationships of RNA is used as labeled training data to generate a secondary structure prediction model; afterward, the base sequence of a target RNA is input into the secondary structure prediction model, generating the symmetric matrix output from the secondary structure prediction model as an initial prediction matrix; and in the initial prediction matrix, base pairs that can form hydrogen bonds and whose corresponding element values are greater than or equal to a threshold value among the base pairs included in the base sequence of the target RNA are set to 1, while all other elements are set to 0, thereby generating the final secondary structure prediction matrix.
Need to check novelty before this filing date? Find Prior Art

Description

Method and system for predicting the secondary structure of RNA

[0001] The present invention relates to prediction of the structure of RNA (Ribonucleic Acid), and more particularly, to a method and system for predicting the secondary structure of RNA.

[0002] Deoxyribonucleic acid (DNA), which carries genetic information, is a double-helix polymer composed of nucleotide polymers. The genetic information contained in DNA is expressed as specific proteins through a complex process.

[0003] The base sequence contained in DNA contains information on the amino acids that make up proteins. For a specific protein to be expressed, a process called transcription takes place, copying the specific base sequence of DNA required for protein expression into messenger ribonucleic acid (mRNA). The copied mRNA then binds to a ribosome, necessary for protein synthesis, and the amino acids matching the mRNA base sequence are synthesized, forming a protein (translation). The protein produced through this process is transported to an appropriate location within the cell or outside the cell, ultimately completing gene expression.

[0004] At this time, reducing the activity of mRNA copied from DNA can interfere with gene expression. Drugs with a new mechanism of action, utilizing RNA, which carries genetic information during the cellular protein formation process, are called RNA therapeutics.

[0005] These RNA therapeutics are emerging as a new solution for incurable diseases for which there are no existing treatments, and research on RNA therapeutics that can treat incurable diseases such as genetic diseases, cardiovascular diseases, and cancer is actively underway.

[0006] In general, in order to develop an RNA therapeutic agent for a specific disease, it is necessary to identify the mRNA associated with the synthesis of the protein causing the specific disease, and then find an RNA base sequence that can effectively reduce the activity of the specific mRNA.

[0007] However, RNA bases do not generally exist as a single long strand, but rather have a secondary structure such as a hairpin structure by forming hydrogen bonds between the bases.

[0008] Therefore, in order to find an RNA base sequence that can effectively reduce the activity of the specified mRNA, a process of accurately identifying the secondary structure of the specified mRNA is necessary.

[0009] Accordingly, the technical problem of the present invention was conceived from this point of view, and one object of the present invention is to provide a method and system capable of effectively predicting the secondary structure of RNA.

[0010] In order to achieve the above-described object of the present invention, in a method for predicting a secondary structure of RNA according to an embodiment of the present invention, a learning data generation unit uses a secondary structure database that stores in advance the folded secondary structures of a plurality of RNAs included in a human body, and generates learning data using a symmetric matrix as a label for the input data, wherein the learning data has the same number of rows and columns as the length of the base sequence of the RNA corresponding to the input data, and when the p-th base and the q-th base in the base sequence of the RNA are hydrogen bonded to each other, the element corresponding to the p-th row and the q-th column and the element corresponding to the q-th row and the p-th column have a value of 1, and when the r-th base and the s-th base are not hydrogen bonded to each other, the element corresponding to the r-th row and the s-th column and the element corresponding to the s-th row and the r-th column have a value of 0, and when the secondary structure prediction unit inputs the base sequence of the RNA into an artificial neural network using the learning data, the artificial neural network generates its own hydrogen bonds between the bases included in the base sequence of the input RNA. The artificial neural network is trained to output a symmetric matrix representing a relationship to generate a secondary structure prediction model, and the secondary structure prediction unit inputs a base sequence of a target RNA received from the outside into the secondary structure prediction model to generate the symmetric matrix output from the secondary structure prediction model as an initial prediction matrix, and a post-processing unit analyzes the base sequence of the target RNA and, in the initial prediction matrix, changes at least some of the elements at positions corresponding to base pairs capable of hydrogen bonding with each other among base pairs included in the base sequence of the target RNA and at positions where the element value of the corresponding position is greater than or equal to a predetermined threshold value to 1, and changes the remaining elements of the initial prediction matrix to 0 to generate a secondary structure prediction matrix.

[0011] In order to achieve the above-described object of the present invention, a system for predicting a secondary structure of RNA according to an embodiment of the present invention includes a learning data generation unit, a secondary structure prediction unit, and a post-processing unit. The learning data generation unit uses a secondary structure database that stores folded secondary structures of a plurality of RNAs included in a human body in advance, and generates learning data using a symmetric matrix as a label for the input data, wherein the base sequence of each of the plurality of RNAs is input data, and the matrix has the same number of rows and columns as the length of the base sequence of the RNA corresponding to the input data, and when the p-th base and the q-th base in the base sequence of the RNA are hydrogen bonded to each other, the element corresponding to the p-th row and the q-th column and the element corresponding to the q-th row and the p-th column have a value of 1, and when the r-th base and the s-th base are not hydrogen bonded to each other, the element corresponding to the r-th row and the s-th column and the element corresponding to the s-th row and the r-th column have a value of 0. The secondary structure prediction unit trains the artificial neural network to output a symmetric matrix representing its own hydrogen bonding relationship between bases included in the base sequence of the input RNA when inputting the base sequence of RNA into the artificial neural network using the training data, thereby generating a secondary structure prediction model, and then inputs the base sequence of a target RNA received from the outside into the secondary structure prediction model to generate the symmetric matrix output from the secondary structure prediction model as an initial prediction matrix. The post-processing unit analyzes the base sequence of the target RNA, and changes at least some of the elements at positions corresponding to base pairs capable of hydrogen bonding with each other among base pairs included in the base sequence of the target RNA in the initial prediction matrix to 1, and changes the remaining elements of the initial prediction matrix to 0, thereby generating a secondary structure prediction matrix.

[0012] The RNA secondary structure prediction system and RNA secondary structure prediction method according to embodiments of the present invention generate an initial prediction matrix through a machine learning model that predicts the secondary structure of a target RNA in response to a base sequence of the target RNA, and then performs post-processing on the initial prediction matrix to generate a secondary structure prediction matrix, thereby enabling more effective and accurate prediction of the secondary structure of the target RNA.

[0013] FIG. 1 is a diagram showing a system for predicting the secondary structure of RNA according to one embodiment of the present invention.

[0014] Figure 2 is a flowchart showing a method for predicting the secondary structure of RNA according to one embodiment of the present invention.

[0015] FIG. 3 is a diagram for explaining an example of the operation of a learning data generation unit included in the secondary structure prediction system of RNA of FIG. 1.

[0016] Hereinafter, preferred embodiments of the present invention will be described in more detail with reference to the attached drawings. Identical components in the drawings are designated by the same reference numerals, and redundant descriptions of identical components are omitted.

[0017] FIG. 1 is a diagram showing a system for predicting the secondary structure of RNA according to one embodiment of the present invention.

[0018] Typically, most RNAs exist folded into complex structures. For example, RNA strands fold themselves by forming hydrogen bonds between base pairs.

[0019] The RNA secondary structure prediction system (10) illustrated in FIG. 1 represents a system that, when receiving a base sequence (R_SEQ) of a target RNA (Ribonucleic Acid), predicts a secondary structure in which the target RNA is folded and generates a secondary structure prediction matrix (SSG_MAT) representing the predicted secondary structure of the target RNA.

[0020] Figure 2 is a flowchart showing a method for predicting the secondary structure of RNA according to one embodiment of the present invention.

[0021] The method for predicting the secondary structure of RNA illustrated in FIG. 2 can be performed through the RNA secondary structure prediction system (10) of FIG. 1.

[0022] Hereinafter, the configuration and operation of the RNA secondary structure prediction system (10) and the RNA secondary structure prediction method performed by the RNA secondary structure prediction system (10) will be described in detail with reference to FIGS. 1 and 2.

[0023] Referring to FIG. 1, the RNA secondary structure prediction system (10) includes a training data generator (100), a secondary structure predictor (200), and a post-processor (300).

[0024] The learning data generation unit (100) generates learning data (L_DATA) for generating a machine learning model that predicts the secondary structure of RNA using a secondary structure database (SS_DB) (110) that stores in advance the folded secondary structures of multiple RNAs included in the human body (step S100).

[0025] In one embodiment, the secondary structure database (110) may be an ArchiveII database, an RNAStrAlign database, or the like.

[0026] Although FIG. 1 illustrates that the secondary structure database (110) is included within the RNA secondary structure prediction system (10), depending on the embodiment, the secondary structure database (110) may exist outside the RNA secondary structure prediction system (10).

[0027] In one embodiment, the learning data generation unit (100) may generate learning data (L_DATA) using the base sequences of each of the plurality of RNAs stored in the secondary structure database (110) as input data, and using a square matrix having the same number of rows and columns as the length of the base sequence of the RNA corresponding to the input data and representing the hydrogen bonding relationship between bases of the RNA corresponding to the input data as a label for the input data.

[0028] FIG. 3 is a diagram for explaining an example of the operation of a learning data generation unit included in the secondary structure prediction system of RNA of FIG. 1.

[0029] In Fig. 3, RNA corresponding to the input data is illustrated as having a base sequence corresponding to GGGAAACGUUCCG, as an example.

[0030] As illustrated in Fig. 3, RNA corresponding to the input data can exist in a folded state by forming hydrogen bonds between complementary base pairs.

[0031] In the case of RNA corresponding to the input data illustrated in FIG. 3, the second base, Guanine, and the 12th base, Cytosine, are shown to form hydrogen bonds with each other, the third base, Guanine, and the 11th base, Cytosine, are shown to form hydrogen bonds with each other, and the fourth base, Adenine, and the 10th base, Uracil, are shown to form hydrogen bonds with each other.

[0032] In this case, the learning data generation unit (100) can determine a square matrix (S_MAT) that has the same number of rows and columns as the length of the base sequence of the RNA corresponding to the input data and represents the hydrogen bonding relationship between the bases of the RNA corresponding to the input data formed by the secondary structure in which the RNA corresponding to the input data is folded, as a label for the input data.

[0033] Specifically, as illustrated in FIG. 3, the learning data generation unit (100) can generate learning data (L_DATA) using a symmetric matrix (S_MAT) as a label for the input data, the base sequence of each of the plurality of RNAs stored in the secondary structure database (110) as input data, and having the same number of rows and columns as the length of the base sequence of the RNA corresponding to the input data, and in which the p-th base and the q-th base in the base sequence of the RNA corresponding to the input data are hydrogen bonded to each other, the element corresponding to the p-th row and the q-th column and the element corresponding to the q-th row and the p-th column have a value of 1, and in which the r-th base and the s-th base are not hydrogen bonded to each other, the element corresponding to the r-th row and the s-th column and the element corresponding to the s-th row and the r-th column have a value of 0. Here, p, q, r, and s represent natural numbers.

[0034] Therefore, the symmetric matrix (S_MAT) corresponding to the label for the input data generated by the learning data generation unit (100) can represent the folded secondary structure of RNA corresponding to the input data.

[0035] In this way, the learning data generation unit (100) can generate learning data (L_DATA) using the base sequence of each of the plurality of RNAs stored in the secondary structure database (110) as the input data and the symmetric matrix (S_MAT) representing the folded secondary structure of the RNA corresponding to the input data as the label for the input data.

[0036] Referring again to FIG. 1, the learning data (L_DATA) generated by the learning data generation unit (100) can be provided to the secondary structure prediction unit (200).

[0037] The secondary structure prediction unit (200) trains the artificial neural network to output a symmetric matrix (S_MAT) representing its own hydrogen bonding relationship between bases included in the base sequence of the input RNA when the base sequence of RNA is input to the artificial neural network using the learning data (L_DATA) provided from the learning data generation unit (100), thereby generating a secondary structure prediction model (step S200).

[0038] In one embodiment, the artificial neural network includes one input layer into which a base sequence of RNA is input, one or more hidden layers, and one output layer that outputs a symmetric matrix (S_MAT) representing hydrogen bonding relationships between bases in the base sequence of the input RNA, and the secondary structure prediction unit (200) can generate the secondary structure prediction model by performing supervised learning on the artificial neural network using learning data (L_DATA).

[0039] Specifically, the secondary structure prediction unit 1 (200) can generate the secondary structure prediction model by performing supervised learning on the artificial neural network so that, for each base pair included in the base sequence of RNA input to the artificial neural network, if a specific base pair is predicted to hydrogen bond with each other, the elements at two symmetric positions corresponding to the specific base pair in the symmetric matrix (S_MAT) are classified as 1, and if the specific base pair is predicted not to hydrogen bond with each other, the elements at two symmetric positions corresponding to the specific base pair in the symmetric matrix (S_MAT) are classified as 0.

[0040] Therefore, each element of the symmetric matrix (S_MAT) output from the secondary structure prediction model in response to the base sequence of RNA input to the secondary structure prediction model may have a real number value greater than or equal to 0 and less than or equal to 1.

[0041] After generating the secondary structure prediction model, the secondary structure prediction unit (200) can receive the base sequence (R_SEQ) of the target RNA for which the secondary structure is to be predicted from the outside.

[0042] In this case, the secondary structure prediction unit (200) can input the base sequence (R_SEQ) of the target RNA into the secondary structure prediction model and generate a symmetric matrix (S_MAT) output from the secondary structure prediction model as an initial prediction matrix (IG_MAT) (step S300).

[0043] Therefore, each element of the initial prediction matrix (IG_MAT) output from the above secondary structure prediction model can have a real value greater than or equal to 0 and less than or equal to 1.

[0044] The initial prediction matrix (IG_MAT) generated from the secondary structure prediction unit (200) can be provided to the post-processing unit (300).

[0045] The post-processing unit (300) analyzes the base sequence (R_SEQ) of the target RNA, and changes at least some of the elements at positions corresponding to base pairs capable of hydrogen bonding with each other among the base pairs included in the base sequence (R_SEQ) of the target RNA in the initial prediction matrix (IG_MAT) to 1, and changes the remaining elements of the initial prediction matrix (IG_MAT) to 0 to generate a secondary structure prediction matrix (SSG_MAT) (step S400).

[0046] As described above with reference to FIG. 3, for general RNA, the square matrix (S_MAT) representing the folded secondary structure of the RNA satisfies the following five conditions.

[0047] (1) In the base sequence of RNA, when a specific base pair is hydrogen bonded to each other, the elements at the two symmetric positions corresponding to the specific base pair have a value of 1, and when the specific base pair is not hydrogen bonded to each other, the elements at the two symmetric positions corresponding to the specific base pair have a value of 0, so all elements of the square matrix (S_MAT) have a value of 0 or 1.

[0048] (2) Since the elements at two symmetric positions in a square matrix (S_MAT) have the same values, the square matrix (S_MAT) is a symmetric matrix.

[0049] (3) In the base sequence of RNA, only the Watson-Crick base pairs corresponding to the base pairs of adenine and uracil and guanine and cytosine, and the wobble base pairs corresponding to the base pairs of guanine and uracil can form hydrogen bonds with each other, while the base pairs of adenine and cytosine cannot form hydrogen bonds with each other. Therefore, in the square matrix (S_MAT), the elements at the two symmetric positions corresponding to the base pairs of adenine and cytosine always have a value of 0.

[0050] (4) General RNA exists in a folded state by forming hydrogen bonds between base pairs on their own, but since it cannot form a rapidly folded loop, hydrogen bonds cannot be formed between bases that are too close together. Specifically, in a general RNA base sequence, hydrogen bonds cannot be formed between base pairs that are located within three bases of each other because the distance between the atoms participating in the hydrogen bond is too far. Therefore, in the base sequence of RNA, even if a specific base pair is a Watson-Crick base pair or a wobble base pair, if the specific base pair is located within three bases of each other, the elements at the two symmetrical positions corresponding to the specific base pair in the square matrix (S_MAT) always have a value of 0. That is, when (|uv|<=3), the element corresponding to the u-th row and the v-th column in the square matrix (S_MAT) and the element corresponding to the v-th row and the u-th column always have a value of 0. Here, u and v represent natural numbers.

[0051] (5) In a typical RNA base sequence, no base can form hydrogen bonds with two or more other bases simultaneously. That is, if a specific base forms a hydrogen bond with another base, that specific base cannot form a hydrogen bond with another base. Therefore, in the RNA base sequence, each base either does not form a hydrogen bond or forms a hydrogen bond with only one other base. Therefore, in a square matrix (S_MAT), the sum of the elements in each row is 0 or 1, and the sum of the elements in each column is 0 or 1.

[0052] Accordingly, when the post-processing unit (300) receives the initial prediction matrix (IG_MAT) from the secondary structure prediction unit (200), the post-processing unit (300) can change the element values ​​of the initial prediction matrix (IG_MAT) so that the initial prediction matrix (IG_MAT) satisfies the above five conditions based on the base sequence (R_SEQ) of the target RNA to generate the secondary structure prediction matrix (SSG_MAT).

[0053] As described below, the post-processing unit (300) can generate a secondary structure prediction matrix (SSG_MAT) by changing the element values ​​of the initial prediction matrix (IG_MAT) so that the initial prediction matrix (IG_MAT) satisfies the above five conditions by utilizing the concept of an optimization algorithm such as the Hungarian Algorithm.

[0054] Below, the process of generating a secondary structure prediction matrix (SSG_MAT) by the post-processing unit (300) is described in detail.

[0055] The post-processing unit (300) can sequentially assign an index number to each base located in the 5'-end to 3'-end direction of the base sequence (R_SEQ) of the target RNA.

[0056] For example, if the length of the base sequence (R_SEQ) of the target RNA is L, the post-processing unit (300) can sequentially assign a natural number from 1 to L as the index number to each base located in the direction from the 5'-end to the 3'-end.

[0057] Thereafter, the post-processing unit (300) can determine a first index number set having as elements the index numbers indicating the order of adenine and guanine in the base sequence (R_SEQ) of the target RNA.

[0058] Additionally, the post-processing unit (300) can determine a first index number set having as elements the index numbers indicating the order of Uracil and Cytosine in the base sequence (R_SEQ) of the target RNA.

[0059] For example, if the base sequence (R_SEQ) of the target RNA is GGGAAACGUUCCG as shown in FIG. 3, the post-processing unit (300) can determine the first index number set as {1, 2, 3, 4, 5, 6, 8, 13} and the second index number set as {7, 9, 10, 11, 12}.

[0060] Thereafter, the post-processing unit (300) can determine a plurality of ordered pair sets (J) that have ordered pairs composed of elements of the first index number set and elements of the second index number set as elements and satisfy the conditions below.

[0061]

[0062]

[0063] Here, IN1 represents the first index number set, IN2 represents the second index number set, J represents the set of ordered pairs, a, i1, and i2 represent elements of the first index number set, and b, j1, and j2 represent elements of the second index number set.

[0064] Accordingly, the plurality of ordered pair sets (J) may correspond to all ordered pair sets that can be generated by using each of the elements of the first index number set and each of the elements of the second index number set at most once, while having as elements an ordered pair composed of an element of the first index number set and an element of the second index number set.

[0065] For example, if the first index number set is {1, 2, 3, 4, 5, 6, 8, 13} and the second index number set is {7, 9, 10, 11, 12}, the post-processing unit (300) can generate (8C1*5C1) ordered pair sets (J) each having 1 element, (8C2*5C2* 2!) ordered pair sets (J) each having 2 elements, (8C3*5C3* 3!) ordered pair sets (J) each having 3 elements, (8C4*5C4* 4!) ordered pair sets (J) each having 4 elements, and (8C5*5C5* 5!) ordered pair sets (J) each having 5 elements.

[0066] Thereafter, the post-processing unit (300) can determine a gain for each of the ordered pairs included in the ordered pair set (J) based on the base sequence (R_SEQ) of the target RNA and the element value at the position corresponding to each of the ordered pairs included in the ordered pair set (J) in the initial prediction matrix (IG_MAT), for each of the plurality of ordered pair sets (J).

[0067] In one embodiment, when the sequence pair (i, j) is included in the sequence pair set (J), the post-processing unit (300) may determine the gain for the sequence pair (i, j) as 0 when the i-th base in the base sequence (R_SEQ) of the target RNA is adenine and the j-th base is cytosine, when the difference between i and j is 3 or less, and when the element value at the position corresponding to the i-th row and the j-th column in the initial prediction matrix (IG_MAT) is smaller than the threshold, and for all other cases, the gain for the sequence pair (i, j) may be determined as the element value at the position corresponding to the i-th row and the j-th column in the initial prediction matrix (IG_MAT).

[0068] For example, in order to consider that base pairs corresponding to positions of specific element values ​​are likely to form hydrogen bonds with each other only when the specific element value of the initial prediction matrix (IG_MAT) generated from the secondary structure prediction model is at least greater than or equal to 0.5, the threshold value may have a value greater than or equal to 0.5 and less than 1.

[0069] Accordingly, the post-processing unit (300) determines the gain for the ordered pair (i, j) as 0 when the base pair corresponding to the ordered pair (i, j) is neither a Watson-Crick base pair nor a Wobble base pair, when the base pair corresponding to the ordered pair (i, j) is located at a distance of three bases or less, and when the element value at the position corresponding to the i-th row and j-th column in the initial prediction matrix (IG_MAT) is smaller than the threshold value and thus the base pair corresponding to the ordered pair (i, j) is judged to have a low possibility of forming a hydrogen bond with each other, and for all other cases, the element value at the position corresponding to the i-th row and j-th column in the initial prediction matrix (IG_MAT) can be determined as the gain for the ordered pair (i, j).

[0070] In another embodiment, when the sequence pair (i, j) is included in the sequence pair set (J), the post-processing unit (300) determines the gain for the sequence pair (i, j) as 0 when the i-th base in the base sequence (R_SEQ) of the target RNA is adenine and the j-th base is cytosine, when the difference between i and j is 3 or less, and when the element value at the position corresponding to the i-th row and the j-th column in the initial prediction matrix (IG_MAT) is smaller than the threshold value, and in other cases when the i-th base in the base sequence (R_SEQ) of the target RNA is guanine and the j-th base is uracil, the gain for the sequence pair (i, j) is determined as a value obtained by multiplying the element value at the position corresponding to the i-th row and the j-th column in the initial prediction matrix (IG_MAT) by a predetermined decay rate, and in all other cases, the sequence pair (i, The above gain for j) can be determined as the element value at the position corresponding to the i-th row and j-th column in the initial prediction matrix (IG_MAT).

[0071] For example, the attenuation ratio may have a value greater than 0.5 and less than 1.

[0072] Accordingly, the post-processing unit (300) determines the gain for the ordered pair (i, j) as 0 if the base pair corresponding to the ordered pair (i, j) is not a Watson-Crick base pair and a Wobble base pair, if the base pair corresponding to the ordered pair (i, j) is located at a distance of three bases or less, and if the element value at the position corresponding to the i-th row and the j-th column in the initial prediction matrix (IG_MAT) is smaller than the threshold value and thus the base pair corresponding to the ordered pair (i, j) is unlikely to form a hydrogen bond with each other, and in addition, if the base pair corresponding to the ordered pair (i, j) is a Wobble base pair, since the bonding strength is weaker than in the case of a Watson-Crick base pair, the value obtained by multiplying the element value at the position corresponding to the i-th row and the j-th column in the initial prediction matrix (IG_MAT) by the decay rate having a value less than 1 is determined as the gain for the ordered pair The above gain for (i, j) is determined, and for all other cases, the element value at the position corresponding to the i-th row and j-th column in the initial prediction matrix (IG_MAT) can be determined as the above gain for the ordered pair (i, j).

[0073] Meanwhile, both guanine-cytosine base pairs and adenine-uracil base pairs are Watson-Crick base pairs, but guanine-cytosine base pairs form three hydrogen bonds with each other, while adenine-uracil base pairs form two hydrogen bonds with each other. Therefore, the binding strength of adenine-uracil base pairs is weaker than that of guanine-cytosine base pairs. In addition, the binding strength of guanine-uracil base pairs, which are wobble base pairs, is weaker than that of Watson-Crick base pairs.

[0074] Accordingly, in another embodiment, when the sequence pair (i, j) is included in the sequence pair set (J), the post-processing unit (300) determines the gain for the sequence pair (i, j) as 0 when the i-th base in the base sequence (R_SEQ) of the target RNA is adenine and the j-th base is cytosine, when the difference between i and j is 3 or less, and when the element value at the position corresponding to the i-th row and the j-th column in the initial prediction matrix (IG_MAT) is less than the threshold, and otherwise, when the i-th base in the base sequence (R_SEQ) of the target RNA is guanine and the j-th base is cytosine, the gain for the sequence pair (i, j) is determined as the element value at the position corresponding to the i-th row and the j-th column in the initial prediction matrix (IG_MAT), and otherwise, when the i-th base in the base sequence (R_SEQ) of the target RNA is In the case where the i-th base is adenine and the j-th base is uracil, the gain for the ordered pair (i, j) may be determined as a value obtained by multiplying the element value at the position corresponding to the i-th row and the j-th column in the initial prediction matrix (IG_MAT) by a predetermined first decay rate. In all other cases, in the case where the i-th base is guanine and the j-th base is uracil in the base sequence (R_SEQ) of the target RNA, the gain for the ordered pair (i, j) may be determined as a value obtained by multiplying the element value at the position corresponding to the i-th row and the j-th column in the initial prediction matrix (IG_MAT) by a predetermined second decay rate having a value smaller than the first decay rate.

[0075] For example, the first attenuation rate and the second attenuation rate may have values ​​greater than 0.5 and less than 1.

[0076] Accordingly, the post-processing unit (300) determines the gain for the ordered pair (i, j) as 0 if the base pair corresponding to the ordered pair (i, j) is not a Watson-Crick base pair and a Wobble base pair, if the base pair corresponding to the ordered pair (i, j) is located at a distance of three bases or less, and if the element value at the position corresponding to the i-th row and j-th column in the initial prediction matrix (IG_MAT) is smaller than the threshold value and thus the base pair corresponding to the ordered pair (i, j) is judged to have a low possibility of forming a hydrogen bond with each other, and in addition, if the base pair corresponding to the ordered pair (i, j) is a Guanine-Cytosine base pair among the Watson-Crick base pairs, the element value at the position corresponding to the i-th row and j-th column in the initial prediction matrix (IG_MAT) is applied as is to the ordered pair (i, j). In the case where the base pair corresponding to the ordered pair (i, j) is an adenine-uracil base pair among Watson-Crick base pairs, the gain for the ordered pair (i, j) is determined by multiplying the element value at the position corresponding to the i-th row and j-th column in the initial prediction matrix (IG_MAT) by the first decay rate having a value less than 1, and in all other cases where the base pair corresponding to the ordered pair (i, j) is a wobble base pair, the gain for the ordered pair (i, j) is determined by multiplying the element value at the position corresponding to the i-th row and j-th column in the initial prediction matrix (IG_MAT) by the second decay rate having a value less than the first decay rate.

[0077] As described above, the post-processing unit (300) determines that, for each of a plurality of ordered pair sets (J), the base pair corresponding to the ordered pair (i, j) included in the ordered pair set (J) in the base sequence (R_SEQ) of the target RNA has no or low possibility of forming a hydrogen bond with each other, and determines the gain for the ordered pair (i, j) as 0.

[0078] Accordingly, the post-processing unit (300) can determine the gain for each of all ordered pairs included in the ordered pair set (J) for each of the plurality of ordered pair sets (J), and then determine the ordered pair sets that include only the ordered pairs whose gain is greater than 0 among the plurality of ordered pair sets (J) as the hydrogen bond candidate ordered pair sets.

[0079] Thereafter, the post-processing unit (300) can calculate the sum of the gains for all the ordered pairs included in each of the hydrogen bond candidate ordered pair sets, and determine the total gain for each of the hydrogen bond candidate ordered pair sets.

[0080] As described above, the gain for the sequence pair (i, j) included in the sequence pair set (J) has a larger value as the probability that the base pairs corresponding to the sequence pair (i, j) form hydrogen bonds with each other increases. In addition, as the number of hydrogen bonds formed between bases in the base sequence (R_SEQ) of the target RNA increases, the target RNA has a more stable structure.

[0081] Therefore, the post-processing unit (300) can determine the hydrogen bond candidate ordered pair set with the maximum total gain among the hydrogen bond candidate ordered pair sets as the final ordered pair set.

[0082] Thereafter, the post-processing unit (300) can output a matrix in which, for each of all ordered pairs (xi, xj) included in the final ordered pair set, the element corresponding to the xi-th row and the xj-th column and the element corresponding to the xj-th row and the xi-th column in the initial prediction matrix (IG_MAT) is changed to 1, and all other remaining elements of the initial prediction matrix (IG_MAT) are changed to 0, as a second-order structure prediction matrix (SSG_MAT).

[0083] Therefore, the secondary structure prediction matrix (SSG_MAT) generated by the post-processing unit (300) in response to the base sequence (R_SEQ) of the target RNA satisfies all of the above five conditions.

[0084] Specifically, since all elements of the secondary structure prediction matrix (SSG_MAT) have values ​​of 0 or 1, the secondary structure prediction matrix (SSG_MAT) satisfies the first condition above.

[0085] The secondary structure prediction matrix (SSG_MAT) is a symmetric matrix because it corresponds to a matrix in which, for each of all ordered pairs (xi, xj) included in the final ordered pair set, the elements corresponding to the xi-th row and the xj-th column and the elements corresponding to the xj-th row and the xi-th column in the initial prediction matrix (IG_MAT) are changed to 1, and all other elements of the initial prediction matrix (IG_MAT) are changed to 0. Therefore, the secondary structure prediction matrix (SSG_MAT) satisfies the second condition.

[0086] When a sequence pair (i, j) is included in a sequence pair set (J), the post-processing unit (300) determines the gain for the sequence pair (i, j) to be 0 when the i-th base in the base sequence (R_SEQ) of the target RNA is adenine and the j-th base is cytosine, and since the final sequence pair set includes only sequence pairs having the gain greater than 0, the elements corresponding to the i-th row and j-th column, which are two symmetric positions corresponding to the base pairs of adenine and cytosine, and the elements corresponding to the j-th row and i-th column always have a value of 0. Therefore, the secondary structure prediction matrix (SSG_MAT) satisfies the third condition.

[0087] When an ordered pair (i, j) is included in an ordered pair set (J), the post-processing unit (300) determines the gain for the ordered pair (i, j) to be 0 if the difference between i and j is 3 or less, and since the final ordered pair set includes only ordered pairs whose gain is greater than 0, in the secondary structure prediction matrix (SSG_MAT), when (|ij|<=3), the element corresponding to the i-th row and j-th column and the element corresponding to the j-th row and i-th column always have a value of 0. Therefore, the secondary structure prediction matrix (SSG_MAT) satisfies the fourth condition.

[0088] The plurality of ordered pair sets (J) correspond to all ordered pair sets that can be generated by using each element of the first index number set and each element of the second index number set at most once, while having as elements an ordered pair composed of an element of the first index number set and an element of the second index number set, and the final ordered pair set is a set selected from the plurality of ordered pair sets (J). Therefore, in the secondary structure prediction matrix (SSG_MAT), the sum of the elements included in each row is 0 or 1, and the sum of the elements included in each column is 0 or 1. Therefore, the secondary structure prediction matrix (SSG_MAT) satisfies the fifth condition.

[0089] In addition, the post-processing unit (300) determines the hydrogen bond candidate ordered pair set with the maximum total gain among the hydrogen bond candidate ordered pair sets as the final ordered pair set, and generates a secondary structure prediction matrix (SSG_MAT) in which the elements at positions corresponding to each of the ordered pairs included in the final ordered pair set have a value of 1 and all other elements have a value of 0, so that the post-processing unit (300) can generate a secondary structure prediction matrix (SSG_MAT) corresponding to the most stable secondary structure while satisfying the five conditions above.

[0090] As described above with reference to FIGS. 1 to 3, the RNA secondary structure prediction system (10) and the RNA secondary structure prediction method according to embodiments of the present invention generate an initial prediction matrix (IG_MAT) through a machine learning model that predicts the secondary structure of the target RNA in response to the base sequence (R_SEQ) of the target RNA, and then performs post-processing on the initial prediction matrix (IG_MAT) to generate a secondary structure prediction matrix (SSG_MAT), so that the secondary structure of the target RNA can be predicted more effectively and accurately.

[0091] The present invention can be usefully used to predict the secondary structure in which RNA is folded.

[0092] As described above, although the present invention has been described with reference to preferred embodiments thereof, it will be understood by those skilled in the art that various modifications and changes may be made to the present invention without departing from the spirit and scope of the present invention as set forth in the claims below.

Claims

1. A step of generating learning data using a secondary structure database in which folded secondary structures of a plurality of RNAs included in a human body are stored in advance, the secondary structure database being used as input data for the base sequence of each of the plurality of RNAs, and having a number of rows and columns equal to the length of the base sequence of the RNA corresponding to the input data, wherein if the p-th base and the q-th base in the base sequence of the RNA are hydrogen bonded to each other, the element corresponding to the p-th row and the q-th column and the element corresponding to the q-th row and the p-th column have a value of 1, and if the r-th base and the s-th base are not hydrogen bonded to each other, the element corresponding to the r-th row and the s-th column and the element corresponding to the s-th row and the r-th column have a value of 0, as a label for the input data; A step of generating a secondary structure prediction model by training an artificial neural network to output a symmetric matrix representing its own hydrogen bonding relationship between bases included in the base sequence of the input RNA when the secondary structure prediction unit inputs the base sequence of RNA into the artificial neural network using the learning data; A step of inputting the base sequence of the target RNA received from the outside into the secondary structure prediction model and generating the symmetric matrix output from the secondary structure prediction model as an initial prediction matrix; and A method for predicting the secondary structure of RNA, comprising the step of a post-processing unit analyzing the base sequence of the target RNA, changing at least some of the elements of positions corresponding to base pairs capable of hydrogen bonding with each other among the base pairs included in the base sequence of the target RNA in the initial prediction matrix and having an element value greater than or equal to a predetermined threshold value to 1, and changing the remaining elements of the initial prediction matrix to 0 to generate a secondary structure prediction matrix.

2. A method for predicting the secondary structure of RNA in the first paragraph, wherein each of the elements of the symmetric matrix output from the secondary structure prediction model has a real number value greater than or equal to 0 and less than or equal to 1.

3. In the first paragraph, the step of generating the second structure prediction matrix by the post-processing unit is as follows: A step of determining a first index number set having as elements index numbers indicating the order of adenine and guanine in the base sequence of the target RNA; A step of determining a second index number set having as elements index numbers representing the order of Uracil and Cytosine in the base sequence of the target RNA; A step of determining a plurality of ordered pair sets that have ordered pairs composed of elements of the first index number set and elements of the second index number set as elements and satisfy the conditions below. (wherein IN1 represents the first index number set, IN2 represents the second index number set, J represents the set of ordered pairs, a, i1, and i2 represent elements of the first index number set, and b, j1, and j2 represent elements of the second index number set); For each of the plurality of ordered pair sets, a step of determining a gain for each of the ordered pairs included in the ordered pair set based on the base sequence of the target RNA and the element value of the position corresponding to each of the ordered pairs included in the ordered pair set in the initial prediction matrix; A step of determining, among the plurality of ordered pair sets, ordered pair sets that include only ordered pairs having a gain greater than 0 as hydrogen bond candidate ordered pair sets; For each of the above hydrogen bond candidate ordered pair sets, a step of calculating the sum of the gains for all ordered pairs included in the above hydrogen bond candidate ordered pair set to determine the total gain for each of the above hydrogen bond candidate ordered pair sets; A step of determining a hydrogen bond candidate ordered pair set having the maximum total gain among the above hydrogen bond candidate ordered pair sets as a final ordered pair set; and A method for predicting the secondary structure of RNA, comprising the step of outputting, as the secondary structure prediction matrix, a matrix in which, for each of all ordered pairs (xi, xj) included in the final ordered pair set, the element corresponding to the xi-th row and the xj-th column and the element corresponding to the xj-th row and the xi-th column in the initial prediction matrix are changed to 1, and all other remaining elements of the initial prediction matrix are changed to 0.

4. In the third paragraph, the step of determining the gain for each of the ordered pairs included in the set of ordered pairs by the post-processing unit is as follows: When the ordered pair (i, j) is included in the set of ordered pairs, A step of determining the gain for the ordered pair (i, j) as 0 when the i-th base in the base sequence of the target RNA is adenine and the j-th base is cytosine, when the difference between i and j is 3 or less, and when the element value at the position corresponding to the i-th row and the j-th column in the initial prediction matrix is ​​less than a predetermined threshold; and A method for predicting the secondary structure of RNA, comprising, for all other cases, determining the gain for the ordered pair (i, j) as the element value at the position corresponding to the i-th row and j-th column in the initial prediction matrix.

5. A method for predicting the secondary structure of RNA in the fourth paragraph, wherein the threshold value has a value greater than or equal to 0.5 and less than 1.

6. In the third paragraph, the step of determining the gain for each of the ordered pairs included in the set of ordered pairs by the post-processing unit is as follows: If the ordered pair (i, j) is included in the set of ordered pairs, A step of determining the gain for the ordered pair (i, j) as 0 when the ith base in the base sequence of the target RNA is adenine and the jth base is cytosine, when the difference between i and j is 3 or less, and when the element value at the position corresponding to the ith row and the jth column in the initial prediction matrix is ​​less than a predetermined threshold; In addition, when the i-th base in the base sequence of the target RNA is Guanine and the j-th base is Uracil, a step of determining the gain for the ordered pair (i, j) as a value obtained by multiplying the element value at the position corresponding to the i-th row and the j-th column in the initial prediction matrix by a predetermined attenuation rate; and A method for predicting the secondary structure of RNA, comprising, for all other cases, determining the gain for the ordered pair (i, j) as the element value at the position corresponding to the i-th row and j-th column in the initial prediction matrix.

7. A method for predicting the secondary structure of RNA in clause 6, wherein the decay rate has a value greater than 0.5 and less than 1.

8. In the third paragraph, the step of determining the gain for each of the ordered pairs included in the set of ordered pairs by the post-processing unit is as follows: If the ordered pair (i, j) is included in the set of ordered pairs, A step of determining the gain for the ordered pair (i, j) as 0 when the ith base in the base sequence of the target RNA is adenine and the jth base is cytosine, when the difference between i and j is 3 or less, and when the element value at the position corresponding to the ith row and the jth column in the initial prediction matrix is ​​less than a predetermined threshold; In addition, when the ith base in the base sequence of the target RNA is Guanine and the jth base is Cytosine, a step of determining the gain for the ordered pair (i, j) as the element value at the position corresponding to the ith row and the jth column in the initial prediction matrix; In addition, when the i-th base in the base sequence of the target RNA is adenine and the j-th base is uracil, a step of determining the gain for the ordered pair (i, j) as a value obtained by multiplying a predetermined first attenuation rate by the element value at the position corresponding to the i-th row and the j-th column in the initial prediction matrix; and In addition, a method for predicting the secondary structure of RNA, comprising the step of determining the gain for the ordered pair (i, j) as a value obtained by multiplying the element value at the position corresponding to the i-th row and the j-th column in the initial prediction matrix by a predetermined second decay rate having a value smaller than the first decay rate, when the ith base in the base sequence of the target RNA is guanine and the jth base is uracil.

9. A method for predicting the secondary structure of RNA in the 8th paragraph, wherein the first decay rate and the second decay rate have values ​​greater than 0.5 and less than 1.

10. A learning data generation unit that uses a secondary structure database that stores in advance the folded secondary structures of a plurality of RNAs included in a human body, and generates learning data using a symmetric matrix as a label for the input data, wherein the base sequence of each of the plurality of RNAs is input data, and the number of rows and columns is the same as the length of the base sequence of the RNA corresponding to the input data, and in which, if the p-th base and the q-th base in the base sequence of the RNA are hydrogen bonded to each other, the element corresponding to the p-th row and the q-th column and the element corresponding to the q-th row and the p-th column have a value of 1, and if the r-th base and the s-th base are not hydrogen bonded to each other, the element corresponding to the r-th row and the s-th column and the element corresponding to the s-th row and the r-th column have a value of 0; A secondary structure prediction unit that trains the artificial neural network to output a symmetric matrix representing its own hydrogen bonding relationship between bases included in the base sequence of the input RNA when the base sequence of RNA is input to the artificial neural network using the above learning data, and then generates a secondary structure prediction model by inputting the base sequence of a target RNA received from the outside into the secondary structure prediction model and generating the symmetric matrix output from the secondary structure prediction model as an initial prediction matrix; and A secondary structure prediction system for RNA, comprising a post-processing unit that analyzes the base sequence of the target RNA, changes at least some of the elements at positions corresponding to base pairs capable of hydrogen bonding with each other among the base pairs included in the base sequence of the target RNA in the initial prediction matrix to 1 and changes the remaining elements of the initial prediction matrix to 0 to generate a secondary structure prediction matrix.

Citation Information

Patent Citations

  • Predicting method of RNA two-grade structure

    CN110010194A

  • Rna secondary structure prediction system and method

    JP1996154677A

  • Systems and methods to determine RNA structure and uses thereof

    WO2022246473A1