RNA sequence design algorithm based on RNA secondary structure substructure solution

By identifying the key substructures in the RNA secondary structure and performing deep-first search and enumeration adjustment, the shortcomings of existing algorithms in designing complex RNA secondary structures are solved, and the success rate and efficiency of RNA sequence design is improved. It is suitable for RNA molecular sensors, gene regulation and RNA drug design.

CN120279979APending Publication Date: 2025-07-08SHENZHEN INST OF ADVANCED TECH CHINESE ACAD OF SCI
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510350006.X
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-03-24
Publication Date
2025-07-08

AI Technical Summary

Technical Problem

The existing RNA sequence design algorithms are difficult to effectively solve when designing complex RNA secondary structures, such as loop structures and symmetric structures at close distances, and cannot meet all target RNA secondary structures in the EteRNA 100 test set.

Method used

By identifying key substructures in the RNA secondary structure, using a deep-first search and exhaustive enumeration method, the initial RNA sequence is generated, and the base pairs and individual sites are adjusted through the optimization process to meet the target RNA secondary structure until the maximum number of iterations is reached.

Benefits of technology

It improves the success rate of identifying and designing complex RNA secondary structures, can better adapt to the close-range loop structure and symmetric structure, improves the efficiency and interpretability of RNA sequence design, and is suitable for RNA molecular sensors, gene regulation and RNA drug design.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120279979A_ABST
    Figure CN120279979A_ABST
Patent Text Reader

Abstract

The invention discloses an RNA sequence design algorithm based on RNA secondary structure substructure solution, and belongs to the technical field of bioinformatics. The RNA sequence design algorithm comprises the following steps: 1, inputting a target RNA secondary structure; 2, identifying a key substructure in the RNA secondary structure; 3, solving a key substructure in the RNA secondary structure; 4, generating an initial RNA sequence by using the satisfactory sequence of the key substructure; and 5, judging whether a target secondary structure is met or not, and if not, repeating the optimization process until a satisfactory RNA sequence is output. The RNA sequence designed on the basis of the target RNA secondary structure can be widely applied to many biological scenes such as RNA molecular sensor design, gene regulation and control and RNA medicine design.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention belongs to the technical field of bioinformatics, and particularly relates to an RNA sequence design algorithm based on the solution of RNA secondary structure substructures. Background Art

[0002] The secondary structure of RNA molecules is closely related to the biological functions of RNA molecules. Designing RNA sequences that conform to specific target secondary structures is an important issue in RNA biology and synthetic biology, and can be widely used in many biological scenarios such as designing RNA molecule sensors, gene regulation, and designing RNA drugs. Currently, existing RNA sequence design algorithms all have certain deficiencies in designing RNA secondary structures that are difficult to meet, and there is no algorithm that can solve all the target RNA secondary structures in the EteRNA 100 test set, which is widely used to test the RNA sequence design ability.

[0003] The problem of designing RNA sequences based on target RNA secondary structures is an important technical issue in RNA biology, especially in RNA synthetic biology. The latest algorithms in the prior art, such as NEMO (Portela F. An unexpectedly effective Monte Carlo technique for the RNA inverse folding problem. bioRxiv 2018:345587), SAMFEO (Zhou T, Dai N, Li S, Ward M, Mathews DH, Huang L. RNA design via structure-aware multifrontier ensemble optimization. Bioinformatics. 2023 Jun 30;39(39 Suppl 1):i563-i571. doi:10.1093 / bioinformatics / btad252.), etc. These algorithms first generate initial RNA sequences, and perform single-site and base-pair transformations on the RNA sequences. Subsequently, according to different objective functions in these algorithms, relatively optimal RNA sequences are selected, and the above RNA sequence optimization process is continuously repeated until an RNA sequence that can meet the target secondary structure is obtained or the maximum number of iterations is reached. However, these RNA sequence design algorithms based on specific target secondary structures have the following disadvantages: they are not applicable when designing complex RNA substructures, including loop structures with close distances, RNA substructures with symmetric structures, etc. Summary of the Invention

[0004] Aiming at the defects of the above-mentioned existing technologies, the purpose of the present invention is to design and provide an algorithm for identifying and solving substructures in RNA secondary structures, obtaining RNA sequences that can satisfy the substructures, and using the RNA sequences that can satisfy the substructures in the design of RNA sequences that meet specific target secondary structures. RNA sequences with more complex secondary structures can be designed, and RNA sequences that meet specific target RNA secondary structures can be designed more efficiently, thereby improving the ability and efficiency of RNA sequence design in various situations such as nucleic acid drugs, gene editing, and gene regulation in synthetic biology. It is still applicable to complex situations such as ring structures with relatively close distances and RNA substructures with symmetric structures.

[0005] To achieve the above object, the present invention adopts the following technical solutions:

[0006] On the one hand, the present invention provides an RNA sequence design algorithm based on the solution of RNA secondary structure substructures, including the following steps:

[0007] (1) Input the target RNA secondary structure;

[0008] (2) Identify the key substructures in the RNA secondary structure;

[0009] (3) Solve the key substructures in the RNA secondary structure;

[0010] (4) Generate an initial RNA sequence using the sequences that can satisfy the key substructures;

[0011] (5) If the initial RNA sequence satisfies the target RNA secondary structure, output the RNA sequence that can be satisfied; if the initial RNA sequence does not satisfy the target RNA secondary structure, perform an optimization process, that is, select a better sequence according to the objective function, perform base pair, single-site, and substructure position sequence transformations on the better sequence, repeat the optimization process until the target RNA secondary structure is satisfied or the maximum number of iterations is reached, and output the RNA sequence that can be satisfied.

[0012] In the described RNA sequence design algorithm based on the solution of RNA secondary structure substructures, the specific process of identifying the key substructures in the RNA secondary structure in step (2) is as follows:

[0013] (a) Input whether to screen the positive and negative of the ring structure energy and the ring structure spacing limit;

[0014] (b) Obtain the positions of all RNA ring structures in the target RNA secondary structure;

[0015] (c) Calculate the distance matrix between ring structures according to the positions of all RNA ring structures;

[0016] (d) If it is necessary to screen the positive and negative of the energy of the loop structure, use RNA secondary structure prediction software for determination. Fill the base pairs of the loop structure to be determined with G-C, fill the unpaired bases with A, and obtain the energy value of this loop structure;

[0017] (e) Starting from each loop structure as the starting point, perform a depth-first search. If the distance between two loop structures is less than or equal to the loop structure spacing limit, it is regarded as connected, and a depth-first search path starting from each loop structure is obtained;

[0018] (f) If the length of the screened depth-first search path is greater than 2, then the entire loop structure on this path is used as the key substructure in the RNA secondary structure.

[0019] In the described RNA sequence design algorithm based on the solution of RNA secondary structure substructures, all RNA loop structures described in step (b) include hairpin loops, internal loops, multi-branch loops, and external loops.

[0020] In the described RNA sequence design algorithm based on the solution of RNA secondary structure substructures, the specific process of solving the key substructures in the RNA secondary structure in step (3) is as follows:

[0021] (ⅰ) Input the target RNA secondary structure and the list of key substructures in the RNA secondary structure:

[0022] (ⅱ) Respectively take the base with the smallest number in each key substructure in the RNA secondary structure in the list as the index to sort the loop structures from small to large, and obtain an ordered list of loop structures;

[0023] (ⅱ) Take the loop structure with the smallest index in the ordered list of loop structures as the initial current substructure to be solved;

[0024] (ⅲ) Connect the next loop structure in the ordered list of loop structures to the current substructure to be solved;

[0025] (ⅳ) Fill base pairs, or base pairs and hairpin structures in front of or behind the base pairs connecting the current substructure to be solved and the non-key substructures in the RNA secondary structure; The criterion for judging "the base pairs connecting the non-key substructures in the RNA secondary structure" is: sort the sites of the current substructure to be solved that have not been filled to obtain a list of sites, such as [3, 4, 5, 100, 101]. When mapping these sites to the representation of the initially input target secondary structure, if "(().)" is obtained, then the middle () part is only base pairs rather than hairpins and needs to be filled.

[0026] (ⅴ) If the number of ring structures in the structure obtained in the above step (ⅳ) is 2, then exhaust all the enumeration sites, set the unpaired sites as A, and randomly fill the paired sites in the filling area with G-C or C-G base pairs to obtain a satisfiable sequence of the current sub-structure to be solved;

[0027] (ⅵ) If the number of ring structures in the structure obtained in the above step (ⅳ) is greater than 2, then calculate the enumeration sites of the newly connected ring structures and the sites where the existing sequences are located, exhaust the enumeration sites of the newly connected ring structures, fill the sites where the existing sequences are located with the existing sequences, set the unpaired sites as A, and randomly fill the paired sites in the filling area with G-C or C-G base pairs to obtain a satisfiable sequence of the current sub-structure to be solved;

[0028] (ⅶ) Perform secondary structure prediction on the satisfiable sequence of the current sub-structure to be solved obtained in step (ⅴ) or (ⅵ), and determine whether the newly connected ring structure is the last item in the ordered list of ring structures. If so, output the satisfiable sequence of the RNA sub-structure to be solved; if not, repeat step (ⅲ) until the newly connected ring structure is the last item in the ordered list of ring structures, and output the satisfiable sequence of the RNA sub-structure to be solved.

[0029] In the described RNA sequence design algorithm based on solving RNA secondary structure sub-structures, the numbering method in step (ⅰ) is as follows: number the bases in ascending order from the 5'-end to the 3'-end of the RNA.

[0030] In the described RNA sequence design algorithm based on solving RNA secondary structure sub-structures, the enumeration sites in step (ⅴ) are base pairs, a pair of unpaired bases adjacent to the base pair in the ring structure, and the unpaired bases contained in a convex ring with a length of 1.

[0031] In the described RNA sequence design algorithm based on solving RNA secondary structure sub-structures, the exhaustive search in step (ⅴ) means that the values of base pairs are successively G-C, C-G, A-U, U-A, U-G, G-U, and the values of unpaired bases are successively A, U, C, G.

[0032] In the second aspect, the present invention provides an RNA sequence design device based on solving RNA secondary structure sub-structures, and the device includes: a memory and a processor;

[0033] The memory is used to store program instructions;

[0034] The processor is used to call the program instructions, and when the program instructions are executed, it is used to execute any one of the described RNA sequence design algorithms based on solving RNA secondary structure sub-structures.

[0035] In a third aspect, the present invention provides an RNA sequence design system based on solving substructures of RNA secondary structures, including:

[0036] An input unit for inputting a target RNA secondary structure;

[0037] An identification unit for identifying key substructures in the RNA secondary structure;

[0038] A solving unit for solving key substructures in the RNA secondary structure;

[0039] A generating unit for generating an initial RNA sequence using a satisfiable sequence of the key substructure;

[0040] An output unit for outputting a satisfiable RNA sequence if the initial RNA sequence satisfies the target RNA secondary structure; if the initial RNA sequence does not satisfy the target RNA secondary structure, an optimization process is performed, that is, a better sequence is selected according to the objective function, base pair, single-site, and substructure position sequence transformations are performed on the better sequence, and the optimization process is repeated until the target RNA secondary structure is satisfied or the maximum number of iterations is reached, and a satisfiable RNA sequence is output.

[0041] In a fourth aspect, the present invention provides a computer-readable storage medium storing a computer program;

[0042] When the computer program is executed by a processor, it implements any one of the RNA sequence design algorithms based on solving substructures of RNA secondary structures.

[0043] Compared with the prior art, the present invention has the following beneficial effects:

[0044] The present invention is an algorithm for designing a satisfiable RNA sequence based on a target RNA secondary structure. The present invention first identifies key substructures in the target RNA secondary structure, designs RNA sequences that can satisfy each substructure, and uses the RNA sequences that can satisfy each substructure for the design of an RNA sequence that can satisfy the target RNA secondary structure. Through the above design algorithm of the present invention, key substructures in the RNA secondary structure can be better identified and located, and the solution of key RNA substructures can be carried out. The success rate of solving a satisfiable sequence of an RNA secondary structure with complex RNA secondary structure substructures using the RNA sequence solving algorithm of the present invention is higher. Due to the identification and location of key substructures in the RNA secondary structure, the interpretability of the RNA sequence design algorithm used in the present invention is stronger. The design of RNA sequences based on the target RNA secondary structure in the present invention can be widely used in many biological scenarios such as designing RNA molecular sensors, gene regulation, and designing RNA drugs. BRIEF DESCRIPTION OF THE DRAWINGS

[0045] Figure 1 This is the overall algorithm flow of the present invention;

[0046] Figure 2 This is the recognition process of the key substructures of the RNA secondary structure;

[0047] Figure 3 This is the process of solving the key substructures in the RNA secondary structure;

[0048] Figure 4 This is an example of filling base pairs and hairpin structures in the RNA substructure to be solved;

[0049] Figure 5 This is the illustration of the target secondary structure of "Short String 4";

[0050] Figure 6 This is the illustration of the key substructures in the target RNA secondary structure;

[0051] Figure 7 This is the graph of sorting the loop structures in the key substructures from small to large;

[0052] Figure 8 This is the comparison of the performance of the present invention and other RNA sequence design algorithms on the EteRNA 100 test set. Detailed implementation manners

[0053] The present invention will be further described below with reference to the accompanying drawings and embodiments.

[0054] Embodiment 1:

[0055] (1) The overall algorithm flow adopted by the present invention is as shown in the accompanying drawings. The input form of the RNA secondary structure is the dot-bracket representation of the RNA secondary structure, such as "((((....))))". After inputting the target RNA secondary structure, all loop structures in the target RNA secondary structure are obtained, including hairpin loops, internal loops, multi-branch loops, external loops, etc. Subsequently, according to the needs of the algorithm, the determination criteria for the key substructures in the RNA secondary structure are input, and the key substructures in the RNA secondary structure are identified based on this criterion, and the key substructures are solved. Figure 1 (2) The recognition process of the key substructures in the RNA secondary structure is as follows

[0056] (2) The recognition process of the key substructures in the RNA secondary structure is as follows Figure 2As shown, according to whether to screen the positive and negative of the loop structure energy and the restriction of the loop structure distance of the input, all loop structures in the target RNA secondary structure are first obtained, including hairpin loops, internal loops, multi-branched loops, external loops, etc. If it is necessary to screen the loop structures with positive or negative energy, an RNA secondary structure prediction software is used for determination. The base pairs of the loop structure to be determined are filled with G-C, and the unpaired bases are filled with A, and the energy value of this loop structure is obtained. According to the positions of all loop structures, that is, the position numbers of all bases in the loop structure, the distance matrix between loop structures is calculated.

[0057] The base numbers are numbered in ascending order from the 5' end to the 3' end of the RNA in sequence. The distance between loop structures is defined as follows: if two loop structures are adjacent, the distance between the two loop structures is the absolute value of the difference between the two closest base numbers in the two loop structures; otherwise, it is infinity. The element in the i-th row and j-th column of the loop structure distance matrix represents the distance between the i-th loop structure and the j-th loop structure. The order of the loop structures is sorted from small to large according to the base with the smallest number in each loop structure as the index.

[0058] Using depth-first search with each loop structure as the starting point according to the loop structure distance matrix, where loop structures with a distance less than or equal to the loop structure spacing limit between two loop structures are regarded as connected, and the corresponding depth search path is obtained. If the length of this depth-first search path is greater than or equal to 2 (that is, the number of loop structures is greater than or equal to 2), it means that there are multiple loop structures adjacent to each other, and the loop structures on this depth-first search path are used as the key substructures of the RNA to be solved.

[0059] (3) The solution process of the key substructure in the RNA secondary structure is as shown in the appendix Figure 3 As shown, first, the loop structures are sorted from small to large according to the base with the smallest number in each loop structure as the index, and an ordered list of loop structures is obtained. The loop structures in the ordered list of loop structures are connected in sequence, and a new loop structure is connected at each stage. And base pairs or a combination of base pairs and hairpin structures are filled in front of or behind the bases that are partially connected between the current substructure to be solved and the non-key substructures. An example of filling is Figure 4 As shown, where the part within the yellow box is the connected loop structures, and the part within the red box is the filled base pairs or the combination of base pairs and hairpin structures.

[0060] The enumerated sites in the RNA substructure to be solved are defined as base pairs, a pair of unpaired bases adjacent to the base pair in the loop structure (if any), and the unpaired bases included in a convex loop with a length of 1. The exhaustive enumeration of the enumerated sites means that for the base pair part, the values are taken in sequence as G-C, C-G, A-U, U-A, U-G, G-U, and for the unpaired bases, the values are taken in sequence as A, U, C, G.

[0061] For the RNA secondary structure after filling base pairs and hairpin structures, if the number of loop structures in the structure is 2 in a certain solution stage, then all the enumerated sites are exhausted. For the remaining sites, if they are base pairs, random G-C or C-G base pairs are used for filling, and non-paired bases are filled with A to obtain an RNA sequence.

[0062] If the number of loop structures in the structure is greater than 2 in a certain solution stage, then the satisfiable sequences of the previous solution stage are enumerated in the existing loop structure part, all the enumerated sites are enumerated in the newly connected loop structure part. For the remaining sites, if they are base pairs, random G-C or C-G base pairs are used for filling, and non-paired bases are filled with A to obtain an RNA sequence.

[0063] Use Vienna RNA to predict the secondary structure of this RNA sequence. If the prediction result is the same as the current sub-structure to be solved after filling, it is determined that the base combination of the enumerated sites satisfies the current sub-structure to be solved. All the RNA sequences that satisfy the current sub-structure to be solved are used for the solution of the next solution stage, or as the final satisfiable sequence of the key sub-structure. If there is no such RNA sequence, the solution fails, and the above steps are repeated to connect the subsequent loop structures to the current sub-structure to be solved.

[0064] (4) Generate an initial RNA sequence according to the satisfiable sequences of each RNA sub-structure obtained. For all the sequences at the positions of the key sub-structures in the recognized RNA secondary structure in the initial RNA sequence, the satisfiable sequences in the key sub-structures are randomly selected. The other non-paired regions in the initial RNA sequence are filled with A, and the paired regions are filled with random G-C base pairs and C-G base pairs. Subsequently, the objective function values of each RNA sequence in the initial RNA sequence are calculated, and a better RNA sequence is selected from them. Single bases, base pairs, and the sequence of the RNA sub-structure positions are changed to generate new RNA sequences. The above optimization process is continuously repeated until an RNA sequence that satisfies the target secondary structure is obtained, or the maximum number of iterations is reached.

[0065] Example 2:

[0066] Select the RNA secondary structure "Short String4" that could not be successfully solved by other methods in the EteRNA100 test set. Its secondary structure is "...((.((..((..(...(.((.((....))))..)...(.((..((....))..))..)...).))..)))).....................". Use this method to design an RNA sequence that satisfies the secondary structure of the sub-goal. The illustration of its target secondary structure is as Figure 5 shown, and the specific process is as follows:

[0067] 1. First, input the target RNA secondary structure, set the distance limit of the loop structure to 1, and do not limit the positive or negative energy of the loop structure.

[0068] 2. Obtain a set of loop structures that meet the requirements according to the distance limit of the loop structure. This is the part where the middle multi-branched loop is connected to the 3 inner loops, and it is solved as the key sub-structure in the target RNA secondary structure. For example, Figure 6 as shown in the part within the red box:

[0069] 3. Sort the loop structures in the key sub-structure from smallest to largest to obtain an ordered list of loop structures as follows Figure 7 the loop structures corresponding to 1, 2, 3, and 4 below;

[0070] 4. Solve the key sub-structure, which is the following steps:

[0071] (1) The 1st loop structure is used as the current sub-structure to be solved;

[0072] (2) Connect the 1st loop structure with the 2nd loop structure and perform filling. The filled secondary structure is "(((((((..(...(((((((((....)))))))))...(((((((((....)))))))))...).)))))))". Enumerate all the enumerated sites to obtain 6 satisfiable sequences in the current solution stage;

[0073] (3) Connect the existing part with the 3rd loop structure and perform filling. The filled secondary structure is "(((((((..(...(.(((((((((....)))))))))..)...(((((((((....)))))))))...).)))))))". Enumerate the enumerated sites in the newly connected 3rd loop structure to obtain 23 satisfiable sequences in the current solution stage;

[0074] (4) Connect the existing part with the 4th loop structure and perform filling. The filled secondary structure is "(((((((..(...(.(((((((((....)))))))))..)...(.(((((((((....)))))))))..)...).)))))))", obtaining 54 satisfiable sequences in the current solution stage. Since the 4th loop structure is the last item in the ordered list of loop structures, the satisfiable sequences obtained in this step are the satisfiable sequences of the key sub-structure.

[0075] 5. Generate an initial RNA sequence using a satisfiable sequence with key substructures, such as "AAAAAACGAGGAACUGACGAAGAGCACCAAAAGGGUGACAAAGAGCAAGCAAAA GCAAGUGACAAAGAGGAACCCGAAAAAAAAAAAAAAAAAAAAA"

[0076] 6. Continuously optimize and adjust the initial RNA sequence to obtain a sequence that satisfies the target RNA secondary structure, such as "AAAAAACGAGGGACUGACGAAGAGCAUCUAAUGAGUGACAAAGAGCGAGGAAAA CCGAGUGACAAAGAGGGACCCGAAAAAAAAAAAAAAAAAAAAA".

[0077] Literature of the EteRNA100 test set: Anderson-Lee J, Fisker E, Kosaraju V, Wu M, Kong J, Lee J, Lee M, Zada M, Treuille A, Das R; Eterna Players. Principles for Predicting RNA Secondary Structure Design Difficulty. J Mol Biol. 2016 Feb 27;428(5 Pt A):748-757.

[0078] Literature source of other existing methods that failed to solve the RNA secondary structure selected in Example 2: Rohan V. Koodli, Boris Rudolfs, Hannah K. Wayment-Steele, Eterna Structure Designers, Rhiju Das Redesigning the Eterna100 for the Vienna 2 folding engine bioRxiv 2021.08.26.457839.

[0079] In summary, the present invention can solve more RNA secondary structures with complex substructures, which cannot be achieved by the latest technologies such as the NEMO algorithm and the SAMFEO algorithm.

[0080] Example 3:

[0081] The four RNA sequence design algorithms of the present invention, namely the NEMO algorithm, the SAMFEO algorithm, the LEARNA algorithm, and the EteRNABrain algorithm, were tested on the EteRNA 100 test set (a test set widely used to evaluate RNA sequence design algorithms). Among them, the specific process of the NEMO algorithm can be found in Portela F. An unexpectedly effective Monte Carlo technique for the RNA inverse folding problem. bioRxiv 2018:345587; the specific process of the SAMFEO algorithm can be found in Zhou T, Dai N, Li S, Ward M, Mathews DH, Huang L. RNA design via structure-aware multifrontier ensemble optimization. Bioinformatics. 2023 Jun 30;39(39 Suppl 1):i563-i571. doi:10.1093 / bioinformatics / btad252.. The specific process of the LEARNA algorithm can be found in Runge F, Stoll D, Falkner S, Hutter F. Learning to Design RNA. ICLR 2019. 2018. The specific process of the EteRNABrain algorithm can be found in Koodli RV, Keep B, Coppess KR, Portela F, Eterna participants, Das R. EternaBrain: and Automated RNA design through movesets and strategies from an Internet-scale RNA videogame. PLoS Comput Biol.

[0082] 2019;15:e1007059. doi:10.1371 / journal.pcbi.1007059.

[0083] The specific performance results are as Figure 8 shown. Its performance is on par with the best NEMO algorithm and has obvious advantages compared with some other RNA sequence design algorithms. This shows that the present invention provides an RNA secondary structure design algorithm with excellent performance and is feasible.

Claims

1. An RNA sequence design algorithm based on solving the sub-structure of RNA secondary structure, characterized in that, It includes the following steps: (1) Input the target RNA secondary structure; (2) Identify the key substructures in the RNA secondary structure; (3) Solve the key substructures in the RNA secondary structure; (4) Generate an initial RNA sequence using the satisfiable sequences of the key substructures; (5) If the initial RNA sequence satisfies the target RNA secondary structure, output the satisfiable RNA sequence; if the initial RNA sequence does not satisfy the target RNA secondary structure, perform an optimization process, that is, select a better sequence according to the objective function, perform base pair, single-site, and substructure position sequence transformations on the better sequence, repeat the optimization process until the target RNA secondary structure is satisfied or the maximum number of iterations is reached, and output the satisfiable RNA sequence.

2. The RNA sequence design algorithm based on the solution of RNA secondary structure substructures as described in claim 1, wherein, The specific process of identifying the key substructures in the RNA secondary structure described in step (2) is as follows: (a) Input whether to screen the positive and negative of the loop structure energy and the loop structure spacing limit; (b) Obtain the positions of all RNA loop structures in the target RNA secondary structure; (c) Calculate the distance matrix between loop structures based on the positions of all RNA loop structures; (d) If it is necessary to screen the positive and negative of the loop structure energy, use RNA secondary structure prediction software for determination. Fill the base pairs of the loop structure to be determined with G-C, and fill the unpaired bases with A, and obtain the energy value of this loop structure; (e) Conduct a depth-first search starting from each loop structure. If the distance between two loop structures is less than or equal to the loop structure spacing limit, it is regarded as connected, and obtain the depth-first search path starting from each loop structure; (f) If the length of the screened depth-first search path is greater than 2, regard the overall loop structures on this path as the key substructures in the RNA secondary structure.

3. The RNA sequence design algorithm based on the solution of RNA secondary structure substructures as described in claim 2, characterized in that All the RNA loop structures described in step (b) include hairpin loops, internal loops, multi-branch loops, and external loops.

4. An RNA sequence design algorithm based on the solution of RNA secondary structure substructures as claimed in claim 1, wherein, The specific process of solving the key substructures in the RNA secondary structure described in step (3) is as follows: (ⅰ) Input the target RNA secondary structure and the list of key substructures in the RNA secondary structure: (ⅱ) Sort the loop structures from small to large by using the base with the smallest number in each key substructure in the RNA secondary structure in the list as the index, and obtain an ordered list of loop structures; (ⅱ) Regard the loop structure with the smallest index in the ordered list of loop structures as the initial current substructure to be solved; (ⅲ) Connect the next loop structure in the ordered list of loop structures to the current substructure to be solved; (ⅳ) Fill in base pairs, or base pairs and hairpin structures in front of or behind the base pairs connecting the current substructure to be solved and the non-key substructures in the RNA secondary structure; (ⅴ) If the number of loop structures in the structure obtained in the above step (ⅳ) is 2, enumerate all enumeration sites, set the unpaired sites as A, and randomly fill the paired sites in the filling area with G-C or C-G base pairs to obtain the satisfiable sequence of the current substructure to be solved. (ⅵ) If the number of ring structures in the structure obtained in the above step (ⅳ) is greater than 2, calculate the enumeration sites of the newly connected ring structures and the sites where the existing sequences are located. Exhaustively enumerate the enumeration sites of the newly connected ring structures, fill the sites where the existing sequences are located with the existing sequences, set the unpaired sites as A, and randomly fill the paired sites in the filled area with G-C or C-G base pairs to obtain the satisfiable sequences of the currently to-be-solved sub-structures; (ⅶ) For the satisfiable sequences of the currently to-be-solved sub-structures obtained in step (ⅴ) or (ⅵ) Perform secondary structure prediction, and determine whether the newly connected ring structure is the last item in the ordered list of ring structures. If so, output the satisfiable sequences of the to-be-solved RNA sub-structures; if not, repeat step (ⅲ) until the newly connected ring structure is the last item in the ordered list of ring structures, and output the satisfiable sequences of the to-be-solved RNA sub-structures.

5. The RNA sequence design algorithm based on the solution of RNA secondary structure substructures as claimed in claim 4, wherein The numbering method described in step (ⅰ) is: number the bases in ascending order from the 5'-end to the 3'-end of the RNA.

6. The RNA sequence design algorithm based on the solution of RNA secondary structure substructures according to claim 4, wherein The enumeration sites described in step (ⅴ) are base pairs, a pair of unpaired bases adjacent to the base pair in the ring structure, and the unpaired bases contained in the bulge loop with a length of 1.

7. The RNA sequence design algorithm based on the solution of RNA secondary structure substructures as described in claim 4, characterized in that The exhaustion described in step (ⅴ) is that the values of the base pairs are successively G-C, C-G, A-U, U-A, U-G, G-U, and the values of the unpaired bases are successively A, U, C, G.

8. An RNA sequence design device based on the solution of RNA secondary structure substructures, characterized in that, The device includes: a memory and a processor; The memory is used to store program instructions; The processor is used to call the program instructions. When the program instructions are executed, it is used to execute the RNA sequence design algorithm for solving RNA secondary structure sub-structures as described in any one of claims 1-7.

9. An RNA sequence design system based on the solution of RNA secondary structure substructures, characterized in that, Including: An input unit for inputting the target RNA secondary structure; An identification unit for identifying the key sub-structures in the RNA secondary structure; A solving unit for solving the key sub-structures in the RNA secondary structure; A generating unit for generating an initial RNA sequence using the satisfiable sequences of the key sub-structures; An output unit for outputting the satisfiable RNA sequences if the initial RNA sequence satisfies the target RNA secondary structure; If the initial RNA sequence does not satisfy the target RNA secondary structure, perform an optimization process, that is, select a better sequence according to the objective function, perform base pair, single-site, and sub-structure position sequence transformations on the better sequence, and repeat the optimization process until the target RNA secondary structure is satisfied or the maximum number of iterations is reached, and output the satisfiable RNA sequences.

10. A computer-readable storage medium, characterized in that, The computer-readable storage medium stores a computer program; When the computer program is executed by the processor, it implements the RNA sequence design algorithm for solving RNA secondary structure sub-structures as described in any one of claims 1-7.