Design method and application of dsRNA / siRNA target based on thermodynamic parameters and computer program product

By analyzing and scoring dsRNA/siRNA target sequences based on thermodynamic parameters, the problem of low design efficiency in the existing technology is solved, and a fast and efficient target sequence design is achieved, which is suitable for the application of RNA interference technology.

CN120126561AActive Publication Date: 2025-06-10SILICON GENE TECH (SHANGHAI) CO LTD

Patent Information

Application Number
CN202510525341.9
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-04-25
Publication Date
2025-06-10
Estimated Expiration
2045-04-25

AI Technical Summary

Technical Problem

The prior art relies on random methods when designing dsRNA target sequences, resulting in low design efficiency and time-consuming and difficult to quickly obtain efficient target sequences.

Method used

The dsRNA/siRNA target sequence design method based on thermodynamic parameters was used to standardize the coding sequence of candidate genes, and the base pair thermodynamic parameters were analyzed using the individual nearest neighbor (INN) model, and the final thermodynamic parameters were calculated in combination with the corrected parameters, scored and sorted to determine the optimal target sequence.

Benefits of technology

It realizes the rapid design of efficient dsRNA/siRNA interference target sequences, shortens the research cycle, saves time and resources, and improves the efficiency and accuracy of target design.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120126561A_ABST
    Figure CN120126561A_ABST
Patent Text Reader

Abstract

The invention relates to the technical field of biological information, and provides a thermodynamic parameter-based dsRNA / siRNA target sequence design method, which comprises the following steps of: standardizing coding sequences of candidate genes; presetting a reading frame, and segmenting the coding sequence of the candidate gene through translation of the reading frame to obtain a plurality of candidate dsRNA / siRNA target sequences; the candidate dsRNA / siRNA target sequences are read and analyzed, and base pairs of all the candidate dsRNA / siRNA target sequences are obtained; presetting thermodynamic parameters of all the base pairs, calculating the sum of the thermodynamic parameters of all the base pairs, and obtaining final thermodynamic parameters of the candidate target sequence in combination with the correction parameters; and calculating the score of each candidate dsRNA / siRNA target sequence according to the thermodynamic parameters, and determining the candidate dsRNA / siRNA target sequence with the highest score as the optimal dsRNA / siRNA target sequence.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the field of bioinformatics technology, and in particular to an RNA target design method, application and computer program product based on thermodynamic parameters. Background Art

[0002] RNA interference is a highly conserved gene silencing mechanism in eukaryotes. The development process of RNA interference technology has undergone a profound transformation from basic research to practical application. In the agricultural field, RNA interference technology shows great application potential as an innovative pest management tool. By introducing specific dsRNA molecules into pests, key survival genes of pests can be precisely targeted for gene silencing, thereby inhibiting the survival or reproduction of pests and effectively reducing the pest population. Since dsRNA molecules degrade rapidly in the environment, the potential impact on non-target organisms and ecosystems can be significantly reduced. The development of this RNA interference-based biopesticide has the characteristics of high specificity, low toxicity and environmental friendliness. In the future, agriculture is expected to achieve a more precise and environmentally friendly pest control method, reduce the use of chemical pesticides, improve the safety of agricultural products, and promote the sustainable development of agriculture.

[0003] Research shows that for the same gene, the interference efficiency of double-stranded RNA (dsRNA) synthesized from different fragments varies significantly. There are many factors affecting RNA interference efficiency, including dsRNA fragment length, position on the gene, and base sequence characteristics. Currently, the design of dsRNA target sequences mostly relies on random methods and verification, which requires a large amount of time and economic costs. Therefore, how to quickly design efficient dsRNA target sequences has become a key problem to be solved urgently. Summary of the Invention

[0004] In order to solve the above technical problems, the present invention provides a dsRNA / siRNA target sequence design method based on thermodynamic parameters, and the design method includes:

[0005] (1) Standardize the coding sequence of the candidate gene; the standardization process includes: sequence inspection to analyze whether the coding sequence is symmetric; removing redundant information to ensure that the coding sequence only contains DNA sequences of A / G / C / T; completing case conversion to unify the coding sequence into uppercase format; converting T in the DNA sequence to U; wherein, the coding sequence of the candidate gene is directly obtained by querying databases such as NCBI;

[0006] (2) Preset the length of a reading frame as a and the translation step size as b, and divide the coding sequence of the candidate gene by translating the reading frame to obtain multiple candidate dsRNA / siRNA target sequences;

[0007] (3) Read and analyze the candidate dsRNA / siRNA target sequences based on the individual nearest neighbor (INN) model to obtain the base pairs of all candidate dsRNA / siRNA target sequences;

[0008] (4) Preset the thermodynamic parameters of all base pairs, calculate the sum of the thermodynamic parameters of all base pairs, and combine the correction parameters to obtain the final thermodynamic parameters of the candidate target sequence;

[0009] (5) Calculate the score of each candidate dsRNA / siRNA target sequence based on the final thermodynamic parameters, sort the scores, and determine one or more candidate dsRNA / siRNA target sequences with the highest scores as the optimal dsRNA / siRNA target sequences.

[0010] Furthermore, the specific steps of step (2) are:

[0011] (21) On the full-length sequence of the candidate gene, a reading frame of length a is set starting from the mth base. The sequence contained in the reading frame is the initially generated candidate dsRNA / siRNA target sequence. According to literature reports, it is generally believed that the first 100 bases are not important for target design, so m can be selected as a number after 100.

[0012] (22) The reading frame is translated, and the step length of the translation is set to b. Each time the translation is performed, the sequence contained in the reading frame becomes a new candidate dsRNA / siRNA target sequence. The translation operation is continued until the reading frame can no longer be moved;

[0013] (23) All generated candidate dsRNA / siRNA target sequences were summarized into one set.

[0014] Furthermore, the specific steps of obtaining base pairs are: defining two adjacent bases as a base pair, analyzing the candidate dsRNA / siRNA target sequence in the 5' to 3' direction, and obtaining all possible base pairs.

[0015] Furthermore, the specific steps for obtaining base pairs are as follows: two adjacent bases in a reading frame are a base pair; and all base pairs are obtained by translation of the reading frame.

[0016] Furthermore, the calculation formula of the final thermodynamic parameter FE in step (4) is:

[0017] FE=G nn + G sym + G t1 + G t2 + G in ;

[0018] Among them, Gnn is the sum of the base pair energy parameters calculated based on the nearest neighbor model, reflecting the energy contribution generated by the interaction of adjacent base pairs;

[0019] G sym is used to measure the structural symmetry of the candidate dsRNA target sequence. If the structure of the candidate dsRNA target sequence exhibits symmetric characteristics, G sym takes 0.43, otherwise it takes 0;

[0020] G t1 judges the first base of the candidate dsRNA target sequence. If the first base is A or U, G t1 takes 0.45, otherwise it takes 0;

[0021] G t2 judges the last base of the candidate dsRNA target sequence. If the last base is A or U, G t2 takes 0.45, otherwise it takes 0;

[0022] G in is the compensation value for all dsRNA target sequences, used to appropriately correct and adjust the overall energy parameters.

[0023] In this way, the influence of multiple factors on the energy of the candidate sequence can be comprehensively considered, so as to provide a more accurate basis for subsequent screening and evaluation.

[0024] Furthermore, the specific scoring method in step (5) is as follows:

[0025] All candidate dsRNA / siRNA target sequences are sorted in descending order according to the value of the final thermodynamic parameter FE. The target sequence ranked first gets a score of 100, and the target sequence ranked last gets a score of 50. The score calculation formula for the remaining target sequences is:

[0026] Score = 50 + (100 - 50) / (FE max - FE min ) * (fe - FE min );

[0027] where FE max is the FE value of the target sequence ranked first, FE min is the FE value of the target sequence ranked last, and fe is the FE value of the candidate target for which the Score value is to be calculated.

[0028] The second aspect of the present invention provides an application of a method for designing dsRNA target sequences based on thermodynamic parameters in RNA interference and in studying the binding of exogenous dsRNA / siRNA to target positions.

[0029] In a third aspect of the present invention, there is provided an application of a method for designing dsRNA target sequences based on thermodynamic parameters in the design of RNA interference targets in the field of animal and plant protection or the field of pest control and prevention.

[0030] In a fourth aspect of the present invention, there is provided a computer program product for designing dsRNA target sequences based on thermodynamic parameters, and the computer program product is used to execute the above-mentioned design method; the structure of Tar of the computer program consists of three parts: input data, execution program, and output data.

[0031] Furthermore, the input data is the coding sequence of a candidate gene; the execution program is used to process the input data. By presetting the length of a reading frame as a and the translation step size as b, the coding sequence of the candidate gene is segmented by translation of the reading frame to obtain a plurality of candidate dsRNA / siRNA target sequences. Based on the individual nearest neighbor (INN) model, the candidate dsRNA / siRNA target sequences are read and analyzed to obtain the base pairs of all candidate dsRNA target sequences; the sum of the thermodynamic parameters of all base pairs is calculated, and the final thermodynamic parameters of the candidate target sequences are obtained by combining correction parameters; the score of each candidate dsRNA / siRNA target sequence is calculated according to the final thermodynamic parameters, and the scores are sorted, and one or more candidate dsRNA / siRNA target sequences with the highest scores are determined as the optimal dsRNA / siRNA target sequences; the output data is used to output the optimal dsRNA / siRNA target sequences.

[0032] In a fifth aspect of the present invention, there is provided a computer system, including a memory, a processor, and a computer program stored on the memory, and the processor executes the above-mentioned computer program product.

[0033] The present invention has the following beneficial effects:

[0034] (1) In the present invention, by presetting the length of a reading frame as a and the translation step size as b, the entire coding sequence of the candidate gene is traversed one by one to complete the design work of the target, so as to ensure that the generated candidate targets can completely cover all regions of the gene;

[0035] (2) In the present invention, through the individual nearest neighbor model (INN model), thermodynamic analysis is carried out on the candidate target sequences. By presetting the thermodynamic parameters of all base pairs and combining correction parameters to obtain the final thermodynamic parameters of the candidate target sequences, scientific evaluation and screening of the candidate target sequences are realized, and finally efficient dsRNA / siRNA interference target sequences are quickly designed;

[0036] (3) In the present invention, the scoring is achieved by adding the sum of the base pair energy parameters calculated by the nearest neighbor model, the score assigned for judging whether the structure of the dsRNA target sequence is symmetric, the score assigned for judging whether the first and last bases are A or U, and the compensation value of the dsRNA target sequence. By comprehensively considering the influence of multiple factors on the energy of the candidate sequence, it ensures that the selected candidate target sequence has relative advantages in terms of thermodynamic properties, etc., thereby improving the efficiency of the candidate target and providing a more accurate basis for screening the optimal dsRNA / siRNA target sequence;

[0037] (4) The present invention has flexibility in its operation mode. It can either run independently to meet the needs of researchers for direct target design, or be called by other software, facilitating integration into more complex research processes to achieve data sharing and expansion of analysis;

[0038] (5) The present invention directly performs thermodynamic energy calculations without relying on other complex calculation tools, simplifies the research process, and improves the calculation efficiency;

[0039] (6) It only takes a few seconds from the input of the candidate sequence to the completion of the target design in the present invention. It can complete the design of the best dsRNA / siRNA target sequence of the candidate target gene in a short time, greatly shortening the research cycle, saving the time and resources required for researching efficient targets, providing an accelerating impetus for the research process of excellent nucleic acid pesticide targets, and is expected to bring new breakthroughs to the field of agricultural pest control. BRIEF DESCRIPTION OF THE DRAWINGS

[0040] Figure 1 is the flow chart of dsRNA / siRNA target sequence design in the present invention.

[0041] Figure 2 is the statistical chart of the results of different target microinjections in the examples.

[0042] Figure 3 is the analysis chart of the transcriptional level of the target gene FucT6 in the examples. DETAILED DESCRIPTION OF THE INVENTION

[0043] The technical solutions of the present invention will be further described in detail below in conjunction with specific embodiments. However, these embodiments are not intended to limit the present invention. Any similar structures and their similar changes adopted in the present invention should be included in the protection scope of the present invention. The commas in the present invention all represent the relationship of "and". The English letters in the present invention are case-sensitive.

[0044] The thermodynamic properties of RNA sequences are crucial for revealing the structure-function relationships of RNA. During RNA interference, the specific binding of small interfering RNAs (siRNAs) cleaved from double-stranded RNA (dsRNA) by dicer to target messenger ribonucleic acid (mRNA) is a complex thermodynamic process involving multiple energy changes, including the energy change for opening the target binding site, the energy released by the binding of siRNA to mRNA, etc. Therefore, analyzing the energy changes at the target positions on the mRNA of the target gene is of great significance for studying the binding efficiency of exogenous dsRNA / siRNA to the target positions.

[0045] As Figure 1 shown, the present invention provides a method for designing dsRNA / siRNA target sequences based on thermodynamic parameters, and the design method includes:

[0046] S1, performing normalization processing on the coding sequence of the candidate gene; the normalization processing includes sequence inspection to analyze whether the coding sequence is symmetric; removing redundant information to ensure that the coding sequence only contains DNA sequences of A / G / C / T; completing case conversion to unify the coding sequence into an uppercase format; converting T of the DNA sequence to U; wherein, the coding sequence of the candidate gene is directly obtained by querying from databases such as NCBI;

[0047] S2, presetting the length of a reading frame as a and the translation step size as b, and dividing the coding sequence of the candidate gene by translating the reading frame to obtain multiple candidate dsRNA / siRNA target sequences; the specific steps are:

[0048] S21, on the full-length sequence of the candidate gene, setting a reading frame with a length of a starting from the m-th base, and the sequence contained in the reading frame is the initially generated candidate dsRNA / siRN target sequence; according to literature reports, generally the first 100 bases are considered unimportant for target design, so m can be selected as a position after 100.

[0049] S22, performing a translation operation on the reading frame, setting the translation step size as b, and each time after translation, the sequence contained in the reading frame is the new candidate dsRNA / siRNA target sequence, and continuously performing the translation operation until the reading frame can no longer move;

[0050] S23, summarizing all the generated candidate dsRNA / siRNA target sequences into a set.

[0051] S3, reading and analyzing the candidate dsRNA / siRNA target sequences based on the individual nearest neighbor (INN) model to obtain the base pairs of all candidate dsRNA / siRNA target sequences;

[0052] The specific steps to obtain base pairs are as follows: Define two adjacent bases as a base pair, and analyze them sequentially from the 5' to 3' direction of the candidate dsRNA / siRNA target sequence to obtain all possible base pairs; alternatively, two adjacent bases within the reading frame can be defined as a base pair; by shifting the reading frame, all base pairs can be obtained.

[0053] S4. Preset the thermodynamic parameters of all base pairs. Among them, AA = -0.93, UU = -0.93, AU = -1.10, UA = -1.33, CU = -2.08, AG = -2.08, CA = -2.11, UG = -2.11, GU = -2.24, AC = -2.24, GA = -2.35, UC = -2.35, CG = -2.36, GG = -3.26, CC = -3.26, GC = -3.42; Calculate the sum of the thermodynamic parameters of all base pairs, and combine the correction parameters to obtain the final thermodynamic parameter FE of the candidate target sequence: FE = G nn + G sym + G t1 + G t2 + G in ;

[0054] Among them, G nn is the sum of the base pair energy parameters calculated based on the nearest neighbor model, reflecting the energy contribution generated by the interaction of adjacent base pairs;

[0055] G sym is used to measure the structural symmetry of the candidate dsRNA target sequence. If the structure of the candidate dsRNA target sequence shows symmetric characteristics, G sym takes 0.43, otherwise takes 0;

[0056] G t1 is to judge the first base of the candidate dsRNA target sequence. If the first base is A or U, G t1 takes 0.45, otherwise takes 0;

[0057] G t2 is to judge the last base of the candidate dsRNA target sequence. If the last base is A or U, G t2 takes 0.45, otherwise takes 0;

[0058] G in is the compensation value of all dsRNA target sequences, used to appropriately correct and adjust the overall energy parameters.

[0059] In this way, the impacts of multiple factors on the energy of candidate sequences can be comprehensively considered, thus providing a more accurate basis for subsequent screening and evaluation.

[0060] S5 calculates the scores of each candidate dsRNA / siRNA target sequence according to the final thermodynamic parameters, sorts the scores, and determines one or more candidate dsRNA / siRNA target sequences with the highest scores as the optimal dsRNA / siRNA target sequences;

[0061] The specific scoring method is as follows: all candidate dsRNA / siRNA target sequences are sorted in descending order according to the value of the final thermodynamic parameter FE. The score of the target sequence ranked 1st is 100, and the score of the target sequence ranked last is 50. The score calculation formula for the remaining target sequences is:

[0062] Score = 50 + (100 - 50) / (FE max - FE min ) * (fe - FE min );

[0063] Among them, FE max is the FE value of the target sequence ranked 1st, FE min is the FE value of the target sequence ranked last, and fe is the FE value of the candidate target for which the Score value is to be calculated.

[0064] The present invention also provides an application of the above dsRNA / siRNA target sequence design method based on thermodynamic parameters in RNA interference and in studying the binding of exogenous dsRNA / siRNA to target positions.

[0065] The present invention also provides an application of the above dsRNA / siRNA target sequence design method based on thermodynamic parameters in the design of RNA interference targets in the fields of animal and plant protection or pest and disease control.

[0066] The present invention also provides a computer program product for designing RNA target sequences based on thermodynamic parameters, which is used to execute the above design method; the structure of the Tar of the computer program consists of three parts: input data, execution program, and output data. The input data is the coding sequence of candidate bases; the execution program is used to process the input data. By presetting the length of a reading frame as a and the translation step size as b, the coding sequence of the candidate gene is segmented through the translation of the reading frame to obtain multiple candidate dsRNA / siRNA target sequences. Based on the individual nearest neighbor (INN) model, the candidate dsRNA / siRNA target sequences are read and analyzed to obtain the base pairs of all candidate dsRNA / siRNA target sequences; calculate the sum of the thermodynamic parameters of all base pairs, and combine the correction parameters to obtain the final thermodynamic parameters of the candidate target sequences; calculate the score of each candidate dsRNA / siRNA target sequence according to the thermodynamic parameters, sort the scores, and determine one or more candidate dsRNA / siRNA target sequences with the highest scores as the optimal dsRNA target sequences; the output data is used to output the optimal dsRNA / siRNA target sequences. The present invention also provides a computer system, including a memory, a processor, and a computer program stored on the memory, and the processor executes the above computer program product.

[0067] Example

[0068] To verify the effectiveness of the present invention, a dsRNA target sequence was designed using the present invention.

[0069] Target Design

[0070] In this embodiment, the gene of Nilaparvata lugens α-1,6-fucosyltransferase (FucT6) was selected as the research object, aiming to design RNA interference targets for genes related to the development of Nilaparvata lugens and verify the control efficiency of different scored targets against Nilaparvata lugens. Specifically, all possible target fragments of the FucT6 gene were comprehensively designed using the present invention, and corresponding scores were assigned to each target fragment. Subsequently, these target fragments were ranked in descending order according to the scores. One target fragment was selected from the ranking results at the front, middle, and rear positions respectively. Subsequently, experimental verification work on the RNA interference efficiency will be carried out for these 3 selected target fragments. Specific relevant information is shown in Table 1.

[0071] Table 1

[0072]

[0073] Three targets representing high, medium, and low scores were selected from all candidate targets, namely the high-score target T1 (ds1), the medium-score target T2 (ds21), and the low-score target T3 (ds13);

[0074] Synthesize the target dsRNA for subsequent RNA interference experiment verification. The specific target design information is as follows:

[0075] Target T1, sense strand sequence (SEQ ID NO.1):

[0076] AGCCAGAUCCUCAGACAGCGCAAAGAAUUUCACAAGCUUUCAAAGACCUUCAAAUGCUACAUCAGCAGAGCAUAGAGAUCAAUCAGCUUUUAUCUGACUACAGCGUUAGUGAUGCAACAUUCAGACAAGUGUUCAAAGACAUGGUUGCUAGAAAUUCUGACAAAAAUUUUCAAGGUGUCAAGCUGCCAUUCGGGAAGAAAUUGGAAGGCGAAAGCAGUAAUAUUGGCGAAUGGAGGCAAGAGUCAGGUUUGGACUAUGAGAAGGUGAGACGACGCCUGAAUAUGGGCGUGGAAGAGAUGUGGUACUUCAUCAGUUCCCAAAUAAAGAUGGCCAAAAAGCAAGCUGGCGACUCCGCGCCUCAUGUCUCCAAAUCGCUCGAAAAGAUUCUUGACGAGGGCCUCGAGCAUAAGAGGUGGCUUUUGAAAGACCUUGGCCACCUGGCGGAAGUGGAUGGCCACGCAAAAUGGCGGGAGAGAGAGGCCGAAGAUCUGUCUCGUCUU。

[0077] Antisense strand sequence (SEQ ID NO.2):

[0078] AAGACGAGACAGAUCUUCGGCCUCUCUCUCCCGCCAUUUUGCGUGGCCAUCCACUUCCGCCAGGUGGCCAAGGUCUUUCAAAAGCCACCUCUUAUGCUCGAGGCCCUCGUCAAGAAUCUUUUCGAGCGAUUUGGAGACAUGAGGCGCGGAGUCGCCAGCUUGCUUUUUGGCCAUCUUUAUUUGGGAACUGAUGAAGUACCACAUCUCUUCCACGCCCAUAUUCAGGCGUCGUCUCACCUUCUCAUAGUCCAAACCUGACUCUUGCCUCCAUUCGCCAAUAUUACUGCUUUCGCCUUCCAAUUUCUUCCCGAAUGGCAGCUUGACACCUUGAAAAUUUUUGUCAGAAUUUCUAGCAACCAUGUCUUUGAACACUUGUCUGAAUGUUGCAUCACUAACGCUGUAGUCAGAUAAAAGCUGAUUGAUCUCUAUGCUCUGCUGAUGUAGCAUUUGAAGGUCUUUGAAAGCUUGUGAAAUUCUUUGCGCUGUCUGAGGAUCUGGCU。

[0079] Target T2 (score rank 2), sense strand sequence (SEQ ID NO.3):

[0080] CCUCCGUCAGCUAUUGGCCGGUAGGUACGGCUUCGACCCAGGUGGUGCAGUUGCCGAUAAUCGACAGUGUGGCUCCUCGACCGCCGUACCUGCCGCAGUCGGUGCCGGCCGAUCUGGCCGACCGCAUCUCGCGGCUGCACGGAGCACCCUUUGUCUGGUGGGUCGGCCAGUUCUUUAAGUUCCUCCUCCGGCCGCAGCCGGCCACCGCCUCCAUGCUCAAUGAGACGGCCGCCAGGAUGAACUUCAAGCGACCCAUUGUCGGGGUGCACAUUCGGCGCACAGACAAGAUCGGCACUGAGGCGGCGUUCCACUCGGUCGACGAGUACAUGUCGAAAGUGGUCGAGUACUACGACCAGCUGGCACUUCUCGGCGGCUCGACUUCCCGACGCCGUGUCUAUCUGGCAUCAGACGAUCCCAAGGUAUAUCUUUGUGCAACUGGGUGUAAAUAUGGUGGUAUUGUCUGCAGGGUGGCGUACGAGAUAAUGCAGUCACUGCAUGUU。

[0081] Antisense strand sequence (SEQ ID NO.4):

[0082] AACAUGCAGUGACUGCAUUAUCUCGUACGCCACCCUGCAGACAAUACCACCAUAUUUACACCCAGUUGCACAAAGAUAUACCUUGGGAUCGUCUGAUGCCAGAUAGACACGGCGUCGGGAAGUCGAGCCGCCGAGAAGUGCCAGCUGGUCGUAGUACUCGACCACUUUCGACAUGUACUCGUCGACCGAGUGGAACGCCGCCUCAGUGCCGAUCUUGUCUGUGCGCCGAAUGUGCACCCCGACAAUGGGUCGCUUGAAGUUCAUCCUGGCGGCCGUCUCAUUGAGCAUGGAGGCGGUGGCCGGCUGCGGCCGGAGGAGGAACUUAAAGAACUGGCCGACCCACCAGACAAAGGGUGCUCCGUGCAGCCGCGAGAUGCGGUCGGCCAGAUCGGCCGGCACCGACUGCGGCAGGUACGGCGGUCGAGGAGCCACACUGUCGAUUAUCGGCAACUGCACCACCUGGGUCGAAGCCGUACCUACCGGCCAAUAGCUGACGGAGG。

[0083] Target T3 (ranked 3rd), sense strand sequence (SEQ ID NO.5):

[0084] ACACCUGCCUGGACCCCUCCGGCGCCUCCGUCAGCUAUUGGCCCGGUACGGCUUCGACCCAGGUGGUGCAGUUGCCGAUAAUCGACAGUGUGGCUCCUCGACCGCCGUACCUGCCGCAGUCGGUGCCGGCCGAUCUGGCCGACCGCAUCUCGCGGCUGCACGGAGCACCCUUUGUCUGGUGGGUCGGCCAGUUCUUUAAGUUCCUCCUCCGGCCGCAGCCGGCCACCGCCUCCAUGCUCAAUGAGACGGCCGCCAGGAUGAACUUCAAGCGACCCAUUGUCGGCAACAAUGUCAAUAUGCUGGGACCAACUAAGGGUUGCGGCUACGGGUGCCAACUGCACCACAUAGUGUACUGCAUGCUGGUGGCGUAUGGCACCGAACGCACCAUGGUGCUCAAGUCGAAAGGAUGGCGCUAUCACCGGGCUGGCUGGGAGGACGUCUAUCAGCCCGUCAGCGACACCUGCCUGGACCCCUCCGGCGCCUCCGUCAGCUAUUGGCCG。

[0085] Antisense strand sequence (SEQ ID NO.6):

[0086] CGGCCAAUAGCUGACGGAGGCGCCGGAGGGGUCCAGGCAGGUGUCGCUGACGGGCUGAUAGACGUCCUCCCAGCCAGCCCGGUGAUAGCGCCAUCCUUUCGACUUGAGCACCAUGGUGCGUUCGGUGCCAUACGCCACCAGCAUGCAGUACACUAUGUGGUGCAGUUGGCACCCGUAGCCGCAACCCUUAGUUGGUCCCAGCAUAUUGACAUUGUUGCCGACAAUGGGUCGCUUGAAGUUCAUCCUGGCGGCCGUCUCAUUGAGCAUGGAGGCGGUGGCCGGCUGCGGCCGGAGGAGGAACUUAAAGAACUGGCCGACCCACCAGACAAAGGGUGCUCCGUGCAGCCGCGAGAUGCGGUCGGCCAGAUCGGCCGGCACCGACUGCGGCAGGUACGGCGGUCGAGGAGCCACACUGUCGAUUAUCGGCAACUGCACCACCUGGGUCGAAGCCGUACCGGGCCAAUAGCUGACGGAGGCGCCGGAGGGGUCCAGGCAGGUGU。

[0087] Target dsRNA synthesis

[0088] For subsequent experiments, the total RNA of Nilaparvata lugens was extracted by the Trizol method. The specific operation is as follows: Collect about 5 third-instar nymphs of Nilaparvata lugens, place them in a RNase-free centrifuge tube, quickly freeze them in liquid nitrogen, and then grind them into powder using a ball mill. Add 1 mL of Trizol reagent to the ground sample and shake it well to mix. Then, add 0.2 mL of chloroform to the lysate and shake it vigorously to ensure thorough mixing. After that, place the centrifuge tube in a low-temperature centrifuge at 4°C and centrifuge it at 12,000 rpm for 15 min. After centrifugation, carefully aspirate the upper colorless aqueous phase and transfer it to a new RNase-free centrifuge tube. Subsequently, add an equal volume of isopropanol, gently invert the centrifuge tube to mix the solution, and promote the precipitation of RNA. Place the centrifuge tube at 4°C again and centrifuge it at 12,000 rpm for 10 min. Discard the supernatant and retain the RNA precipitate at the bottom of the tube. To remove residual impurities, wash the RNA precipitate with 75% ethanol. After the precipitate dries at room temperature, add 30 μL of DEPC water to dissolve it.

[0089] Using the extracted total RNA as a template, cDNA was synthesized using a reverse transcription kit (Vazyme, HiScript III All-in-one RT SuperMix Perfect for qPCR). The synthesized cDNA was used as a template, and Taq enzyme was used to perform PCR amplification of the target fragment. The PCR reaction program was set as follows: First, pre-denaturation was carried out at 95°C for 2 minutes, followed by the cycling stage. Each cycle included denaturation at 95°C for 15 seconds, annealing at 55°C for 30 seconds, and extension at 72°C for 30 seconds. A total of 35 cycles were performed, and finally, extension was carried out at 72°C for 5 minutes. Finally, using an in vitro transcription kit, with the PCR product as a template, dsRNA products were synthesized.

[0090] Verification of target effect

[0091] While synthesizing the candidate target dsRNA, dsRNA of green fluorescent protein (GFP) was designed and synthesized as a negative control. The concentration of all dsRNAs was measured using nanodrop, and their concentration was diluted to 1000 ng / μL using ultrapure water.

[0092] Third-instar nymphs of Nilaparvata lugens were selected as experimental subjects. After anesthesia with carbon dioxide, microinjection was performed in the area between the bases of the second pair of legs on the thorax of Nilaparvata lugens, and about 50 ng of dsRNA was introduced into the body of Nilaparvata lugens. After microinjection, the Nilaparvata lugens were left standing for 2 hours, and the individuals with the best vitality were selected and placed into glass test tubes and cultured using rice seedlings as a food source. The injected Nilaparvata lugens nymphs were evenly divided into two parts. Half was used for statistical analysis of survival, and the other half was sampled to extract RNA 24 hours after injection for gene knockdown analysis.

[0093] The survival of Nilaparvata lugens is as Figure 2 shown. Generally speaking, target T1 showed the best control effect. After 7 days of microinjection, the survival rate of Nilaparvata lugens was less than 20%; target T3 had the worst control effect on Nilaparvata lugens nymphs. After 7 days of microinjection, the survival rate was about 43.3%; the control effect of target T2 was between T1 and T3, and the survival rate of Nilaparvata lugens was about 33.3%.

[0094] Based on the above results of the microinjection experiment, the control effects of the three candidate targets on Nilaparvata lugens were highly consistent with the predictions of the target design algorithm. This result fully demonstrates that the algorithm proposed in the present invention can accurately predict the target effect based on the energy parameters of genes, thereby realizing the rapid design of highly efficient targets.

[0095] Gene knockdown analysis of dsRNA targets

[0096] To further verify the effect of the target double-stranded ribonucleic acid (dsRNA), about 5 brown planthoppers were collected and quickly frozen in liquid nitrogen 24 hours after the microinjection operation was completed. For each treatment, 3 biological replicates were set to ensure the reliability and repeatability of the experimental results. RNA in the samples was extracted using the Trizol method. After reverse transcription to synthesize cDNA, quantitative polymerase chain reaction (qPCR) technology was used to analyze gene expression.

[0097] The experimental results are as Figure 3 shown. Compared with the negative control, after injecting 3 target dsRNAs, significant knockdown of the FucT6 gene occurred in the brown planthoppers. Specifically, after injecting the dsRNA of target T1, the relative expression level of the FucT6 gene decreased to 9.3%; after injecting the dsRNA of target T2, the relative expression level was 21.1%; after injecting the dsRNA of target T3, the relative expression level was 27.9%.

[0098] Generally speaking, the knockdown effect of the target gene is highly consistent with the change trend of the survival rate of brown planthoppers after microinjection. This result fully indicates that the RNA interference mechanism was successfully activated after injecting dsRNA, leading to a significant decrease in the expression level of the target gene, and then causing the death of brown planthoppers, ultimately achieving effective control of brown planthoppers.

[0099] Based on the above experimental results, it can be seen that the 3 targets representing high, medium, and low efficiencies designed by the algorithm of the present invention can all effectively cause the death of brown planthopper nymphs through the RNA interference mechanism. In the microinjection and gene knockdown analysis experiments, the actual efficiencies of these 3 targets are highly consistent with the algorithm prediction results, strongly proving that the RNA interference target design algorithm proposed by the present invention can successfully and quickly complete the design of highly efficient target sequences.

[0100] Based on this algorithm, the RNA interference target sequence design program and method of the present invention show significant advantages. The traditional target sequence design work usually has to go through cumbersome batch design and a large number of screening and verification processes, which not only takes a long time but also consumes a large amount of resources. In sharp contrast, the present invention only takes a few seconds from inputting candidate sequences to completing target design, and can complete the design of the best dsRNA target sequence of candidate target genes in an extremely short time, greatly shortening the research cycle and saving the time cost required for screening highly efficient targets. This innovative achievement is of crucial significance for accelerating the research process of excellent targets for nucleic acid pesticides. It is expected to bring new technological breakthroughs and development opportunities to the field of agricultural pest control, and promote the development of this field towards a more efficient and precise direction.

[0101] Although the preferred embodiments of the present application have been described, additional changes and modifications can be made to these embodiments by those skilled in the art once they learn of the basic inventive concept. Therefore, the appended claims are intended to be construed to include the preferred embodiments as well as all changes and modifications that fall within the scope of the present application.

Claims

1. A dsRNA / siRNA target sequence design method based on thermodynamic parameters, characterized in that: The design method comprises: (1) Standardize the coding sequence of the candidate gene; standardization includes: sequence inspection, analysis of whether the coding sequence is symmetrical; removal of redundant information, ensuring that the coding sequence only contains DNA sequences of A / G / C / T; completion of uppercase and lowercase conversion, unifying the coding sequence into uppercase format; conversion of T in the DNA sequence to U; (2) Preset a reading frame length of a and a translation step length of b, and segment the coding sequence of the candidate gene by translation of the reading frame to obtain multiple candidate dsRNA / siRNA target sequences; (3) Read and analyze the candidate dsRNA / siRNA target sequences based on the individual nearest neighbor model to obtain the base pairs of all candidate dsRNA / siRNA target sequences; (4) Preset the thermodynamic parameters of all base pairs, calculate the sum of the thermodynamic parameters of all base pairs, and combine the correction parameters to obtain the final thermodynamic parameters of the candidate dsRNA / siRNA target sequence; (5) Calculate the score of each candidate dsRNA / siRNA target sequence based on the final thermodynamic parameters, sort the scores, and determine one or more candidate dsRNA / siRNA target sequences with the highest scores as the optimal dsRNA / siRNA target sequences.

2. The method for designing dsRNA / siRNA target sequences based on thermodynamic parameters according to claim 1, characterized in that: The specific steps of step (2) are: (21) On the full-length sequence of the candidate gene, a reading frame of length a is set starting from the mth base, and the sequence contained in the reading frame is the initially generated candidate dsRNA / siRNA target sequence; (22) The reading frame is translated, and the step length of the translation is set to b. Each time the translation is performed, the sequence contained in the reading frame becomes a new candidate dsRNA / siRNA target sequence. The translation operation is continued until the reading frame can no longer be moved; (23) All generated candidate dsRNA / siRNA target sequences were summarized into one set.

3. The method for designing dsRNA / siRNA target sequences based on thermodynamic parameters according to claim 1, characterized in that: In step (3), the specific steps of obtaining base pairs are: define two adjacent bases as a base pair, analyze the candidate dsRNA / siRNA target sequence from 5' to 3' direction, and obtain all possible base pairs.

4. The method for designing dsRNA / siRNA target sequences based on thermodynamic parameters according to claim 1, characterized in that: In step (3), the specific steps for obtaining base pairs are as follows: two adjacent bases in the reading frame are one base pair; and all base pairs are obtained by translation of the reading frame.

5. The method for designing dsRNA target sequences based on thermodynamic parameters according to claim 1, characterized in that: The calculation formula of the final thermodynamic parameter FE in step (4) is: FE= G nn + G sym + G t1 + G t2 + G in ; Among them, G nn It is the sum of base pair energy parameters calculated based on the nearest neighbor model; G sym It is used to measure the structural symmetry of the candidate dsRNA target sequence. If the structure of the candidate dsRNA target sequence presents symmetric characteristics, G sym Take 0.43, otherwise take 0; G t1 The first base of the candidate dsRNA target sequence is used for judgment. If the first base is A or U, G t1 Take 0.45, otherwise take 0; G t2 The last base of the candidate dsRNA target sequence is used for judgment. If the last base is A or U, G t2 Take 0.45, otherwise take 0; G in is the compensation value for all dsRNA target sequences and is used to make appropriate corrections and adjustments to the overall energy parameters.

6. The method for designing dsRNA target sequences based on thermodynamic parameters according to claim 1, characterized in that: The specific scoring method of step (5) is: All candidate dsRNA / siRNA target sequences are ranked from high to low according to the value of the final thermodynamic parameter FE. The score of the first-ranked target sequence is 100, and the score of the last-ranked target sequence is 50. The score calculation formula for the remaining target sequences is: Score = 50 + (100 - 50) / (FE max - FE min ) * (fe - FE min ); Among them, FE max is the FE value of the first-ranked target sequence, FE min is the FE value of the last ranked target sequence, and fe is the FE value of the candidate target for which the Score value is to be calculated.

7. The method for designing dsRNA / siRNA target sequences based on thermodynamic parameters according to any one of claims 1 to 6 is used in RNA interference and in studying the binding of exogenous dsRNA / siRNA to target sites.

8. Application of the dsRNA / siRNA target sequence design method based on thermodynamic parameters as described in any one of claims 1 to 6 in RNA interference target design in the field of animal and plant protection or the field of pest control.

9. A computer program product for RNA target sequence design based on thermodynamic parameters, characterized in that: The computer program product is used to execute the design method described in any one of claims 1 to 6; the Tar structure of the computer program consists of three parts: input data, execution program, and output data.

10. The computer program product for RNA target sequence design based on thermodynamic parameters according to claim 9, characterized in that: The input data is the coding sequence of the candidate gene; the execution program is used to process the input data, by presetting a reading frame length a and a translation step length b, the coding sequence of the candidate gene is segmented by translation of the reading frame to obtain multiple candidate dsRNA / siRNA target sequences, and the candidate dsRNA / siRNA target sequences are read and analyzed based on the individual nearest neighbor model to obtain the base pairs of all candidate dsRNA / siRNA target sequences; the final thermodynamic parameters of the candidate target sequences are obtained by calculating the sum of the thermodynamic parameters of all base pairs and combining the correction parameters; the score of each candidate dsRNA / siRNA target sequence is calculated according to the thermodynamic parameters, and the scores are sorted, and one or more candidate dsRNA / siRNA target sequences with the highest scores are determined as the optimal dsRNA / siRNA target sequences; the output data is used to output the optimal dsRNA / siRNA target sequence.

11. A computer system comprising a memory, a processor and a computer program stored in the memory, characterized in that: The processor executes the computer program product of claim 9 or 10.

Citation Information

Patent Citations

  • Novel method for forecasting siRNA interference efficiency based on ARM (Advanced RISC Machines) microprocessor

    CN103020489A

  • Method for screening siRNA on basis of bioinformatics

    CN104419702A

  • Rule weight distribution siRNA design method based on grid search

    CN112951322A

  • Method and system for screening siRNA sequences to reduce off-target effect

    CN116798513A

  • SiRNA inhibition rate prediction method based on antisense strand multi-dimensional embedding characterization

    CN116844644A

Cited By

  • Method and device for generating siRNA precursor double-stranded RNA sequence aiming at target organism and application of siRNA precursor double-stranded RNA sequence

    CN120877845A

  • Double-target dsRNA molecule for preventing and treating tomato brown wrinkled fruit virus and application of double-target dsRNA molecule

    CN122445648A

  • Double-target dsrna molecule for preventing and treating tomato brown rugose fruit virus and application thereof

    CN122445648B

  • Double-stranded RNA for synergistically controlling plutella xylostella with chemical insecticides and application thereof

    CN122629059A