dsRNA / siRNA target design method based on thermodynamic parameters, application and computer program product
Through the dsRNA target sequence design method based on thermodynamic parameters, the time-consuming and cost-effective design of dsRNA target sequence in the prior art is solved, fast and efficient target sequence screening is achieved, RNA interference efficiency is improved, and the accuracy and environmental protection of agricultural pest control are promoted.
Patent Information
- Application Number
- CN202510525341.9
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2025-04-25
- Publication Date
- 2025-08-19
- Estimated Expiration
- 2045-04-25
AI Technical Summary
In the prior art, dsRNA target sequence design depends on random methods, which is time-consuming and costly, making it difficult to quickly design efficient dsRNA target sequences.
The dsRNA target sequence design method based on thermodynamic parameters was used to screen out the optimal target sequence through standardized processing of candidate genes, reading frame translation, individual nearest neighbor model analysis and thermodynamic parameter calculation.
Rapidly and accurately design efficient dsRNA target sequences, shorten research cycles, save resources, and improve RNA interference efficiency, and is suitable for agricultural pest control.
Smart Images

Figure CN120126561B_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the field of bioinformatics, and in particular to an RNA target design method based on thermodynamic parameters, an application thereof, and a computer program product. Background Art
[0002] RNA interference is a highly conserved gene silencing mechanism in eukaryotes. The development of RNA interference technology has undergone a profound transformation from basic research to practical application. In the agricultural field, RNA interference technology has shown great application potential as an innovative pest management tool. By introducing specific dsRNA molecules into the pest body, the pest's key survival genes can be precisely targeted and gene silenced, thereby inhibiting the pest's survival or reproduction and effectively reducing the pest population. Because dsRNA molecules degrade rapidly in the environment, the potential impact on non-target organisms and ecosystems can be significantly reduced. This type of RNA interference-based biopesticide development has the characteristics of high specificity, low toxicity and environmental friendliness. In the future, agriculture is expected to achieve more precise and environmentally friendly pest control methods, reduce the use of chemical pesticides, improve the safety of agricultural products, and promote the sustainable development of agriculture.
[0003] Research has shown that the interference efficiency of double-stranded RNA (dsRNA) synthesized from different fragments targeting the same gene varies significantly. Numerous factors influence RNAi efficiency, including dsRNA fragment length, gene location, and base sequence characteristics. Currently, the design and verification of dsRNA target sequences mostly rely on random methods, a process that is time-consuming and costly. Therefore, how to rapidly design highly effective dsRNA target sequences has become a critical issue that needs to be addressed. Summary of the Invention
[0004] In order to solve the above technical problems, the present invention provides a dsRNA / siRNA target sequence design method based on thermodynamic parameters, the design method comprising:
[0005] (1) Standardize the coding sequence of the candidate gene; standardization includes: sequence inspection, analysis of whether the coding sequence is symmetrical; removal of redundant information, ensuring that the coding sequence contains only DNA sequences of A / G / C / T; completion of uppercase and lowercase conversion, unifying the coding sequence into uppercase format; conversion of T in the DNA sequence to U; the coding sequence of the candidate gene is directly obtained from databases such as NCBI;
[0006] (2) Preset a reading frame length a and a translation step length b, and segment the coding sequence of the candidate gene by translation of the reading frame to obtain multiple candidate dsRNA / siRNA target sequences;
[0007] (3) Read and analyze the candidate dsRNA / siRNA target sequences based on the individual nearest neighbor (INN) model to obtain the base pairs of all candidate dsRNA / siRNA target sequences;
[0008] (4) Preset the thermodynamic parameters of all base pairs, calculate the sum of the thermodynamic parameters of all base pairs, and combine the correction parameters to obtain the final thermodynamic parameters of the candidate target sequence;
[0009] (5) Calculate the score of each candidate dsRNA / siRNA target sequence based on the final thermodynamic parameters, sort the scores, and determine one or more candidate dsRNA / siRNA target sequences with the highest scores as the optimal dsRNA / siRNA target sequences.
[0010] Furthermore, the specific steps of step (2) are:
[0011] (21) On the full-length sequence of the candidate gene, a reading frame of length a is set starting from the mth base. The sequence contained in the reading frame is the initially generated candidate dsRNA / siRNA target sequence. According to literature reports, it is generally believed that the first 100 bases are not important for target design, so m can be selected from the number of digits after 100.
[0012] (22) The reading frame is translated, and the step length of the translation is set to b. Each time the translation is performed, the sequence contained in the reading frame becomes a new candidate dsRNA / siRNA target sequence. The translation operation is continued until the reading frame can no longer be moved;
[0013] (23) All generated candidate dsRNA / siRNA target sequences were aggregated into a set.
[0014] Furthermore, the specific steps of obtaining base pairs are: defining two adjacent bases as a base pair, and analyzing the candidate dsRNA / siRNA target sequence in the 5' to 3' direction to obtain all possible base pairs.
[0015] Furthermore, the specific steps for obtaining base pairs are as follows: two adjacent bases in a reading frame are a base pair; and all base pairs are obtained by translation of the reading frame.
[0016] Furthermore, the calculation formula of the final thermodynamic parameter FE in step (4) is:
[0017] FE=G nn + G sym + G t1 + G t2 + G in ;
[0018] Among them, Gnn It is the sum of the base pair energy parameters calculated based on the nearest neighbor model, reflecting the energy contribution generated by the interaction between adjacent base pairs;
[0019] G sym It is used to measure the structural symmetry of the candidate dsRNA target sequence. If the structure of the candidate dsRNA target sequence presents symmetric characteristics, G sym Take 0.43, otherwise take 0;
[0020] G t1 It is based on the first base of the candidate dsRNA target sequence. If the first base is A or U, G t1 Take 0.45, otherwise take 0;
[0021] G t2 It is based on the last base of the candidate dsRNA target sequence. If the last base is A or U, G t2 Take 0.45, otherwise take 0;
[0022] G in It is the compensation value for all dsRNA target sequences and is used to make appropriate corrections and adjustments to the overall energy parameters.
[0023] In this way, the impact of multiple factors on the energy of candidate sequences can be comprehensively considered, thereby providing a more accurate basis for subsequent screening and evaluation.
[0024] Furthermore, the specific scoring method of step (5) is:
[0025] All candidate dsRNA / siRNA target sequences are ranked from high to low according to the value of the final thermodynamic parameter FE. The first-ranked target sequence has a score of 100, and the last-ranked target sequence has a score of 50. The scores of the remaining target sequences are calculated as follows:
[0026] Score = 50 + (100 - 50) / (FE max -FE min ) * (fe - FE min );
[0027] Among them, FE max is the FE value of the first-ranked target sequence, FE min is the FE value of the last-ranked target sequence, and fe is the FE value of the candidate target for which the Score value is to be calculated.
[0028] The second aspect of the present invention provides a dsRNA target sequence design method based on thermodynamic parameters for use in RNA interference and the study of binding of exogenous dsRNA / siRNA to target sites.
[0029] The third aspect of the present invention provides an application of a dsRNA target sequence design method based on thermodynamic parameters in the design of RNA interference targets in the field of animal and plant protection or the field of pest and disease control.
[0030] The fourth aspect of the present invention provides a computer program product for dsRNA target sequence design based on thermodynamic parameters, which is used to execute the above-mentioned design method; the Tar structure of the computer program consists of three parts: input data, execution program, and output data.
[0031] Furthermore, the input data is the coding sequence of the candidate gene; the execution program is used to process the input data, by presetting a reading frame length a and a translation step length b, segmenting the coding sequence of the candidate gene by translation of the reading frame to obtain multiple candidate dsRNA / siRNA target sequences, reading and analyzing the candidate dsRNA / siRNA target sequences based on an individual nearest neighbor (INN) model to obtain base pairs of all candidate dsRNA target sequences; calculating the sum of the thermodynamic parameters of all base pairs, and combining the correction parameters to obtain the final thermodynamic parameters of the candidate target sequence; calculating the score of each candidate dsRNA / siRNA target sequence based on the final thermodynamic parameters, and sorting the scores, and determining one or more candidate dsRNA / siRNA target sequences with the highest scores as the optimal dsRNA / siRNA target sequence; the output data is used to output the optimal dsRNA / siRNA target sequence.
[0032] A fifth aspect of the present invention provides a computer system comprising a memory, a processor, and a computer program stored in the memory, wherein the processor executes the computer program product.
[0033] The present invention has the following beneficial effects:
[0034] (1) The present invention completes the target design work by presetting a reading frame length a and a translation step length b, and traversing the entire coding sequence of the candidate gene one by one, thereby ensuring that the generated candidate target can completely cover all regions of the gene;
[0035] (2) The present invention uses the individual nearest neighbor model (INN model) to conduct thermodynamic analysis on candidate target sequences, preset the thermodynamic parameters of all base pairs, and combine the correction parameters to obtain the final thermodynamic parameters of the candidate target sequence to achieve scientific evaluation and screening of the candidate target sequence, and ultimately quickly design efficient dsRNA / siRNA interference target sequences;
[0036] (3) The scoring in the present invention is achieved by adding the sum of the base pair energy parameters calculated by the nearest neighbor model, the score for judging whether the structure of the dsRNA target sequence is symmetrical, the score for judging whether the first and last bases are A or U, and the compensation value of the dsRNA target sequence. This comprehensively considers the impact of multiple factors on the energy of the candidate sequence, ensures that the selected candidate target sequence has relative advantages in terms of thermodynamic properties, etc., thereby improving the efficiency of the candidate target and providing a more accurate basis for screening the optimal dsRNA / siRNA target sequence.
[0037] (4) The present invention is flexible in its operation mode. It can be run independently to meet the needs of researchers to directly design targets; it can also be called by other software to facilitate integration into more complex research processes, realizing data sharing and analysis expansion;
[0038] (5) The present invention directly performs thermodynamic energy calculations without relying on other complex calculation tools, which simplifies the research process and improves calculation efficiency;
[0039] (6) The present invention takes only a few seconds from candidate sequence input to target design completion, and can complete the optimal dsRNA / siRNA target sequence design of the candidate target gene in a short period of time, greatly shortening the research cycle and saving the time and resources required for the research of efficient targets. It provides an accelerated driving force for the research process of excellent targets for nucleic acid pesticides and is expected to bring new breakthroughs in the field of agricultural pest control. BRIEF DESCRIPTION OF THE DRAWINGS
[0040] Figure 1 This is a flow chart for dsRNA / siRNA target sequence design in the present invention.
[0041] Figure 2 It is a statistical diagram of the microinjection results of different targets in the examples.
[0042] Figure 3 This is a transcription level analysis diagram of the target gene FucT6 in the example. DETAILED DESCRIPTION
[0043] The technical solution of the present invention is further described in detail below in conjunction with specific embodiments, but this embodiment is not intended to limit the present invention. All similar structures and similar variations of the present invention should be included in the scope of protection of the present invention. The semicolons in the present invention represent the relationship of and, and the English letters in the present invention are case-sensitive.
[0044] The thermodynamic properties of RNA sequences are crucial for revealing the structure-function relationship of RNA. During RNA interference (RNAi), the specific binding of small interfering RNA (siRNA) to its target messenger RNA (mRNA) by dicer cleavage is a complex thermodynamic process involving multiple energy changes, including those required to open the target binding site and the energy released by siRNA binding to mRNA. Therefore, analyzing the energy changes at the target site on the target gene mRNA is crucial for studying the binding efficiency of exogenous dsRNA / siRNA to the target site.
[0045] like Figure 1 As shown, the present invention provides a dsRNA / siRNA target sequence design method based on thermodynamic parameters, the design method comprising:
[0046] S1, standardize the coding sequence of the candidate gene; standardization includes sequence inspection to analyze whether the coding sequence is symmetrical; remove redundant information to ensure that the coding sequence contains only DNA sequences of A / G / C / T; complete uppercase and lowercase conversion to unify the coding sequence into uppercase format; convert T in the DNA sequence to U; the coding sequence of the candidate gene is directly obtained from databases such as NCBI;
[0047] S2, presetting a reading frame length a and a translation step length b, segmenting the coding sequence of the candidate gene by translation of the reading frame to obtain multiple candidate dsRNA / siRNA target sequences; the specific steps are:
[0048] S21. On the full-length sequence of the candidate gene, a reading frame of length a is set starting from the m-th base. The sequence contained in the reading frame is the initially generated candidate dsRNA / siRN target sequence. According to literature reports, it is generally believed that the first 100 bases are not important for target design, so m can be selected from the number of digits after 100.
[0049] S22 performs a translation operation on the reading frame, and the step length of the translation is set to b. Each time the reading frame is translated, the sequence contained in the reading frame becomes a new candidate dsRNA / siRNA target sequence. The translation operation is continued until the reading frame cannot be moved any further.
[0050] S23, summarizing all generated candidate dsRNA / siRNA target sequences into a set.
[0051] S3, reads and analyzes the candidate dsRNA / siRNA target sequences based on the individual nearest neighbor (INN) model to obtain the base pairs of all candidate dsRNA / siRNA target sequences;
[0052] The specific steps for obtaining base pairs are: defining two adjacent bases as a base pair, analyzing the candidate dsRNA / siRNA target sequence from 5' to 3' direction, and obtaining all possible base pairs; alternatively, two adjacent bases within the reading frame can be defined as a base pair; and all base pairs can be obtained by translation of the reading frame.
[0053] S4, preset the thermodynamic parameters of all base pairs, where AA = -0.93, UU = -0.93, AU = -1.10, UA = -1.33, CU = -2.08, AG = -2.08, CA = -2.11, UG = -2.11, GU = -2.24,AC = -2.24, GA = -2.35, UC = -2.35, CG = -2.36, GG = -3.26, CC = -3.26, GC = -3.42; calculate the sum of the thermodynamic parameters of all base pairs and combine them with the correction parameters to obtain the final thermodynamic parameter FE of the candidate target sequence: FE = G nn + G sym + G t1 + G t2 + G in ;
[0054] Among them, G nn It is the sum of the base pair energy parameters calculated based on the nearest neighbor model, reflecting the energy contribution generated by the interaction between adjacent base pairs;
[0055] G sym It is used to measure the structural symmetry of the candidate dsRNA target sequence. If the structure of the candidate dsRNA target sequence presents symmetric characteristics, G sym Take 0.43, otherwise take 0;
[0056] G t1 It is based on the first base of the candidate dsRNA target sequence. If the first base is A or U, G t1 Take 0.45, otherwise take 0;
[0057] G t2 It is based on the last base of the candidate dsRNA target sequence. If the last base is A or U, G t2 Take 0.45, otherwise take 0;
[0058] G in It is the compensation value for all dsRNA target sequences and is used to make appropriate corrections and adjustments to the overall energy parameters.
[0059] In this way, the impact of multiple factors on the energy of candidate sequences can be comprehensively considered, thereby providing a more accurate basis for subsequent screening and evaluation.
[0060] S5: calculating a score for each candidate dsRNA / siRNA target sequence based on the final thermodynamic parameters, ranking the scores, and determining one or more candidate dsRNA / siRNA target sequences with the highest scores as the optimal dsRNA / siRNA target sequences;
[0061] The specific scoring method is: all candidate dsRNA / siRNA target sequences are ranked from high to low according to the value of the final thermodynamic parameter FE. The first-ranked target sequence scores 100, and the last-ranked target sequence scores 50. The scoring formula for the remaining target sequences is:
[0062] Score = 50 + (100 - 50) / (FE max -FE min ) * (fe - FE min );
[0063] Among them, FE max is the FE value of the first-ranked target sequence, FE min is the FE value of the last-ranked target sequence, and fe is the FE value of the candidate target for which the Score value is to be calculated.
[0064] The present invention also provides the application of the dsRNA / siRNA target sequence design method based on thermodynamic parameters in RNA interference and the study of the binding of exogenous dsRNA / siRNA to the target position.
[0065] The present invention also provides the application of the dsRNA / siRNA target sequence design method based on thermodynamic parameters in the design of RNA interference targets in the field of animal and plant protection or the field of pest and disease control.
[0066] The present invention also provides a computer program product for the aforementioned thermodynamic parameter-based RNA target sequence design, which is configured to execute the aforementioned design method. The Tar structure of the computer program comprises three parts: input data, an execution program, and output data. The input data is the coding sequence of a candidate base. The execution program processes the input data by presetting a reading frame length a and a translation step length b, segmenting the coding sequence of the candidate gene by translation of the reading frame to obtain multiple candidate dsRNA / siRNA target sequences. The candidate dsRNA / siRNA target sequences are read and analyzed based on an individual nearest neighbor (INN) model to obtain base pairs for all candidate dsRNA / siRNA target sequences. The sum of the thermodynamic parameters of all base pairs is calculated, and combined with correction parameters, the final thermodynamic parameters of the candidate target sequences are obtained. A score for each candidate dsRNA / siRNA target sequence is calculated based on the thermodynamic parameters, and the scores are ranked. The one or more candidate dsRNA / siRNA target sequences with the highest scores are determined as optimal dsRNA target sequences. The output data is used to output the optimal dsRNA / siRNA target sequence. The present invention also provides a computer system comprising a memory, a processor and a computer program stored in the memory, wherein the processor executes the computer program product.
[0067] Example
[0068] In order to verify the effectiveness of the present invention, a dsRNA target sequence was designed using the present invention.
[0069] Target design
[0070] This example uses the brown planthopper (Nilaparvata lugens) α-1,6-fucosyltransferase (FucT6) gene as a research target. The goal is to design RNAi targets against developmental genes in the brown planthopper and verify the control efficacy of targets with different scores. Specifically, using the present invention, all possible target segments of the FucT6 gene were comprehensively designed and each was assigned a score. These target segments were then ranked in descending order based on their scores. From the ranking results, one target segment was selected from the top, middle, and bottom positions. Subsequent experiments will be conducted to verify the RNAi efficacy of these three selected target segments. Detailed information is provided in Table 1.
[0071] Table 1
[0072]
[0073] Three targets representing high, medium, and low scores were selected from all candidate targets, namely, high-scoring target T1 (ds1), intermediate-scoring target T2 (ds21), and low-scoring target T3 (ds13);
[0074] Synthesize target dsRNA for subsequent RNA interference experiment verification. The specific target design information is as follows:
[0075] Target T1, positive chain sequence (SEQ ID NO.1):
[0076] AGCCAGAUCCUCAGACAGCGCAAAGAAUUUCACAAGCUUUCAAAGACCUUCAAAUGCUACAUCAGCAGAGCAUAGAGAUCAAUCAGCUUUUAUCUGACUACAGCGUUAGUGAUGCAACAUUCAGA CAAGUGUUCAAAGACAUGGUUGCUAGAAAUUCUGACAAAAAUUUUCAAGGUGUCAAGCUGCCAUUCGGGAAGAAAUUGGAAGGCGAAAGCAGUAAUAUUGGCGAAUGGAGGCAAGAGUCAGGUUU GGACUAUGAGAAGGUGAGACGACGCCUGAAUAUGGGCGUGGAAGAGAUGUGGUACUUCAUCAGUUCCCAAAUAAAGAUGGCCAAAAAGCAAGCUGGCGACUCCGCGCCUCAUGUCUCCAAAUCGC UCGAAAAGAUUCUUGACGAGGGCCUCGAGCAUAAGAGGUGGCUUUUGAAAGACCUUGGCCACCUGGCGGAAGUGGAUGGCCACGCAAAAUGGCGGGAGAGAGGCCGAAGAUCUGUCUCGUCUU.
[0077] Antisense strand sequence (SEQ ID NO.2):
[0078] AAGACGAGACAGAUCUUCGGCCUCUCUCUCCCGCCAUUUUGCGUGGCCAUCCACUUCCGCCAGGUGGCCAAGGUCUUUCAAAAGCCACCUCUUAUGCUCGAGGCCCUCGUCAAGAAUCUUUUCGAGCGAUUUGGAGACAUGAGGCGCGGAGUCGCCAGCUUGCUUUUUGGCCAUCUUUAUUUGGGAACUGAUGAAGUACCACAUCUCUUCCACGCCCAUAUUCAGGCGUCGUCUCACCUUCUCAUAGUCCAAACCUGACUCUUGCCUCCAUUCGCCAAUAUUACUGCUUUCGCCUUCCAAUUUCUUCCCGAAUGGCAGCUUGACACCUUGAAAAUUUUUGUCAGAAUUUCUAGCAACCAUGUCUUUGAACACUUGUCUGAAUGUUGCAUCACUAACGCUGUAGUCAGAUAAAAGCUGAUUGAUCUCUAUGCUCUGCUGAUGUAGCAUUUGAAGGUCUUUGAAAGCUUGUGAAAUUCUUUGCGCUGUCUGAGGAUCUGGCU。
[0079] Target T2 (ranked 2nd), sense strand sequence (SEQ ID NO.3):
[0080] CCUCCGUCAGCUAUUGGCCGGUAGGUACGGCUUCGACCCAGGUGGUGCAGUUGCCGAUAAUCGACAGUGUGGCUCCUCGACCGCCGUACCUGCCGCAGUCGGUGCCGGCCGAUCUGGCCGACCGCAUCUCGCGGCUGCACGGAGCACCCUUUGUCUGGUGGGUCGGCCAGUUCUUUAAGUUCCUCCUCCGGCCGCAGCCGGCCACCGCCUCCAUGCUCAAUGAGACGGCCGCCAGGAUGAACUUCAAGCGACCCAUUGUCGGGGUGCACAUUCGGCGCACAGACAAGAUCGGCACUGAGGCGGCGUUCCACUCGGUCGACGAGUACAUGUCGAAAGUGGUCGAGUACUACGACCAGCUGGCACUUCUCGGCGGCUCGACUUCCCGACGCCGUGUCUAUCUGGCAUCAGACGAUCCCAAGGUAUAUCUUUGUGCAACUGGGUGUAAAUAUGGUGGUAUUGUCUGCAGGGUGGCGUACGAGAUAAUGCAGUCACUGCAUGUU。
[0081] Antisense strand sequence (SEQ ID NO.4):
[0082] AACAUGCAGUGACUGCAUUAUCUCGUACGCCACCCUGCAGACAAUACCACCAUAUUUACACCCAGUUGCACAAAGAUAUACCUUGGGAUCGUCUGAUGCCAGAUAGACACGGCGUCGGGAAGUCGAGCCGCCGAGAAGUGCCAGCUGGUCGUAGUACUCGACCACUUUCGACAUGUACUCGUCGACCGAGUGGAACGCCGCCUCAGUGCCGAUCUUGUCUGUGCGCCGAAUGUGCACCCCGACAAUGGGUCGCUUGAAGUUCAUCCUGGCGGCCGUCUCAUUGAGCAUGGAGGCGGUGGCCGGCUGCGGCCGGAGGAGGAACUUAAAGAACUGGCCGACCCACCAGACAAAGGGUGCUCCGUGCAGCCGCGAGAUGCGGUCGGCCAGAUCGGCCGGCACCGACUGCGGCAGGUACGGCGGUCGAGGAGCCACACUGUCGAUUAUCGGCAACUGCACCACCUGGGUCGAAGCCGUACCUACCGGCCAAUAGCUGACGGAGG。
[0083] Target T3 (ranked 3rd), sense strand sequence (SEQ ID NO.5):
[0084] ACACCUGCCUGGACCCCUCCGGCGCCUCCGUCAGCUAUUGGCCCGGUACGGCUUCGACCCAGGUGGUGCAGUUGCCGAUAAUCGACAGUGUGGCUCCUCGACCGCCGUACCUGCCGCAGUCGGUGCCGGCCGAUCUGGCCGACCGCAUCUCGCGGCUGCACGGAGCACCCUUUGUCUGGUGGGUCGGCCAGUUCUUUAAGUUCCUCCUCCGGCCGCAGCCGGCCACCGCCUCCAUGCUCAAUGAGACGGCCGCCAGGAUGAACUUCAAGCGACCCAUUGUCGGCAACAAUGUCAAUAUGCUGGGACCAACUAAGGGUUGCGGCUACGGGUGCCAACUGCACCACAUAGUGUACUGCAUGCUGGUGGCGUAUGGCACCGAACGCACCAUGGUGCUCAAGUCGAAAGGAUGGCGCUAUCACCGGGCUGGCUGGGAGGACGUCUAUCAGCCCGUCAGCGACACCUGCCUGGACCCCUCCGGCGCCUCCGUCAGCUAUUGGCCG。
[0085] Antisense strand sequence (SEQ ID NO.6):
[0086] CGGCCAAUAGCUGACGGAGGCGCCGGAGGGGUCCAGGCAGGUGUCGCUGACGGGCUGAUAGACGUCCUCCCAGCCAGCCCGGUGAUAGCGCCAUCCUUUCGACUUGAGCACCAUGGUGCGUUCGG UGCCAUACGCCACCAGCAUGCAGUACACUAUGUGGUGCAGUUGGCACCCGUAGCCGCAACCCUUAGUUGGUCCCAGCAUAUUGACAUUGUUGCCGACAAUGGGUCGCUUGAAGUUCAUCCUGGCG GCCGUCUCAUUGAGCAUGGAGGCGGUGGCCGGCUGCGGCCGGAGGAGGAACUUAAAGAACUGGCCGACCCACCAGACAAAGGGGUCCUCCGUGCAGCCGCGAGAUGCGGUCGGCCAGAUCGGCCGG CACCGACUGCGGCAGGUACGGCGGUCGAGGAGCCACACUGUCGAUUAUCGGCAACUGCACCACCUGGGUCGAAGCCGUACCGGGCCAAUAGCUGACGGAGGCGCCGGAGGGGUCCAGGCAGGUGU.
[0087] Target dsRNA synthesis
[0088] For subsequent experiments, this study used the Trizol method to extract total RNA from brown planthoppers. The procedure was as follows: approximately five third-instar nymphs of brown planthoppers were collected and placed in an RNase-free centrifuge tube. They were quickly frozen in liquid nitrogen and then ground into a powder using a ball mill. 1 mL of Trizol reagent was added to the ground sample and thoroughly mixed. Next, 0.2 mL of chloroform was added to the lysate and the solution was vigorously shaken to ensure thorough mixing. The tube was then centrifuged at 12,000 rpm for 15 minutes at 4°C. After centrifugation, the colorless upper aqueous phase was carefully aspirated and transferred to a fresh RNase-free centrifuge tube. An equal volume of isopropanol was then added, and the tube was gently inverted to mix the solution and induce RNA precipitation. The tube was again centrifuged at 12,000 rpm at 4°C for 10 minutes. The supernatant was discarded, retaining the RNA pellet at the bottom of the tube. To remove residual impurities, the RNA pellet was washed with 75% ethanol. After the precipitate was dried at room temperature, 30 μL of DEPC water was added to dissolve it.
[0089] cDNA was synthesized using the extracted total RNA as a template using a reverse transcription kit (HiScript III All-in-one RT SuperMix Perfect for qPCR). The synthesized cDNA was used as a template for PCR amplification of the target fragment using Taq enzyme. The PCR reaction program was as follows: initial denaturation at 95°C for 2 minutes, followed by 35 cycles of denaturation at 95°C for 15 seconds, annealing at 55°C for 30 seconds, and extension at 72°C for 30 seconds, with a final extension at 72°C for 5 minutes. Finally, dsRNA was synthesized using an in vitro transcription kit using the PCR product as a template.
[0090] Target effect verification
[0091] While synthesizing candidate target dsRNAs, a dsRNA encoding green fluorescent protein (GFP) was designed as a negative control. All dsRNA concentrations were determined using a nanodrop and diluted to 1000 ng / μL using ultrapure water.
[0092] Third-instar nymphs of brown planthoppers were used as experimental subjects. After carbon dioxide anesthesia, microinjection of approximately 50 ng of dsRNA was performed between the bases of the second pair of legs on the thorax. After microinjection, the nymphs were left to rest for 2 hours. The most vigorous individuals were selected and placed in glass test tubes and cultured with rice seedlings as a food source. The injected nymphs were divided equally into two groups: one half was used for survival analysis, and the other half was sampled 24 hours after injection for RNA extraction and gene knockdown analysis.
[0093] The survival of brown planthoppers Figure 2 Overall, target T1 showed the best control effect, with a survival rate of less than 20% of brown planthoppers 7 days after microinjection; target T3 had the worst control effect on brown planthopper nymphs, with a survival rate of approximately 43.3% 7 days after microinjection; target T2 had a control effect between T1 and T3, with a brown planthopper survival rate of approximately 33.3%.
[0094] Based on the microinjection results, the control efficacy of the three candidate targets against brown planthoppers was highly consistent with the predictions of the target design algorithm. This result fully demonstrates that the algorithm proposed in this paper can accurately predict target efficacy based on gene energy parameters, thereby enabling the rapid design of efficient targets.
[0095] Gene knockdown analysis of dsRNA targets
[0096] To further validate the efficacy of the target double-stranded RNA (dsRNA), approximately five brown planthoppers were collected 24 hours after microinjection and quickly frozen in liquid nitrogen. Three biological replicates were performed for each treatment to ensure the reliability and reproducibility of the experimental results. RNA from the samples was extracted using Trizol, reverse-transcribed into cDNA, and gene expression was analyzed using quantitative polymerase chain reaction (qPCR).
[0097] The experimental results are as follows Figure 3 As shown in the figure, compared with the negative control, injection of the three target dsRNAs significantly knocked down the FucT6 gene in brown planthoppers. Specifically, after injection of target T1 dsRNA, the relative expression level of the FucT6 gene dropped to 9.3%; after injection of target T2 dsRNA, the relative expression level was 21.1%; and after injection of target T3 dsRNA, the relative expression level was 27.9%.
[0098] Overall, the knockdown effect of the target gene was highly consistent with the survival trend of the brown planthopper after microinjection. This result fully demonstrates that the dsRNA injection successfully activated the RNA interference mechanism, causing a significant decrease in the expression level of the target gene, which in turn led to the death of the brown planthopper and ultimately achieved effective control of the brown planthopper.
[0099] The above experimental results show that the three targets designed using the algorithm of this invention, representing high, medium, and low efficiencies, all effectively kill brown planthopper nymphs through the RNAi mechanism. In microinjection and gene knockdown assays, the actual efficiencies of these three targets closely matched the algorithm's predictions, strongly demonstrating that the proposed RNAi target design algorithm can successfully and rapidly design highly effective target sequences.
[0100] Based on this algorithm, the RNA interference target sequence design program and method of the present invention show significant advantages. Traditional target sequence design work usually has to go through tedious batch design and a large number of screening and verification processes, which is not only time-consuming but also consumes a lot of resources. In sharp contrast, the present invention only takes a few seconds from inputting the candidate sequence to completing the target design. It can complete the optimal dsRNA target sequence design of the candidate target gene in a very short time, greatly shortening the research cycle and saving the time cost required to screen efficient targets. This innovative achievement is of vital importance to accelerating the research process of excellent targets for nucleic acid pesticides. It is expected to bring new technological breakthroughs and development opportunities to the field of agricultural pest control, and promote the development of this field in a more efficient and precise direction.
[0101] Although the preferred embodiments of the present application have been described, those skilled in the art may make additional changes and modifications to these embodiments once they have learned the basic creative concept. Therefore, the appended claims are intended to be interpreted as including the preferred embodiments and all changes and modifications that fall within the scope of the present application.
Claims
1. A method for designing dsRNA / siRNA target sequences based on thermodynamic parameters, characterized in that: The design method includes: (1) Standardize the coding sequence of the candidate gene; standardization includes: sequence inspection, analysis of whether the coding sequence is symmetrical; removal of redundant information, ensuring that the coding sequence contains only DNA sequences of A / G / C / T; completion of uppercase and lowercase conversion, unifying the coding sequence into uppercase format; conversion of T in the DNA sequence to U; (2) Preset a reading frame length a and a translation step length b, and segment the coding sequence of the candidate gene by translation of the reading frame to obtain multiple candidate dsRNA / siRNA target sequences; (3) Read and analyze the candidate dsRNA / siRNA target sequences based on the individual nearest neighbor model to obtain the base pairs of all candidate dsRNA / siRNA target sequences; (4) Preset the thermodynamic parameters of all base pairs, calculate the sum of the thermodynamic parameters of all base pairs, and combine the correction parameters to obtain the final thermodynamic parameters of the candidate dsRNA / siRNA target sequence; The calculation formula of the final thermodynamic parameter FE is: FE= G nn + G sym + G t1 + G t2 + G in ; Among them, G nn It is the sum of the base pair energy parameters calculated based on the nearest neighbor model; G sym It is used to measure the structural symmetry of the candidate dsRNA target sequence. If the structure of the candidate dsRNA target sequence presents symmetric characteristics, G sym Take 0.43, otherwise take 0; G t1 It is based on the first base of the candidate dsRNA target sequence. If the first base is A or U, G t1 Take 0.45, otherwise take 0; G t2 It is based on the last base of the candidate dsRNA target sequence. If the last base is A or U, G t2 Take 0.45, otherwise take 0; G in It is the compensation value of all dsRNA target sequences, which is used to make appropriate corrections and adjustments to the overall energy parameters; (5) Calculate the score of each candidate dsRNA / siRNA target sequence based on the final thermodynamic parameters, sort the scores, and determine one or more candidate dsRNA / siRNA target sequences with the highest scores as the optimal dsRNA / siRNA target sequences.
2. The method for designing dsRNA / siRNA target sequences based on thermodynamic parameters according to claim 1, wherein: The specific steps of step (2) are: (21) On the full-length sequence of the candidate gene, a reading frame of length a is set starting from the mth base, and the sequence contained in the reading frame is the initially generated candidate dsRNA / siRNA target sequence; (22) The reading frame is translated, and the step length of the translation is set to b. Each time the translation is performed, the sequence contained in the reading frame becomes a new candidate dsRNA / siRNA target sequence. The translation operation is continued until the reading frame can no longer be moved; (23) All generated candidate dsRNA / siRNA target sequences were aggregated into a set.
3. The method for designing dsRNA / siRNA target sequences based on thermodynamic parameters according to claim 1, wherein: In step (3), the specific steps for obtaining base pairs are as follows: define two adjacent bases as a base pair, analyze the candidate dsRNA / siRNA target sequence from 5' to 3' direction, and obtain all possible base pairs.
4. The method for designing dsRNA / siRNA target sequences based on thermodynamic parameters according to claim 1, wherein: In step (3), the specific steps for obtaining base pairs are as follows: two adjacent bases in the reading frame are one base pair; all base pairs are obtained by translation of the reading frame.
5. The method for designing dsRNA / siRNA target sequences based on thermodynamic parameters according to claim 1, wherein: The specific scoring method for step (5) is: All candidate dsRNA / siRNA target sequences are ranked from high to low according to the value of the final thermodynamic parameter FE. The first-ranked target sequence has a score of 100, and the last-ranked target sequence has a score of 50. The scores of the remaining target sequences are calculated as follows: Score = 50 + (100 - 50) / (FE max - FE min ) * (fe - FE min ); Among them, FE max is the FE value of the first-ranked target sequence, FE min is the FE value of the last-ranked target sequence, and fe is the FE value of the candidate target for which the Score value is to be calculated.
6. Use of the dsRNA / siRNA target sequence design method based on thermodynamic parameters according to any one of claims 1 to 5 in RNA interference and the study of binding of exogenous dsRNA / siRNA to target sites.
7. Use of the dsRNA / siRNA target sequence design method based on thermodynamic parameters according to any one of claims 1 to 5 in RNA interference target design in the field of plant protection or pest control.
8. A computer program product for RNA target sequence design based on thermodynamic parameters, characterized in that: The computer program product is used to execute the design method described in any one of claims 1 to 5; the Tar structure of the computer program consists of three parts: input data, execution program, and output data.
9. The computer program product for RNA target sequence design based on thermodynamic parameters according to claim 8, characterized in that The input data is the coding sequence of the candidate gene; the execution program is used to process the input data, by presetting a reading frame length a and a translation step length b, segmenting the coding sequence of the candidate gene by translation of the reading frame to obtain multiple candidate dsRNA / siRNA target sequences, reading and analyzing the candidate dsRNA / siRNA target sequences based on an individual nearest neighbor model, and obtaining base pairs of all candidate dsRNA / siRNA target sequences; calculating the sum of thermodynamic parameters of all base pairs and combining them with correction parameters to obtain final thermodynamic parameters of the candidate target sequences; calculating a score for each candidate dsRNA / siRNA target sequence based on the thermodynamic parameters, sorting the scores, and determining one or more candidate dsRNA / siRNA target sequences with the highest scores as optimal dsRNA / siRNA target sequences; and the output data is used to output the optimal dsRNA / siRNA target sequence.
10. A computer system comprising a memory, a processor, and a computer program stored in the memory, wherein: The processor executes the computer program product according to claim 8 or 9.
Citation Information
Patent Citations
Novel method for forecasting siRNA interference efficiency based on ARM (Advanced RISC Machines) microprocessor
CN103020489A
Method for screening siRNA on basis of bioinformatics
CN104419702A