A microRNA prediction method and system based on transcriptome data
By using transcriptome data and non-redundant protein databases to screen non-protein coding sequences, combining existing microRNA mature sequences and constructing secondary structures, the accurate prediction of target biological microRNAs is achieved, solving the problem of inaccurate microRNA prediction in the existing technology, and providing a powerful tool for biological research.
Patent Information
- Application Number
- CN202311189496.7
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2023-09-15
- Publication Date
- 2025-05-23
- Estimated Expiration
- 2043-09-15
AI Technical Summary
The prior art is difficult to use transcriptomic data to make efficient microRNA accurate predictions, resulting in the discarding of non-protein coding sequences that are not annotated as protein-encoded genes.
By obtaining transcriptome data of the target organism at different growth stages, non-redundant protein databases are used to screen non-protein coding sequences, setting distance windows and intercepting windows to slide the microRNA predicted candidate precursor sequences, combining existing microRNA mature sequences for labeling, and constructing secondary structures to screen target microRNA precursor sequences.
Accurate microRNA prediction of target organisms at different growth stages has been achieved, providing a powerful tool for biological research and the exploration of gene regulation networks.
Smart Images

Figure CN117198409B_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the field of biotechnology, and in particular to a microRNA prediction method and system based on transcription data. Background Art
[0002] With the continuous deepening of biological research and the rapid development of high-throughput sequencing technology, transcriptome data plays an increasingly important role in revealing gene regulatory networks, biological characteristics and disease mechanisms. At present, in transcriptome data analysis, protein-coding genes are usually studied, but most transcriptome data cannot be annotated as protein-coding genes. However, these non-protein coding sequences play an equally important role in the life activities of organisms. Since most species lack genomic data and there is little non-coding RNA information in public databases, non-protein coding sequences in transcriptome data that cannot be annotated as protein-coding genes are often discarded.
[0003] MicroRNA is a type of endogenous non-coding single-stranded RNA molecule with a length of 18 to 25 nt (nucleotides), which participates in post-transcriptional gene expression regulation in plants and animals. Currently, microRNA prediction is usually obtained from genomic data and small RNA libraries. There is little research on the process of predicting microRNA from transcriptome data, and there is no effective microRNA prediction based on transcriptome data. Therefore, there is an urgent need for a microRNA prediction method and system based on transcriptome data to accurately predict microRNA. Summary of the invention
[0004] In response to the shortcomings of current technology and the needs of practical applications, in the first aspect, the present invention provides a microRNA prediction method based on transcriptome data, which aims to use transcriptome data to accurately predict the microRNA of a target organism. The microRNA prediction method based on transcriptome data provided by the present invention includes the following steps: obtaining transcriptome data of the target organism at one or more growth stages; obtaining a non-redundant protein database, and using the non-redundant protein database to screen out non-protein coding sequences from the transcriptome data set; setting a distance window and a capture window, and using the length of the distance window as a sliding unit, using the capture window to slide on the non-protein coding sequence to capture multiple microRNA prediction candidate precursor sequences; obtaining an existing microRNA mature sequence, and combining the microRNA prediction candidate precursor sequence with the existing microRNA mature sequence to obtain a microRNA mature sequence. Marking; based on the existing microRNA mature body sequence, screen out the microRNA precursor sequence in the microRNA prediction candidate precursor sequence; construct the secondary structure of the microRNA precursor sequence, and obtain the minimum free energy and minimum free energy coefficient of the secondary structure; set the minimum free energy threshold and the minimum free energy coefficient threshold, and combine the minimum free energy and the minimum free energy coefficient to screen out the target microRNA precursor sequence from the microRNA precursor sequence; use the microRNA mature body sequence marker to match the target microRNA precursor sequence to obtain the microRNA mature sequence of the target organism at different growth stages. The present invention fully combines the known protein coding sequence and microRNA mature body sequence using transcriptome data to achieve accurate prediction of microRNA at different growth stages of the target organism, providing a powerful tool for biological research and exploration of gene regulatory networks.
[0005] Optionally, the target organism includes Plutella xylostella. This option is aimed at the agricultural pest Plutella xylostella. Accurately predicting its microRNA through the method of the present invention can bring new ideas and methods for controlling these pests, which is helpful to create a more environmentally friendly and sustainable agricultural production model.
[0006] Further optionally, the growth stage includes the diamondback moth egg stage, the diamondback moth larva stage, the diamondback moth pupa stage and the diamondback moth adult stage. This option combines the transcriptome data of the diamondback moth at different growth stages to accurately predict microRNAs, further bringing new ideas and methods for controlling these pests.
[0007] Optionally, the use of the non-redundant protein database to screen out non-protein coding sequences from the transcriptome data set includes the following steps: assembling the transcription data in the transcriptome data set to obtain multiple non-repetitive continuous sequences; comparing the known protein coding sequences in the non-redundant protein database with the non-repetitive continuous sequences to obtain the coding regions in the non-repetitive continuous sequences that are similar to the known protein coding sequences, and calculating the similarity between the sequences in the coding regions and the known protein coding sequences; setting a similarity threshold, and judging whether the non-repetitive continuous sequences are non-protein coding sequences or protein coding sequences by comparing the similarity with the similarity threshold. This option uses a non-redundant protein database to identify non-protein coding sequences from transcriptome data through sequence alignment and similarity calculation, ensuring the effective acquisition of potential microRNA precursor sequences and optimizing the accuracy and efficiency of the prediction method.
[0008] Optionally, the length of the distance window ranges from 18 nt to 25 nt, and the length of the interception window is at least 120 nt. The design of the length of the distance window and the interception window in this option helps to capture potential microRNA precursor sequences and improve the accuracy and comprehensiveness of the prediction method.
[0009] Further optionally, the length of the distance window is set to 25 nt, and the length of the interception window is set to 120 nt; the length of the distance window is used as a sliding unit, and the interception window is used to slide on the non-protein coding sequence to intercept multiple microRNA prediction candidate precursor sequences, and any microRNA prediction candidate precursor sequence satisfies the following counting model: L i (25i-24,25i+95), where i∈N * , N * represents a positive integer, N represents the total number of nucleotides contained in the non-protein coding sequence, L i (25i-24, 25i+95) represents the i-th microRNA prediction candidate precursor sequence obtained from the non-protein coding sequence, and the i-th microRNA prediction candidate precursor sequence includes the 25i-24th nucleotide to the 25i+95th nucleotide in the non-protein coding sequence. The setting of this option helps to efficiently capture the microRNA prediction candidate precursor sequence.
[0010] Optionally, the step of obtaining a microRNA mature sequence tag by combining the microRNA prediction candidate precursor sequence with the existing microRNA mature sequence comprises the following steps: aligning the existing microRNA mature sequence with the microRNA prediction candidate precursor sequence to obtain a microRNA prediction candidate precursor sequence having a comparison site; and marking a microRNA mature sequence similar to the comparison site sequence as a microRNA mature sequence tag according to the microRNA prediction candidate precursor sequence having a comparison site. This option compares the existing microRNA mature sequence with the microRNA prediction candidate precursor sequence, identifies the candidate precursor sequence having a comparison site, and marks the portion similar to the comparison site as a microRNA mature sequence, thereby adding a mature sequence tag to the prediction result, thereby improving the reliability and accuracy of the prediction. This option compares with the existing microRNA mature sequence, marks the microRNA mature sequence similar to the comparison site sequence, provides an accurate basis for the microRNA mature sequence tag, and improves the reliability of the prediction result.
[0011] Optionally, the obtaining of the minimum free energy and minimum free energy coefficient of the secondary structure comprises the following steps: constructing a minimum free energy model and a minimum free energy coefficient model respectively; using the minimum free energy model and the minimum free energy coefficient model to obtain the minimum free energy and the minimum free energy coefficient corresponding to the secondary structure respectively. This option uses the minimum free energy model and the minimum free energy coefficient model to calculate and obtain the minimum free energy and the minimum free energy coefficient of the secondary structure of the microRNA precursor sequence, providing a solid theoretical support for accurate microRNA prediction.
[0012] Optionally, the use of the microRNA mature body sequence marker to match the target microRNA precursor sequence to obtain the microRNA mature sequence of the target organism at different growth stages includes the following steps: using the microRNA mature body sequence corresponding to the microRNA mature body sequence marker to match the sequence of the stem region in the secondary structure of the target microRNA precursor sequence; and obtaining the microRNA mature sequence of the target organism at different growth stages based on the matching result. This option further enhances the credibility and accuracy of microRNA prediction by matching the microRNA mature body sequence marker with the stem region in the secondary structure of the target microRNA precursor sequence to obtain the microRNA mature sequence of the target organism at different growth stages.
[0013] In the second aspect, in order to better implement the above-mentioned microRNA prediction method based on transcriptome data, the present invention also provides a microRNA prediction system based on transcriptome data. The microRNA prediction system based on transcriptome data proposed by the present invention includes an input device, a processor, a memory and an output device, wherein the input device, the processor, the memory and the output device are interconnected, wherein the memory is used to store a computer program, the computer program includes program instructions, and the processor is configured to call the program instructions to execute the microRNA prediction method based on transcriptome data described in the first aspect of the present invention. The microRNA prediction system based on transcriptome data provided by the present invention stores and executes computer programs through the interconnection of input devices, processors, memories and output devices, effectively realizing the aforementioned microRNA prediction method based on transcriptome data, and providing a convenient and efficient tool for biological research and agricultural pest management. BRIEF DESCRIPTION OF THE DRAWINGS
[0014] Figure 1 A flow chart of a microRNA prediction method based on transcriptome data provided by an embodiment of the present invention;
[0015] Figure 2 This is a structural diagram of a microRNA prediction system based on transcriptome data provided in an embodiment of the present invention. DETAILED DESCRIPTION
[0016] The specific embodiments of the present invention will be described in detail below. It should be noted that the embodiments described herein are only for illustration and are not intended to limit the present invention. In the following description, a large number of specific details are set forth in order to provide a thorough understanding of the present invention. However, it is obvious to those of ordinary skill in the art that these specific details do not need to be adopted to implement the present invention. In other examples, in order to avoid confusing the present invention, known circuits, software or methods are not specifically described.
[0017] Throughout the specification, references to "one embodiment," "an embodiment," "an example," or "an example" mean that a particular feature, structure, or characteristic described in conjunction with the embodiment or example is included in at least one embodiment of the present invention. Therefore, the phrases "in one embodiment," "in an embodiment," "an example," or "an example" appearing in various places throughout the specification do not necessarily all refer to the same embodiment or example. In addition, particular features, structures, or characteristics may be combined in one or more embodiments or examples in any suitable combination and / or subcombination. In addition, it should be understood by those of ordinary skill in the art that the figures provided herein are for illustrative purposes and that the figures are not necessarily drawn to scale.
[0018] In an alternative embodiment, see Figure 1 , Figure 1 Flow chart of the microRNA prediction method based on transcriptome data provided by the embodiment of the present invention. Figure 1 As shown, the microRNA prediction method based on transcriptome data comprises the following steps:
[0019] S01. Obtain transcriptome data of a target organism at one or more growth stages.
[0020] The target organism described in the present invention refers to the organism concerned by the present invention, which can be an animal, a plant or other organism. It can be understood that through the microRNA prediction method based on transcriptome data described in the present invention, the expression of microRNA at different growth stages can be understood based on a specific target organism.
[0021] The growth stage refers to a specific period in the life cycle of the target organism, that is, different developmental stages from embryonic development to maturity. It is understandable that different growth stages of an organism are often accompanied by changes in gene expression, including the expression of microRNA.
[0022] The transcription data refers to the collection of RNA transcription products of all genes of the target organism at a specific growth stage. Transcriptome data can be obtained by high-throughput sequencing technology (such as RNA-Seq), and transcriptome data shows the transcription level of each gene in a specific growth stage.
[0023] In order to predict the expression preference of microRNA of diamondback moth at different growth stages and provide new ideas for effectively preventing and controlling the population outbreak of diamondback moth, in an optional embodiment, the target organism of concern is diamondback moth.
[0024] Further, in this embodiment, the growth stages of interest include the diamondback moth egg stage (Egg), the diamondback moth larva stage (Larva), the diamondback moth pupa stage (Pupa) and the diamondback moth adult stage (Adult).
[0025] In this embodiment, the step S01 of obtaining transcriptome data of a target organism at one or more growth stages includes the following steps:
[0026] S011. Collect samples of the target organism at one or more growth stages, and extract RNA information of the target organism at different growth stages based on the samples at different growth stages.
[0027] Step S011 collects samples of diamondback moth at four growth stages: Egg, Larva, Pupa and Adult; grinds the samples at the four stages using liquid nitrogen; and extracts total RNA of diamondback moth from the grinded samples at different growth stages using RNA extraction reagent Trizol.
[0028] S012. Constructing a total RNA library according to the RNA information, and obtaining a transcriptome data set by combining the total RNA library with a high-throughput sequencing technology.
[0029] Since RNA is easily degraded during analysis and experiments, while DNA is more stable and easier to handle, RNA is subjected to ligation reaction and reverse transcription to generate corresponding cDNA (complementary DNA) for subsequent analysis.
[0030] It is understood that the cDNA is a DNA molecule synthesized by a reverse transcription process, and its sequence is complementary to the corresponding part of the RNA molecule. Further, the high-throughput sequencing technology is RNA sequencing (RNA-Seq) technology.
[0031] It should be understood that through steps S011 to S012, the transcriptome data of Plutella xylostella at four growth stages, namely Egg, Larva, Pupa and Adult, can be obtained, and the transcriptome data of any growth stage includes protein coding sequence information and non-protein coding sequence information.
[0032] In one or some other embodiments, the step S01 of obtaining transcriptome data of the target organism at one or more growth stages can be completed through an existing database.
[0033] Further, in a specific embodiment, the transcriptome data of Plutella xylostella at four growth stages, namely Egg, Larva, Pupa and Adult, were downloaded from the SRA database (Sequence Read Archive) of NCBI (National Center for Biotechnology Information), and their numbers are SRR179062, SRR179508, SRR179509 and SRR179510, respectively.
[0034] S02. Obtain a non-redundant protein database, and use the non-redundant protein database to screen out non-protein coding sequences from the transcriptome data set.
[0035] The non-redundant protein database described in the present invention refers to a database that stores known protein coding sequences from various biological studies, for example, the nr database (Non-Redundant Protein Database) of NCBI.
[0036] Further, based on the transcriptome data of Plutella xylostella at four growth stages of Egg, Larva, Pupa and Adult from the SRA database (Sequence Read Archive) of NCBI in the above embodiment, the step S02 of using the non-redundant protein database to screen out non-protein coding sequences from the transcriptome data set includes the following steps:
[0037] S021. Assemble the transcription data in the transcriptome data set to obtain multiple non-repetitive continuous sequences.
[0038] In step S021, Trinity software may be used to assemble sequences in the transcription data, wherein Trinity is an open source software for transcriptome assembly, and is used to reconstruct transcripts of multiple genes from RNA-Seq (transcriptome sequencing) data.
[0039] In this example, the Trinity software (default parameters) was used to assemble the transcriptome data of the four growth stages, respectively, to obtain a plurality of non-repetitive continuous sequences.
[0040] S022. Compare the known protein coding sequence in the non-redundant protein database with the non-repetitive continuous sequence, obtain the coding region in the non-repetitive continuous sequence that is similar to the known protein coding sequence, and calculate the similarity between the sequence in the coding region and the known protein coding sequence.
[0041] In step S022, Blastx software may be used to compare the assembled multiple non-repetitive sequences with the non-redundant protein database (nr database) respectively; wherein Blastx is a sequence similarity search tool based on comparison, which is used to compare nucleic acid sequences with protein databases to determine the similarity and matching degree between sequences.
[0042] In this embodiment, in addition to using Blastx software to compare the assembled non-repetitive sequence with the non-redundant protein database (nr database), Blastx software is also used to identify the coding region in the non-repetitive continuous sequence that is similar to the known protein coding sequence, and Blastx software is used to calculate the similarity between the sequence in the coding region and the known protein coding sequence.
[0043] S023. Setting a similarity threshold, and determining whether the non-repetitive continuous sequence is a non-protein coding sequence or a protein coding sequence by comparing the similarity with the similarity threshold.
[0044] In this example, based on the comparison results obtained by the above-mentioned Blastx software, the similarity threshold is set to 0.00001 to determine whether the non-repetitive continuous sequence is a non-protein coding sequence or a protein coding sequence:
[0045] When the similarity between the sequence in the coding region and the known protein coding sequence is greater than or equal to 0.00001, the non-repetitive continuous sequence corresponding to the coding region is the protein coding sequence.
[0046] When the similarity between the sequence in the coding region and the known protein coding sequence is less than 0.00001, the non-repetitive continuous sequence corresponding to the coding region is a non-protein coding sequence.
[0047] In this embodiment, after step S021 to step S023, the number of sequences corresponding to each stage of this embodiment is shown in the following table:
[0048]
[0049] Among them, the data corresponding to the original sequence refers to the number of sequences in the diamondback moth transcriptome data set at different growth stages; the data corresponding to the non-repetitive sequence refers to the number of sequences obtained after the original sequence is assembled at different growth stages; the data corresponding to the annotated sequence refers to the number of known protein coding sequences in the non-redundant protein database at different growth stages; the data corresponding to the non-coding sequence refers to the number of non-protein coding sequences at different growth stages; the longest coding sequence refers to the number of nucleotides in the longest non-coding sequence at different growth stages; the shortest coding sequence refers to the number of nucleotides in the shortest non-protein coding sequence among the non-coding sequences at different growth stages; the data corresponding to the non-coding sequence ratio (%) refers to the ratio of non-coding sequences to non-repetitive sequences at different growth stages.
[0050] S03, setting a distance window and a capture window, and using the length of the distance window as a sliding unit, and using the capture window to slide on the non-protein coding sequence to capture multiple microRNA prediction candidate precursor sequences.
[0051] Since the length of microRNA is within the range of 18 to 25 nt, in order to ensure that the typical microRNA length range is covered during the scanning of non-protein coding sequences, the length range of the distance window includes 18 nt to 25 nt, and the length of the interception window is at least 120 nt.
[0052] In a specific embodiment, in step S03, the length of the distance window is set to 25 nt, and the length of the interception window is set to 120 nt. That is, with 25 nt as the sliding unit, the interception window is used to slide on the non-protein coding sequence to intercept multiple microRNA prediction candidate precursor sequences, and any microRNA prediction candidate precursor sequence satisfies the following counting model: L i (25i-24,25i+95), where i∈N * , N * represents a positive integer, represents N divided by 25 and rounded up, where N represents the total number of nucleotides contained in the non-protein coding sequence, and L i (25i-24, 25i+95) represents the i-th microRNA prediction candidate precursor sequence obtained from the non-protein coding sequence, and the i-th microRNA prediction candidate precursor sequence includes the 25i-24th nucleotide to the 25i+95th nucleotide in the non-protein coding sequence.
[0053] Furthermore, the candidate precursor sequences of microRNA prediction corresponding to any non-protein coding sequence of length N include: 1 (1,120), L 2 (26,145),…,L i (25i-24,25i+95),…, Where i∈N * , N * represents a positive integer, N represents the total number of nucleotides contained in the non-protein coding sequence, L i (25i-24,25i+95) represents the i-th microRNA prediction candidate precursor sequence obtained from the non-protein coding sequence, and the i-th microRNA prediction candidate precursor sequence includes the 25i-24th nucleotide to the 25i+95th nucleotide in the non-protein coding sequence. Among them, L(1,120) represents the first microRNA prediction candidate precursor sequence obtained from the non-protein coding sequence, and the first microRNA prediction candidate precursor sequence includes the 1st nucleotide to the 120th nucleotide in the non-protein coding sequence. L 2 (26,145) represents the second microRNA prediction candidate precursor sequence obtained from the non-protein coding sequence, wherein the second microRNA prediction candidate precursor sequence includes the 26th nucleotide to the 145th nucleotide in the non-protein coding sequence, ...,
[0054] L i(25i-24,25i+95) represents the i-th microRNA prediction candidate precursor sequence obtained from the non-protein coding sequence, wherein the i-th microRNA prediction candidate precursor sequence includes the 25i-24th nucleotide to the 25i+95th nucleotide in the non-protein coding sequence, ..., Indicates the first The candidate precursor sequence of microRNA prediction The candidate precursor sequences of microRNA prediction include the first nucleotides to the Nth nucleotide.
[0055] S04. Obtain an existing microRNA mature sequence, and obtain a microRNA mature sequence marker by combining the microRNA predicted candidate precursor sequence with the existing microRNA mature sequence.
[0056] The existing mature microRNA sequence of the present invention refers to a mature microRNA sequence that has been confirmed and disclosed in a database or literature. Further, the existing mature microRNA sequence can be obtained through a relevant database, for example, the miRBase database.
[0057] In an optional embodiment, the step S04 of obtaining a microRNA mature sequence tag by combining the microRNA predicted candidate precursor sequence with the existing microRNA mature sequence comprises the following steps:
[0058] S041. Compare the existing microRNA mature sequence with the microRNA predicted candidate precursor sequence to obtain a microRNA predicted candidate precursor sequence having a comparison site.
[0059] In this embodiment, Seqmap can be used to align the mature microRNA sequence with the predicted candidate microRNA precursor sequence. Seqmap (Sequence Mapping and Assembly Program) is a computational tool for sequence alignment and assembly, which is designed to quickly and accurately map and assemble sequences from high-throughput sequencing data.
[0060] S042. Predict candidate precursor sequences according to the microRNA having the alignment site, and mark the microRNA mature sequence similar to the alignment site sequence as a microRNA mature sequence marker.
[0061] In this example, the 1.0.13 version of Seqmap was used to align the mature microRNA sequence with the predicted candidate precursor microRNA sequence, and sequences with alignment sites were obtained according to the standards of 2_1_1 and 3_1_1 (mismatch_insertion_deletion), respectively, wherein the mature microRNA sequence was from the miRBase database.
[0062] S05. Screening out a microRNA precursor sequence from the microRNA prediction candidate precursor sequences according to the existing microRNA mature sequence.
[0063] In an optional embodiment, in step S05, triplet-SVM algorithm-related software may be used to identify characteristic sequences in the microRNA prediction candidate precursor sequence having the alignment site to determine whether the microRNA prediction candidate precursor sequence having the alignment site is a microRNA.
[0064] Furthermore, the triplet-SVM algorithm related software is a kind of software based on machine learning algorithm, which is an extended form of support vector machine (SVM) and is mainly used to process the sorting and ranking problems of triplet data.
[0065] It can be understood that the triplet-SVM algorithm-related software uses known microRNA mature sequences and non-microRNA sequence data as training sets (and verification) to build a learning model for classifying predicted sequences, and can preliminarily determine whether a candidate precursor sequence with a comparison site microRNA prediction is a microRNA.
[0066] S06. Constructing the secondary structure of the microRNA precursor sequence, and obtaining the minimum free energy and the minimum free energy coefficient of the secondary structure.
[0067] In an optional embodiment, step S06 can use RNAfold to predict the secondary structure, and use RNAfold software to obtain the minimum free energy (MFE) of the secondary structure and the corresponding minimum free energy index (MFEI). Wherein, RNAfold is a computational tool for predicting the secondary structure of RNA molecules.
[0068] Furthermore, in this embodiment, the minimum free energy and minimum free energy coefficient of any secondary structure of the microRNA precursor sequence constructed in step S06 can be calculated using the above-mentioned RNAfold software.
[0069] In one or some other embodiments, based on the constructed secondary structure, obtaining the minimum free energy and the minimum free energy coefficient of the secondary structure comprises the following steps:
[0070] S061. Build the minimum free energy model and the minimum free energy coefficient model respectively.
[0071] In this embodiment, the minimum free energy and the minimum free energy coefficient satisfy the following models respectively:
[0072]
[0073] Wherein, MFE represents the minimum free energy of the secondary structure, MFEI represents the minimum free energy coefficient of the secondary structure, i and j both represent the base positions in the microRNA precursor sequence, w(i,j) represents the energy between the base at the i-th position and the base at the j-th position in the microRNA precursor sequence, δ(i,j) represents the pairing indicator function between the base at the i-th position and the base at the j-th position in the microRNA precursor sequence, w(i) represents the energy of the base at the i-th position in the microRNA precursor sequence, δ(i) represents the stability indicator function of the base at the i-th position in the microRNA precursor sequence, R represents the ideal gas constant, T represents the absolute temperature, l represents the number of bases in the predicted secondary structure, m(G&C) represents the number of bases G and base C in the microRNA precursor sequence, and m(G&C&A&U) represents the number of bases G, base C, base A, and base U in the microRNA precursor sequence.
[0074] Furthermore, for the pairing indicator function δ(i,j): when the base at the i-th position in the microRNA precursor sequence can pair with the base at the j-th position, δ(i,j) = 1, when the base at the i-th position in the microRNA precursor sequence cannot pair with the base at the j-th position, δ(i,j) = 0. For the stability indicator function δ(i): when the base at the i-th position in the microRNA precursor sequence is in a stable state, δ(i) = 1, when the base at the i-th position in the microRNA precursor sequence is in an unstable state, δ(i) = 0.
[0075] S062. Using the minimum free energy model and the minimum free energy coefficient model, respectively obtaining the minimum free energy and the minimum free energy coefficient corresponding to the secondary structure can be calculated by the following minimum free energy model and minimum free energy coefficient model.
[0076] S07. Setting a minimum free energy threshold and a minimum free energy coefficient threshold, and combining the minimum free energy and the minimum free energy coefficient to screen out a target microRNA precursor sequence from the microRNA precursor sequences.
[0077] In an optional embodiment, based on the secondary structure predicted by the RNAfold software, and the minimum free energy and the minimum free energy coefficient calculated by the RNAfold software, a corresponding minimum free energy threshold and a minimum free energy coefficient threshold are set: the minimum free energy threshold MFEMFE≤
[0078] -25Kcal / mol, minimum free energy coefficient threshold MFEI ≥ 0.85.
[0079] Furthermore, when the minimum free energy and the minimum free energy coefficient of the secondary structure constructed by the RNAfold software meet the above corresponding thresholds, the microRNA precursor sequence corresponding to the secondary structure is the target microRNA precursor sequence; otherwise, the microRNA precursor sequence corresponding to the secondary structure is not the target microRNA precursor sequence.
[0080] S08. Using the microRNA mature sequence marker to match the target microRNA precursor sequence, obtain the microRNA mature sequence of the target organism at different growth stages.
[0081] In an optional embodiment, the step S07 of using the mature microRNA sequence marker to match the target microRNA precursor sequence to obtain the mature microRNA sequence of the target organism at different growth stages includes the following steps:
[0082] S081. Using the mature microRNA sequence corresponding to the mature microRNA sequence tag, match the sequence of the stem region in the secondary structure of the target microRNA precursor sequence.
[0083] In this embodiment, the stem region refers to the portion of the sequence in the microRNA candidate precursor sequence that forms the stem-loop structure. Specifically, the stem region is composed of two complementary sequences that form a stable double helix structure through complementary pairing.
[0084] S082. Obtain the mature microRNA sequences of the target organism at different growth stages according to the matching results.
[0085] Specifically, if the mature microRNA sequence corresponding to the mature microRNA sequence tag can completely match the sequence of the stem region in the secondary structure of the target microRNA precursor sequence, the mature microRNA sequence corresponding to the mature microRNA sequence tag is the mature microRNA sequence of the target organism. Further, redundant sequences with repeated positions or consistent precursor sequences are removed from the prediction results.
[0086] In a specific embodiment, based on the transcriptome data of Plutella xylostella at four growth stages of Egg, Larva, Pupa and Adult from the SRA database (Sequence Read Archive) of NCBI in the above embodiment, through the above steps S02 to S08, the distribution numbers of microRNAs predicted from the transcriptome data of Plutella xylostella in Egg, Larva, Pupa and Adult are 62, 35, 69 and 76 respectively.
[0087] In an optional embodiment, in order to better implement the above-mentioned microRNA prediction method based on transcriptome data, the present invention also provides a microRNA prediction system based on transcriptome data, see Figure 2 , Figure 2 This is a structural diagram of a microRNA prediction system based on transcriptome data provided in an embodiment of the present invention.
[0088] like Figure 2 As shown, the microRNA prediction system based on transcriptome data proposed in the present invention includes an input device, a processor, a memory and an output device, wherein the input device, the processor, the memory and the output device are interconnected, wherein the memory is used to store a computer program, the computer program includes program instructions, and the processor is configured to call the program instructions to execute the microRNA prediction method based on transcriptome data provided by the present invention.
[0089] Furthermore, the input device may be a keyboard, a mouse, a touch screen, etc., for the user to interact with the system. For example, researchers may provide transcriptome data to be analyzed through an input device so that the system can perform subsequent processing.
[0090] Furthermore, the processor is a core component of the system, which is used to execute computer programs and process data. It should be understood that the processor is responsible for calling computer program instructions stored in the memory to execute the microRNA prediction method based on transcriptome data.
[0091] Further, the memory is used to store computer programs, data and intermediate results. In the present invention, the memory stores the computer program instructions required for executing the microRNA prediction method. This may include transcriptome data, non-redundant protein databases, existing microRNA mature body sequences, etc.
[0092] Furthermore, the output device is used to display the results of the system processing to the user. For example, the system can display the predicted microRNA mature sequence results to researchers through the output device so that they can analyze and study.
[0093] In this embodiment, the microRNA prediction system based on transcriptome data of the present invention uses an input device to obtain transcriptome data, executes a computer program via a processor, uses a memory to store the required program and data, and finally presents the prediction result to the user through an output device. The microRNA prediction system based on transcriptome data provided by the present invention can efficiently and effectively implement the aforementioned microRNA prediction method based on transcriptome data, providing a convenient and efficient tool for biological research and agricultural pest management.
[0094] Finally, it should be noted that the above embodiments are only used to illustrate the technical solutions of the present invention, rather than to limit it. Although the present invention has been described in detail with reference to the aforementioned embodiments, those skilled in the art should understand that they can still modify the technical solutions described in the aforementioned embodiments, or replace some or all of the technical features therein by equivalents. These modifications or replacements do not make the essence of the corresponding technical solutions deviate from the scope of the technical solutions of the embodiments of the present invention, and they should all be included in the scope of the claims and specification of the present invention.
Claims
1. A microRNA prediction method based on transcriptome data, It is characterized in that The microRNA prediction method based on transcriptome data comprises the following steps: obtaining transcriptome data of a target organism at one or more growth stages; Obtaining a non-redundant protein database, and using the non-redundant protein database to screen out non-protein coding sequences from a transcriptome data set; Setting a distance window and a capture window, and using the length of the distance window as a sliding unit, sliding the capture window on the non-protein coding sequence to capture multiple microRNA prediction candidate precursor sequences; Obtaining an existing microRNA mature sequence, and obtaining a microRNA mature sequence marker by combining the microRNA predicted candidate precursor sequence with the existing microRNA mature sequence; According to the existing microRNA mature sequence, screening out microRNA precursor sequences from the microRNA prediction candidate precursor sequences; Constructing the secondary structure of the microRNA precursor sequence, and obtaining the minimum free energy and minimum free energy coefficient of the secondary structure; Setting a minimum free energy threshold and a minimum free energy coefficient threshold, and combining the minimum free energy and the minimum free energy coefficient to screen out a target microRNA precursor sequence from the microRNA precursor sequences; Using the microRNA mature sequence marker to match the target microRNA precursor sequence to obtain the microRNA mature sequence of the target organism at different growth stages; The step of obtaining the minimum free energy and the minimum free energy coefficient of the secondary structure comprises the following steps: The minimum free energy model and the minimum free energy coefficient model are constructed respectively, and the minimum free energy model and the minimum free energy coefficient model are respectively: , , in, represents the minimum free energy of the secondary structure, represents the minimum free energy coefficient of the secondary structure, i and j both represent the base position in the microRNA precursor sequence, represents the energy between the base at the i-th position and the base at the j-th position in the microRNA precursor sequence, represents the pairing indicator function between the base at the i-th position and the base at the j-th position in the microRNA precursor sequence. When the base at the i-th position in the microRNA precursor sequence can pair with the base at the j-th position, =1, when the base at the i-th position in the microRNA precursor sequence cannot pair with the base at the j-th position, =0, represents the base energy at the ith position in the microRNA precursor sequence, represents the stability indicator function of the base at the i-th position in the microRNA precursor sequence. When the base at the i-th position in the microRNA precursor sequence is in a stable state, =1, when the base at position i in the microRNA precursor sequence is in an unstable state, =0, R represents the ideal gas constant, T represents the absolute temperature, indicates the number of bases in the predicted secondary structure, Indicates the bases in the microRNA precursor sequence and bases The number of , represents the number of base G, base C, base A, and base U in the microRNA precursor sequence; Using the minimum free energy model and the minimum free energy coefficient model, respectively obtaining the minimum free energy and the minimum free energy coefficient corresponding to the secondary structure; The length range of the distance window includes 18 nt to 25 nt, and the length of the interception window is at least 120 nt; Set the distance window length to 25 nt, and set the interception window length to 120 nt; The length of the distance window is used as a sliding unit, and the interception window is used to slide on the non-protein coding sequence to intercept multiple microRNA prediction candidate precursor sequences. Any microRNA prediction candidate precursor sequence satisfies the following counting model: ,in, , , represents a positive integer, represents the total number of nucleotides contained in the non-protein coding sequence. represents the i-th microRNA prediction candidate precursor sequence obtained from the non-protein coding sequence, wherein the i-th microRNA prediction candidate precursor sequence includes the i-th microRNA prediction candidate precursor sequence in the non-protein coding sequence. Nucleotides to nucleotides; The method of using the non-redundant protein database to screen out non-protein coding sequences from a transcriptome data set comprises the following steps: Assemble the transcription data in the transcriptome dataset to obtain multiple non-repetitive continuous sequences; Comparing the known protein coding sequence in the non-redundant protein database with the non-repetitive continuous sequence, obtaining a coding region in the non-repetitive continuous sequence that is similar to the known protein coding sequence, and calculating the similarity between the sequence in the coding region and the known protein coding sequence; Setting a similarity threshold, and determining whether the non-repetitive continuous sequence is a non-protein coding sequence or a protein coding sequence by comparing the similarity with the similarity threshold; The method of obtaining a microRNA mature body sequence tag by combining the microRNA predicted candidate precursor sequence with the existing microRNA mature body sequence comprises the following steps: Comparing the existing microRNA mature sequence with the microRNA predicted candidate precursor sequence to obtain a microRNA predicted candidate precursor sequence having a comparison site; Predicting candidate precursor sequences according to the microRNA with the alignment site, marking a microRNA mature body sequence similar to the alignment site sequence as a microRNA mature body sequence marker; The method of using the mature microRNA sequence marker to match the target microRNA precursor sequence to obtain the mature microRNA sequence of the target organism at different growth stages comprises the following steps: Using the mature microRNA sequence corresponding to the mature microRNA sequence marker to match the sequence of the stem region in the secondary structure of the target microRNA precursor sequence; According to the matching results, the mature sequences of microRNAs of the target organisms at different growth stages are obtained.
2. The microRNA prediction method based on transcriptome data according to claim 1, It is characterized in that The target organisms include Plutella xylostella.
3. The microRNA prediction method based on transcriptome data according to claim 2, It is characterized in that The growth stages include the diamondback moth egg stage, the diamondback moth larva stage, the diamondback moth pupa stage and the diamondback moth adult stage.
4. A microRNA prediction system based on transcriptome data, It is characterized in that The microRNA prediction system based on transcriptome data comprises an input device, a processor, a memory and an output device, wherein the input device, the processor, the memory and the output device are interconnected, wherein the memory is used to store a computer program, the computer program comprises program instructions, and the processor is configured to call the program instructions to execute the microRNA prediction method based on transcriptome data as described in any one of claims 1 to 3.
Citation Information
Patent Citations
Biological genome sequence-based microRNA predicting method
CN107523631A
Method for extracting immunotherapy new antigen based on tumor specific transcriptional region assembled by new transcript and application
CN111627497A
Transcriptome analysis method and system without reference genome sequence
CN112397149A