Quantitative method of SRBSDV based on rice infected tissue transcriptome sequencing
Through high-throughput transcriptome sequencing technology and FPKB standardized function, the problem of difficulty in evaluating SRBSDV infection levels in the prior art is solved, and accurate quantities of viral gene expression and in-depth understanding of viral infectious dynamics are achieved, thus reducing detection costs.
Patent Information
- Application Number
- CN202510433196.1
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-04-08
- Publication Date
- 2025-07-22
- Estimated Expiration
- 2045-04-08
AI Technical Summary
The prior art is difficult to quickly and comprehensively evaluate the infection level of Southern Rice Black Strip Dwarf Vaccine (SRBSDV). Traditional detection methods have limitations and cannot provide comprehensive viral load information and high-throughput detection capabilities.
High-throughput transcriptome sequencing technology based on rice infected tissues is adopted, and the viral gene expression level is accurately determined by constructing a viral genome index module, obtaining high-quality sequencing data, and using FPKB standardized function to analyze the viral gene expression level.
It improves the accuracy and efficiency of virus quantification, provides more comprehensive viral load information, enhances the understanding of the dynamics of virus infection, and supports virus research and prevention and control.
Smart Images

Figure CN120356523A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the fields of biotechnology and bioinformatics, and particularly to a method for rapidly quantifying the gene expression level of Southern rice black-streaked dwarf virus using high-throughput sequencing technology. Background Art
[0002] Southern rice black-streaked dwarf disease, caused by Southern rice black-streaked dwarf virus (SRBSDV), poses a major threat to the global rice industry. The virus is transmitted by vector insects such as planthoppers, resulting in symptoms such as dwarfing of rice plants, increased tillering, and black streaks on leaves, seriously affecting the growth and yield of rice. Due to its rapid spread and difficult-to-control characteristics, SRBSDV has caused huge economic losses to the rice industry globally, especially in Asian regions.
[0003] In the prior art, there are certain limitations in the methods for detecting and quantifying SRBSDV. Serological detection is simple to operate, but is limited by the need for specific antibodies and has limited detection ability for low-abundance viruses. Electron microscopy observation can identify the morphology of virus particles, but requires expensive equipment and professional skills and cannot perform quantitative analysis. PCR or RT-PCR techniques detect the virus by amplifying the specific gene sequence of the virus, with high sensitivity and specificity, but can only design precise primers for specific fragments, with high requirements and little information obtained; at the same time, the detection results rely on the electrophoresis detection results in molecular biology experiments, with a high error rate. Chip hybridization technology uses specific probes to hybridize with virus RNA or DNA, with a high cost. These methods usually cannot provide comprehensive virus load information, such as simultaneously quantifying the transcriptional expression of multiple genes; and cannot standardize the detection task of large samples, restricting the in-depth understanding of the virus infection dynamics. Therefore, there is an urgent need for a new technology to capture a more comprehensive SRBSDV virus infection situation in a relatively standardized form.
[0004] As a breakthrough in the field of biotechnology, high-throughput transcriptome sequencing technology can rapidly and comprehensively perform high-throughput analysis of the transcriptomes of a large number of samples. With the reduction of sequencing costs and the improvement of data processing capabilities, transcriptome sequencing has become an important tool for studying gene expression, gene regulatory networks, and disease mechanisms. In terms of virus detection, transcriptome sequencing exhibits high sensitivity and specificity, and is particularly suitable for the high-throughput detection of viruses such as SRBSDV with few commercial detection products but with genomic data. However, due to the complexity of high-throughput transcriptome sequencing data, existing technical analysis methods often have difficulty effectively processing and interpreting the expression level of SRBSDV, such as comparing the expression situation of time series, due to the lack of a systematic bioinformatics analysis process and a dedicated expression level quantification method.
[0005] Therefore, how to provide an innovative method for more accurately and with higher throughput to evaluate the infection level of SRBSDV has become an urgent problem to be solved currently. Summary of the Invention
[0006] To solve the above technical problems, the present invention provides a method for quantifying Southern rice black-streaked dwarf virus (SRBSDV) based on transcriptome sequencing of infected rice tissues. The FPKB normalization method proposed by the present invention provides a new perspective for the analysis of viral gene expression levels. This innovative method helps to more accurately evaluate the infection level of the virus; at the same time, by making full use of high-throughput transcriptome sequencing data, more virological information can be mined, maximizing the utilization of data, and providing support for virus research and prevention and control. Moreover, it not only improves the accuracy of virus quantification, but also enhances the understanding of the dynamics of virus infection by providing more comprehensive virus load information, which is of great significance for early diagnosis and epidemic monitoring.
[0007] To achieve the above object, one of the objects of the present invention is to provide a method for quantifying SRBSDV based on transcriptome sequencing of infected rice tissues, comprising the following steps:
[0008] 1) Obtain the genome assembly and gene annotation files of the SRBSDV virus, and the raw data of high-throughput transcriptome sequencing of virus-infected rice tissues;
[0009] 2) Construct an SRBSDV virus genome indexing module;
[0010] 3) Obtain the non-redundant length of each gene in the SRBSDV virus genome annotation file;
[0011] 4) Process the raw data of high-throughput transcriptome sequencing to obtain high-quality transcriptome sequencing data;
[0012] 5) Obtain the total number of short reads of each sample in the high-quality transcriptome sequencing data;
[0013] 6) Based on the SRBSDV virus genome indexing module, align the high-quality transcriptome sequencing data with the SRBSDV virus genome to obtain an individual-level BAM file;
[0014] 7) Based on the BAM file and the gtf format of the SRBSDV virus genome annotation file, obtain the sum of short read fragments aligned to the viral gene region, and obtain the original expression levels of different viral genes in each sample;
[0015] 8) Obtain a new gene expression level normalization function FPKB for normalizing the original expression levels.
[0016] Preferably, step 1) specifically includes:
[0017] 11) Download the genome assembly and gene annotation files of the SRBSDV virus from the public database NCBI;
[0018] 12) Extract the total RNA from the rice infected tissues, isolate and obtain mRNA, construct a library for high-throughput sequencing, and obtain the raw data of high-throughput transcriptome sequencing of the virus-infected rice tissues.
[0019] Furthermore, the link for downloading in step 11) from the public database NCBI is: https: / / ftp.ncbi.nlm.nih.gov / genomes / all / GCA / 031 / 528 / 665 / GCA_031528665.1_ASM3152866v1 / .
[0020] The beneficial effects of adopting the above technical solutions at least include: the above steps can make full use of the already published genome assembly, improving the universality of the invention and the unity of the output results; compared with the time-consuming de novo assembly, the above steps can improve the quantification efficiency.
[0021] Preferably, the high-throughput sequencing in step 12) is performed using the Illumina Hiseq or DNBSE Q-T7 platform for paired-end 150bp sequencing.
[0022] The beneficial effects of adopting the above technical solutions at least include: the two sequencing platforms supported by the above steps are the most widely used in commercial applications, laying a foundation for the subsequent popularization and utilization of the present invention.
[0023] Preferably, step 2) is specifically: unzip the genome assembly file of the SRBSDV virus, and use the hisat2-build software to construct the SRBSDV virus genome index module.
[0024] The beneficial effects of adopting the above technical solutions at least include: the above index can significantly improve the running speed and accuracy of subsequent step 6).
[0025] Preferably, step 3) is specifically:
[0026] 31) Unzip the gene annotation file GCA_031528665.1_ASM3152866v1_genomic.gff.gz of the SRBSDV virus;
[0027] The above steps can make full use of the already published gene annotation data, improving the universality of the invention and the unity of the output results; compared with the time-consuming de novo assembly, the above steps can improve the quantification efficiency;
[0028] 32) Use the gffread software to convert the annotation file format from gff3 to gtf;
[0029] 33) Obtain the non-redundant length of each gene.
[0030] The above steps utilize the true length of the gene, solving the problem that the small number of viral genes is likely to cause inaccurate standardization of expression levels.
[0031] Preferably, it is characterized in that the processing of the original high-throughput transcriptome sequencing data in step 4) means: using the fastp software to remove adapters, primers, and low-quality short read fragments.
[0032] Preferably, it is characterized in that step 5) uses the seqkit software to obtain the total number of short reads for each sample in the high-quality transcriptome sequencing data;
[0033] Adopting the above technical solution has at least the following technical effects: Through large-scale parallel design in the calculation logic, this tool can accurately and rapidly calculate the total number of short reads for each sample.
[0034] Preferably, it is characterized in that the high-quality transcriptome sequencing data in step 6) is aligned with the SRBSD V virus genome using the hisat2 software to obtain a BAM file at the individual level.
[0035] Preferably, it is characterized in that step 7) uses the featureCounts tool to count the total sum of short read fragments aligned to the viral gene region.
[0036] Adopting the above technical solution has at least the following technical effects: This tool can quickly extract the effective aligned read fragments in the BAM file, which has an overall improvement effect on the running speed of the present invention.
[0037] Preferably, step 8) is specifically: using the total number of short reads of all high-quality sequencing data in each individual as a measure of the sequencing library size;
[0038] A gene expression level normalization function FPKB = total gene fragments count / (total sequenced paired reads count (Billion) * gene length (KB));
[0039] Where total gene fragments count is obtained from step S7,
[0040] total sequenced paired reads count is obtained from step S5,
[0041] The gene length is obtained in step S3.
[0042] Adopting the above technical solution has at least the following technical effects: Based on the normalization formula of eukaryotes, the true gene expression length of the SRBSDV virus is introduced for correction, improving the accuracy and reliability of the quantitative results of the expression level.
[0043] Furthermore, considering that the original expression level is affected by the size of the sequencing library and the gene length, data transformation for normalization is required. Several existing gene expression level normalization formulas sum up the alignment fragments of all genes as a measure of the size of the sequencing library, but this is not applicable to the current project because the virus has only 10 genes in total, which is asymmetric compared to the total number of 40,000 to 50,000 genes in rice. Therefore, the present invention uses the total number of short reads of all high-quality sequencing data in each individual as a measure of the size of the sequencing library.
[0044] In summary, the technical effects that the present invention can achieve at least include:
[0045] (1) By means of an automated high-throughput sequencing data analysis process, the efficiency of the quantification of Southern rice black-streaked dwarf virus is significantly improved;
[0046] (2) Using the specific gene expression level normalization method FPKB, the accuracy and reliability of the quantitative results are improved;
[0047] (3) Compared with the existing detection methods, the present invention reduces the dependence on specific reagents and expensive equipment, and reduces the detection cost;
[0048] (4) The present invention can provide comprehensive information on the virus load, which helps to deeply understand the dynamics of virus infection; furthermore, the present invention is mainly based on bioinformatics sequencing and analysis integration, and the final result can obtain more diverse and comprehensive information, including the expression of different genes of this virus;
[0049] (5) The FPKB normalization method proposed by the present invention provides a new perspective for the analysis of virus gene expression levels, which helps to more accurately evaluate the infection level of the virus; especially for the problem that the original expression quantitative results cannot be compared and there is a lack of a dedicated original data normalization (standardization) algorithm, the FPKB formula proposed by the present invention solves this problem;
[0050] (6) By making full use of high-throughput transcriptome sequencing data, the present invention can mine more virological information, provide support for virus research and prevention and control, and maximize the utilization of data as much as possible. Description of the Drawings
[0051] To more clearly illustrate the technical solutions in the embodiments of the present invention or the prior art, the following will briefly introduce the accompanying drawings required for the description of the embodiments or the prior art. Obviously, the accompanying drawings in the following description are only the embodiments of the present invention. For those of ordinary skill in the art, without creative efforts, other accompanying drawings can also be obtained based on the provided accompanying drawings.
[0052] Figure 1 It is a schematic flowchart of the method for quantifying the gene expression level of Southern rice black-streaked dwarf virus in the embodiments of the present invention;
[0053] Figure 2 It is the quantification result of the gene expression level of Southern rice black-streaked dwarf virus identified in the embodiments of the present invention. Detailed implementation manners
[0054] The following will clearly and completely describe the technical solutions in the embodiments of the present invention in conjunction with the accompanying drawings in the embodiments of the present invention. Obviously, the described embodiments are only a part of the embodiments of the present invention, rather than all of the embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those of ordinary skill in the art without creative efforts belong to the scope of protection of the present invention.
[0055] Example 1
[0056] Taking rice (Oryza sativa) as an example, the method for quantifying the gene expression level of Southern rice black-streaked dwarf virus provided by the following method is as follows:
[0057] 1) Virus infection treatment: Indoor, use virus-carrying white-backed planthoppers to bite rice. The treatment day is the 0th day, and then collect the diseased tissues at the base of the flag leaf on the 5th and 11th days after treatment. Set 3 biological replicates for each time point;
[0058] 2) Extract total RNA from rice tissues, then isolate mRNA from the total RNA, and then construct a sequencing library with a specification of 150 base pairs;
[0059] 3) Perform transcriptome sequencing on the library using the Illumina Hiseq sequencing platform;
[0060] 4) Download the genome assembly and gene annotation files of Southern rice black-streaked dwarf virus from the public database NCBI;
[0061] 5) After decompressing the genome file downloaded in step 4), use the hisat2-build software to construct an index of the Southern rice black-streaked dwarf virus genome;
[0062] 6) After decompressing the gene annotation file downloaded in step 4), use the gffread software to convert the file format from gff3 to gtf.
[0063] 7) Based on the gtf file obtained in step 6), calculate the non-redundant length of each gene.
[0064] 8) Based on the original high-throughput transcriptome sequencing data obtained in step 3), which has been uploaded to the public database (https: / / ngdc.cncb.ac.cn / bioproject / browse / PRJCA034117), use the fastp software for filtering, including removing adapters, primers, and low-quality short read fragments to obtain high-quality transcriptome sequencing data.
[0065] 9) Based on the high-quality sequencing data obtained in step 8), use the seqkit software to count the total number of short reads in each sample, which is the library size.
[0066] 10) Based on the genomic index obtained in step 5), use the hisat2 software to align the high-quality sequencing data obtained in step 8) to the genome of Southern rice black-streaked dwarf virus to generate an individual-level BAM file.
[0067] 11) Based on the BAM file obtained in step 10) and the gene annotation file (gtf format) obtained in step 6), use the featureCounts tool to count the total number of short read fragments aligned to the viral gene region to obtain the raw expression levels of different viral genes in each sample.
[0068] 12) Develop a gene expression normalization function FPKB (Fragments Per Kilobase of genemodelperBillion sequencedpaired reads) to normalize the raw expression levels to account for the effects of the sequencing library size and gene length. FPKB = total gene fragments count / (total sequenced pairedreads count (Billion) * gene length (KB)). Where total gene fragments count is obtained in step S7, total sequenced paired reads count is obtained in step S5, and gene length is obtained in step S3.
[0069] 13) Based on the function obtained in step 12), normalize the raw expression levels in step 11) to obtain the FPKB values of different viral genes in each individual.
[0070] It must be noted that, unless otherwise specified, technical terms or scientific terms in this application document shall be interpreted according to the general understanding of those skilled in the art. The specific steps, numerical formulas, and numerical values mentioned in these embodiments shall not limit the scope of protection of the present invention unless otherwise stated. In other words, the specific numerical values in all examples shall be regarded as exemplary in nature and not as a limitation on the present invention. Therefore, other possible numerical ranges shall also be regarded as embodiments of the present invention.
[0071] The various embodiments in this specification are described in a progressive manner. Each embodiment focuses on the differences from other embodiments. For the same or similar parts among the various embodiments, reference can be made to each other.
[0072] The above description of the disclosed embodiments enables those skilled in the art to implement or use the present invention. Various modifications to these embodiments will be obvious to those skilled in the art. The general principles defined herein can be implemented in other embodiments without departing from the spirit or scope of the present invention. Therefore, the present invention will not be limited to these embodiments shown herein, but rather will conform to the widest scope consistent with the principles and novel features disclosed herein.
Claims
1. A quantitative method for SRBSDV based on transcriptome sequencing of infected rice tissues, characterized in that, It includes the following steps: 1) Obtain the genome assembly and gene annotation files of SRBSDV virus, and the raw data of high-throughput transcriptome sequencing of rice tissues infected with the virus; 2) Construct an SRBSDV virus genome indexing module; 3) Obtain the non-redundant length of each gene in the SRBSDV virus genome annotation file; 4) Process the raw data of high-throughput transcriptome sequencing to obtain high-quality transcriptome sequencing data; 5) Obtain the total number of short reads of each sample in the high-quality transcriptome sequencing data; 6) Based on the SRBSDV virus genome indexing module, align the high-quality transcriptome sequencing data with the SRBSDV virus genome to obtain an individual-level BAM file; 7) Based on the BAM file and the SRBSDV virus genome annotation file in gtf format, sum up the short read fragments aligned to the viral gene region to obtain the original expression levels of different genes of the virus in each sample; 8) Obtain a new gene expression level normalization function FPKB for normalizing the original expression levels.
2. The quantitative method of SRBSDV based on the transcriptome sequencing of infected rice tissues according to claim 1, wherein Step 1) specifically includes: 11) Download the genome assembly and gene annotation files of SRBSDV virus from the public database NCBI; 12) Extract the total RNA of the rice infected tissues, isolate and obtain mRNA, construct a library for high-throughput sequencing, and obtain the raw data of high-throughput transcriptome sequencing of rice tissues infected with the virus.
3. The quantitative method of SRBSDV based on transcriptome sequencing of infected rice tissues according to claim 2, wherein, In step 12), the high-throughput sequencing is performed using the Illumina Hiseq or DNBSEQ-T7 platform for paired-end 150bp sequencing.
4. The quantitative method of SRBSDV based on the transcriptome sequencing of infected tissues of rice according to claim 1, characterized in that, Step 2) is specifically: Unzip the genome assembly file of SRBSDV virus, and use the hisat2-build software to construct an SRBSDV virus genome indexing module.
5. The quantitative method of SRBSDV based on the transcriptome sequencing of infected rice tissues according to claim 1, wherein Step 3) is specifically: 31) Unzip the gene annotation file GCA_031528665.1_ASM3152866v1_genomic.gff.gz of SRBSDV virus; 32) Use the gffread software to convert the annotation file format from gff3 to gtf; 33) Obtain the non-redundant length of each gene.
6. The quantitative method of SRBSDV based on transcriptome sequencing of infected rice tissues according to claim 1, wherein In step 4), the processing of the raw data of high-throughput transcriptome sequencing refers to: Using the fastp software to remove adapters, primers, and low-quality short read fragments.
7. The quantitative method of SRBSDV based on transcriptome sequencing of infected rice tissues according to claim 1, characterized in that In step 5), use the seqkit software to obtain the total number of short reads of each sample in the high-quality transcriptome sequencing data.
8. The quantitative method of SRBSDV based on transcriptome sequencing of infected rice tissues according to claim 1, characterized in that, In step 6), the high-quality transcriptome sequencing data is aligned with the SRBSDV virus genome using the hisat2 software to obtain an individual-level BAM file.
9. The quantitative method of SRBSDV based on transcriptome sequencing of infected rice tissues according to claim 1, characterized in that In step 7), use the featureCounts tool to count the total sum of short read fragments aligned to the viral gene region.
10. The quantitative method of SRBSDV based on the transcriptome sequencing of infected rice tissues according to claim 1, characterized in that Step 8) is specifically: Use the total number of short reads of all high-quality sequencing data in each individual as a measure of the sequencing library size; A gene expression normalization function FPKB = total gene fragments count (total number of gene fragments) / (total sequenced paired reads count (in billions) * gene length (in KB)); where total gene fragments count is obtained from step S7, total sequenced paired reads count is obtained from step S5, and gene length is obtained from step S3.
Citation Information
Patent Citations
Screening method for improving non-parameter transcriptome microsatellite marker polymorphism
CN107604047A
Whole genome association analysis method based on comparison of multiple genomes and next-generation sequencing data
CN113628685A
Method for detecting and analyzing virus expression quantity in single cell transcriptome sequencing data
CN115512767A
Repeat-aware profiling of cell-free RNA
WO2024010875A1
Copy number variation detection method based on high-throughput transcriptome sequencing data of single sample
WO2025039433A1