A method for quantifying srbdv based on sequencing of the transcriptome of infected tissues of rice

By using high-throughput transcriptome sequencing and FPKB normalization functions, the problem of difficulty in quantifying SRBSDV gene expression levels in existing technologies has been solved, achieving more accurate and efficient viral quantification, providing comprehensive viral load information, and supporting virus research and prevention and control.

CN120356523BActive Publication Date: 2025-11-07RICE RES INST GUANGDONG ACADEMY OF AGRI SCI
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202510433196.1
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2025-04-08
Publication Date
2025-11-07
Estimated Expiration
2045-04-08

AI Technical Summary

Technical Problem

Existing technologies are insufficient for rapidly and comprehensively quantifying gene expression levels in Southern Rice Black-Streaked Dwarf Virus (SRBSDV), and existing methods cannot provide comprehensive information on viral load, limiting a deeper understanding of the virus infection dynamics.

Method used

High-throughput transcriptome sequencing based on infected rice tissue was employed. By constructing an SRBSDV viral genome index module and obtaining high-quality transcriptome sequencing data, gene expression analysis was performed using the FPKB normalization function. The process included steps such as high-throughput sequencing, data processing, alignment, and expression level normalization.

Benefits of technology

It improves the accuracy and efficiency of viral quantification, provides more comprehensive viral load information, enhances the understanding of viral infection dynamics, reduces detection costs, and is suitable for large-sample testing.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120356523B_ABST
    Figure CN120356523B_ABST
Patent Text Reader

Abstract

The application discloses a kind of based on rice infected tissue transcriptome sequencing of SRBSDV Quantitative method of virus, comprising: 1) obtain the genome assembly and gene annotation file of SRBSDV virus, and the high-throughput transcriptome sequencing original data of rice virus infected tissue;2) construct SRBSDV virus genome index module;3) obtain the non-redundant length of each gene in SRBSDV virus genome annotation file;4) obtain high-quality transcriptome sequencing data;5) obtain the total number of short reads of each sample in step 4);6) obtain individual level BAM file;7) obtain the original expression of different genes in each sample Virus;8) obtain gene expression standardization function FPKB.The application can more accurately evaluate the infection level of virus by means of standardized bioinformatics analysis framework;And more virology information can be mined, the maximization of data utilization is realized, support is provided for virus research and prevention and control, and the accuracy of virus quantification is improved.
Need to check novelty before this filing date? Find Prior Art

Description

TECHNICAL FIELD

[0001] The present application relates to the field of biotechnology and bioinformatics, and particularly relates to a method for rapidly quantifying the expression level of Southern rice black-streaked dwarf virus gene by using high-throughput sequencing technology. BACKGROUND

[0002] Southern rice black-streaked dwarf disease, caused by Southern rice black-streaked dwarf virus (SRBSDV), poses a significant threat to the rice industry. The virus is transmitted by planthoppers and other insect vectors, leading to stunting, increased tillering, and black stripes on leaves in rice plants, severely affecting the growth and yield of rice. Due to its rapid spread and difficulty in control, SRBSDV has caused significant economic losses to the rice industry.

[0003] In the prior art, there are certain limitations in detecting and quantifying SRBSDV. Serological detection is simple to operate, but is limited by the need for specific antibodies and has limited detection capability for low-abundance viruses. Electron microscopy can identify the morphology of virus particles, but requires expensive equipment and professional skills, and cannot perform quantitative analysis. PCR or RT-PCR technology detects viruses by amplifying specific viral gene sequences, with high sensitivity and specificity, but requires precise primer design for specific fragments, which is demanding and provides limited information; the detection results depend on the electrophoresis detection results in molecular biology experiments, with a high error rate. Chip hybridization technology uses specific probes to hybridize with viral RNA or DNA, which is costly. These methods usually cannot provide comprehensive viral load information, such as simultaneous quantification of multiple gene transcription expressions; and cannot standardize large sample detection tasks, limiting the in-depth understanding of viral infection dynamics. Therefore, there is an urgent need for a new technology to capture a more comprehensive SRBSDV viral infection in a more standardized form.

[0004] High-throughput transcriptome sequencing technology, as a breakthrough in the field of biotechnology, can rapidly and comprehensively analyze the transcriptome of a large number of samples. With the reduction of sequencing cost and the improvement of data processing capacity, transcriptome sequencing has become an important tool for studying gene expression, gene regulation networks, and disease mechanisms. In the field of virus detection, transcriptome sequencing has shown high sensitivity and specificity, and is particularly suitable for high-throughput detection of viruses such as SRBSDV, which have less commercial detection products but have genomic data. However, due to the complexity of high-throughput transcriptome sequencing data, existing technical analysis methods often fail to effectively process and interpret the expression level of SRBSDV, such as comparing the expression over time, due to the lack of a systematic bioinformatics analysis process and specialized expression quantification methods.

[0005] Therefore, how to provide an innovative method capable of more accurately and higher throughput evaluating the infection level of SRBSDV has become a current urgent problem to be solved. SUMMARY

[0006] To solve the above technical problems, the present application provides a method for quantifying southern rice black-streaked dwarf virus (SRBSDV) based on transcriptome sequencing of infected rice tissues, the FPKB normalization method provided by the present application provides a new perspective for viral gene expression analysis, which helps to more accurately evaluate the infection level of the virus; at the same time, by fully utilizing the high-throughput transcriptome sequencing data, more virology information can be mined, and the maximization of data utilization is realized, which provides support for virus research and prevention and control. And not only improve the accuracy of virus quantification, but also provide more comprehensive viral load information, enhance the understanding of viral infection dynamics, which has important significance for early diagnosis and epidemic monitoring.

[0007] To achieve the above purpose, one of the objects of the present application provides a method for quantifying SRBSDV based on transcriptome sequencing of infected rice tissues, comprising the following steps:

[0008] 1) Obtain the genome assembly and gene annotation file of SRBSDV virus, and the high-throughput transcriptome sequencing raw data of the virus-infected rice tissue;

[0009] 2) Construct a SRBSDV virus genome index module;

[0010] 3) Obtain the non-redundant length of each gene in the SRBSDV virus genome annotation file;

[0011] 4) Process the high-throughput transcriptome sequencing raw data to obtain high-quality transcriptome sequencing data;

[0012] 5) Obtain the total number of short reads of each sample in the high-quality transcriptome sequencing data;

[0013] 6) Based on the SRBSDV virus genome index module, the high-quality transcriptome sequencing data is aligned with the SRBSDV virus genome to obtain the individual level BAM file;

[0014] 7) Based on the BAM file and the SRBSDV virus genome annotation file gtf format, the total number of short reads aligned to the viral gene region is obtained, and the original expression of each sample in the different genes of the virus is obtained;

[0015] 8) Obtain a new gene expression normalization function FPKB for normalizing the original expression.

[0016] Preferably, step 1) specifically comprises:

[0017] 11) Download the genome assembly and gene annotation file of SRSBDV virus from the public database NCBI;

[0018] 12) Extract total RNA from rice infected tissues, isolate mRNA, construct a library for high-throughput sequencing, and obtain high-throughput transcriptome sequencing raw data of rice infected virus tissues.

[0019] Further, the link for downloading from the public database NCBI in step 11) is: https: / / ftp.ncbi.nlm.nih.gov / genomes / all / GCA / 031 / 528 / 665 / GCA_031528665.1_ASM3152866v1 / .

[0020] The beneficial effects of the above technical solutions at least include: the above steps can make full use of the published genome assembly, improve the universality of the application and the uniformity of the output results; compared with the time-consuming de novo assembly, the above steps can improve the quantitative efficiency.

[0021] Preferably, the high-throughput sequencing in step 12) is carried out by Illumina Hiseq or DNBSE Q-T7 platform for double-end 150bp sequencing.

[0022] The beneficial effects of the above technical solutions at least include: the two sequencing platforms supported by the above steps are most widely used in commercial applications, laying a foundation for the subsequent popularization and utilization of the application.

[0023] Preferably, step 2) is specifically: decompressing the genome assembly file of SRSBDV virus, and constructing an SRSBDV virus genome index module by using hisat2-build software.

[0024] The beneficial effects of the above technical solutions at least include: the above index can significantly improve the running speed and accuracy of subsequent step 6).

[0025] Preferably, step 3) is specifically:

[0026] 31) decompressing the gene annotation file GCA_031528665.1_ASM3152866v1_genomic.gff.gz of SRSBDV virus;

[0027] The above steps can make full use of the published gene annotation data, improve the universality of the application and the uniformity of the output results; compared with the time-consuming de novo assembly, the above steps can improve the quantitative efficiency;

[0028] 32) converting the annotation file format from gff3 to gtf by using gffread software;

[0029] 33) Obtain the non-redundant length of each gene.

[0030] The above steps utilize the actual length of the gene, and solve the problem that the small number of viral genes easily causes inaccurate standardization of expression.

[0031] Preferably, the processing of the high-throughput transcriptome sequencing raw data in step 4) refers to the removal of adapters, primers and low-quality short read fragments by using fastp software.

[0032] Preferably, step 5) uses seqkit software to obtain the total number of short reads of each sample in the high-quality transcriptome sequencing data.

[0033] The above technical solution at least includes the following technical effects: the tool is designed in large-scale parallelization in computing logic, and can accurately and rapidly calculate the total number of short reads of each sample.

[0034] Preferably, the high-quality transcriptome sequencing data and the SRBSD V viral genome in step 6) are aligned by using hisat2 software to obtain the individual-level BAM file.

[0035] Preferably, step 7) uses the featureCounts tool to count the total number of short reads aligned to the viral gene region.

[0036] The above technical solution at least has the following technical effects: the tool can quickly extract the effective alignment read fragments in the BAM file, and has a whole level of improvement on the running speed of the present application.

[0037] Preferably, step 8) is specifically: using the total number of short reads of all high-quality sequencing data in each individual as the measurement standard of the sequencing library size.

[0038] A gene expression standardization function FPKB = total gene fragments count / (total sequenced paired reads count (Billion) * gene length (KB));

[0039] total gene fragments count is obtained in step S7,

[0040] total sequenced paired reads count is obtained in step S5,

[0041] The gene length is obtained in step S3.

[0042] The technical scheme has at least the following technical effects: the formula is based on the standardization formula of eukaryotes, the real length of gene expression of SRSBDV virus is introduced for correction, and the accuracy and reliability of the quantitative result of expression are improved.

[0043] Further, since the original expression is affected by the sequencing library size and the gene length, data transformation needs to be standardized. The existing several gene expression standardization formulas all add the aligned fragments of all genes as the measurement standard of the sequencing library size, but this is not applicable to the present subject matter, because the virus has only 10 genes in total, which is asymmetric compared to the total number of 40,000 to 50,000 genes of rice, and therefore the present application uses the total number of short reads of all high-quality sequencing data in each individual as the measurement standard of the sequencing library size.

[0044] In summary, the technical effects that can be achieved by the present application include at least:

[0045] (1) The efficiency of quantification of Southern rice stripe virus is significantly improved by an automated high-throughput sequencing data analysis process;

[0046] (2) The accuracy and reliability of the quantitative result are improved by using the specific gene expression standardization method FPKB;

[0047] (3) Compared with the existing detection methods, the present application reduces the dependence on specific reagents and expensive equipment, and reduces the detection cost;

[0048] (4) The present application can provide comprehensive information on viral load, which helps to better understand the dynamic of viral infection; further, the present application is mainly based on bioinformatics sequencing and analysis integration, and the final result can obtain more diverse and comprehensive information, including the expression of different genes of the virus;

[0049] (5) The FPKB standardization method proposed by the present application provides a new perspective for viral gene expression analysis, which helps to more accurately evaluate the infection level of the virus; especially for the original expression quantitative result which cannot be compared, lacking a dedicated original data normalization (standardization) algorithm, the FPKB formula proposed by the present application solves this problem;

[0050] (6) By fully utilizing the high-throughput transcriptome sequencing data, the present application can mine more virology information, provide support for virus research and prevention, and maximize the utilization of data as much as possible. BRIEF DESCRIPTION OF DRAWINGS

[0051] In order to more clearly illustrate the technical solutions in the embodiments of the present application or the prior art, the accompanying drawings needed to be used in the embodiments or prior art description will be briefly introduced. Obviously, the accompanying drawings in the following description only represent some embodiments of the present application, and all other drawings obtained by those of ordinary skill in the art without creative effort based on the provided drawings also belong to the protection scope of the present application.

[0052] Figure 1 A flowchart of the method for quantifying the expression amount of the Southern rice black-streaked dwarf virus gene in the embodiments of the present application is shown.

[0053] Figure 2 The quantification results of the expression amount of the Southern rice black-streaked dwarf virus gene identified in the embodiments of the present application are shown. DETAILED DESCRIPTION

[0054] The technical solutions in the embodiments of the present application will be described clearly and completely in combination with the accompanying drawings in the embodiments of the present application. Obviously, the described embodiments only represent some embodiments of the present application, rather than all the embodiments. Based on the embodiments in the present application, all other embodiments obtained by those of ordinary skill in the art without creative effort also belong to the protection scope of the present application.

[0055] Embodiment 1

[0056] For example, the Southern rice black-streaked dwarf virus gene expression amount quantification method provided by the following method is used for rice (Oryza sativa), and the specific steps are as follows:

[0057] 1) Virus infection treatment: indoor white-backed planthoppers with virus bite rice, the treatment day is the 0th day, and then the diseased tissues collected on the 5th day and the 11th day after treatment, i.e. the base of the flag leaf, are set up with 3 biological replicates at each time point;

[0058] 2) Extract total RNA from rice tissues, then separate total RNA to obtain mRNA, and then construct a 150 base pair standard sequencing library;

[0059] 3) Based on the library, use Illumina Hiseq sequencing platform for transcriptome sequencing;

[0060] 4) Download the genome assembly and gene annotation file of the Southern rice black-streaked dwarf virus from the public database NCBI;

[0061] 5) After decompressing the genome file downloaded in step 4), use hisat2-build software to construct the index of the Southern rice black-streaked dwarf virus genome;

[0062] 6) After decompressing the gene annotation file downloaded in step 4), use gffread software to convert the file format from gff3 to gtf.

[0063] 7) Based on the gtf file obtained in step 6), calculate the non-redundant length of each gene.

[0064] 8) Based on the high-throughput transcriptome sequencing raw data obtained in step 3), which has been uploaded to the public database (https: / / ngdc.cncb.ac.cn / bioproject / browse / PRJCA034117), use fastp software for filtering, including removing adapters, primers and low-quality short read fragments, to obtain high-quality transcriptome sequencing data.

[0065] 9) Based on the high-quality sequencing data obtained in step 8), use seqkit software to count the total number of short reads in each sample, i.e. the library size.

[0066] 10) Based on the genome index obtained in step 5), use hisat2 software to align the high-quality sequencing data obtained in step 8) to the genome of Southern rice black-streaked dwarf virus, generating individual-level BAM files.

[0067] 11) Based on the BAM file obtained in step 10) and the gene annotation file (gtf format) obtained in step 6), use the featureCounts tool to count the total number of short read fragments aligned to the viral gene region, obtaining the raw expression of different viral genes in each sample.

[0068] 12) Develop a gene expression normalization function FPKB (Fragments Per Kilobase of genemodel per Billion sequenced paired reads) to normalize the raw expression, taking into account the effects of sequencing library size and gene length. FPKB = total gene fragments count / (total sequenced paired reads count (Billion) * gene length (KB)). Where total gene fragments count is obtained in step S7, total sequenced paired reads count is obtained in step S5, and gene length is obtained in step S3.

[0069] 13) Based on the function obtained in step 12), normalize the raw expression of step 11) to obtain the FPKB value of different viral genes in each individual.

[0070] It must be noted that, unless otherwise specified, the technical or scientific terms used in this application should be interpreted according to the conventional understanding of those skilled in the art. The specific steps, numerical formulas, and values ​​mentioned in these embodiments, unless otherwise specified, should not constitute a limitation on the scope of protection of this invention. In other words, the specific numerical values ​​in all examples should be considered exemplary and not limiting of the invention; therefore, other possible numerical ranges should also be considered as embodiments of the invention.

[0071] The various embodiments in this specification are described in a progressive manner, with each embodiment focusing on the differences from other embodiments. The same or similar parts between the various embodiments can be referred to each other.

[0072] The above description of the disclosed embodiments enables those skilled in the art to make or use the invention. Various modifications to these embodiments will be readily apparent to those skilled in the art, and the general principles defined herein may be implemented in other embodiments without departing from the spirit or scope of the invention. Therefore, the invention is not to be limited to the embodiments shown herein, but is to be accorded the widest scope consistent with the principles and novel features disclosed herein.

Claims

1. A method for quantifying SRBSDV based on sequencing of the transcriptome of infected tissues of rice, characterized in that, The method comprises the following steps: 1) obtaining a genome assembly and gene annotation file of Srbdv virus and high-throughput transcriptome sequencing raw data of rice infected virus tissue; 2) constructing a Srbdv virus genome index module; 3) obtaining the non-redundant length of each gene in the Srbdv virus genome annotation file; 4) processing the high-throughput transcriptome sequencing raw data to obtain high-quality transcriptome sequencing data; 5) obtaining the total number of short reads of each sample in the high-quality transcriptome sequencing data; 6) based on the Srbdv virus genome index module, the high-quality transcriptome sequencing data is aligned with the Srbdv virus genome to obtain the individual level BAM file; 7) based on the BAM file and the Srbdv virus genome annotation file gtf format, the total number of short read fragments aligned to the virus gene region is obtained, and the original expression of each sample in the virus different gene is obtained; 8) obtaining a new gene expression normalization function FPKB for normalizing the original expression; Step 8) is: using the total number of short reads of all high-quality sequencing data in each individual as the measurement standard of sequencing library size; A gene expression normalization function FPKB = total gene fragments count gene fragment total number / (total sequenced paired reads count paired sequencing short read total number (Billion) * genelength gene length (KB)); Where total gene fragments count is obtained in step 7), total sequenced paired reads count is obtained in step 5), gene length is obtained in step 3).

2. The method for quantifying SRBSDV based on sequencing of the transcriptome of infected tissues of rice according to claim 1, characterized in that, Step 1) specifically comprises: 11) downloading the genome assembly and gene annotation file of Srbdv virus from the public database NCBI; 12) extracting total RNA of rice infected tissue, and separating and obtaining mRNA to construct a library for high-throughput sequencing to obtain high-throughput transcriptome sequencing raw data of rice infected virus tissue.

3. The method for quantifying SRBSDV based on sequencing of the transcriptome of infected tissues of rice according to claim 2, characterized in that, Step 12) uses Illumina Hiseq or DNBSEQ-T7 platform for double-end 150bp sequencing.

4. The method for quantifying SRBSDV based on sequencing of the transcriptome of infected tissues of rice according to claim 1, wherein, Step 2) is: decompressing the genome assembly file of Srbdv virus, and constructing the Srbdv virus genome index module by using hisat2-build software.

5. The method for quantifying SRBSDV based on sequencing of the transcriptome of infected tissues of rice according to claim 1, wherein, Step 3) is specifically: 31) decompressing the gene annotation file GCA_031528665.1_ASM3152866v1_genomic.gff.gz of Srbdv virus; 32) converting the annotation file format from gff3 to gtf by using gffread software; 33) obtaining the non-redundant length of each gene. 6.The method for quantifying SRBSDV based on sequencing of the transcriptome of infected tissues of rice according to claim 1, wherein, Step 4) The processing of the raw data of high-throughput transcriptome sequencing refers to the removal of adapters, primers and short read fragments with low quality by using fastp software.

7. The method for quantifying SRBSDV based on sequencing of the transcriptome of infected tissues of rice according to claim 1, wherein, Step 5) The total number of short read fragments of each sample in the high-quality transcriptome sequencing data is obtained by using seqkit software. 8.The method for quantifying SRBSDV based on sequencing of the transcriptome of infected tissues of rice according to claim 1, wherein, Step 6) The high-quality transcriptome sequencing data is aligned with the SRBSDV viral genome by using hisat2 software to obtain the individual level BAM file. 9.The method of quantifying SRBSDV based on sequencing of the transcriptome of infected tissues of rice according to claim 1, wherein, Step 7) The total number of short read fragments aligned to the viral gene region is counted by using featureCounts tool.

Citation Information

Patent Citations

  • Screening method for improving non-parameter transcriptome microsatellite marker polymorphism

    CN107604047A

  • Whole genome association analysis method based on comparison of multiple genomes and next-generation sequencing data

    CN113628685A