Digital quality control product for monitoring metagenome biological information analysis process detection
By using digital quality control products in the metagenomic biological information analysis process, including specific base sequences, the problem of inaccuracy of detection results of the second-generation sequencing method is solved, and higher-precision analysis results and process optimization are achieved.
Patent Information
- Application Number
- CN202510194498.8
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-02-21
- Publication Date
- 2025-06-20
AI Technical Summary
In the existing biological product detection, the second-generation sequencing method of metagenomic technology has high sensitivity, high background noise and sequencing randomness, making it difficult to judge whether the test results are true or false, and lacks a recognized quality control method to monitor the accuracy of the biotechnology analysis process.
It provides a digital quality control product containing specific base sequences (SEQ ID NO.1-10) to monitor the metagenomic bioinformatics analysis process, and evaluate the accuracy and stability of the analysis process by incorporating sequencing data to avoid false positive and false negative results.
It improves the detection accuracy of metagenomic biological information analysis, ensures the reliability and specificity of the analysis results, avoids interference with gene composition, and provides a recognized quality control method to optimize the analysis process.
Smart Images

Figure BDA0005281051700000031 
Figure BDA0005281051700000041 
Figure BDA0005281051700000061
Abstract
Description
Technical Field
[0001] The present invention belongs to the field of bioinformatics, and particularly relates to a digital quality control product for monitoring the detection process of metagenomic bioinformatics analysis. Background Art
[0002] Metagenomics extracts nucleic acid substances of all microorganisms from different samples, then constructs a metagenomic library, and further analyzes the genetic composition of all microorganisms in the sample and the potential functions of their communities using genomic research strategies. This omics technology does not rely on the isolation and pure culture technology of single bacteria, and to a large extent solves the problem that most microorganisms are difficult to study because they cannot be isolated and cultured. At the same time, it can also reflect the true situation of the microbial composition in the research ecological environment. In the microecological research on human health, quantifying microorganisms based on metagenomic sequencing data is the basis for studying related laws such as their community composition, species interaction, and exploring the relationship between their occurrence and development of diseases. With the progress of scientific research, more and more studies have shown that it is increasingly important to accurately annotate lower taxonomic units of specific species, namely strains. If simply studying the relationship between bacteria and diseases at a higher taxonomic level, it is very likely to add together categories that are positively correlated, uncorrelated, or even negatively correlated with the development of diseases, which is clearly fallacious both biologically and statistically; and the existing research results also urgently need to be corrected by improving the accuracy of microbial quantification or conducting more in-depth mechanism research.
[0003] Currently, in the biological product detection industry, more and more people use the metagenomic technology route for exogenous virus detection. Through high-throughput metagenomic technology, it is expected to detect all microorganisms (bacteria, viruses, fungi, etc.) in a sample at one time, but its subsequent verification work needs to be carried out as soon as possible to popularize this detection technology.
[0004] The metagenomic technology developed in recent years avoids the traditional microbial isolation and culture methods and directly extracts total nucleic acids from environmental samples. By constructing and screening metagenomic libraries, new functional genes and bioactive substances can be obtained. The metagenomic library includes both culturable and unculturable microbial genetic information, so it increases the chance of obtaining new bioactive substances. Its quantification method has a relatively high resolution, that is, it can annotate to the species or strain level, so it is currently widely used in pathogen detection.
[0005] In the field of biological product detection, currently the industry mainly still uses traditional methods for detection, such as PCR, electrophoresis, and culture methods. However, the throughput of these traditional methods is relatively low, and they cannot detect a large number of different types of viruses at the same time. Therefore, methods based on the next-generation sequencing (NGS) technology route are gradually developing.
[0006] Although next-generation sequencing can make up for the shortcoming of low throughput, its characteristics of high sensitivity, high background noise, and random sequencing lead to difficulty in judging the true and false positives of the detection results. At the same time, the sample preparation of next-generation sequencing, the extraction of DNA and RNA, and the subsequent library construction process are prone to introducing environmental or consumable contamination. Therefore, while promoting the use of metagenomics for biological product safety detection, it is necessary to develop quality control products to test the entire metagenomic workflow, including the correctness of data during bioinformatics analysis in wet experiments. Due to the diverse methods and analysis software tools in dry experiments (bioinformatics processes), there is currently no generally recognized and most reliable detection method and tool for everyone to use. Therefore, it is very necessary to construct digital quality control products to monitor the bioinformatics analysis process. Summary of the Invention
[0007] To solve the above problems, the present invention provides a digital quality control product for testing whether the metagenomic bioinformatics workflow is correct.
[0008] The technical solution adopted by the present invention is: a digital quality control product, including a quality control sequence; the quality control sequence is selected from at least one of the base sequences SEQ ID NO.1-10.
[0009] The present invention is mainly applied to the monitoring of the bioinformatics analysis process for the detection of exogenous viruses in biological samples. Therefore, during the construction of the digital quality control product library, the nucleic acid sequences of the quality control products need to avoid being the same as the nucleic acid sequences of viruses, and at the same time, avoid the nucleic acid sequences of the quality control products being the same as the nucleic acid sequences of the hosts of the detected biological samples. Through gene Blast, the quality control sequences shown in base sequences such as SEQ ID NO.1-10 are selected as the basic sequences of the digital quality control products.
[0010] Preferably, the quality control sequence is selected from at least one of the base sequences SEQ ID NO.1-5. Further, the present invention combines the quality control sequences shown in base sequences such as SEQ ID NO.1-5 as the digital quality control product for the metagenomic bioinformatics workflow. That is, more preferably, the quality control sequence includes the base sequences SEQ ID NO.1-5.
[0011] Preferably, the quality control sequence is selected from at least one of the base sequences SEQ ID NO.6-10. Further, the present invention combines the quality control sequences shown in base sequences such as SEQ ID NO.6-10 as the digital quality control product for the metagenomic bioinformatics workflow. That is, more preferably, the quality control sequence includes the base sequences SEQ ID NO.6-10.
[0012] Preferably, the digital quality control product further includes a high-GC sequence; the high-GC sequence is selected from at least one of the base sequences SEQ ID NO.11 to 13.
[0013] Preferably, the digital quality control product further includes a high-AT sequence; the high-AT sequence is the base sequence SEQ ID NO.14. Further, to avoid the interference of the particularity of the gene composition, such as the proportion of bases G, C, T, A and repetitive sequences, etc., on the bioinformatics analysis process, the present invention also adds special gene sequences such as high-GC sequences and high-AT sequences to the digital quality control product for supplementation.
[0014] Preferably, the digital quality control product includes a quality control sequence, a high-GC sequence, and a high-AT sequence; the quality control sequence includes the base sequences SEQ ID NO.1 to 5; the high-GC sequence is selected from at least one of the base sequences SEQ ID NO.11 to 13; the high-AT sequence is the base sequence SEQ ID NO.14.
[0015] Preferably, the digital quality control product includes a quality control sequence, a high-GC sequence, and a high-AT sequence; the quality control sequence includes the base sequences SEQ ID NO.6 to 10; the high-GC sequence is selected from at least one of the base sequences SEQ ID NO.11 to 13; the high-AT sequence is the base sequence SEQ ID NO.14.
[0016] Table 1. Digital quality control product
[0017]
[0018]
[0019] The present invention also provides the application of the digital quality control product in monitoring the detection of the metagenomic bioinformatics analysis process of biological samples.
[0020] Preferably, the detection includes: detecting exogenous viruses in biological samples.
[0021] Preferably, the application includes: incorporating the digital quality control product into sequencing data for metagenomic bioinformatics analysis; if the number of reads detected in the species genome is within a preset fluctuation range, the result of the bioinformatics analysis is reliable.
[0022] Preferably, the fluctuation range is 50% to 150%.
[0023] Preferably, the incorporation quantity of the digital quality control product is 50 - 1000000.
[0024] The beneficial effects of the present invention:
[0025] In the bioinformatics analysis process of the present invention, digital quality control products with different proportions are incorporated into the sequencing data. By analyzing the detection conditions of the digital quality control products at different gradients and their influence on the detection results of other species, the accuracy, stability, and reliability of the entire bioinformatics analysis process are evaluated and monitored, thereby providing a basis for optimizing the analysis process parameters and improving the detection accuracy. Moreover, the digital quality control products do not overlap with the detection target and the host sequence, thus effectively avoiding false positive or false negative results caused by sequence similarity and making the detection results more real and reliable. In addition, to avoid the interference of the particularity of the gene composition, such as the proportion of bases G, C, T, A and repetitive sequences, etc., on the bioinformatics analysis process, the present invention also adds special gene sequences such as high-GC sequences and high-AT sequences to the digital quality control products for supplementation. Specific Embodiments
[0026] The following specific embodiments illustrate the implementation manners of the present invention. Those skilled in the art can easily understand other advantages and effects of the present invention from the content disclosed in this specification. The present invention can also be implemented or applied through other different specific implementation manners, and various details in this specification can also be modified or changed based on different viewpoints and applications without departing from the spirit of the present invention. It should be noted that, without conflict, the following embodiments and the features in the embodiments can be combined with each other.
[0027] In the embodiments of the present invention, two groups of digital quality control products, namely combination 1 and combination 2, are provided.
[0028] Combination 1 includes: quality control sequences with base sequences as shown in SEQ ID NO.1 - 5, a high-GC sequence with a base sequence as shown in SEQ ID NO.11, and a high-AT sequence with a base sequence as shown in SEQ ID NO.14.
[0029] Combination 2 includes: quality control sequences with base sequences as shown in SEQ ID NO.6 - 10, a high-GC sequence with a base sequence as shown in SEQ ID NO.11, and a high-AT sequence with a base sequence as shown in SEQ ID NO.14.
[0030] Perform exogenous virus detection on the sample poxvirus - related virus seed (cultured through chicken embryos) to obtain sequencing data, and perform bioinformatics analysis on the sequencing data; at the same time, incorporate combination 1 and combination 2 into the sequencing data at 50, 500, 5000, and 1000000 reads respectively for bioinformatics analysis.
[0031] The analysis results are shown in Table 2. From the sample sequencing data, the viruses Vaccinia virus, Cytomegalovirus humanbeta5, Monkeypox virus, and Camelpox virus can be detected, and both digital quality control product combinations 1 and 2 are detected to varying degrees. Since the sample is derived from virus seeds cultured in chicken embryos, there are genomic residues of Gallusgallus and Homo sapiens.
[0032] In this embodiment, digital quality control product combination 1 or combination 2 is used to monitor whether there are false positive data or false negative data caused by mutual interference between genomes in the bioinformatics analysis process. Therefore, the sequencing data is analyzed using the bioinformatics analysis process without incorporating digital quality control products to obtain the corresponding value S1_num, and the sequencing data is analyzed using the bioinformatics analysis process with different quantity gradients of digital quality control products incorporated to obtain the corresponding values S1_qc50_num, S1_qc500_num, S1_qc5000_num, and S1_qc1000000_num. By comparing the two sets of data, it can be found that the total number of reads belonging to Gallusgallus detected after adding different amounts of quality control sequences does not change, indicating that there is no interference between genes between the added quality control product sequences and the sequences in the genome of Gallusgallus during the analysis using the bioinformatics analysis process. The results show that the total number of reads detected for viruses such as Vaccinia virus, Cytomegalovirus humanbeta5, and Camelpox virus decreases, while the total number of reads detected for Homo sapiens increases, indicating that there may be some interference between the added quality control product sequences and the sequences in the genomes of the above species during the analysis using the bioinformatics analysis process. However, as long as the total number of reads detected for the above species still fluctuates within the methodological range (for example, the reference ELISA methodological range is 80% - 120%, and in this embodiment, the fluctuation range is 50% - 150%), the results of this bioinformatics analysis process are considered reliable, that is, the bioinformatics process in this embodiment exhibits good specificity and can distinguish target sequences from non-target sequences to avoid false positives or false negatives. It can be understood that the above fluctuation range can be adjusted by those skilled in the art according to the actual analysis results. If the total number of reads detected for a species exceeds the set fluctuation range, the bioinformatics analysis process should be optimized, such as adjusting the Cut-off value.
[0033] Therefore, the digital quality control product described in the present invention can be understood as spiking in the experiment to determine whether some key information is missing or added during the analysis of actual sample data by the bioinformatics analysis process, thereby detecting whether there are false positives or false negatives caused by mutual interference between genomes in the bioinformatics analysis process.
[0034] Table 2. Analysis results after digital quality control products were incorporated into sequencing data
[0035]
[0036] Note: S1_num is the number of reads detected in the input sample. S1_qc50_num, S1_qc500_num, S1_qc5000, and S1_qc1000000_num are the numbers of reads detected after adding digital quality control products with 50, 500, 5000, and 1000000 reads respectively to the sample sequencing data.
[0037] The embodiments described above are only descriptions of the preferred embodiments of the present invention and do not limit the scope of the present invention. Without departing from the design spirit of the present invention, various deformations and improvements made by those of ordinary skill in the art to the technical solutions of the present invention shall fall within the protection scope of the present invention.
Claims
1. Digital quality control products, characterized by: It includes a quality control sequence; the quality control sequence is selected from at least one of the base sequences SEQ ID NO.1-10.
2. The digital quality control product according to claim 1, characterized in that: The quality control sequence is selected from at least one of the base sequences SEQ ID NO.1-5.
3. The digital quality control product according to claim 1, characterized in that: The quality control sequence is selected from at least one of the base sequences SEQ ID NO.6-10.
4. The digital quality control product according to any one of claims 1 to 3, characterized in that: It also includes a high GC sequence; the high GC sequence is selected from at least one of the base sequences SEQ ID NO.11-13.
5. The digital quality control product according to any one of claims 1 to 3, characterized in that: It also includes a high AT sequence; the high AT sequence is the base sequence SEQ ID NO.
14.
6. The digital quality control product according to claim 1, characterized in that: Including quality control sequences, high GC sequences, and high AT sequences; The quality control sequence includes base sequences SEQ ID NO.1-5; The high GC sequence is selected from at least one of the base sequences SEQ ID NO.11 to 13; The high AT sequence is the base sequence SEQ ID NO.
14.
7. The digital quality control product according to claim 1, characterized in that: Including quality control sequences, high GC sequences, and high AT sequences; The quality control sequence includes base sequences SEQ ID NO.6-10; The high GC sequence is selected from at least one of the base sequences SEQ ID NO.11 to 13; The high AT sequence is the base sequence SEQ ID NO.
14.
8. Use of the digital quality control product as described in any one of claims 1 to 7 in monitoring the bioinformatics analysis process of biological sample metagenomes.
9. The use according to claim 8, characterized in that The application includes: incorporating the digital quality control product into sample sequencing data to perform metagenomic bioinformatics analysis; if the number of reads detected in the species genome is within a preset fluctuation range, the result of the metagenomic bioinformatics analysis is reliable.
10. The use according to claim 9, characterized in that The number of digital quality control products added is 50-1,000,000; and / or the fluctuation range is 50% to 150%.