Digital standards for verifying the accuracy of metagenomic biological information detection.
Patent Information
- Application Number
- JP2026513046
- Authority / Receiving Office
- JP · JP
- Patent Type
- Applications
- Current Assignee / Owner
- Priority Date
- 2023-08-25
- Filing Date
- 2024-08-26
- Publication Date
- 2026-09-08
AI Technical Summary
【0046】 [発明の効果] 本発明の範囲内で、本発明の上記の各技術的特徴と以下(例えば、実施例)に具体的に説明される各技術的特徴との間を、互いに組み合わせることにより、新しいまたは好ましい技術的解決策を構成することができることを理解されたい。スペースに限りがあるため、ここでは繰り返さない。
Smart Images

Figure 2026530471000001_ABST
Abstract
Description
Detailed Description of the Invention
[0001] [Technical Field] The present invention relates to the field of biological information. Specifically, the present invention relates to a digital standard product for verifying the accuracy of metagenomic biological information detection.
[0002] [Background Art] Metagenomics adopts a genomic research strategy by extracting DNA from all microorganisms in a microbial ecosystem sample and constructing a metagenomic library, to analyze the genetic composition of all microorganisms in the sample and the potential functions of their community. Since this omics technology does not rely on the isolation and pure culture technology of single bacteria, it can greatly solve the problem that most microorganisms are difficult to study due to the difficulty in isolation and culture, and at the same time, it can also reflect the real situation composed of microorganisms in the ecological environment that is the research object. In the research of microbial ecosystems related to human health, quantification of microorganisms based on metagenomic sequencing data is the basis for studying relevant rules such as the composition of the community, interspecific interactions, and the association with the occurrence and development of diseases. With the progress of scientific research, more and more studies have shown that accurate annotation of low-level taxonomic units, that is, strains, for specific species has become increasingly important. Only researching the relationship between bacteria and diseases at a high taxonomic level will easily lump together categories that show positive correlation, irrelevant correlation or negative correlation with the development of diseases, which has obvious defects both biologically and statistically. Existing research results need to be corrected by improving the accuracy of microbial quantification, or implemented through more in-depth mechanism research.
[0003] Currently, in the biological product detection industry, more and more people use metagenomic technology pathways to detect exogenous viruses. High-throughput metagenomic technology has the potential to detect all microorganisms (bacteria, viruses, fungi, etc.) in a sample, but subsequent verification work needs to be carried out urgently so that this detection technology can be popularized.
[0004] Recent advancements in metagenomics technology bypass conventional microbial isolation and culture methods, directly extracting total DNA from environmental samples to construct and screen metagenomic libraries, thereby obtaining novel functional genes and bioactive substances. These metagenomic libraries contain genetic information not only from culturable microorganisms but also from inculpable ones, thus increasing opportunities to obtain new bioactive substances. Furthermore, their quantitative methods offer higher resolution, allowing annotation down to the species or strain level, and are therefore widely applied in pathogen detection.
[0005] In the field of biological product detection, the industry currently primarily uses conventional methods such as PCR, electrophoresis, and culture. However, because these conventional methods have lower throughput and make it difficult to analyze large quantities of different types of viruses simultaneously, methods based on next-generation sequencing (NGS) technology are gradually being developed.
[0006] Next-generation sequencing can overcome the drawback of low throughput, but its high sensitivity, high background noise, and random sequencing characteristics make it difficult to distinguish between true and false positives in detection results. Simultaneously, sample preparation for next-generation sequencing, including DNA and RNA extraction and subsequent library construction processes, is susceptible to environmental or consumable contamination. Therefore, while promoting the safe detection of biological products using metagenomics, it is necessary to develop standards to validate the entire metagenomic workflow, including dry and wet experiments. Here, the methods and tools for dry experiments (biological information processing) are diverse, and currently, there are no universally accepted and highly reliable detection methods and tools. Therefore, it is essential to validate the biological information analysis process by constructing digital standards.
[0007] Therefore, in this field, there is a need for the development of digital standards to verify the accuracy of metagenomic biological information detection. [Overview of the prefecture] [Problems the invention aims to solve] The objective of this invention is to provide a digital standard for verifying the accuracy of metagenomic biological information detection.
[0008] [Means for solving the problem] A first aspect of the present invention provides a method for manufacturing a digital standard for detecting genomic biological information, comprising the following steps: (i) Provide n reference genomes of microorganisms, where n is an integer selected from 10 to 24, preferably 15 to 24, more preferably 20 to 24. (ii) Simulation community design: A simulation community is generated based on the reference genome and taxonomic information, (iii) Read simulation: A read simulator is used to sample each genome of the simulation community, thereby generating a simulation sample, and a simulation dataset is obtained by mixing the simulation samples. (iv) Standard assembly: Assemble the simulation dataset to obtain the digital standard.
[0009] In another preferred example, the microorganism is a virus. In another preferred example, the microorganisms include Betapolyomavirus macacae (monkey virus 40), Bluetongue virus, Bovine polyomavirus 3, Human betaherpesvirus 6A, Human betaherpesvirus 6B, Fowl aviadenovirus A (exogenous triadenovirus type 1), Simian retrovirus 8, Simian foamy virus, Human gammaherpesvirus 4 (human EB virus) (Epstein-Barr virus), Human polyomavirus 12, Human circovirus VS6600022, Human betaherpesvirus 5, Human immunodeficiency virus 1, and Murine The group consists of n microorganisms selected from the following: respirovirus (Sendai virus), Murid betaherpesvirus 8 (mouse salivary gland virus / mouse cytomegalovirus), Nipah henipavirus, Suid alphaherpesvirus 1 (pseudorabies virus), Swinepox virus, Porcine reproductive and respiratory syndrome virus, Suid betaherpesvirus 2 (porcine cytomegalovirus), Porcine circovirus 2 (porcine circovirus), Human gammaherpesvirus 8 (human herpesvirus-8), Bat rotavirus (mouse rotavirus), and Murid betaherpesvirus 1 (mouse salivary gland virus / mouse cytomegalovirus).
[0010] In another preferred example, the n microorganisms are Betapolyomavirus macacae (monkey virus 40), Bluetongue virus, Bovine polyomavirus 3 (bovine polyomavirus), Human betaherpesvirus 6A (human herpesvirus-6), Human betaherpesvirus 6B (human herpesvirus-6), Fowl aviadenovirus A (exogenous triadenovirus type 1), Simian retrovirus 8 (monkey retrovirus), Simian foamy virus (sulfomyvirus), Human gammaherpesvirus 4 (human EB virus) (Epstein-Barr virus), Human polyomavirus 12 (human polyomavirus), Human circovirus VS6600022 (human circovirus), Human betaherpesvirus 5 (human cytomegalovirus), Human immunodeficiency virus 1 (human immunodeficiency virus), Murine It is composed of microorganisms such as respirovirus (Sendai virus), Murid betaherpesvirus 8 (mouse salivary gland virus / mouse cytomegalovirus), Nipah henipavirus, Suid alphaherpesvirus 1 (pseudorabies virus), Swinepox virus, Porcine reproductive and respiratory syndrome virus, Suid betaherpesvirus 2 (porcine cytomegalovirus), Porcine circovirus 2 (porcine circovirus), Human gammaherpesvirus 8 (human herpesvirus-8), Bat rotavirus (mouse rotavirus), and Murid betaherpesvirus 1 (mouse salivary gland virus / mouse cytomegalovirus).
[0011] In another preferred example, the reference genome is obtained from the NCBI genome database. In another preferred example, steps (ii) to (iv) are performed using CAMISIM software.
[0012] In another preferred example, the simulation dataset described in step (ii) is generated by a method selected from the group consisting of de novo simulation, dataset simulation, or a combination thereof.
[0013] In another preferred example, the de novo simulation includes the following steps: A1) Provide specified taxonomic information, which includes the name of the taxonomic unit and the abundance of the taxonomic unit. A2) Provide the genome classification name of the reference genome in a known database, and match the classification unit name with the reference genome by mapping the classification unit name to the genome classification name in a known database (e.g., NCBI database classification ID). A3) Based on the novelty level of the reference genome and the corresponding classification unit name, m genomes are extracted from the reference genome to generate the simulation community.
[0014] In another preferred example, the predetermined taxonomic information is a CAMI file or a BIOM file. In another preferred example, the classification unit name of the species refers to the classification name of the operational unit (OTU).
[0015] In another preferred example, in step A2), the OUT classification name is mapped to the NCBI classification ID. In another preferred example, in step A2), the relationship between the classification unit and the mapped genome satisfies the following conditions: The shortest path between the classification unit name and the genome classification name is calculated based on the de Bruin diagram, and The aforementioned classification unit name and genome classification name belong to at least the same family.
[0016] In another preferred example, if no genome that meets the criteria exists in step A2), the classification unit is assigned to a custom-defined new "random" genome.
[0017] In another preferred example, step A3) includes the following steps: 1) Based on the total number of genomes m, the same number of genomes are extracted from each novelty category (new strain, new species, new genus, new family, new order), 2) If there is a shortage of extractable genomes in a particular novelty category, extract more genomes in the same quantity from other categories. 3) Repeat steps 1) and 2) until m genomes are extracted.
[0018] In another preferred example, the novelty category is determined based on the similarity between a reference genome sequence and a currently known viral sequence. In another preferred example, m is 15 to 40, preferably 20 to 25.
[0019] In another preferred example, the simulation community further includes additional genomes for stepwise simulations, the additional genomes being generated by the next step, a) Select a specific genome from the reference genome and extract a strain number from the geometric distribution, the strain number representing the number of related strains generated by the specific genome, b) Repeat this step until the total number of specified strains is reached. c) Additional genomes for the simulation are obtained by randomly selecting a specified strain number from the generated related strains.
[0020] In another preferred example, 20 to 50 related strains, preferably 30 to 40 related strains, are generated for each reference genome. In another preferred embodiment, the genetic distance between each of said related strains and the starting strain is increased stepwise with a step length of 0.1%, and the total distance is 2% to 5%, preferably 3% to 4%.
[0021] In another preferred embodiment, the data set simulation comprises the following steps: B1) sampling from a log-normal distribution to construct an abundance distribution sample set, B2) generating the simulated community based on a reference genome and the abundance distribution sample set.
[0022] In another preferred embodiment, in step B1), using default data of CAMISIM software, the parameter mu of the log-normal distribution is set to 1, and sigma is set to 2.
[0023] In another preferred embodiment, in step B1), using the same parameters, ≧2 abundance distribution samples, preferably ≧5, more preferably ≧10, are independently extracted from the log-normal distribution.
[0024] In another preferred embodiment, in step B1), for the (x+1)-th abundance distribution sample, a new value is obtained by sampling from the log-normal distribution, and the (x+1)-th abundance distribution sample is obtained by adding the new value to the value of the x-th abundance distribution sample and dividing the sum by 2.
[0025] In another preferred embodiment, in step (iii), simulated reads are generated based on specified parameters, and said parameters comprise read length, insert size, error rate, or a combination thereof.
[0026] In another preferred embodiment, the read length is 50 to 500 bp, preferably 100 to 350 bp, more preferably 150 to 250 bp. In another preferred example, the insertion size is 120 bp to 400 bp, preferably 220 bp to 320 bp, and more preferably 270 bp to 280 bp.
[0027] In another preferred example, the error rate is ≤10%, preferably ≤5%, and more preferably ≤1%. In another preferred example, in step (iii), samples are taken from each genome of the simulation population according to a ratio based on a specified genome abundance.
[0028] In another preferred example, the simulation dataset includes 10 to 500 simulation samples, preferably 50 to 200 simulation samples, and more preferably 100 to 150 simulation samples.
[0029] In another preferred example, the specified abundance in each sample is different. In another preferred example, step (iii) further includes generating a second read simulation using the same genomic abundance and different average insertion sizes, and mixing it with the previously obtained simulation sample.
[0030] In another preferred example, in step (iii), the size of each sample in the simulation dataset includes simulation reads of 1 to 10 G, preferably 3 to 5 G.
[0031] In another preferred example, in step (iii), the generated simulation dataset is for each sample, Simulation read sequences such as FASTQ files, and This includes alignment information between simulation reads such as BAM files and a reference genome.
[0032] In another preferred example, step (iv) includes the following step: From the IVA read simulation, the position of each read on the reference genome is obtained, and the coverage area of each region is calculated. ivb) Gold standard contig sequences are obtained by degrading the reference genome sequence at a zero coverage site.
[0033] In another preferred example, step (iv) includes the following step: The challenge dataset is obtained by shuffling and anonymizing the array of the IVC simulation dataset.
[0034] In another preferred example, the method further comprises, after step (iv), step (v), filtering the data based on sequencing data bases and sequence quality of the digital standard, and aligning and identifying a database constructed based on the digital standard and a reference genome using sequence alignment software, and removing sequences aligned to non-target viral genomes.
[0035] In another preferred example, the digital standard includes reference genome species and corresponding abundance information contained within each simulation sample. A second aspect of the present invention provides a digital standard for metagenomic bioinformation detection, the digital standard comprising p samples, each of which comprises genome sequences derived from n different microbial reference genomes. Here, the p samples have different genomic abundances.
[0036] In another preferred example, the microorganism is a virus. In another preferred example, the microorganisms include Betapolyomavirus macacae (monkey virus 40), Bluetongue virus, Bovine polyomavirus 3, Human betaherpesvirus 6A, Human betaherpesvirus 6B, Fowl aviadenovirus A (exogenous triadenovirus type 1), Simian retrovirus 8, Simian foamy virus, Human gammaherpesvirus 4 (human EB virus) (Epstein-Barr virus), Human polyomavirus 12, Human circovirus VS6600022, Human betaherpesvirus 5, Human immunodeficiency virus 1, and Murine The group consists of n microorganisms selected from the following: respirovirus (Sendai virus), Murid betaherpesvirus 8 (mouse salivary gland virus / mouse cytomegalovirus), Nipah henipavirus, Suid alphaherpesvirus 1 (pseudorabies virus), Swinepox virus, Porcine reproductive and respiratory syndrome virus, Suid betaherpesvirus 2 (porcine cytomegalovirus), Porcine circovirus 2 (porcine circovirus), Human gammaherpesvirus 8 (human herpesvirus-8), Bat rotavirus (mouse rotavirus), and Murid betaherpesvirus 1 (mouse salivary gland virus / mouse cytomegalovirus).
[0037] In another preferred example, the n microorganisms are Betapolyomavirus macacae (monkey virus 40), Bluetongue virus, Bovine polyomavirus 3 (bovine polyomavirus), Human betaherpesvirus 6A (human herpesvirus-6), Human betaherpesvirus 6B (human herpesvirus-6), Fowl aviadenovirus A (exogenous triadenovirus type 1), Simian retrovirus 8 (monkey retrovirus), Simian foamy virus (sulfomyvirus), Human gammaherpesvirus 4 (human EB virus) (Epstein-Barr virus), Human polyomavirus 12 (human polyomavirus), Human circovirus VS6600022 (human circovirus), Human betaherpesvirus 5 (human cytomegalovirus), Human immunodeficiency virus 1 (human immunodeficiency virus), Murine It is composed of microorganisms such as respirovirus (Sendai virus), Murid betaherpesvirus 8 (mouse salivary gland virus / mouse cytomegalovirus), Nipah henipavirus, Suid alphaherpesvirus 1 (pseudorabies virus), Swinepox virus, Porcine reproductive and respiratory syndrome virus, Suid betaherpesvirus 2 (porcine cytomegalovirus), Porcine circovirus 2 (porcine circovirus), Human gammaherpesvirus 8 (human herpesvirus-8), Bat rotavirus (mouse rotavirus), and Murid betaherpesvirus 1 (mouse salivary gland virus / mouse cytomegalovirus).
[0038] In another preferred example, p is an integer selected from 10 to 500, preferably 50 to 200, and more preferably 100 to 150. In another preferred example, n is an integer selected from 10 to 24, preferably 15 to 24, and more preferably 20 to 24.
[0039] In another preferred example, the digital standard includes reference genome species and corresponding abundance information contained within each sample. In another preferred example, the digital standard is filtered, which means aligning and identifying a database built on the digital standard and a reference genome using sequence alignment software to remove sequences aligned to non-target viral genomes.
[0040] In another preferred example, the digital standard is constructed by the method described in the first aspect of the present invention. A third aspect of the present invention provides an apparatus for manufacturing digital standards for metagenomic biological information detection, the apparatus comprising: 1) An input module used to input n reference genomes of microorganisms, where n is an integer selected from 10 to 24, preferably 15 to 24, more preferably 20 to 24, and 2) Generation of a simulated community based on the above-mentioned reference genome and taxonomic information, A read simulator is used to sample from each genome of the simulated community, thereby generating a simulation sample, and the simulation dataset is obtained by mixing the simulation samples. A simulation module used to obtain the digital standard by assembling the aforementioned simulation dataset, and 3) Output module: The output module is used to output the digital standard.
[0041] In another preferred example, the microorganism is a virus. In another preferred example, the microorganisms include Betapolyomavirus macacae (monkey virus 40), Bluetongue virus, Bovine polyomavirus 3, Human betaherpesvirus 6A, Human betaherpesvirus 6B, Fowl aviadenovirus A (exogenous triadenovirus type 1), Simian retrovirus 8, Simian foamy virus, Human gammaherpesvirus 4 (human EB virus) (Epstein-Barr virus), Human polyomavirus 12, Human circovirus VS6600022, Human betaherpesvirus 5, Human immunodeficiency virus 1, and Murine The group consists of n microorganisms selected from the following: respirovirus (Sendai virus), Murid betaherpesvirus 8 (mouse salivary gland virus / mouse cytomegalovirus), Nipah henipavirus, Suid alphaherpesvirus 1 (pseudorabies virus), Swinepox virus, Porcine reproductive and respiratory syndrome virus, Suid betaherpesvirus 2 (porcine cytomegalovirus), Porcine circovirus 2 (porcine circovirus), Human gammaherpesvirus 8 (human herpesvirus-8), Bat rotavirus (mouse rotavirus), and Murid betaherpesvirus 1 (mouse salivary gland virus / mouse cytomegalovirus).
[0042] In another preferred example, the n microorganisms are Betapolyomavirus macacae (monkey virus 40), Bluetongue virus, Bovine polyomavirus 3 (bovine polyomavirus), Human betaherpesvirus 6A (human herpesvirus-6), Human betaherpesvirus 6B (human herpesvirus-6), Fowl aviadenovirus A (exogenous triadenovirus type 1), Simian retrovirus 8 (monkey retrovirus), Simian foamy virus (sulfomyvirus), Human gammaherpesvirus 4 (human EB virus) (Epstein-Barr virus), Human polyomavirus 12 (human polyomavirus), Human circovirus VS6600022 (human circovirus), Human betaherpesvirus 5 (human cytomegalovirus), Human immunodeficiency virus 1 (human immunodeficiency virus), Murine It is composed of microorganisms such as respirovirus (Sendai virus), Murid betaherpesvirus 8 (mouse salivary gland virus / mouse cytomegalovirus), Nipah henipavirus, Suid alphaherpesvirus 1 (pseudorabies virus), Swinepox virus, Porcine reproductive and respiratory syndrome virus, Suid betaherpesvirus 2 (porcine cytomegalovirus), Porcine circovirus 2 (porcine circovirus), Human gammaherpesvirus 8 (human herpesvirus-8), Bat rotavirus (mouse rotavirus), and Murid betaherpesvirus 1 (mouse salivary gland virus / mouse cytomegalovirus).
[0043] A fourth aspect of the present invention provides a method for verifying the accuracy of metagenomic biological information detection. I) Providing a method for detecting metagenomic biological information to be tested, II) A step of obtaining the detection result of the detection method by detecting each sample in the digital standard described in the second aspect of the present invention using the detection method, III) The step of determining the accuracy of the detection method by comparing the detection results of the detection method with the reference genome species and corresponding abundance information of each sample in the digital standard.
[0044] In another preferred example, the method includes comparing whether the abundance information of the same virus in a standard sample matches the detection results of the biological information detection method under test.
[0045] A fifth aspect of the present invention provides an apparatus for verifying the accuracy of a metagenomic biological information detection method, the apparatus comprising: I) An input module used to input the metagenomic biological information detection method to be tested, II) A calculation module used to obtain the detection result of the method under test by detecting the digital standard described in the second aspect of the present invention using the detection method described above, III) A comparison module used to determine the accuracy of the biological information detection method by comparing the detection results of the detection method with the reference genome species and corresponding abundance information of each sample in the digital standard, IV) Includes an output module used to output the accuracy determination results of the biological information detection method.
[0046] [Effects of the invention] It should be understood that, within the scope of the present invention, new or preferred technical solutions can be constructed by combining the above-described technical features of the present invention with the technical features specifically described below (e.g., in the examples). Due to space limitations, this will not be repeated here. [Brief explanation of the drawing]
[0047] The following drawings are used to illustrate specific embodiments of the present invention and are not intended to limit the scope of the invention as defined by the claims. [Figure 1] This diagram shows the process of constructing simulation data for a specific reference genome using CAMISIM. [Figure 2] This example demonstrates the results of verifying the accuracy of bioinformation detection alignment using digital standards. [Figure 3] This shows the results of overlapping virus types in the standard and detection results, for which a digital standard was constructed using 59 types of viruses. [Figure 4] This shows an example of the accuracy results of standard validation bioinformation detection alignment, which was constructed using 59 types of viruses to create digital standards. [Modes for carrying out the invention]
[0048] Through extensive and meticulous research, the inventors have developed the first digital standards for verifying the accuracy of the metagenomic bioinformation detection process. Specifically, the present invention selected 24 particularly noteworthy viral genomes from the Chinese Pharmacopoeia and constructed digital standards for verifying the accuracy of the metagenomic bioinformation detection process, including qualitative and quantitative methods. The present invention constructed a total of 100 samples, each containing the genome sequences of 24 different viruses, with randomly generated abundances, resulting in a different distribution of viral abundances in each sample. Through these digital standards, the accuracy of the bioinformation process constructed based on the metagenomic can be effectively verified. Based on this, the present invention was completed.
[0049] Standard product The digital standard of the present invention is constructed by the method of the first aspect of the present invention. Specifically, it is constructed based on a known viral genome selected from Table 1.
[0050] [Table 1] TIFF2026530471000003.tif250170TIFF2026530471000004.tif70170
[0051] The size of a viral genome is at least several kilobytes to tens of kilobytes. The difference in base pairs between a viral subtype and the original is only 100 to 200 base pairs, representing a very small proportion of the entire genome. Since next-generation sequencing read lengths are typically 250 to 300 bp, it is difficult to reliably extract sequences containing different base pairs during the random library construction process. Current viral alignment methods still rely on sequencing accuracy, and conventional methods tolerate a single base mismatch. In subtype classification, even a single base difference can represent a virus of a different subtype, potentially leading to alignment errors in the final sequencing results. Furthermore, current viral libraries largely contain sequences that are highly conserved between viruses, and no sequence library exists that can completely distinguish between subtypes of viruses.
[0052] This invention involves selecting particularly important viruses from the Chinese Pharmacopoeia and screening approximately 100 viruses that are important in the Chinese Pharmacopoeia. The results show that these viruses contain many viral subtypes, and their genomes contain highly similar viral subtype sequences. As a result, the manufactured digital standards contain a large number of similar sequences, and when the accuracy of a bioinformation detection method is evaluated using such digital standards, interference occurs, making it difficult to obtain truly reliable evaluation results.
[0053] As a result of extensive testing and screening, the present invention has selected 24 viruses shown in Table 1. These viruses have diverse origins, including bovine, mouse, human, and porcine, and do not contain or contain very few similar viral sequences. Therefore, they do not interfere with the test results of digital standards and can be accurately used to evaluate bioinformation detection methods. At the same time, all of these viruses are important viruses in the Chinese Pharmacopoeia, and the digital standards composed of them have extremely high clinical significance and application value.
[0054] CAMISIM The present invention provides a method for producing digital standards using a microbial community and metagenomic simulator, CAMISIM. The CAMISIM software can simulate various microbial abundance profiles, multiple sample time series, and differential abundance studies, including actual and simulated strain-level diversity, and can generate second and third-generation sequencing data.
[0055] CAMISIM allows customization of many attributes of the generated communities and datasets, such as the total number of genomes, species diversity, genome abundance distribution, sample size, number of replicates, and the sequencing technology used. Sequencing data simulated with CAMISIM software exhibits high agreement with actual data. Based on conventional techniques, two simulated multi-sample datasets of human and mouse gut microbiota generated with CAMISIM software show results that demonstrate high agreement with actual data. Simultaneously, using thousands of small datasets generated with CAMISIM, we study the effects of different genomic diversity, sequencing depth, and read errors on two widely used metagenomic sequence assembly software, MEGAHIT and metaSPAdes.
[0056] CAMISIM's data simulation is divided into the following three stages: 1) This includes community design, selection of community members and their genomes, and assignment of their relative abundances. Genome selection is based on a truncated geometric distribution, and abundance is based on a log-normal distribution.
[0057] 2) Metagenome sequencing data simulation, 3) Includes post-processing, binning, and assembly stages. Kraken2 In one embodiment, the digital standard manufacturing process of the present invention can be aligned with known sequences, thereby improving the accuracy of final validation by removing simulated sequences from non-target viral genomes. This alignment can be performed by Kraken software.
[0058] Kraken and Kraken2 (an improved version of Kraken) are the most common software for microbiome analysis, and compared to other software with similar features, they are characterized by their speed and high accuracy. The principle is that it is a highly accurate metagenomic sequence classification software based on the k-mer algorithm, which can rapidly classify sequenced (or simulated) reads into species.
[0059] Using Kraken software involves two stages: preparation (library building) and identification. Stage 1: Preparation (Library Construction) • Building a taxon database corresponding to k-mer (Box1) • Mapping of database and index files into memory Stage 2: Identification • Cleavage of the k-mer of the sequence to be identified. The LCA_taxon(Box2) obtained by aligning k-mer data in the database, and the number of times the alignment was performed.
[0060] After constructing a classification tree from the above data, calculate the sum of all weights up to each root-to-leaf point, and select the sequence that is the largest in the classification tree of that sequence.
[0061] Methods for building digital standards The present invention's method for constructing digital standards requires 1) providing a CAMISIM tool and 2) providing a reference genome of a specific microorganism.
[0062] In one embodiment, the present invention's construction method includes: 1) installing the CAMISIM tool; 2) downloading a reference genome of a specific microorganism; 3) performing a standard data simulation using the reference genome according to the CAMISIM tool's tutorial process (the process is shown in Figure 1); and 4) performing data alignment on the simulated data using Kraken2 (reference) and removing sequences aligned to non-target viral genomes.
[0063] In the digital standard construction process of the present invention, good simulation array accuracy is obtained using a specific amount of simulation data, simulation read length, and insertion fragment size. The read length used in the construction process of the present invention may be 50 to 500 bp, preferably 100 to 350 bp, more preferably 150 to 250 bp, the insertion fragment size may be 120 to 400 bp, preferably 220 to 320 bp, more preferably 270 to 280 bp, and the size of each sample in the simulation dataset may be a simulation read containing 1 to 10 G, preferably 3 to 5 G. In a preferred embodiment, the construction process of the present invention uses a read length of 150 bp, the insertion fragment size is set to 270, and the amount of data for each simulation sample is set to a 5 G r read.
[0064] Virus detection is performed based on simulation data, and due to sequence similarity issues, alignment with viruses from other non-reference genomes is possible. Therefore, the method of the present invention may also include a step of filtering the digital standard after obtaining it by simulation. Specifically, the data is filtered based on the sequencing data bases and sequence quality of the digital standard, and the database constructed based on the digital standard and the reference genome is aligned and identified using sequence alignment software to remove sequences aligned to non-target viral genomes. The digital standard obtained by the present invention removes sequences that easily misalign with sequences not present in the reference viral genome, and misalignment situations hardly occur even when using the standard of the present invention.
[0065] The main advantages of this invention are as follows: 1) The present invention can verify the accuracy of metagenomic bioinformatics processes in the detection of specific viruses. It helps establish the accuracy and versatility of next-generation sequencing technology in the detection of biological products and is advantageous for the use of next-generation sequencing in the field of biological products. 2) The present invention develops metagenomic data containing 24 viruses from the Chinese Pharmacopoeia, which have known virus species and viral abundances, and can therefore be used to evaluate the sensitivity and specificity of the metagenomic bioinformation detection process. At the same time, the accuracy of the process quantification can be verified.
[0066] 3) The viruses selected in this invention have diverse origins, including those from cattle, mice, humans, and pigs, and do not contain or contain very few similar viral sequences. This significantly reduces the interference of viral subtypes on the test results of the digital standard. At the same time, all of these viruses are important viruses in the Chinese Pharmacopoeia, and the digital standard composed of them has extremely high clinical significance and application value.
[0067] 4) The present invention screens for advantageous simulation data volume, simulation read length, and insertion fragment size by constructing digital standards, thereby obtaining good accuracy for simulation sequences.
[0068] The present invention will be further described below in conjunction with specific examples. These examples are used solely to illustrate the present invention and should not be used to limit its scope. In the following examples, experimental methods that do not specify conditions typically follow conventional conditions, such as those described in Sambrook et al., Molecular Cloning: An Experimental Manual (New York: Cold Spring Harbor Laboratory Press, 1989), or conditions proposed by the manufacturer. Unless otherwise specified, percentages and parts are calculated as weight percentages and weight parts.
[0069] Example 1. Construction of a metagenomic biological information detection standard. 1.1. Installing the CAMISIM tool Using Git, download version 71a1d86 from https: / / github.com / CAMI-challenge / CAMISIM and install it on your server following the tutorial.
[0070] 1.2. Downloading a reference genome of a specific microorganism Download the genomes of 24 particularly noteworthy viruses from the Chinese Pharmacopoeia using the NCBI Genome Database (November 19, 2022 edition).
[0071] 1.3. Metagenome Simulation Method 1: de novo metagenomic simulation 1.3.1. Input of the reference genome For de novo metagenomic simulations, ownership sequences of less than 1000 base pairs are removed from the provided genome, and the validation sequences contain only valid characters. The input genome can be a draft genome in fasta format, which may include ambiguous DNA characters such as "RYWSMKHBVDN" in addition to the bases ACGT. Here, all remaining non-base characters are removed except for those with probability "N".
[0072] 1.3.2. Taxonomic information If the input taxonomic information is in CAMI or BIOM format, its taxonomic units are matched with closely related complete genomes (e.g., novel genomes) in NCBI RefSeq or other reference files, and the simulated metagenome (genome and its abundance) is made to reflect the input profile as accurately as possible. This step is performed by matching scientific names in the input configuration file. The NCBI classification ID can be estimated from the classification names of Operational Taxonomic Units (OTUs) in the BIOM configuration file. The classification ID is also required for all reference genomes.
[0073] In the next step, the phylogenetic relationships between OTUs and genomes are compared, and the OTUs are mapped to genomes with the shortest pathway between the OTU classification ID and the genome classification ID, the shortest pathway being calculated based on the de Bruin diagram. Because the shortest pathway can be very long, a threshold is set at the family level, ensuring that OTUs are not mapped to genomes belonging to different families; i.e., each mapped genome belongs to at least one family. Typically, for this reason, some OTUs are not mapped if the scientific name of the OTU does not correspond to an NCBI classification ID, or if a complete genome does not exist for the classification ID. Unmapped OUTs can be assigned to "random" genomes in the reference set, as this would increase complexity through additional selection.
[0074] 1.3.3. Crowd design For de novo community design, a specified number of sequences are selected from all input sequences (e.g., circular elements (i.e., sequence segments representing a specific microorganism) or genomes) based on a specific standard to generate a specific dataset, where the specified number represents the number of generated viral sequence types to be simulated. The novelty level is determined based on the similarity between combinations of genome sequences and currently known viral sequences. Random selection of genomes in a specific dataset is performed based on the novelty level and members in the OTU. To increase taxonomic spread, for a specified total number of genomes to be included in a specific dataset, the selection algorithm attempts to extract the same number of genomes from each novelty level (new strain, new species, new genus, new family, new order). If there are not enough genomes in a category to reach the specified number, more genomes are extracted in equal amounts from other categories, if possible. This operation is repeated until the specified number of genomes are extracted.
[0075] 1.3.4. Adding a Genome To increase strain-level diversity in the dataset, we simulated many additional genomes using sgEvolver12, which represent strains associated with the reference genome selected above.
[0076] Therefore, a single genome is randomly selected, and a strain number is extracted from the geometric distribution (p=0.3), which represents the number of associated strains created for that particular genome. This process is repeated until the total number of specified strains generated is reached. For each genome, sgEvolver simulates 40 strains, and the genetic distance between each strain and the preceding strain increases stepwise (with a step length of 0.1%) until the total distance reaches 4%. From the 40 strains, a specified strain number is randomly selected and added to the genome set.
[0077] 1.4. Metagenome Simulation Method 2: Dataset Simulation For dataset simulations, each dataset requires the genome and the abundance distribution of circular elements.
[0078] For a single sample dataset, the abundance distribution is constructed by sampling from a log-normal distribution. By default, the log-log mu of the log distribution is set to 1, and sigma is set to 2. The higher the sigma, the flatter the normal distribution becomes, with smaller peaks. To reflect differential abundance experiments, two or more abundance samples can be independently extracted from the log-normal distribution using the same parameter settings.
[0079] For continuous samples in a time series, the abundance of the next sample is calculated by sampling new values from a log-normal distribution, adding these values to the previous values, and dividing by 2. The absolute abundance of circular elements can be reproduced by setting it to 15 times the microbial genome abundance. For each sample, the taxonomic spectrum of higher-level taxa is calculated based on the genome abundance value and its classification ID.
[0080] 1.5. Lead Simulation Based on the genome abundance distribution, genome size, and specified metagenomic sample output size, coverage in the simulation sample is calculated for each genome. Next, the selected read simulator generates reads using user-specified techniques and attributes (e.g., read length, insertion size, or error rate (if applicable)) and samples each genome proportionally based on the abundance specified in the population. The read length is set to 150 bp and the insertion size to 270 bp. The use of mixed techniques can be reproduced with the same sample by running the read simulation again on the dataset using different average insertion sizes and the same genome abundance as before. The size of each sample in each dataset is set to 5G reads.
[0081] The output at this stage targets each sample in the dataset, along with FASTQ and BAM files, which specify the position of each read in the input sequence.
[0082] 1.6. Standard Assembly Using SAMtools, gold standard assembly is performed on each sample and all samples (integrated gold standard) in each dataset, including each sequence location with at least one coverage. Specifically, the SAM file obtained from read simulation specifies the location of each read on the reference genome and is used with SAMtools to calculate the coverage of each region. Next, the reference genome sequence is decomposed into gold standard contig sequences at the zero-coverage sites. Then, the contig sequence is marked with a suffix of the unique dataset and sample identification prefix and ascending index. In the final step, the sequences of the simulation dataset are shuffled and anonymized using the unix command "shuf" to create the challenge dataset.
[0083] Example 2. Construction of a metagenomic bioinformation analysis process A digital standard is constructed using the de novo simulation method according to the method described in Example 1.
[0084] 2.1. Construction of a proprietary metagenomic biological information data analysis process a) Filter the data based on the sequencing data bases and sequence quality of the digital standard using common sequence quality control software such as FastQC or FastP.
[0085] b) Download viral sequence alignment software (e.g., Kraken2, Centrifuge, etc.) and build a viral database using these 24 viral genomes.
[0086] c) Using the above software and preset parameters, perform viral sequence alignment and identification on the sample, and remove sequences aligned to non-target viral genomes.
[0087] 2.2. Accuracy testing of biological information detection methods The virus detection results are verified based on the types and amounts of viruses contained in each digital standard provided in this invention. The results are shown in Figure 2, and the accuracy of the biological information detection method can be determined using the digital standards of this invention.
[0088] Comparative Example Due to the limitations of current sequence alignment methods and virus libraries, it is extremely difficult to perfectly align subtype viruses using alignment methods. Before testing the 24 viruses in this study, the inventors manufactured and validated digital standards using 59 particularly noteworthy viruses from the Chinese Pharmacopoeia. As a result, it was found that many false positives occurred when using standards constructed with 59 viruses. As shown in Figure 3, although theoretically only about 59 viruses should have been detected, a total of 136 virus types were detected after detection using the bioinformatics process. At the same time, the results in Figure 4 also show that false positives and missed detections exist for some virus types in the standards.
[0089] Analysis suggests that this phenomenon may be caused by interference from subtypes present within these viruses, and that false positives occur because the genome sequences in digital standards are very similar to other sequences not present in the target viruses.
[0090] All documents referenced in this invention are cited as references in this application, as if each document were cited individually. Furthermore, after reading the above teachings of this invention, persons skilled in the art can make various changes or modifications to the invention, and these equivalent forms are also included within the scope defined by the claims appended to this application.
Claims
1. A method for manufacturing a digital standard for the detection of metagenomic biological information, Including the following stage, (i) Provide n reference genomes of microorganisms, where n is an integer selected from 10 to 24, preferably 15 to 24, more preferably 20 to 24. (ii) Simulation community design: A simulation community is generated based on the reference genome and taxonomic information. (iii) Size simulation: Using a read simulator, sampling is performed from each genome of the simulation community to generate simulation samples, and a simulation dataset is obtained by mixing the simulation samples. (iv) Standard assembly: A method for manufacturing a digital standard, characterized by assembling the simulation dataset and thereby obtaining the digital standard.
2. The aforementioned n microorganisms are Betapolyomavirus macacae (monkey virus 40), Bluetongue virus, Bovine polyomavirus 3 (bovine polyomavirus), Human betaherpesvirus 6A (human herpesvirus-6), Human betaherpesvirus 6B (human herpesvirus-6), Fowl aviadenovirus A (exogenous triadenovirus type 1), Simian retrovirus 8 (monkey retrovirus), Simian foamy virus (sulfomyvirus), Human gammaherpesvirus 4 (Human EB virus) (Epstein-Barr virus), Human polyomavirus 12 (Human polyomavirus), Human circovirus VS6600022 (Human circovirus), Human betaherpesvirus 5 (Human cytomegalovirus), Human immunodeficiency virus 1 (Human immunodeficiency virus), Murine respirovirus (Sendai virus), Murid betaherpesvirus 8 (Mouse salivary gland virus - Mouse cytomegalovirus), Nipah henipavirus (Nipah virus), Suid alphaherpesvirus It is characterized by being composed of the following microorganisms: 1 (pseudorabies virus), Swinepox virus, Porcine reproductive and respiratory syndrome virus, Suid betaherpesvirus 2 (porcine cytomegalovirus), Porcine circovirus 2 (porcine circovirus), Human gammaherpesvirus 8 (human herpesvirus-8), Bat rotavirus (mouse rotavirus), and Murid betaherpesvirus 1 (mouse salivary gland virus - mouse cytomegalovirus). A method for manufacturing a digital standard according to claim 1.
3. The aforementioned steps (ii) to (iv) are characterized by being performed using CAMISIM software. A method for manufacturing a digital standard according to claim 1.
4. The simulation dataset comprises 10 to 500 simulation samples, preferably 50 to 200 simulation samples, more preferably 100 to 150 simulation samples, and is characterized in that the specified abundance in each sample is different. A method for manufacturing a digital standard according to claim 1.
5. The method further comprises step (v) after step (iv): filtering the data based on the sequencing data bases and sequence quality of the digital standard, and aligning and identifying the database constructed based on the digital standard and the reference genome using sequence alignment software, and removing sequences aligned to non-target viral genomes. A method for manufacturing a digital standard according to claim 1.
6. A digital standard for the detection of metagenomic biological information, The aforementioned digital standard comprises p samples, where each sample comprises genome sequences derived from n different microbial reference genomes. Here, the p samples have different genomic abundances, and p is an integer selected from 10 to 500, preferably 50 to 200, and more preferably 100 to 150. The digital standard is characterized in that n is an integer selected from 10 to 24, preferably 15 to 24, and more preferably 20 to 24.
7. The aforementioned n microorganisms are Betapolyomavirus macacae (monkey virus 40), Bluetongue virus, Bovine polyomavirus 3 (bovine polyomavirus), Human betaherpesvirus 6A (human herpesvirus-6), Human betaherpesvirus 6B (human herpesvirus-6), Fowl aviadenovirus A (exogenous triadenovirus type 1), Simian retrovirus 8 (monkey retrovirus), Simian foamy virus (sulfomyvirus), Human gammaherpesvirus 4 (Human EB virus) (Epstein-Barr virus), Human polyomavirus 12 (Human polyomavirus), Human circovirus VS6600022 (Human circovirus), Human betaherpesvirus 5 (Human cytomegalovirus), Human immunodeficiency virus 1 (Human immunodeficiency virus), Murine respirovirus (Sendai virus), Murid betaherpesvirus 8 (Mouse salivary gland virus - Mouse cytomegalovirus), Nipah henipavirus (Nipah virus), Suid alphaherpesvirus It is characterized by being composed of the following microorganisms: 1 (pseudorabies virus), Swinepox virus, Porcine reproductive and respiratory syndrome virus, Suid betaherpesvirus 2 (porcine cytomegalovirus), Porcine circovirus 2 (porcine circovirus), Human gammaherpesvirus 8 (human herpesvirus-8), Bat rotavirus (mouse rotavirus), and Murid betaherpesvirus 1 (mouse salivary gland virus - mouse cytomegalovirus). The digital standard product according to claim 6.
8. The digital standard is filtered, and the filtering is characterized by aligning and identifying a database constructed based on the digital standard and a reference genome using sequence alignment software, and removing sequences aligned to non-target viral genomes. The digital standard product according to claim 6.
9. The aforementioned digital standard is characterized by being constructed by the method described in claim 1. The digital standard product according to claim 6.
10. A manufacturing apparatus for digital standards for metagenomic biological information detection, The aforementioned device is 1) An input module used to input n reference genomes of microorganisms, where n is an integer selected from 10 to 24, preferably 15 to 24, more preferably 20 to 24, and 2) Generation of a simulated community based on the above-mentioned reference genome and taxonomic information, A read simulator is used to sample from each genome of the simulated community, thereby generating a simulation sample, and the simulation dataset is obtained by mixing the simulation samples. A simulation module used to obtain the digital standard by assembling the aforementioned simulation dataset, 3) The manufacturing apparatus, characterized by including an output module used to output the digital standard.
11. The aforementioned n microorganisms are Betapolyomavirus macacae (monkey virus 40), Bluetongue virus, Bovine polyomavirus 3 (bovine polyomavirus), Human betaherpesvirus 6A (human herpesvirus-6), Human betaherpesvirus 6B (human herpesvirus-6), Fowl aviadenovirus A (exogenous triadenovirus type 1), Simian retrovirus 8 (monkey retrovirus), Simian foamy virus (sulfomyvirus), Human gammaherpesvirus 4 (Human EB virus) (Epstein-Barr virus), Human polyomavirus 12 (Human polyomavirus), Human circovirus VS6600022 (Human circovirus), Human betaherpesvirus 5 (Human cytomegalovirus), Human immunodeficiency virus 1 (Human immunodeficiency virus), Murine respirovirus (Sendai virus), Murid betaherpesvirus 8 (Mouse salivary gland virus - Mouse cytomegalovirus), Nipah henipavirus (Nipah virus), Suid alphaherpesvirus It is characterized by being composed of the following microorganisms: 1 (pseudorabies virus), Swinepox virus, Porcine reproductive and respiratory syndrome virus, Suid betaherpesvirus 2 (porcine cytomegalovirus), Porcine circovirus 2 (porcine circovirus), Human gammaherpesvirus 8 (human herpesvirus-8), Bat rotavirus (mouse rotavirus), and Murid betaherpesvirus 1 (mouse salivary gland virus - mouse cytomegalovirus). The manufacturing apparatus according to claim 10.
12. A method for verifying the accuracy of metagenomic biological information detection, I) A step of providing a method for detecting metagenomic biological information to be tested, II) A step of obtaining the detection result of the detection method by detecting each sample in the digital standard described in any one of claims 6 to 10 using the detection method, III) A method for verifying the accuracy of metagenomic biological information detection, comprising the step of determining the accuracy of the detection method by comparing the detection result of the detection method with the reference genome species and corresponding abundance information of each sample in the digital standard.
13. A device for verifying the accuracy of a metagenomic biological information detection method, The aforementioned device is I) An input module used to input the metagenomic biological information detection method for the subject of the test, II) A calculation module used to obtain the detection result of the method under test by detecting the digital standard described in any one of claims 6 to 10 using the detection method described above, III) A comparison module used to determine the accuracy of the biological information detection method by comparing the detection results of the detection method with the reference genome species and corresponding abundance information of each sample in the digital standard, IV) An apparatus for verifying the accuracy of a metagenomic biological information detection method, comprising an output module used to output the accuracy determination results of the biological information detection method.