A bioinformatics denoising analysis method and system based on hidden subgroups

By using the Hugo-DNA method, hidden subgroups can be identified by utilizing the ratio of test samples to negative and bioinformatics controls, thus solving the problem of noise removal in metagenomic sequencing and enabling more accurate data interpretation.

CN115719614BActive Publication Date: 2026-05-19HUGOBIOTECH BEIJING CO LTD +1
View PDF 1 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
HUGOBIOTECH BEIJING CO LTD
Filing Date
2021-08-27
Publication Date
2026-05-19

AI Technical Summary

Technical Problem

Existing metagenomic sequencing methods have shortcomings in noise removal. Conventional methods cannot effectively distinguish between real and noise signals, leading to inaccurate data interpretation.

Method used

A noise reduction analysis method based on hidden subgroups (Hugo-DNA) was adopted. By comparing the proportion of taxonomic units of the test sample with negative control and bioinformatics control, a model was constructed to remove noise signals. Computer graph theory was used to identify hidden subgroups and remove noise.

Benefits of technology

It improves the sensitivity and specificity of real signals, effectively removes noise signals introduced by experiments and bioinformatics analysis, and enhances the accuracy of data interpretation.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN115719614B_ABST
    Figure CN115719614B_ABST
Patent Text Reader

Abstract

The application relates to a bioinformatics analysis method and system for removing background noise of sequencing data. The method is based on the fact that introduced noise signals form hidden subgroups, analyzes the connection between components in the sample, and thus more effectively removes the introduced noise signals while retaining the real signals to the greatest extent. The application does not rely on any prior knowledge of the real signals and noise signals to be measured, and can efficiently remove the noise.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application relates to the field of bioinformatics analysis, specifically to a method for background noise removal from sequencing data. Technical Background

[0002] The composition of microbial communities is closely related to environmental ecosystems, human health, and clinical diseases. Metagenomic next-generation sequencing (mNGS) is a method that can detect all nucleic acid molecules in a sample, providing a quantitative description of the microbial community composition. Due to its extremely high sensitivity and extremely low limit of detection (LOD), metagenomic methods can detect minute amounts of real nucleic acid molecules, but they can also detect nucleic acid molecules from contaminants introduced during the experiment (such as environmental microorganisms). More problematic is that, due to the high similarity between microbial species, data from a single microorganism can be annotated as belonging to multiple microorganisms in subsequent bioinformatics analyses, leading to false positives. Without effective analysis and processing, these noise signals—i.e., experimentally introduced contaminants and bioinformatics-introduced false positives—can lead to inaccurate data interpretation.

[0003] Metagenomic sequencing methods involve several steps, including extracting all nucleic acid molecules from a sample, constructing a library, sequencing the library, and data analysis. All of these steps can introduce noise signals. Although some literature reports that contamination can be controlled and prevented through rigorous experimental procedures, these methods have not been very effective. Therefore, the mainstream approach is to use bioinformatics analysis tools to eliminate background noise during later data processing. One widely used method is to remove all components below a specified relative abundance threshold. However, this method relies on the choice of threshold and often removes low-abundance true signals while retaining high-abundance noise signals. A second method uses a negative control for noise removal. In practice, when performing metagenomic sequencing on the sample to be tested, a negative control (i.e., water without nucleic acids) is simultaneously set up, and the same procedures as the sample to be tested are performed concurrently to simulate a background without true signals. Therefore, existing methods often use the composition of negative controls to correct the composition of the test sample, or directly remove all identification results reported by the negative controls, or remove components that are more abundant in the negative controls but less abundant in the test sample after normalization (such as the Z-score method). However, due to the randomness of sequencing, this method often erroneously removes true signals and retains noise signals. The third type of analytical method requires a large amount of prior information about contamination to maintain a "blacklist," and directly removes components from this "blacklist" in the detection of the test sample. However, this prior information is often unclear, and because noise signals vary greatly between different laboratories, historical records cannot fully reflect the specific state of the current experiment. Therefore, removing components from the "blacklist" may lead to false negatives and false positives. The fourth type of analytical method assumes a negative correlation between noise signals and DNA concentration after library preparation, but due to the randomness of sequencing, this method may not be suitable for low levels of contaminants. In summary, existing analytical methods all analyze single components of the sample composition and, through different assumptions and information, judge and remove single components judged as noise signals. However, due to the extremely high sensitivity, extremely low detection limit, and strong randomness of metagenomic sequencing, methods that identify and eliminate single components cannot effectively eliminate background noise signals.

[0004] In view of the above, this application is hereby submitted. Summary of the Invention

[0005] The core problem this application aims to solve is how to utilize the information provided by current experiments to perform targeted and efficient noise removal on the current experimental results.

[0006] As is known in the art, during bioinformatics analysis, the composition of the sample to be tested is typically a superposition of real signals and noise signals. To remove noise signals, this application requires comparing the composition of the sample to be tested with the background composition containing only noise signals, identifying the factors that remain constant, and constructing a model based on these constant factors to achieve effective noise removal.

[0007] First, an empirical observation is that the noise signal in any gene sequencing result (such as metagenomics) is not a single component, but contains an extremely diverse combination of species. This suggests that the noise is introduced in units of a subgroup, rather than a single noise signal introduced from a single pollution source.

[0008] Secondly, by examining the various steps of the sequencing method, this application found that the sources of noise signals can be mainly divided into two categories. The first category is contamination introduced during the experiment, and these noise signals are all real nucleic acid molecules. The second category is erroneous signals introduced during bioinformatics analysis, such as multiple alignments from real nucleic acid molecules, which are not real nucleic acid molecules. According to the definition of a negative control, the negative control itself does not contain any real signals. Theoretically, the composition of the observed negative control only includes contamination introduced during the experiment. Therefore, the negative control can be considered as the basic reference for experimental noise signals during the experiment.

[0009] In bioinformatics analysis, there are generally two main types of methods for obtaining the taxonomic unit composition of the sample: one is based on sequence alignment (such as BLASTN), and the other is not based on sequence alignment (such as the kMER method). Regardless of the method, due to the similarity between taxonomic units and errors introduced by sequencing, sequencing reads from the real components can be misclassified as other taxonomic units, creating bioinformatics noise. Specifically, when grouping taxonomic units based on sequence alignment, bioinformatics noise often aligns simultaneously with the real components onto some conservative sequencing reads, resulting in multiple alignments. In this case, multiple alignments can indicate the source of bioinformatics noise, so the alignment results serve as a good basic reference for bioinformatics noise signals and can be used as a bioinformatics control. When grouping taxonomic units without sequence alignment, reference sequences of the real components can be used to simulate sequencing reads according to the sequencer's error rate distribution. The classification results of the simulated data can then be used to construct a basic reference for bioinformatics noise signals, serving as a bioinformatics control. Therefore, the composition of both the negative control and the bioinformatics control are important references for noise signals.

[0010] Furthermore, noise removal requires consideration of the specific method in which the noise signal was introduced. This application found that during experiments, experimental contamination noise signals introduced by the same operation (e.g., components D, E, and F, each 1 unit) will undergo various transformations in subsequent operations (e.g., simultaneous 2-fold amplification, becoming D, E, and F, each 2 units), but the ratio between noise signals within the same group remains relatively constant (e.g., D:E:F = 1:1:1 before and after transformation). Although there may be no stable relationship between contamination noise signals introduced in different steps, the paired ratios of the same group of contamination signals introduced in the same step at the same time remain constant. Therefore, for experimental noise, whether it is experimental noise can be determined by calculating whether the ratio between classification units remains consistent in the test sample and the negative control.

[0011] This application finds that during the experiment, assuming that the classification unit A' is close to the A sequence and A' is not a bioinformatics false positive of A, then the number of A's should not be less than the number of A's obtained by erroneous alignment after simulation with the current number of A's (or the number of A's obtained by multiple alignment with the current number of A's). As mentioned above, the ratio between bioinformatics noise and true components is similar in the test sample and the bioinformatics control. Therefore, for bioinformatics noise, whether it is bioinformatics noise can be determined by calculating whether the ratio between classification units is consistent in the test sample and the bioinformatics control.

[0012] In summary, whether the contamination introduced during the experimental process or the false detection introduced during the bioinformatics analysis process, it can be determined whether it is a noise signal by judging whether the proportion of the composition within the subgroup remains unchanged (comparing the composition of the test sample with that of the negative control group, and comparing the composition of the test sample with that of the bioinformatics control).

[0013] In this application, we independently developed a hidden subgroup-based de-noising algorithm (Hugo-DNA) to effectively reduce noise signals (experimentally introduced contamination and bioinformatics-introduced false positives). Hugo-DNA is based on the assumption that a pair of noise signals from the same source will maintain their ratio during subsequent processing and transformation. Specifically, for experimentally introduced contamination, the contamination signals introduced by earlier experimental steps will maintain their internal ratios unchanged in subsequent experiments; for bioinformatics-introduced false positives, the proportion of a real signal incorrectly identified as a noise signal in the current data will not change. By comparing the composition data of the test sample with the composition data of the negative control, experimentally introduced contamination noise signals can be eliminated. By comparing the composition data of the test sample with the composition data of the bioinformatics control, bioinformatics-introduced false positive noise signals can be eliminated. Hugo-DNA can efficiently remove contamination without prior knowledge or setting thresholds.

[0014] Therefore, the primary objective of this application is to provide a calculation method and system for distinguishing between real signals and background signals;

[0015] The second objective of this application is to provide an application of hidden subgroups in bioinformatics analysis for denoising sequencing data.

[0016] Based on the above objectives, this application provides the following technical solution:

[0017] This application first provides a bioinformatics analysis method for distinguishing between real signals and background signals, including the following steps:

[0018] Step 1) Sequencing steps for the sample to be tested and the negative control sample;

[0019] Step 2) Grouping the sequencing data of the test sample and negative control sample according to the classification unit;

[0020] Step 3) Statistical steps for the classification units of the test sample and the negative control sample;

[0021] Step 4) Compare the results of the test sample with the negative control and calculate the relationship between the taxonomic units;

[0022] Step 5) Compare the results of the test sample with the negative control to identify the hidden subgroup.

[0023] Furthermore, in step 2), the grouping is performed by grouping the test samples and negative control samples separately based on an alignment method;

[0024] Preferably, the alignment method involves using alignment software that retains non-single alignment results (such as BLASTN software) to perform sequence alignment on the sequence readouts and then group them.

[0025] Furthermore, in step 2), the grouping is performed by grouping the test samples and negative control samples separately using a non-alignment method.

[0026] Preferably, grouping is performed using methods including but not limited to the kmer method, hash table method, or string matching.

[0027] Furthermore, the statistics of the classification units in step 3) include, but are not limited to, the following statistics: the number of sequences supported for sequencing reads for each classification unit, the relative proportion of each classification unit, or the statistics of each classification unit after some normalization.

[0028] Furthermore, in step 4), the calculation of the interrelationship between classification units is a statistical measure of the classification units in step 3). Every two classification units in the test sample are paired, and the ratio of the two classification units in the pair is calculated to remain stable in the test sample and the negative control. If it remains stable, the paired classification units are considered to be related to each other.

[0029] Furthermore, in step 5), the identification of hidden subgroups involves processing and filtering the relationships between the classification units in step 2) and step 4), and identifying the remaining relationships between the classification units as hidden subgroups.

[0030] Preferably, the identification and / or analysis is performed by hidden subgroup analysis without prior information; the identification and / or analysis is performed using analytical methods for analyzing the correlation between two or more elements or the characteristics of the elements themselves.

[0031] More preferably, the identification and / or analysis is performed using graph theory methods from computer science, that is, each classification unit in step 2) is taken as a vertex of the graph, and each related pair in step 4) is taken as an edge of the graph, forming a complete undirected graph; in the undirected graph, classical graph analysis methods are used to find the complete subgraph, which is taken as the hidden subgroup;

[0032] More preferably, the number of vertices in the complete subgraph is 2 or more; for example, 2, 3, 4, 5, or 6.

[0033] Furthermore, step 5) further includes the following steps:

[0034] Step 6) Construct bioinformatics controls and statistically analyze their taxonomic unit steps;

[0035] Step 7) Compare the test sample with the bioinformatics control results and calculate the relationship between taxonomic units;

[0036] Step 8) Compare the test sample with the bioinformatics control results to identify hidden subgroups.

[0037] Furthermore, the bioinformatics control in step 6) is based on the alignment results in step 2), and the alignment results are used as bioinformatics controls according to the alignment of the sequence readouts.

[0038] Furthermore, the bioinformatics control in step 6) is based on the non-alignment results in step 2). Data simulation is performed on the reference genome of each taxonomic unit. According to the error distribution pattern of the sequencer, the sequencing readout sequence of the taxonomic unit is simulated, and the simulated sequencing readout sequence is used to group the data. The grouping result of the simulated data can be used as the bioinformatics control for the taxonomic unit.

[0039] Furthermore, step 7) is a statistical measure of the classification units in step 6), which involves pairing every two classification units in the test sample and calculating whether the proportion of the two classification units in the pair is higher or the same in the test sample than in the bioinformatics control.

[0040] Preferably, if the values ​​are higher or equal, the paired taxa are considered to be related to each other;

[0041] More preferably, the ratio of the two classification units is obtained by dividing the classification unit statistics; or by performing a statistical test on the unit statistics.

[0042] Furthermore, in step 7), after removing the taxonomic units that form hidden subgroups in step 5), the statistical measures of the taxonomic units in step 6) are used to pair up every two taxonomic units in the test sample and calculate whether the ratio of the two taxonomic units in the pair is higher or the same in the test sample than in the bioinformatics control.

[0043] Preferably, if the values ​​are higher or equal, the paired taxa are considered to be related to each other;

[0044] More preferably, the ratio of the two classification units is obtained by dividing the classification unit statistics; or by performing a statistical test on the unit statistics.

[0045] Furthermore, step 8) involves processing and filtering the relationships between the classification units of the sample to be tested in step 2) and the classification units in step 7), and identifying the remaining relationships between the classification units as hidden subgroups.

[0046] Furthermore, step 8) involves identifying and / or analyzing the relationships between the classification units of the sample to be tested in step 2) and the classification units in step 7) after removing the classification units that form hidden subgroups in step 5).

[0047] Furthermore, the identification and / or analysis is performed by hidden subgroup analysis without prior information; the identification and / or analysis is performed using analytical methods for analyzing the correlation between two or more elements or the characteristics of the elements themselves.

[0048] Preferably, the identification and / or analysis is performed using graph theory methods from computer science. That is, each classification unit of the sample to be tested in step 2) is taken as a vertex of the graph, and each related pair in step 7) is taken as an edge of the graph, forming a complete undirected graph. In the undirected graph, classical graph analysis methods are used to find the complete subgraph, which can be used as the hidden subgroup.

[0049] More preferably, the number of vertices in the complete subgraph should be 2 or more; for example, 2, 3, 4, 5, or 6.

[0050] This application also provides a bioinformatics analysis method for background noise removal from sequencing data, including any of the methods described above, and further including the following steps:

[0051] Step 9) Remove background noise from the sample to be tested.

[0052] Preferably, the elimination in step 9) is to eliminate all classification units that form hidden subgroups in step 5) (contamination introduced in the experiment) for the classification units of the test sample in step 2); and / or, for the classification units in step 6), to eliminate all classification units that form hidden subgroups in step 8) and have low statistics in the paired classification units (i.e. false detections introduced by bioinformatics analysis).

[0053] This application also provides a method for calculating the taxonomic unit components of sequencing data, the method comprising any of the above methods, and further comprising the following steps:

[0054] Step 10) Taxonomic unit abundance calibration step;

[0055] Preferably, the abundance calibration step involves regrouping the sequencing readouts according to the classification units preferentially supported by the alignment results after removing noisy classification units, and counting the number of sequences in each group and the proportion of the total number of readouts.

[0056] This application also provides a bioinformatics analysis system for distinguishing between real signals and background signals, characterized by comprising the following modules:

[0057] Module 1) Sequencing module for test samples and negative control samples;

[0058] Module 2) Sequencing data of the test samples and negative control samples are grouped into different taxonomic unit modules;

[0059] Module 3) A module for statistically analyzing the classification units of the test sample and the negative control sample;

[0060] Module 4) Comparison of test samples and negative control results, and calculation of the interrelationships between taxonomic units;

[0061] Module 5) Comparison of test samples with negative control results to identify hidden subgroups.

[0062] Furthermore, it also includes the following modules:

[0063] Module 6) Constructing bioinformatics controls and statistically analyzing their taxonomic units;

[0064] Module 7) Comparison of test samples with bioinformatics control results, and calculation of the interrelationships between classification units;

[0065] Module 8) Comparison of test samples with bioinformatics control results to identify hidden subgroups.

[0066] Preferred options also include:

[0067] Module 9) Remove background noise from the test sample.

[0068] The above modules execute steps 1)-9 respectively.

[0069] Furthermore, in any of the foregoing methods or systems, the sequencing is second-generation sequencing, third-generation sequencing, or fourth-generation sequencing;

[0070] Preferably, the sequencing is performed using Illumina, BGI, ION TORRENT, PacBio, Roche, Helicos, ABI, or nanopore sequencing.

[0071] More preferably, the sequencing is performed using Illumina, BGI, or nanopore sequencing.

[0072] Furthermore, the sequencing data is gene sequencing data;

[0073] Preferably, the sequencing data is genomic sequencing data;

[0074] More preferably, the sequencing data is metagenomic sequencing data;

[0075] More preferably, the sequencing data is low-biomass metagenomic sequencing data.

[0076] This application also provides a computer system including a memory, a processor, and a computer program stored in the memory, wherein the processor executes the computer program to implement the steps of any of the methods described above.

[0077] This application also provides a computer-readable storage medium having a computer program stored thereon, characterized in that the computer program, when executed by a processor, implements the steps of any of the methods described above.

[0078] This application also provides a computer program product, including a computer program, characterized in that, when the computer program is executed by a processor, it implements the steps of any of the methods described above.

[0079] This application also provides an application of hidden subgroups in distinguishing real signals from background signals during bioinformatics analysis. The hidden subgroups are derived from the same source of noise signals introduced during the experiment and / or bioinformatics analysis. The ratio between each pair of elements within the hidden subgroup remains stable under two or more conditions.

[0080] Preferably, the hidden subgroup is a hidden subgroup formed between the test sample and the negative control, and / or, the hidden subgroup is a hidden subgroup formed between the test sample and the bioinformatics control.

[0081] The beneficial technical effects of this application are:

[0082] 1. This application is an innovation in the algorithm of conventional noise reduction methods. It innovatively proposes the hypothesis that "introduced noise signals will form hidden subgroups" to remove noise from the composition data of genome sequencing samples.

[0083] 2. This application, for the first time, determines whether taxonomic units originate from the same subgroup (i.e., noise) by comparing the consistency of the proportions between pairs of components in the sample composition, negative control group composition, and multiple alignment composition. This solves the problem that existing noise reduction methods cannot fully utilize real experimental information and require a large amount of prior knowledge, effectively improving the sensitivity and specificity of the real signal.

[0084] 3. The computational framework of this application is independent of the choice of a specific sequencing platform, and can be applied to sequencing data from multiple platforms such as second-generation sequencing technology and third-generation sequencing technology, and can be applied to detection samples from different sources or different species. Attached Figure Description

[0085] Figure 1 : Schematic diagram of the introduction of metagenomic noise signals.

[0086] Figure 2The noise signals introduced at the same time will form subgroups, and the proportional relationship between each pair of taxa within the subgroup remains unchanged in the sample and the negative control.

[0087] Figure 3 The ratio of real signal to noise signal in the test sample and negative control, and the ratio of noise signal to noise signal in the test sample and negative control.

[0088] Figure 4 : Hugo-DNA process diagram;

[0089] Figure 5 : An undirected graph constructed from the classification units of the sample to be tested.

[0090] Figure 6 A comparison of the noise reduction performance of Hugo-DNA with other methods. Detailed Implementation

[0091] The embodiments of this application will be described in detail below with reference to examples. However, those skilled in the art will understand that the following examples are for illustrative purposes only and should not be considered as limiting the scope of this application. Unless otherwise specified in the examples, conventional conditions or conditions recommended by the manufacturer shall apply. Reagents or instruments whose manufacturers are not specified are all conventional products that can be purchased on the market.

[0092] Definitions of some terms

[0093] Unless otherwise defined below, all technical and scientific terms used in the specific embodiments of this application are intended to have the same meaning as commonly understood by those skilled in the art. While it is believed that the following terms will be well understood by those skilled in the art, the following definitions are set forth to better explain this application.

[0094] As used in this application, the terms “comprising,” “including,” “having,” “containing,” or “involving” are inclusive or open-ended and do not exclude other unlisted elements or method steps. The term “consisting of” is considered a preferred embodiment of the term “comprising.” If a group is defined below as comprising at least a certain number of embodiments, this should also be understood to disclose a group that preferably consists only of those embodiments.

[0095] When referring to a singular noun, the indefinite or definite article used, such as "a" or "a kind of," "the," includes the plural form of the noun.

[0096] The term "approximately" in this application refers to an accuracy range that, as would be understood by those skilled in the art, still guarantees the technical effects of the discussed features. This term typically indicates a deviation from the indicated value of ±10%, preferably ±5%.

[0097] Furthermore, the terms first, second, third, (a), (b), (c), and similar terms used in the specification and claims are for distinguishing similar elements and are not necessary for the order of description or chronological sequence. It should be understood that such terms are interchangeable in appropriate contexts, and the embodiments described herein can be implemented in a different order than that described or illustrated herein.

[0098] The following terms or definitions are provided merely to aid in understanding this application. These definitions should not be construed as having a scope less than that understood by those skilled in the art.

[0099] The term "sequencing readout" in this application refers to a single or group of nucleic acid sequences read out by a sequencing platform.

[0100] The term "alignment" in this application refers to the correspondence between a sequencing read and a reference sequence. A sequencing read can have multiple alignment results simultaneously.

[0101] The term "taxonomic unit" in this application refers to a classification concept at the same level in biology. For example, when studying biological evolutionary relationships, a taxonomic unit is the concept of each group of organisms at the levels of kingdom, phylum, class, order, family, genus, species, and strain. For instance, Protozoa, Primates, Staphylococcus aureus, and Salmonella enterica subsp. enterica are all taxonomic units. In systems biology, a taxonomic unit can be a research unit representing different genes, different genomic elements, different genomic locations, or different sequence characteristics.

[0102] The term "species" in this application refers to a special taxonomic unit that refers to a group of organisms that can mate and reproduce.

[0103] The term "(a set of taxonomic units) mutually exclusive" in this application means that any two taxonomic units A and B in the set of taxonomic units can be selected such that taxonomic unit A neither contains nor is contained by taxonomic unit B. For example, the three taxonomic units "Escherichia coli, Salmonella enterica, and Klebsiella pneumoniae" are mutually exclusive; while the two taxonomic units "Klebsiella pneumoniae" are not mutually exclusive.

[0104] The terms “a sequencing read sequence aligned to a certain taxonomic unit”, “a sequencing read sequence aligned to a certain taxonomic unit”, and “a sequencing read sequence supporting a certain taxonomic unit” in this application refer to the fact that the alignment result of the sequencing read sequence includes a reference sequence from that taxonomic unit.

[0105] The term "single alignment" in this application refers to a sequence for which all alignment results of a sequencing readout correspond to reference sequences from the same taxonomic unit.

[0106] The term "multiple alignment" in this application refers to the alignment result of a sequencing readout containing reference sequences from two or more mutually exclusive taxa.

[0107] In this application, the term "false positive" refers to a sequence alignment result from a sequencing read that actually originates from a taxonomic unit containing a reference sequence from another taxonomic unit that is neither included nor contained by the aforementioned taxonomic unit. It should be noted that the alignment result itself is computationally correct, and the base agreement rate can be very high, but the taxonomic unit aligned to does not actually match the original taxonomic unit of the sample. This type of false alignment generally occurs between taxonomic units that are geographically close.

[0108] The term "experimental contamination" in this application refers to nucleic acid molecules not originating from the sample that are introduced during the sample processing. These contaminants are not nucleic acid molecules from the sample itself, but rather contaminating nucleic acid molecules introduced during the experiment. Possible sources include sampling, instruments, reagents, etc.

[0109] The term "detected signal" in this application refers to all taxonomic units reported after analysis of the sequencing data of a sample. It includes both real and noise signals.

[0110] The term "true signal" in this application refers to the actual nucleic acid molecules to be tested contained in the sample. In the experimental design, this application incorporates specific, known, high-purity nucleic acid molecules into nucleic acid-free water; these nucleic acid molecules constitute the true signal. When the sample to be tested is unknown, the nucleic acid molecules without any treatment constitute the true signal.

[0111] The term "noise signal" in this application refers to non-authentic signals introduced during sample processing. These signals are not nucleic acid molecules of the sample itself, but rather contaminating nucleic acid molecules introduced during the experiment or erroneous comparisons and false detections introduced during bioinformatics analysis.

[0112] The "method for calculating the taxonomic units of sequencing data" in this application typically refers to a method for calculating the presence and proportion of various taxonomic units in sequencing data through bioinformatics analysis. This application preferably uses a method based on "hidden subgroup noise removal" to obtain the taxonomic unit composition of sequencing data. This method effectively removes false positive results in the calculation of taxonomic unit components. It is understood that this application solves the noise removal problem that is difficult to address with conventional taxonomic unit component calculation methods by introducing the calculation of "hidden subgroups," effectively improving the specificity and accuracy of pathogen detection. Furthermore, the calculation framework of this application is independent of the specific sequencing platform and is not limited by it, applicable to sequencing data from various platforms such as second-generation sequencing and third-generation sequencing technologies. Moreover, this calculation framework is also independent of the specific sequencing data source. It is understood in the art that the calculation framework of this application addresses any misalignment of homologous sequences; therefore, the sequence source does not limit the application of this application. Besides the metagenomic data source preferred in this application, other genomic or gene data sources are also suitable for this application.

[0113] As described in the preceding claims, the core of this application is based on the principle that "introduced noise signals form hidden subgroups," analyzing the connections between components in the sample to achieve more effective removal of introduced noise signals while preserving the true signal to the greatest extent possible. This method can be based solely on the hidden subgroups formed between the test sample and the negative control for noise reduction, or it can be based on both the hidden subgroups formed between the test sample and the negative control, and the hidden subgroups formed between the test sample and the bioinformatics control. Therefore, the method for distinguishing between true bioinformatics signals and background signals described in this application can be as follows:

[0114] According to one aspect of this application, the bioinformatics analysis method for distinguishing between real signals and background signals described in this application includes the following steps:

[0115] Step 1) Sequencing steps for the sample to be tested and the negative control sample;

[0116] Step 2) Grouping the sequencing data of the test sample and negative control sample according to the classification unit;

[0117] Step 3) Statistical steps for the classification units of the test sample and the negative control sample;

[0118] Step 4) Compare the results of the test sample with the negative control and calculate the relationship between the taxonomic units;

[0119] Step 5) Compare the results of the test sample with the negative control to identify the hidden subgroup.

[0120] The term "negative control" in this application refers to a blank control set up during sequencing experiments, which undergoes the same operational procedures as the sample to be tested. This negative control has the traditional meaning of a negative control. For example, in a sequencing experiment, it could be a nucleic acid-free water sample. Theoretically, the detection signal of this sample should be an empty set, meaning nothing should be detected. However, due to contamination introduced by the experiment and false positives introduced by bioinformatics analysis, this sample may become the best representative of the noise signal in the same batch of experiments. As another example, in a sequencing experiment, a negative control is a blank system that does not contain the target substance, relative to the target substance. Theoretically, if a target substance is considered to be present in sample one but not in sample two, sample two can be used as a negative control for sample one, thereby finding the element contained only in sample one.

[0121] The "method for calculating the taxonomic units of sequencing data" in this application typically refers to a method for calculating the presence and proportion of various taxonomic units in sequencing data through bioinformatics analysis. This application preferably uses a method based on "hidden subgroup noise removal" to obtain the taxonomic unit composition of sequencing data. This method effectively removes false positive results in the calculation of taxonomic unit components. It is understood that this application solves the noise removal problem that is difficult to address with conventional taxonomic unit component calculation methods by introducing the calculation of "hidden subgroups," effectively improving the specificity and accuracy of pathogen detection. Furthermore, the calculation framework of this application is independent of the specific sequencing platform and is not limited by it, applicable to sequencing data from various platforms such as second-generation sequencing and third-generation sequencing technologies. Moreover, this calculation framework is also independent of the specific sequencing data source. It is understood in the art that the calculation framework of this application addresses any misalignment of homologous sequences; therefore, the sequence source does not limit the application of this application. Besides the metagenomic data source preferred in this application, other genomic or gene data sources are also suitable for this application.

[0122] In this application, the terms "hidden subgroup," "subgroup," or "sub-group" refer to a hypothetical co-occurrence of signals, which may consist of two or more elements. A hidden subgroup requires that its elements be correlated or connected. The hidden subgroup may originate from signals introduced from the same source during the experiment or bioinformatics analysis, and the ratio between any two elements within the hidden subgroup remains stable under two or more conditions. In some implementations, the hidden subgroup is a hidden subgroup formed between the test sample and the positive control; a hidden subgroup formed between the test sample and the negative control; a hidden subgroup formed between the test sample and the bioinformatics control; a hidden subgroup formed between the test sample and other controls; a hidden subgroup formed between the test samples; a hidden subgroup formed between the negative control and the positive control; a hidden subgroup formed between the bioinformatics control and the negative control; a hidden subgroup formed between other controls and the positive control; a hidden subgroup formed between the bioinformatics control and the negative control; a hidden subgroup formed between other controls and the negative control; or a hidden subgroup formed between the bioinformatics control and other controls.

[0123] For example, in bioinformatics noise reduction, the hidden subgroups are noise signals introduced from the same source during the experiment or analysis. These noise signals from the same source undergo various processing simultaneously under different conditions, but their internal proportions remain stable. Therefore, the proportions between any two elements within a hidden subgroup remain stable under two or more conditions. In absolute quantification, high-abundance background signals often appear simultaneously in the negative control, positive control, and test sample. The proportions between these high-abundance background signals may remain stable in the negative control, positive control, and test sample, forming hidden subgroups. These hidden subgroups can serve as a pseudo-internal control (PIC). The PIC can be used to calculate the proportion between the negative control and positive control, and also the proportion between the negative control and the test sample. Through correlation, the proportion between the positive control and the test sample can be calculated, and the absolute quantification of the test sample can be calculated based on the absolute quantification (known) of the positive control. This application specifically refers to the hidden subgroups formed between the test sample and the negative control, and the hidden subgroups formed between the test sample and the bioinformatics control.

[0124] Step 5 above specifically refers to the hidden subgroup formed between the test sample and the negative control.

[0125] The terms “interrelated” or “connected” in this application refer to the proportional relationship between paired taxa that can be rigorously quantified and that remains stable under two or more conditions.

[0126] The term "co-detection" in this application refers to the simultaneous detection of a certain taxonomic unit under different conditions, such as detection in the sample and negative control, or detection in the sample and positive control, or detection in the sample and bioinformatics control, etc.

[0127] The term "detected alone" in this application refers to a classification unit being detected under only a single condition, such as detection in either the sample or the negative control. In its description, it is often specified that either the sample or the negative control was detected alone.

[0128] The term "coexistence" in this application refers to a taxonomic unit that is both a real signal and a noise signal. This situation often occurs when the real signal in a sample contains nucleic acid molecules from environmental microorganisms, which are a common source of contamination.

[0129] The term "existing alone" in this application refers to a classification unit that is either a real signal or a noise signal.

[0130] In some implementations, the grouping in step 2) is performed by grouping the test samples and negative control samples separately based on an alignment method;

[0131] Preferably, the alignment method involves using alignment software that retains non-single alignment results (such as BLASTN software) to perform sequence alignment on the sequence readouts and then group them.

[0132] In some implementations, the grouping in step 2) is performed by grouping the test samples and negative control samples separately using a non-alignment method;

[0133] Preferably, grouping is performed using methods including but not limited to the kmer method, hash table method, or string matching.

[0134] The taxonomic unit statistics described in this application include, but are not limited to, the following statistics: the number of sequences supported for sequencing reads for each taxonomic unit, the relative proportion of each taxonomic unit, or the statistics of each taxonomic unit after some normalization.

[0135] In some implementations, the calculation of the interrelationship between classification units in step 4) is a statistical measure of the classification units in step 3). Every two classification units in the test sample are paired, and the ratio of the two classification units in the pair is calculated to remain stable in the test sample and the negative control. If it remains stable, the paired classification units are considered to be related to each other.

[0136] In some implementations, the identification of hidden subgroups in step 5) involves processing and filtering the relationships between the classification units in step 2) and step 4), and identifying the remaining relationships between the classification units as hidden subgroups.

[0137] Preferably, the identification and / or analysis is performed by hidden subgroup analysis without prior information; the identification and / or analysis is performed using analytical methods for analyzing the correlation between two or more elements or the characteristics of the elements themselves.

[0138] The approach of using hidden subgroups as the research object in this application differs from that of single elements, where only the element's own data characteristics can be used as the research object for subsequent analysis methods. Hidden subgroups include both the elements within a subgroup and the relationships between elements within a subgroup. Therefore, the analysis methods for hidden subgroups can not only focus on the relationships between elements within a subgroup, but also on the summation (or other combinations, such as products, statistics, etc.) of the same data characteristics of all elements within a subgroup. These types of analysis methods are collectively referred to as "analysis methods for analyzing the relationships between two or more elements or the characteristics of the elements themselves."

[0139] More preferably, the identification and / or analysis is performed using graph theory methods from computer science, that is, each classification unit in step 2) is taken as a vertex of the graph, and each related pair in step 4) is taken as an edge of the graph, forming a complete undirected graph; in the undirected graph, classical graph analysis methods are used to find the complete subgraph, which is taken as the hidden subgroup;

[0140] More preferably, the number of vertices in the complete subgraph is 2 or more; for example, 2, 3, 4, 5, or 6.

[0141] The terms "vertices" and "edges" refer to the constituent elements of a "graph" in computer data structures. A vertex is an element in a set, while an edge is a connection between two vertices. In this application, vertices are classification units, and edges are connections between classification units.

[0142] The term "subgraph" refers to a subset of the elements in an undirected graph.

[0143] The term "complete subgraph" refers to a subset of elements in an undirected graph that satisfies the condition that every pair of vertices in the subset is connected by an edge.

[0144] In this application, the term "outlier" refers to an element in an undirected graph that is not connected to any other element.

[0145] It is understood that, based on the negative control, this application can further perform noise reduction through bioinformatics control to achieve a better noise reduction effect. Therefore, in some embodiments of this application, step 5) may be further followed by the following steps:

[0146] Step 6) Construct bioinformatics controls and statistically analyze their taxonomic unit steps;

[0147] Step 7) Compare the test sample with the bioinformatics control results and calculate the relationship between taxonomic units;

[0148] Step 8) Compare the test sample with the bioinformatics control results to identify hidden subgroups.

[0149] The term "bioinformatics control" can encompass multiple components. In bioinformatics analysis, two main categories of methods obtain the taxonomic composition of a sample: sequence alignment-based methods (e.g., BLASTN) and non-sequence alignment-based methods (e.g., the kMER method). Regardless of the method, due to the similarity between taxonomic units and errors introduced by sequencing, sequencing reads from the real components can be misclassified as other taxonomic units, creating bioinformatics noise. Specifically, in the case of sequence alignment-based taxonomic unit grouping, bioinformatics noise often aligns simultaneously with the real components onto some conservative sequencing reads, resulting in multiple alignments. In this case, multiple alignments can indicate the source of bioinformatics noise, so the alignment results serve as a good basic reference for bioinformatics noise signals and can be used as a bioinformatics control. In the case of non-sequence alignment-based taxonomic unit grouping, reference sequences of the real components can be used to simulate sequencing reads according to the sequencer's error rate distribution. The classification results of the simulated data can then be used to construct a basic reference for bioinformatics noise signals, serving as a bioinformatics control. Furthermore, the bioinformatics control in step 6) is based on the alignment results in step 2), and the alignment results are used as bioinformatics controls according to the alignment of the sequence readouts.

[0150] In some implementations, the bioinformatics control in step 6) is based on the non-alignment results of step 2). Data simulation is performed on the reference genome of each taxonomic unit, and the sequencing readout sequence of the taxonomic unit is simulated according to the error distribution pattern of the sequencer. The simulated sequencing readout sequence is then used to group the data, and the grouping result of the simulated data can be used as the bioinformatics control for the taxonomic unit.

[0151] In some implementations, step 7) is a statistical measure of the classification units in step 6), which involves pairing every two classification units in the test sample and calculating whether the proportion of the two classification units in the pair is higher or the same in the test sample than in the bioinformatics control.

[0152] Preferably, if the values ​​are higher or equal, the paired taxa are considered to be related to each other;

[0153] More preferably, the ratio of the two classification units is obtained by dividing the classification unit statistics; or by performing a statistical test on the unit statistics.

[0154] In some implementations, after removing the taxonomic units that form hidden subgroups in step 5), step 7) then pairs every two taxonomic units in the test sample based on the statistics of the taxonomic units in step 6), and calculates whether the proportion of the two taxonomic units in the pair is higher or the same in the test sample than in the bioinformatics control.

[0155] Preferably, if the values ​​are higher or equal, the paired taxa are considered to be related to each other;

[0156] More preferably, the ratio of the two classification units is obtained by dividing the classification unit statistics; or by performing a statistical test on the unit statistics.

[0157] In some implementations, step 8) involves processing and filtering the relationships between the classification units of the sample to be tested in step 2) and the classification units in step 7), and identifying the remaining relationships between the classification units as hidden subgroups.

[0158] In some implementations, step 8) involves identifying and / or analyzing the relationships between the classification units of the sample to be tested in step 2) and the classification units in step 7) after removing the classification units that form hidden subgroups in step 5).

[0159] Furthermore, the identification and / or analysis is performed by hidden subgroup analysis without prior information; the identification and / or analysis is performed using analytical methods for analyzing the correlation between two or more elements or the characteristics of the elements themselves.

[0160] Preferably, the identification and / or analysis is performed using graph theory methods from computer science. That is, each classification unit of the sample to be tested in step 2) is taken as a vertex of the graph, and each related pair in step 7) is taken as an edge of the graph, forming a complete undirected graph. In the undirected graph, classical graph analysis methods are used to find the complete subgraph, which can be used as the hidden subgroup.

[0161] More preferably, the number of vertices in the complete subgraph should be 2 or more; for example, 2, 3, 4, 5, or 6.

[0162] In some embodiments, the above method further includes the following steps:

[0163] Step 9) Remove background noise from the sample to be tested.

[0164] Preferably, the elimination in step 9) is to eliminate all classification units that form hidden subgroups in step 5) (contamination introduced in the experiment) for the classification units of the test sample in step 2); and / or, for the classification units in step 6), to eliminate all classification units that form hidden subgroups in step 8) and have low statistics in the paired classification units (i.e. false detections introduced by bioinformatics analysis).

[0165] In some implementations, the following steps are further included:

[0166] Step 10) Taxonomic unit abundance calibration step;

[0167] Preferably, the abundance calibration step involves regrouping the sequencing readouts according to the classification units preferentially supported by the alignment results after removing noisy classification units, and counting the number of sequences in each group and the proportion of the total number of readouts.

[0168] It is understood that the method of this application is not limited to the sequencing platform, the source of sequencing samples, etc. Therefore, in any of the aforementioned methods, the sequencing can be second-generation sequencing, third-generation sequencing, or fourth-generation sequencing.

[0169] Preferably, the sequencing is performed using Illumina, BGI, ION TORRENT, PacBio, Roche, Helicos, ABI, or nanopore sequencing.

[0170] More preferably, the sequencing is performed using Illumina, BGI, or nanopore sequencing.

[0171] In some implementations, the sequencing data is gene sequencing data;

[0172] Preferably, the sequencing data is genomic sequencing data;

[0173] More preferably, the sequencing data is metagenomic sequencing data;

[0174] More preferably, the sequencing data is low-biomass metagenomic sequencing data.

[0175] To better explain this application, according to some specific embodiments of this application, the bioinformatics analysis method for distinguishing between real signals and background signals can be a specific method including the following steps (this explanation does not limit the scope of the invention):

[0176] Step 1) Sequencing of the test sample and negative control sample; Step 2) Grouping the sequencing data of the test sample and negative control sample into different taxonomic units; Step 3) Statistical analysis of the taxonomic units of the test sample and negative control sample; Step 4) Constructing a bioinformatics control and statistically analyzing its taxonomic units; Step 5) Comparing the results of the test sample and negative control, and calculating the relationships between taxonomic units; Step 6) Comparing the results of the test sample and negative control, and identifying hidden subgroups; May also include: Step 7) Comparing the results of the test sample and bioinformatics control, and calculating the relationships between taxonomic units; Step 8) Comparing the results of the test sample and bioinformatics control, and identifying hidden subgroups; Step 9) Removing background noise from the test sample.

[0177] In some implementations, step 1) involves performing parallel experiments on the test sample and negative control sample, while simultaneously constructing a sequencing library. Parallel experiments require that experimental conditions such as the experimental site, personnel, reagents, instruments, procedures, steps, techniques, and sequencing platform be kept as consistent as possible to ensure that the negative control can systematically and comprehensively capture various background noise signals introduced during the experiment.

[0178] In some embodiments, step 2), based on the sequencing data (generally a base text file) from step 1), uses an alignment-based method for grouping. Alignment software is used to align the sequencing reads, and grouping is performed according to the alignment results. For example, all sequencing reads aligned to the same taxonomic unit are grouped together. In some embodiments, step 2), based on the sequencing data (generally a base text file) from step 1), uses a non-alignment-based method for grouping. The kmer method first divides the sequencing reads into kmers and the alignment database into kmers, and then counts which reference sequence in the database a given sequencing read matches best, thus grouping the sequencing read into the taxonomic unit of that reference sequence. Preferably, for alignment-based grouping methods, alignment software that retains multiple alignment results (such as BLASTN software) is used to perform sequence alignment and then group the sequencing reads. Preferably, for non-alignment-based grouping methods, methods including but not limited to kmer methods, hash table methods, and string matching methods are used. This step performs the same processing on both sample data and negative control data.

[0179] In some implementations, step 3) involves calculating statistics for each taxonomic unit based on the grouping results of the taxonomic units in step 2), specifically for the test sample and the negative control sample. This includes calculating the number of supported sequencing reads for each taxonomic unit, the relative proportion of each taxonomic unit, and the statistics for each taxonomic unit after normalization (e.g., reads per million, RPM). In some implementations, step 3) involves dividing the alignment results of step 2) into two groups based on the alignment of the sequencing reads: single alignment (i.e., each sequencing read aligns to the same taxonomic unit) and multiple alignment (i.e., each sequencing read aligns to different taxonomic units). When calculating the statistics for each taxonomic unit, considering that single alignment results have smaller errors, only the results of single alignments are used to calculate the number and percentage of sequencing reads aligned to each taxonomic unit. In some implementations, step 3) directly uses a non-alignment-based grouping method (such as the kmer method) to obtain statistics such as the number and percentage of supporting sequencing reads for each taxonomic unit, based on the non-alignment results of step 2). This step performs the same processing on the test sample data and the negative control data to obtain the taxonomic unit composition results of the test sample and the negative control.

[0180] In some implementations, step 4) is used to extract bioinformatics controls from the alignment results of the test sample or to simulate bioinformatics controls from the classification unit results of the test sample, for subsequent steps to remove background noise signals introduced by bioinformatics analysis. If a classification unit A' obtained in step 3) is a bioinformatics false positive for classification unit A, then its number must not be higher than the number of A's in the multiple alignment results of the real classification unit A (or the number of A's in the classification units of the simulated data of the real classification unit A). Based on this, bioinformatics controls are constructed one by one for each classification unit obtained in step 3), and the pairing relationship in the bioinformatics controls is limited to classification units with a high number of potential true signals and classification units with a low number of potential noise signals. Here, the number refers to the statistic in step 3). The high or low number is judged according to the significance in the experimental design.

[0181] In some embodiments, step 4) involves using the alignment results from step 2) as bioinformatics controls based on the alignment of the sequenced readouts. In other embodiments, step 4) involves performing data simulation on the reference genome for each taxonomic unit based on the non-alignment results from step 2). The simulated sequenced readouts for that taxonomic unit are then grouped according to the error distribution patterns of the current sequencer, and the grouping results of this simulated data can be used as bioinformatics controls for that taxonomic unit.

[0182] In some implementations, step 5) involves pairing every two taxonomic units in the test sample based on the taxonomic unit statistics obtained in step 3), and calculating whether the ratio of the two taxonomic units in each pair remains stable in both the test sample and the negative control. If it remains stable, the paired taxonomic units are considered to be related. In some specific implementations, the ratio of the two taxonomic units in each pair is obtained by dividing the taxonomic unit statistics, and whether it remains stable is determined by calculating whether the ratio of the pair is close to 1 under two (or more) conditions. For example, if the statistics of two taxonomic units A and B are 10 and 30 under condition one, respectively, with a ratio of 0.333, and the statistics are 1 and 3 under condition two, respectively, with a ratio of 0.333, then the ratio of the pair A and B in condition one (0.333) and condition two (0.333) is (0.333:0.333)1, which is considered stable. In some specific implementations, the ratio of the two classification units in each pair is obtained by performing statistical tests on the unit statistics, such as the p-value of the goodness-of-fit test. For example, if the statistics of two classification units A and B are 10 and 30 respectively under condition one, and their ratio is 0.333, and the statistics are 1 and 3 respectively under condition two, the goodness-of-fit test is performed on these four numbers (two-tailed) and the p-value is calculated. If the p-value is greater than or equal to a certain threshold (such as 0.05), it is considered to be stable.

[0183] In some specific implementations, determining whether the proportion of a certain pairing in the test sample and the negative control is close to 1 requires setting a threshold, such as greater than or equal to 0.5 and less than or equal to 2. This threshold setting is related to the specific application scenario and cannot be uniformly applied across different scenarios; therefore, no specific threshold range is defined. In some specific implementations, statistical tests are used to determine whether the proportion of a certain pairing in the test sample and the negative control remains stable. The threshold setting for this is often a commonly used statistical threshold such as 0.05 or 0.01. However, this threshold setting is also related to the specific application scenario and cannot be uniformly applied across different scenarios; therefore, no specific threshold range is defined.

[0184] In some implementations, step 6) involves identifying and / or analyzing the relationships between the classification units of the test sample in step 2) and the classification units in step 5). In some implementations, hidden subgroup analysis is performed without prior information, such as using graph theory methods from computer science. That is, each classification unit in step 2) is treated as a vertex of the graph, and each related pair in step 5) is treated as an edge, forming a complete undirected graph. In the undirected graph, classical graph analysis methods are used to find complete subgraphs, also called "cliques." These cliques can then be considered as hidden subgroups. Preferably, the number of vertices in the complete subgraph should be 2 or more. More preferably, the number of vertices in the complete subgraph should be 2, 3, 4, 5, or 6.

[0185] In some embodiments, step 7) involves pairing every two classification units in the test sample based on the classification unit statistics from step 4), and calculating whether the ratio of the two classification units in the pair (potential true component to potential noise component) is higher or equal in the test sample than in the bioinformatics control. In some embodiments, step 7) involves removing the classification units that form hidden subgroups in step 6) based on the classification unit statistics from step 4), pairing every two classification units in the test sample, and calculating whether the ratio of the two classification units in the pair (potential true component to potential noise component) is higher or equal in the test sample than in the bioinformatics control. If it is higher or equal, the paired classification units are considered to be related. In some specific embodiments, the ratio of the two classification units in each pair is obtained by dividing the classification unit statistics, and then comparing whether it is higher or equal in the test sample than in the bioinformatics control. In some specific implementations, the ratio of the two classification units in each pair is obtained by statistically testing the unit statistics, such as the p-value of the goodness-of-fit test. For example, if the statistics of two classification units A and B are 10 and 30 respectively in condition one, and their ratio is 0.333, and the statistics are 1 and 3 respectively in condition two, the goodness-of-fit test is performed on these four numbers (one-tailed test) and the p-value is calculated. If the p-value is greater than or equal to a certain threshold (such as 0.05), it is considered that the ratio is higher or the same in the test sample than in the bioinformatics control.

[0186] In some embodiments, step 8) involves identifying and / or analyzing the relationships between the classification units of the test sample in step 2) and the classification units in step 7). In some embodiments, step 8) involves identifying and / or analyzing the relationships between the classification units of the test sample in step 2) after removing the classification units that form hidden subgroups in step 6) and the classification units in step 7). In some embodiments, hidden subgroup analysis is performed without prior information, such as using graph theory methods from computer science. That is, each classification unit of the test sample in step 2) is treated as a vertex of the graph, and each pair of related elements in step 7) is treated as an edge, forming a complete undirected graph. In the undirected graph, classical graph analysis methods are used to find complete subgraphs, also called "cliques." These "cliques" can then be considered as hidden subgroups. Preferably, the number of vertices in the complete subgraph should be 2 or more. More preferably, the number of vertices in a complete subgraph should be 2, 3, 4, 5, or 6.

[0187] In some implementations, in step 9), for the classification units of step 2), all classification units that form hidden subgroups in step 6) are considered contamination introduced by the experimental process, and all classification units that form hidden subgroups in step 8) and have lower statistics in the paired classification units are considered false detections introduced by bioinformatics analysis, marked as noise signals and removed.

[0188] This application also provides a computer system including a memory, a processor, and a computer program stored in the memory, wherein the processor executes the computer program to implement the steps of any of the methods described above.

[0189] Based on the core idea of ​​this application, in some embodiments, this application may also relate to a computer-readable storage medium having a computer program stored thereon, characterized in that the computer program, when executed by a processor, implements the steps of any of the methods described above.

[0190] In some embodiments, this application may also relate to a computer program product, including a computer program, characterized in that the computer program, when executed by a processor, implements the steps of any of the methods described above.

[0191] In some embodiments, this application may also relate to the application of a hidden subgroup in distinguishing real signals from background signals during bioinformatics analysis, wherein the hidden subgroup originates from the same source of noise signals introduced during the experiment and / or bioinformatics analysis, and the ratio between each pair of elements within the hidden subgroup remains stable under two or more conditions.

[0192] Preferably, the hidden subgroup is a hidden subgroup formed between the test sample and the negative control, and / or, the hidden subgroup is a hidden subgroup formed between the test sample and the bioinformatics control.

[0193] The present application will now be described in conjunction with specific embodiments.

[0194] Example 1: Invention Exploration Design

[0195] First, this application examines the entire process of genome (metagenomic) sequencing methods and, after careful analysis, breaks it down as follows. For example... Figure 1 As shown in steps 1-4, during sample preparation, a series of experimental contaminants are simultaneously introduced into both the sample and the negative control. Specifically, the sample originally contains only real nucleic acid molecules (solid dots), while the negative control sample contains nothing. Step 1 introduces a subgroup of contaminated nucleic acid molecules (squares), Step 2 introduces a new subgroup of contaminated nucleic acid molecules (hollow circles), and the proportion of the contaminated nucleic acid molecule subgroup represented by the squares changes (shrinks overall), as does the proportion of the real nucleic acid molecule subgroup represented by the solid dots. In Step 3, a new group of contaminated nucleic acid molecules (small solid dots A) is introduced; this nucleic acid molecule is also a real nucleic acid molecule (solid dot A), i.e., coexisting nucleic acid molecules. Furthermore, Step 3 also involves changes in the proportions of the original nucleic acid molecule subgroups (real nucleic acid molecule subgroup, square contaminated nucleic acid molecules, hollow circle contaminated nucleic acid molecules). In Step 4, cross-contamination between samples (such as aerosols) also occurs.

[0196] During sequencing, all nucleic acid molecules are randomly sampled using a Bernoulli process and their nucleic acid sequences are obtained by a sequencer (e.g., ...). Figure 1 (Sequencing step). After this, the nucleic acid sequences of all real nucleic acid molecules will undergo bioinformatics (such as...) Figure 1The analysis (bioinformatics) yields the final alignment and detection results. However, some misalignments and false positives inevitably occur, as shown in A', A”, C', and C” in the figure. In other words, throughout the process, for the sample, from the initial true signal ABC to the final detected result ABCDEFGHI A'A”C'C, only ABC is the true signal, while DEFGHI A'A”C'C” are noise signals. DEFGHI is experimentally introduced contamination, and A'A”C'C” is a bioinformatics-introduced false positive. Furthermore, for the true signal A, as a coexisting nucleic acid molecule, the final detection result is a superposition of the true nucleic acid molecule A and the contaminating nucleic acid molecule A. For the negative control, which does not contain any real nucleic acid molecules, the final detection results ABDEFGHIJ A'A" are all noise signals. Among them, ABDEFGHIJ are contaminations introduced by the experiment and are real nucleic acid molecules (DEFGHIJ is introduced by steps 1 and 2, A is introduced by step 3, and B is introduced by step 4). A'A" are all false detections introduced by bioinformatics.

[0197] Secondly, this application found that if each detection is examined in isolation, tracing how it was introduced into the sample, the overall transformation process would be lost. However, this application found that the proportional relationship between the various detections is more important than the detected microbial sequence itself. Noise signals from a subgroup introduced in the same step will maintain their intragroup proportions during subsequent processing and transformation. Figure 2 As shown, even if the proportions between different subgroups change significantly, the intragroup proportions remain constant, which is the core concept of this application. Furthermore, through sequencing analysis of samples with known compositions, this application found that the ratio between real and noise signals could not be maintained stably in the test samples and negative controls (stable ratios are represented by dark dots, while unstable ratios are represented by light dots). However, conversely, the ratio between noise signals largely remained stable in the test samples and negative controls (e.g., ...). Figure 3 (As shown).

[0198] Based on this characteristic, this application establishes a noise reduction algorithm (Hugo-DNA) based on hidden subgroup analysis. Figure 4The sequencing reads from the test sample (SP) and negative control (NC) underwent quality control and host removal filtering. The filtered data were then compared. The results of a single alignment (test sample and negative control) served as the basis for taxonomic unit statistics, while the multiple alignment results of the test sample were retained and used subsequently for constructing bioinformatics controls. First, the test sample was compared with the negative control, and the pairwise ratios of the taxonomic units composed of the test sample were calculated to determine if they remained stable in both samples. All taxonomic units of the test sample were used as vertices of an undirected graph, and the pairs with stable ratios were used as edges. Hidden subgroups (i.e., complete subgraphs) were searched within the undirected graph, and all taxonomic units in these hidden subgroups were marked as noise signals. Secondly, for taxonomic units that did not form hidden subgroups (i.e., outliers), a bioinformatics control based on the alignment results was constructed. The test sample was compared with the bioinformatics control. If the proportion in the test sample was not higher than that in the bioinformatics control, it was considered that there was a true signal and a false detection similar to the true signal in the paired taxonomic unit. In this case, the taxonomic unit with fewer sequences supporting sequencing reads was considered a false detection, classified as noise signal, and removed. Finally, all remaining signals were labeled as true signals.

[0199] Example 2: Establishment of the method of this application

[0200] Based on Example 1, the specific steps for establishing the Hugo-DNA method in this application are as follows:

[0201] 1) Constructing a test sample with known composition: This application adds high-purity, artificially synthesized nucleic acid molecules with known sequences to nucleic acid-free water. The true signal of this sample is the artificially synthesized nucleic acid molecule with the known sequence. All other detections are considered noise signals.

[0202] 2) Constructing negative control samples: This application uses nucleic acid-free water as a negative control and processes the negative control in exactly the same way as the test sample to ensure that the negative control can capture noise signals throughout the process.

[0203] 3) Metagenomic Library Construction and Sequencing: Library construction was performed using the VAHTS Universal Plus DNA Library Prep Kit for Illumina (Vazyme, ND617) kit, following the manufacturer's instructions. This application used the Agilent 2100 Bioanalyzer and Agilent 2100 DNA 1000 kit (Agilent, 5067-1504) for library quality checks and a Qubit 4.0 fluorometer for quantification. Each library was normalized to the same concentration and pooled, and finally, paired-end 150bp sequencing was performed on the Nextseq 550Dx sequencing platform.

[0204] 4) Microbial community composition analysis of the data: The genome database used for annotation was extracted from the NCBI nt database and includes archaea, bacteria, fungi, and viruses. This application used BLASTN for sequence alignment and used single-aligned reads for subsequent sample detection. These detections constituted a list of taxonomic units detected in the samples. Both the test samples and negative controls could obtain a taxonomic unit composition to facilitate further elimination of experimentally introduced contamination. For samples, this application also counted multiple-aligned reads. These multiple-aligned reads were used to construct a multiple alignment result matrix, i.e., a matrix of reads in which a taxonomic unit can be simultaneously aligned to other taxonomic units, to facilitate further elimination of bioinformatics-introduced false positives.

[0205] 5) The composition of the taxonomic units in the test sample is compared with that in the negative control to find the subset of taxonomic units that are detected in common. These taxonomic units constitute the vertices in the undirected graph. For these commonly detected taxonomic units, the ratio between each pair is calculated. If the ratio is consistent in the sample and in the negative control, this application considers that there is a connection between the pair of taxonomic units, forming an edge in the undirected graph.

[0206] 6) Whether the proportion of a pair of taxonomic units is consistent in the test sample and the negative control is determined by the following rules:

[0207] a) If their proportions differ by less than 2;

[0208] b) If the p-value of its goodness-of-fit test is higher than 0.05.

[0209] 7) Search for a complete subgraph within the undirected graph above using a depth-first search algorithm. Elements contained within the complete subgraph are labeled as noise signals. Other classification units are labeled as candidate ground truth signals.

[0210] 8) For candidate true signals, compare them with the multiple alignment results. Similarly, calculate whether the proportion between each pair of classification units is consistent with its proportion in the multiple alignment results. The rules for whether the proportions are consistent are the same as in 6) above. If the proportions of a pair of classification units are consistent, this application considers the classification unit with the smaller number of reads to be a false detection of a bioinformatics misalignment and marks it as a noise signal, while all other signals are marked as true signals.

[0211] Example 3: Hidden Subgroup Display in Test Samples

[0212] This application further verifies the existence of hidden subgroups in the core assumption of the Hugo-DNA algorithm. The specific steps are as follows. Based on the method of Example 2 (steps 1-7), this application constructs an undirected graph of the sample to be tested. Figure 5 In this graph, each dot represents a vertex, or classification unit, of the test sample; each gray line represents an edge, indicating that the ratio of the two classification units it connects to remains stable between the test sample and the negative control. Figure 5 As shown, all vertices in the undirected graph except for the true signal (indicated by the arrow) are connected, indicating that all classification units other than the true signal come from different hidden subgroups and are all noise.

[0213] Example 4: Comparison of the method of this application with existing methods

[0214] Furthermore, this embodiment compares the results of Hugo-DNA with other existing methods, and the specific steps are as follows:

[0215] 1) Based on the method of Example 2 (steps 1-4), this example obtains the classification unit composition of the test sample and the negative control. These results will be used as input for subsequent different noise reduction methods. They are also presented as the original data. Figure 5 RAW (Rough / Written). The actual signal is indicated by an arrow.

[0216] 2) For the Hugo-DNA method, this embodiment uses steps 5-8 of Example 2 to perform analysis and obtain annotations of the real signal and noise signal. Figure 6 DNA)

[0217] 3) For Z-score noise reduction, this embodiment normalizes the relative abundance of taxa in the test sample and the negative control using Z-score. Taxa in the test sample whose normalized relative abundance is more than twice that in the negative control are marked as true signals, and the others are noise signals. Figure 6 ZSC)

[0218] 4) For Decontam noise reduction, this embodiment inputs the classification unit results of the test sample and negative control into the Decontam software (version 1.10.0) for frequency-based noise reduction analysis. The selected threshold is 0.5. Additional library concentration information is required during library construction; this information is obtained using Invitrogen. TM Qubit TM The fluidometer is used for measurement before the fluid is used. Figure 6 DCN)

[0219] 5) For noise reduction using the SourceTracker method, this embodiment inputs the classification unit results of the test sample and negative control into the software (version 1.0) for analysis. The SourceTracker method outputs the probability that each classification unit may be contaminated and the probability that it may be true. In this embodiment, classification units with a contaminated probability greater than the true probability are marked as contaminated, and the remaining classification units are marked as true signals. Figure 6 ,STR)

[0220] The results are as follows Figure 6 As shown, the sample to be tested contains approximately 2000 taxonomic units, of which only one is a true signal (red dot), and the rest are noise signals (green dots). Compared with other methods, Hugo-DNA retains both true signals and less noise.

[0221] The above description of the specific embodiments of this application does not limit this application. Those skilled in the art can make various changes or modifications based on this application, and as long as they do not depart from the spirit of this application, they should all fall within the scope of the appended claims.

Claims

1. A bioinformatics analysis method for distinguishing between real signals and background signals, characterized in that, Includes the following steps: Step 1) Sequencing steps for the sample to be tested and the negative control sample; Step 2) Grouping the sequencing data of the test samples and negative control samples according to taxonomic units; Step 3) Statistical steps for the classification units of the test sample and the negative control sample; Step 4) Compare the results of the test sample with the negative control and calculate the relationship between the taxonomic units; Step 5) Compare the results of the test sample with the negative control to identify the hidden subgroup; The classification unit statistics in step 3) include the following statistics: the number of sequences supported for sequencing reads for each classification unit, the relative proportion of each classification unit, or the statistics of each classification unit after some normalization. The calculation of the interrelationship between classification units in step 4) is a statistical measure of the classification units in step 3). Every two classification units in the test sample are paired, and the proportion of the two classification units in the pair is calculated to see if it remains stable in the test sample and the negative control. If the relationship remains stable, the paired taxa are considered to be related to each other. In step 5), the identification of hidden subgroups involves processing and filtering the relationships between the classification units in step 2) and step 4), and identifying the remaining relationships between the classification units as hidden subgroups. The hidden subgroups are derived from signals introduced from the same source during the experiment or bioinformatics analysis, and the ratios between any two elements within the hidden subgroups remain stable under two or more conditions.

2. The bioinformatics analysis method according to claim 1, characterized in that, The grouping in step 2) is based on a comparison method to group the test sample and the negative control sample separately.

3. The bioinformatics analysis method according to claim 1, characterized in that, The grouping in step 2) is based on a non-alignment method to group the test samples and negative control samples separately.

4. The bioinformatics analysis method according to any one of claims 1-3, characterized in that, The identification in step 5) is performed by hidden subgroup analysis without prior information; the identification is performed using analytical methods used to analyze the correlation between two or more elements or the characteristics of the elements themselves.

5. The bioinformatics analysis method according to claim 4, characterized in that, The identification in step 5) is performed using graph theory methods from computer science. That is, each classification unit in step 2) is used as a vertex of the graph, and each related pair in step 4) is used as an edge of the graph to construct a complete undirected graph. In the undirected graph, a complete subgraph is found, which is used as the hidden subgroup.

6. The bioinformatics analysis method according to any one of claims 1-3, characterized in that, Step 5) is further followed by the following steps: Step 6) Construct bioinformatics controls and statistically analyze their taxonomic unit steps; Step 7) Compare the test sample with the bioinformatics control results and calculate the relationship between taxonomic units; Step 8) Compare the test sample with the bioinformatics control results to identify hidden subgroups.

7. The bioinformatics analysis method according to claim 6, characterized in that, The bioinformatics control in step 6) is based on the alignment results in step 2), and the alignment results are used as bioinformatics controls according to the alignment of the sequence readouts.

8. The bioinformatics analysis method according to claim 6, characterized in that, The bioinformatics control in step 6) is based on the non-alignment results in step 2). Data simulation is performed on the reference genome of each taxonomic unit. According to the error distribution pattern of the sequencer, the sequencing readout sequence of the taxonomic unit is simulated, and the simulated sequencing readout sequence is used to group the data. The grouping result of the simulated data can be used as the bioinformatics control for the taxonomic unit.

9. The bioinformatics analysis method according to claim 6, characterized in that, Step 7) involves calculating the statistics of the classification units in step 6). Each pair of classification units in the test sample is paired, and the proportion of the two classification units in the pair is calculated to determine whether it is higher or the same in the test sample than in the bioinformatics control. If it is higher or the same, the paired classification units are considered to be related to each other.

10. The bioinformatics analysis method according to claim 9, characterized in that, In step 7), the ratio of the two classification units is obtained by dividing the classification unit statistics; or by performing a statistical test on the unit statistics.

11. The bioinformatics analysis method according to claim 6, characterized in that, In step 7), after removing the taxonomic units that formed hidden subgroups in step 5), the statistical measures of the taxonomic units in step 6) are used to pair up every two taxonomic units in the test sample and calculate whether the proportion of the two taxonomic units in the pair is higher or the same in the test sample than in the bioinformatics control. If it is higher or the same, it is considered that the paired taxonomic units are related to each other.

12. The bioinformatics analysis method according to claim 6, characterized in that, Step 8) involves processing and filtering the relationships between the classification units of the sample to be tested in step 2) and the classification units in step 7), and identifying the remaining relationships between the classification units as hidden subgroups.

13. The bioinformatics analysis method according to claim 6, characterized in that, Step 8) involves identifying and / or analyzing the relationships between the classification units of the sample to be tested in step 2) and the classification units in step 7) after removing the classification units that form hidden subgroups in step 5).

14. The bioinformatics analysis method according to any one of claims 12-13, characterized in that, The identification is performed through hidden subgroup analysis without prior information; the identification and / or analysis are performed using analytical methods for analyzing the correlation between two or more elements or the characteristics of the elements themselves.

15. A bioinformatics analysis method for background noise removal in sequencing data, characterized in that, The method includes any one of claims 1-14, and further includes the following steps: Step 9) Remove background noise from the sample to be tested.

16. The bioinformatics analysis method according to claim 15, characterized in that, In step 9), the elimination is for the classification units of the sample to be tested in step 2), eliminating all classification units that form hidden subgroups in step 5); and / or, for the classification units in step 6), eliminating all classification units that form hidden subgroups in step 8) and have low statistics in the paired classification units.

17. A method for calculating the taxonomic unit components of sequencing data, characterized in that, The method includes the method of any one of claims 15-16, and further includes the following steps: Step 10) Taxonomic unit abundance calibration step.

18. The method for calculating the components of a classification unit according to claim 17, characterized in that, In step 10), the abundance calibration step involves regrouping the sequencing readouts according to the classification units that are preferentially supported by the alignment results after removing noisy classification units, and counting the number of sequences in each group and the proportion of the total number of readouts.

19. A bioinformatics analysis system for distinguishing between real signals and background signals, characterized in that, Includes the following modules: Module 1) Sequencing module for test samples and negative control samples; Module 2) Sequencing data of the test samples and negative control samples are grouped into different taxonomic unit modules; Module 3) A module for statistically analyzing the classification units of the test sample and the negative control sample; Module 4) Comparison of test samples and negative control results, and calculation of the interrelationships between taxonomic units; Module 5) Comparison of test samples with negative control results to identify hidden subgroups; The classification unit statistics in module 3) include the following statistics: the number of sequences supported for sequencing reads for each classification unit, the relative proportion of each classification unit, or the statistics of each classification unit after some normalization. The calculation of the interrelationship between classification units in module 4) is a statistical measure of the classification units in step 3). Every two classification units in the test sample are paired, and the ratio of the two classification units in the pair is calculated to see if it remains stable in the test sample and the negative control. If the relationship remains stable, the paired taxa are considered to be related to each other. The identification of hidden subgroups in module 5) involves processing and filtering the relationships between the classification units in step 2) and step 4), and identifying the remaining relationships between the classification units as hidden subgroups. The hidden subgroups are derived from signals introduced from the same source during the experiment or bioinformatics analysis, and the ratios between any two elements within the hidden subgroups remain stable under two or more conditions.

20. The bioinformatics analysis system according to claim 19, characterized in that, Further includes the following modules: Module 6) Constructing bioinformatics controls and statistically analyzing their taxonomic units; Module 7) Comparison of test samples with bioinformatics control results, and calculation of the interrelationships between classification units; Module 8) Comparison of test samples with bioinformatics control results to identify hidden subgroups; Module 9) Remove background noise from the test sample.

21. The bioinformatics analysis system according to any one of claims 19-20, characterized in that, The sequencing is second-generation sequencing, third-generation sequencing, or fourth-generation sequencing.

22. The bioinformatics analysis system according to claim 21, characterized in that, The sequencing data were obtained from Illumina, BGI, ION TORRENT, PacBio, Roche, Helicos, ABI, or nanopore sequencing.

23. The bioinformatics analysis system according to claim 21, characterized in that, The sequencing data is gene sequencing data.

24. The bioinformatics analysis system according to claim 23, characterized in that, The gene sequencing data is metagenomic sequencing data.

25. The bioinformatics analysis system according to claim 24, characterized in that, The metagenomic sequencing data mentioned are low-biomass metagenomic sequencing data.

26. A computer system comprising a memory, a processor, and a computer program stored in the memory, wherein, The processor executes the computer program to implement the steps of the method according to any one of claims 1-14.

27. A computer-readable storage medium having a computer program stored thereon, characterized in that, When executed by a processor, the computer program implements the steps of the method according to any one of claims 1-14.

28. A computer program product, comprising a computer program, characterized in that, When executed by a processor, the computer program implements the steps of the method according to any one of claims 1-14.