A bioinformatics noise reduction analysis method and system based on hidden subgroup for bioinformatics noise
By using the Hugo-DNA method, noise signals are identified and removed by analyzing the composition ratio of the test sample with negative and bioinformatics controls. This solves the problem of low noise removal efficiency in existing technologies and achieves more efficient noise removal and preservation of true signals.
Patent Information
- Application Number
- CN202111000227.2
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2021-08-27
- Publication Date
- 2025-11-25
- Estimated Expiration
- 2041-08-27
AI Technical Summary
Existing metagenomic sequencing methods are inefficient in removing noise signals. Conventional methods rely on threshold selection, negative controls, or prior information, which can lead to the incorrect removal of real signals or the retention of noise signals, making it impossible to effectively eliminate background noise.
A noise reduction analysis method based on hidden subgroups (Hugo-DNA) was adopted. By comparing the composition ratio of the test sample with the negative control and bioinformatics control, noise signals were identified and removed. Assuming that the proportion of noise signals from the same source is constant in different steps, a model was constructed for effective removal.
It improves the sensitivity and specificity of real signals, is independent of sequencing platforms and prior knowledge, is applicable to a variety of sequencing technologies and sample sources, and effectively removes noise signals.
Smart Images

Figure CN115732031B_ABST
Abstract
Description
TECHNICAL FIELD
[0001] The present application relates to the field of bioinformatics analysis, in particular to a sequencing data background noise elimination method and system. TECHNICAL BACKGROUND
[0002] The composition of microbial community is closely related to environmental ecosystem, human health and clinical diseases. Metagenomic next generation sequencing (mNGS) is a method that can detect all nucleic acid molecules in the sample, and can quantitatively describe the composition of microbial community in the sample. Due to its high sensitivity and low detection limit (LOD), mNGS can detect very small amounts of real nucleic acid molecule signals, but also detect nucleic acid molecules of environmental microorganisms introduced during the experiment. More troublesome is that due to the high similarity between microbial species, the data of a single microorganism will also be annotated as multiple microorganisms in the subsequent bioinformatics analysis process, resulting in false detection. Without effective analysis and processing, these noise signals, i.e. experimental introduction of pollution and bioinformatics introduction of false detection, will lead to inaccurate data interpretation.
[0003] The metagenomic sequencing method includes several steps such as extraction of all nucleic acid molecules in the sample, library construction, library sequencing and data analysis. The above steps may introduce noise signals. Although the literature has reported that the pollution can be controlled and prevented as much as possible through strict experimental operation, but these schemes have not achieved significant results, therefore the mainstream method is to use bioinformatics analysis tools to eliminate background noise in the later data processing process. Among them, one widely used method is to remove all components below a specified relative abundance threshold, but this method depends on the selection of the threshold, often removing low-abundance true signals and retaining high-abundance noise signals. The second method uses a negative control to remove noise. In practice, while performing metagenomic sequencing on the sample to be tested, a negative control, i.e. nucleic acid-free water, is also set up at the same time, and the same operation as the sample to be tested is performed simultaneously to simulate the background without real signal. Therefore, the existing method often corrects the composition of the sample to be tested using the composition of the negative control, or directly removes all identified results reported by the negative control, or after normalization (such as Z-score method), removes components that are more in the negative control and less in the sample to be tested, but due to the randomness of sequencing, this method often incorrectly removes true signals and retains noise signals. The third type of analysis method requires a large amount of prior information about pollution to maintain a "blacklist", and directly removes components in the "blacklist" from the detection of the sample to be tested. However, this prior information is often unclear, and because the noise signals differ greatly between different laboratories, historical records cannot fully reflect the specific state of this experiment, so removing components from the "blacklist" may cause false negatives and false positives. The fourth type of analysis method assumes that noise signals are negatively correlated with DNA concentration after library preparation, but due to the randomness of sequencing, this method may not be suitable for low-content contaminants. In summary, the existing analysis methods all analyze a single component of the sample composition, and make judgments based on different assumptions and information, and remove single components that are judged to be noise signals. However, due to the high sensitivity, low detection limit and strong randomness of metagenomic sequencing, the method of judging and removing single components cannot effectively eliminate background noise signals.
[0004] Therefore, the present application is proposed. SUMMARY
[0005] The core problem to be solved by the present application is how to use the information available in the current experiment to effectively remove noise from the current experimental results.
[0006] It is known in the art that in bioinformatics analysis, the composition of the sample to be observed is usually the superposition of true signals and noise signals. In order to remove noise signals, the composition of the sample to be observed needs to be compared with the background composition containing only noise signals, and the factors that remain constant are found to construct a model for effective noise removal.
[0007] First of all, an empirical observation is that the composition of noise signals in any gene sequencing results (such as metagenome) is not single, but contains a variety of species combinations, which indicates that the introduction of noise is in the form of sub-group, rather than a single noise signal from a single source of pollution.
[0008] Secondly, by examining the steps of the sequencing method, the application finds that the source of noise signal can be mainly divided into two categories: the first category is the pollution introduced in the experimental process, and these noise signals are real nucleic acid molecules; the second category is the error signal introduced in the bioinformatics analysis process, such as multiple alignment from real nucleic acid molecules, which is not a real nucleic acid molecule. According to the definition of negative control, the negative control itself does not include any real signal, and the composition of the observed negative control theoretically only includes the pollution introduced in the experimental process, so the negative control can be considered as the basic reference of the experimental noise signal in the experimental process.
[0009] In the bioinformatics analysis process, there are usually two categories of methods to obtain the composition of the classification unit of the sample to be observed: one is based on sequence alignment method (such as BLASTN), and the other is not based on sequence alignment method (such as kmer method). No matter which method, due to the similarity between classification units and sequencing errors, sequencing reads from real components will be incorrectly grouped into other classification units, forming bioinformatics noise. Specifically, in the case of classification unit grouping based on sequence alignment, bioinformatics noise is often aligned to some conservative reads together with real components, forming multiple alignment. At this time, multiple alignment can indicate the source of bioinformatics noise, so the alignment result is a good reference for bioinformatics noise signal, which can be used as bioinformatics control; in the case of classification unit grouping without sequence alignment, the reference sequence of real components can be used to simulate sequencing reads according to the error rate distribution of the sequencer, and the classification result of the simulated data can be used to construct the basic reference of bioinformatics noise signal, which can be used as bioinformatics control. Therefore, the composition of negative control and the composition of bioinformatics control are important references for noise signals.
[0010] Again, the removal of noise signals needs to consider the specific way of noise signal introduction. The present application found that in the experimental process, the experimental pollution noise signals introduced by the same operation (such as components D, E, F, all 1 unit) will occur various deformations (such as synchronous amplification 2 times, become D, E, F, all 2 units) in subsequent operations, but the ratio between the same group of noise signals will remain relatively constant (such as D: E: F = 1: 1: 1 before and after deformation). Although there may be no stable relationship between the pollution noise signals introduced by different steps, but the same group of pollution signals introduced at the same time and in the same step maintain constant in the mutual ratio relationship of each other. Therefore, for experimental noise, by calculating whether the ratio relationship between the classification units remains consistent in the test sample and the negative control, it can be judged whether it is an experimental noise.
[0011] The present application found that in the experimental process, assuming that the classification unit A' is close to the sequence of A, and A' is not a bioinformatics false detection of A, then the number of A' should not be lower than the number of A' obtained by simulating the current number of A error comparison (or the number of A' obtained by multiple comparison of the current number of A). According to the above, the ratio relationship between the bioinformatics noise and the real component is similar in the test sample and the bioinformatics control. Therefore, for bioinformatics noise, by calculating whether the ratio relationship between the classification units remains consistent in the test sample and the bioinformatics control, it can be judged whether it is a bioinformatics noise.
[0012] In summary, whether it is pollution introduced in the experimental process or false detection introduced in the bioinformatics analysis process, it can be judged whether it is a noise signal by judging whether the ratio of the composition within the subgroup remains unchanged (the composition of the test sample compared with the composition of the negative control, and the composition of the test sample compared with the bioinformatics control result) such hidden subgroup.
[0013] In the present application, the present application independently develops a hidden subgroup-based de-noising analysis method (Hugo-DNA) to effectively reduce noise signals (experimentally introduced pollution and bioinformatics introduced false positives). Hugo-DNA is based on the following assumptions: a pair of noise signals of the same source will maintain their ratio consistent during subsequent processing and transformation. Specifically, for experimentally introduced pollution, the pollution signals introduced by the first completed experimental step will maintain the internal ratio between each other unchanged during the subsequent experimental process; for bioinformatics introduced false positives, the proportion of a real signal in the current data that is misidentified as a noise signal will not change. By comparing the composition data of the test sample with the composition data of the negative control, the experimentally introduced pollution noise signal can be eliminated. By comparing the composition data of the test sample with the composition data of the bioinformatics control, the bioinformatics introduced false positive noise signal can be eliminated. Hugo-DNA does not require prior knowledge and does not need to set a threshold, and can efficiently remove pollution.
[0014] Therefore, the first object of the present application is to provide a calculation method and system for distinguishing between real signals and background signals;
[0015] The second object of the present application is to provide an application of hidden subgroups in sequencing data de-noising bioinformatics analysis.
[0016] Based on the above purposes, the present application provides the following technical solutions:
[0017] The present application first provides a bioinformatics analysis method for distinguishing between bioinformatics real signals and background signals, comprising the following steps:
[0018] Step 1) sequencing step of the test sample;
[0019] Step 2) grouping the sequencing data of the test sample into different classification units;
[0020] Step 3) classifying the test sample classification units;
[0021] Step 4) constructing a bioinformatics control and counting its classification units;
[0022] Step 5) comparing the results of the test sample and the bioinformatics control, and calculating the relationship between the classification units;
[0023] Step 6) comparing the results of the test sample and the bioinformatics control, and identifying hidden subgroups.
[0024] Further, the hidden subgroup is a form in which signals commonly appear under two or more conditions, and is composed of two or more elements. The ratio between the elements in the hidden subgroup remains stable under two or more conditions;
[0025] Preferably, the hidden subgroups are from the same source introduced signals in the experimental process or bioinformatics analysis process;
[0026] More preferably, the hidden subgroups are selected from the hidden subgroups formed between the test sample and the bioinformatics control.
[0027] Further, the grouping in step 2) is based on alignment method; preferably, the sequencing read sequences are grouped after sequence alignment using alignment software (such as BLASTN software) that retains non-single alignment results;
[0028] Alternatively, the grouping in step 2) is based on non-alignment method; preferably, the grouping is performed using methods including but not limited to kmer method, hash table method or string matching method.
[0029] Further, the statistics obtained in step 3) include but are not limited to the following statistics: the number of supporting sequencing read sequences of each classification unit, the relative proportion of each classification unit or the statistics of each classification unit after certain normalization.
[0030] Further, the bioinformatics control of step 4) is based on the alignment results of step 2);
[0031] Preferably, according to the alignment of the sequencing read sequences, the alignment results are used as the bioinformatics control;
[0032] Alternatively, the bioinformatics control of step 4) is based on the non-alignment results of step 2);
[0033] Preferably, the reference genome of each classification unit is data simulated, the sequencing read sequences of the classification unit are simulated according to the error distribution law of the sequencer, and the simulated sequencing read sequences are grouped, and the grouping results of the simulated data can be used as the bioinformatics control of the classification unit.
[0034] Further, step 5) is to pair each two classification units in the test sample, and calculate whether the proportion of the two classification units in the pair is higher or flat in the test sample than in the bioinformatics control; if higher or flat, it is considered that the classification units in the pair are related to each other;
[0035] Preferably, the proportion of the two classification units is obtained by dividing the classification unit statistics; or is obtained by statistical test on the unit statistics.
[0036] Further, the step 6) is to identify and / or analyze the correlation between the classification units of the sample to be tested in step 3) and the classification units in step 5), and identify the correlation between the remaining classification units as a hidden subgroup.
[0037] Further, the identification and / or analysis is a hidden subgroup analysis by a priori information-free method.
[0038] Preferably, the identification is performed by using an analysis method for analyzing the correlation between two or more elements or the characteristics of the elements themselves.
[0039] In the present application, the hidden subgroup is different from a single element whose data characteristics can be used as the research object of the subsequent analysis method. The hidden subgroup includes not only the elements in the subgroup, but also the correlation between the elements in the subgroup. Therefore, the analysis method for the hidden subgroup can not only use the correlation between the elements in the subgroup as the research object, but also use the sum (or other combination methods such as product, statistical quantity, etc.) of the same data characteristics of all elements in the subgroup as the research object. Such analysis methods are collectively referred to as "analysis methods for analyzing the correlation between two or more elements or the characteristics of the elements themselves".
[0040] More preferably, the identification and / or analysis is performed by using a graph method in computer science, i.e., each classification unit in step 2) is taken as a node of the graph, and each pair with a correlation in step 5) is taken as an edge of the graph to form a complete undirected graph; in the undirected graph, a complete subgraph is found by using a classical graph analysis method, and the complete subgraph can be used as a hidden subgroup.
[0041] Further preferably, the number of nodes of the complete subgraph is 2 or more.
[0042] The present application also provides a bioinformatics analysis method for removing background noise of sequencing data, which comprises any of the above methods and further comprises the following steps:
[0043] Step 7) removing the background noise in the sample to be tested.
[0044] Preferably, the removal in step 7) is to remove all classification units in step 6) that form a hidden subgroup and have a low statistical quantity in the two classification units in the pair, from the classification units in step 2).
[0045] The present application also provides a bioinformatics analysis system for distinguishing between bioinformatics true signals and background signals, which comprises the following modules:
[0046] Module 1) sequencing module for the sample to be tested;
[0047] Module 2) grouping the sequencing data of the sample to be tested into different classification units;
[0048] Module 3) counting the classification units of the sample to be tested;
[0049] Module 4) constructing a bioinformatics control and counting the classification units thereof;
[0050] Module 5) comparing the results of the sample to be tested and the bioinformatics control, and calculating the correlation between the classification units;
[0051] Module 6) comparing the results of the sample to be tested and the bioinformatics control, and identifying a hidden subgroup.
[0052] Preferably, it further comprises:
[0053] Module 7) removing background noise in the sample to be tested;
[0054] The above modules 1) to 7) respectively perform steps 1) to 7) in claims 1 to 8.
[0055] Further, in the above method or process, the sequencing is second-generation sequencing, third-generation sequencing or fourth-generation sequencing; preferably, it is from Illumina, BGI, ION TORRENT, PacBio, Roche, Helicos, ABI or nanopore sequencing; more preferably, it is from Illumina, BGI or nanopore sequencing; the sequencing data is gene sequencing data, preferably genomic sequencing data; more preferably, it is metagenomic sequencing data; further preferably, it is low-biomass metagenomic sequencing data.
[0056] The present application also provides a computer system comprising a memory, a processor and a computer program stored on the memory, wherein the processor executes the computer program to implement the steps of the method of any one of the above.
[0057] The present application also provides a computer-readable storage medium having a computer program stored thereon, which, when executed by a processor, implements the steps of the method of any one of the above.
[0058] The present application also provides a computer program product comprising a computer program, which, when executed by a processor, implements the steps of the method of any one of the above.
[0059] The present application also provides an application of a hidden subgroup in distinguishing a real signal from a background signal in a bioinformatics analysis process, wherein the hidden subgroup is from the same source as noise signals introduced in the bioinformatics analysis process, and the proportion between elements in the hidden subgroup remains stable under two or more conditions; preferably, the hidden subgroup is a hidden subgroup formed between the sample to be tested and the bioinformatics control
[0060] The application has the beneficial effects of:
[0061] 1. The application is an innovation of the algorithm of the conventional noise reduction method. It creatively proposes the hypothesis that the introduced noise signal will form a hidden subgroup, and performs noise elimination on the composition data of the genomic sequencing sample group.
[0062] 2. The application first determines whether the classification units come from the same subgroup, i.e. noise, by comparing whether the proportion between two compositions is consistent in the sample composition and the multiple alignment composition. It solves the problem that the existing noise reduction method cannot fully utilize the current experimental true information and needs a large amount of prior knowledge, and effectively improves the sensitivity and specificity of the true signal.
[0063] 3. The calculation framework of the application is independent of the selection of specific sequencing platforms, and can be applied to sequencing data of various platforms such as second-generation sequencing technology and third-generation sequencing technology, and can be applied to detection samples of different sources or different species. BRIEF DESCRIPTION OF DRAWINGS
[0064] Figure 1 : Schematic diagram of macrogenomic noise signal introduction.
[0065] Figure 2 : The noise signals introduced at the same time will form a subgroup, and the proportion relationship between the classification units in the subgroup remains unchanged in the sample and the negative control.
[0066] Figure 3 : The proportion relationship between the true signal and the noise signal in the sample to be tested and the negative control, and the proportion relationship between the noise signal and the noise signal in the sample to be tested and the negative control.
[0067] Figure 4 : Schematic diagram of Hugo-DNA process;
[0068] Figure 5 : Schematic diagram of noise reduction based on bioinformatics control alone.
[0069] Figure 6 : Undirected graph of the classification unit construction of the sample to be tested.
[0070] Figure 7 : Comparison of the noise reduction effect of Hugo-DNA with other methods. DETAILED DESCRIPTION
[0071] The embodiments of the present application will be described in detail below with examples, but those skilled in the art will understand that the following examples are for illustration only and should not be construed as limiting the scope of the present application. Where specific conditions are not mentioned in the examples, they are carried out under conventional conditions or conditions recommended by the manufacturer. Where the manufacturer of the reagent or instrument is not mentioned, it is a conventional product that can be purchased on the market.
[0072] Definitions of Particular Terms
[0073] Unless otherwise defined, all technical and scientific terms used in the of the application have the same meaning as commonly understood by one of ordinary skill in the art to which this application belongs. Although methods and materials similar or equivalent to those described herein can be used in the practice or testing of the present application, exemplary methods and materials are described below. However, the terminology used in the of the application is for the purpose of describing particular embodiments only and is not intended to limit the scope of the present application. It should be noted that, as used in this specification and the appended claims, the singular form "a", "an" and "the" include plural references
[0074] As used in this application, the terms "comprises", "comprising", "has", "having", "includes" or "including" are inclusive or open-ended and do not exclude additional, unrecited elements or method steps. The term "consisting of is considered to be a preferred embodiment of the term "comprising". If a group of items, definitions or embodiments are disclosed as comprising at least one of a list of elements, it is understood that the group of items, definitions or embodiments can comprise any one of the listed elements, or any combination of the listed elements.
[0075] The indefinite articles "a" or "an", as used herein in referring to a singular noun form, also include the plural referents of the noun, unless the context clearly indicates otherwise.
[0076] The term "about" in the present application denotes an interval of accuracy which the person skilled in the art understands that the technical effect of the feature in question is still guaranteed. This term usually denotes ±10% deviation from the indicated numerical value, preferably ±5%.
[0077] Furthermore, the terms first, second, third, (a), (b), (c), and the like, as used in the specification, are used as identifiers to describe similarity, not as an ordering or time sequence. It is to be understood that the terms so used are interchangeable under appropriate circumstances and that the embodiments of the present application described herein are capable of
[0078] The following terms or definitions are provided solely to aid in the understanding of the present application. These definitions should not be construed to limit the scope or applicability of the claims.
[0079] The term "read" or "reads" in the present application: English for "read" or "reads", refers to one or a set of nucleic acid sequences read by a sequencing platform.
[0080] The term "alignment" in the present application refers to the corresponding result between a sequencing read sequence and a reference sequence. A sequencing read sequence can have multiple alignments.
[0081] The term "taxon" in the present application refers to a classification concept with the same level in biological hierarchy. For example, when studying the evolutionary relationship of organisms, taxon is the concept of each biological group at the level of kingdom, phylum, class, order, family, genus, species, strain, etc. For example, the kingdom of protozoa, the order of primates, Staphylococcus aureus, Salmonella enterica subsp. enterica, etc. are all taxons. When studying systems biology, taxon can be a research unit of different genes, different genomic elements, different genomic positions, different sequence characteristics.
[0082] The term "species" in the present application refers to a special taxon, which is a group of organisms that can mate and reproduce offspring.
[0083] The term "(a group of taxons) mutually exclusive" in the present application refers to that, for any two taxons A and B selected from the group, taxon A does not contain taxon B, and taxon B does not contain taxon A. For example, the three taxons of Escherichia coli, Salmonella enterica, and Klebsiella pneumoniae are mutually exclusive; while the two taxons of Klebsiella and Klebsiella pneumoniae are not mutually exclusive.
[0084] The term "a sequencing read sequence aligns to a taxon", "a sequencing read sequence aligns to a taxon", "a sequencing read sequence supports a taxon" in the present application refers to that the alignment result of the sequencing read sequence contains a reference sequence from the taxon.
[0085] The term "single alignment" in the present application refers to that all alignment results of a sequencing read sequence correspond to reference sequences from the same taxon.
[0086] The term "multiple alignment" in the present application refers to that the alignment result of a sequencing read sequence contains reference sequences from two or more mutually exclusive taxons.
[0087] The term "misidentification" in the present application refers to the alignment result of the sequencing read sequence actually from a certain classification unit containing the reference sequence from another classification unit which neither contains nor is contained by the aforementioned classification unit. It should be noted that the alignment result here is not computationally incorrect, and the base identity rate can be very high, but the classification unit on the alignment does not match the source classification unit of the sample. Such misalignment generally occurs between classification units with relatively close genomic relationships.
[0088] The term "experimental contamination" in the present application refers to non-sample-derived nucleic acid molecules introduced during the processing of the sample after the sample has been processed. These signals are not nucleic acid molecules of the sample itself, but are contaminated nucleic acid molecules introduced during the experiment. Possible sources include sampling, instruments, reagents, etc.
[0089] The term "detected signal" in the present application refers to all classification units reported after analysis of the sequencing data of a sample. It includes true signals and noise signals.
[0090] The term "true signal" in the present application refers to the nucleic acid molecules actually contained in the sample. In the experimental design, the present application will incorporate specific known high-purity nucleic acid molecules into the nucleic acid-free water. These nucleic acid molecules are true signals. When the sample to be tested is an unknown sample, the nucleic acid molecules without any treatment are true signals.
[0091] The term "noise signal" in the present application refers to non-true signals introduced during the processing of the sample after the sample has been processed. These signals are not nucleic acid molecules of the sample itself, but are contaminated nucleic acid molecules introduced during the experiment or false alignment introduced during bioinformatics analysis.
[0092] As described in the above claims, the core of the present application is based on the fact that "introduced noise signals form hidden subgroups", and the connection between the components in the sample is analyzed, so that the introduced noise signals can be more effectively removed while the true signals are maximally preserved. The method can be based solely on the hidden subgroups formed between the sample to be tested and the negative control for noise reduction, or based on the hidden subgroups formed between the sample to be tested and the negative control, and the hidden subgroups formed between the sample to be tested and the bioinformatics control for noise reduction. Therefore, the method for distinguishing bioinformatics true signals from background signals described in the present application can be as follows:
[0093] According to an aspect of the present application, the bioinformatics analysis method for distinguishing true signals from background signals described in the present application comprises the following steps:
[0094] Step 1) sequencing step of the sample to be tested and the negative control sample;
[0095] Step 2) Grouping step of sequencing data of the test sample and the negative control sample according to taxonomic units;
[0096] Step 3) Taxonomic unit statistics step of the test sample and the negative control sample;
[0097] Step 4) Test sample and negative control result comparison step to calculate the correlation between taxonomic units;
[0098] Step 5) Test sample and negative control result comparison step to identify hidden subgroups.
[0099] The term "negative control" in the present application refers to a blank control set during the sequencing experiment, which undergoes the same operation process as the test sample. Such a negative control has the meaning of a traditional negative control. For example, during the sequencing experiment, it can be a nucleic acid-free water sample. In theory, the detection signal of the sample should be empty, i.e., nothing should be detected. However, due to experimental contamination and false positives introduced by bioinformatics analysis, the sample will become the best representative of noise signals in the same batch experiment. For example, during the sequencing experiment, the negative control is a blank system that does not contain the detection target relative to the detection target. In theory, as long as a detection target is considered to be contained in sample 1 but not in sample 2, sample 2 can be used as the negative control of sample 1, so as to find the element contained only in sample 1.
[0100] The "taxonomic unit component calculation method of sequencing data" in the present application generally refers to a calculation method for determining whether various taxonomic units appear in sequencing data and their proportion through bioinformatics analysis of sequencing data. The preferred method of the present application is to obtain the taxonomic unit component of sequencing data after "hidden subgroup noise removal". This method can effectively remove false positive results in taxonomic unit component calculation. It can be understood that the present application solves the noise removal problem that the original conventional taxonomic unit component calculation method cannot solve by introducing the calculation of "hidden subgroups", effectively improving the specificity and accuracy of pathogen detection; at the same time, the calculation framework of the present application is independent of the selection of specific sequencing platforms and is not limited to sequencing platforms, and can be applied to sequencing data of various platforms such as second-generation sequencing technology and third-generation sequencing technology; moreover, the calculation framework is also independent of the specific sequencing data source. It can be understood in the art that the calculation framework of the present application is for any homologous sequence misalignment, therefore, the sequence source does not limit the application of the present application. In addition to the preferred macrogenomic data source of the present application, other genomic or genetic data sources are also suitable for the present application.
[0101] The term "hidden sub-group", "sub-group" or "sub-group" in the present application refers to a form of co-occurrence of signals that are assumed to exist, which can consist of two or more elements. The hidden sub-group requires mutual correlation or connection between its elements. The hidden sub-group can be introduced from the same source in the experimental process or bioinformatics analysis process, and the ratio between the elements in the hidden sub-group remains stable under two or more conditions. In some embodiments, the hidden sub-group is formed between the test sample and the positive control, the hidden sub-group is formed between the test sample and the negative control, the hidden sub-group is formed between the test sample and the bioinformatics control, the hidden sub-group is formed between the test sample and other controls, the hidden sub-group is formed between the test sample and the test sample, the hidden sub-group is formed between the negative control and the positive control, the hidden sub-group is formed between the bioinformatics control and the negative control, the hidden sub-group is formed between the other control and the positive control, the hidden sub-group is formed between the bioinformatics control and the negative control, the hidden sub-group is formed between the other control and the negative control, or the hidden sub-group is formed between the bioinformatics control and the other control.
[0102] For example, in the bioinformatics noise reduction process, the hidden sub-group is the same source of noise signals introduced in the experimental process or analysis process, and these noise signals from the same source undergo various treatments simultaneously under different conditions, but the internal ratio relationship remains stable, so the ratio between the elements in the hidden sub-group remains stable under two or more conditions. In absolute quantification, high-abundance background signals often appear simultaneously in negative controls, positive controls and test samples. The ratio relationship between these high-abundance background signals may remain stable in negative controls, positive controls and test samples, forming a hidden sub-group. These hidden sub-groups can be used as a pseudo-internal control (PIC), which can be used to calculate the ratio between negative controls and positive controls, and also to calculate the ratio between negative controls and test samples. Through the chain relationship, the ratio between positive controls and test samples can be calculated, and the absolute quantification of the test sample can be calculated according to the absolute quantification (known) of the positive control. The present application particularly refers to the hidden sub-group formed between the test sample and the negative control, and the hidden sub-group formed between the test sample and the bioinformatics control.
[0103] Step 5) refers to the hidden sub-group formed between the test sample and the negative control.
[0104] The term "correlation" or "connection" in the present application refers to the proportional relationship between two paired classification units, which can be strictly quantified and remains stable under two or more conditions.
[0105] The term "co-detection" in the present application refers to the detection of a classification unit under different conditions, such as the detection of a sample and a negative control, or the detection of a sample and a positive control, or the detection of a sample and a bioinformatics control, and the like.
[0106] The term "single detection" in the present application refers to the detection of a classification unit under a single condition, such as the detection of a sample or a negative control. In the expression, it is often indicated that the sample is detected alone or the negative control is detected alone.
[0107] The term "coexistence" in the present application refers to a classification unit that is both a true signal and a noise signal. This often occurs when the true signal in the sample contains nucleic acid molecules of environmental microorganisms, which are a common source of contamination.
[0108] The term "single existence" in the present application refers to a classification unit that is only one of a true signal or a noise signal.
[0109] In some embodiments, the grouping in step 2) is based on an alignment method for grouping the test sample and the negative control sample, respectively.
[0110] Preferably, the alignment method is to group the sequencing read sequences after sequence alignment using alignment software that retains non-single alignment results (such as BLASTN software).
[0111] In some embodiments, the grouping in step 2) is based on a non-alignment method for grouping the test sample and the negative control sample, respectively.
[0112] Preferably, the grouping is performed by methods including but not limited to kmer method, hash table method or string matching method.
[0113] The classification unit statistics described in the present application include but are not limited to the following statistical quantities: the number of supporting sequencing read sequences of each classification unit, the relative proportion of each classification unit, or the statistical quantity of each classification unit after certain normalization.
[0114] In some embodiments, the calculation of the correlation between classification units in step 4) is to pair each two classification units in the test sample according to the statistical quantity of the classification units in step 3), and calculate whether the proportion of the two classification units in the pair remains stable in the test sample and the negative control; if it remains stable, it is considered that the paired classification units have a correlation with each other.
[0115] In some embodiments, the identifying the hidden subgroup in step 5) is performed by correlating the classification units of step 2) and the classification units of step 4) with each other, processing and screening the classification unit correlations, and identifying the remaining links between the classification units as the hidden subgroup.
[0116] Preferably, the identifying and / or analyzing is performed by a hidden subgroup analysis without prior information; and the identifying and / or analyzing is performed by using an analysis method for analyzing the correlation between two or more elements or the characteristics of an element itself.
[0117] More preferably, the identifying and / or analyzing is performed by using a graph method in computer science, i.e., each classification unit in step 2) is taken as a node of a graph, and each pair of classification units having a link in step 4) is taken as an edge of the graph, thereby forming a complete undirected graph; in the undirected graph, a complete subgraph is found by using a classical graph analysis method, and the complete subgraph is taken as the hidden subgroup.
[0118] Further preferably, the number of nodes of the complete subgraph is 2 or more, such as 2, 3, 4, 5, or 6.
[0119] The terms "node" and "edge" refer to the constituent elements of a "graph" in computer data structure. A node is an element in a set, and an edge is a connection between two nodes. In the present application, a node is a classification unit, and an edge is a link between classification units.
[0120] The term "subgraph" refers to a subset of elements in an undirected graph.
[0121] The term "complete subgraph" refers to a subset of elements in an undirected graph, which satisfies that there is an edge between any two nodes in the subset.
[0122] The term "outlier" in the present application refers to an element in an undirected graph, which has no connection with other elements.
[0123] It can be understood that, based on the negative control, the present application can further perform noise reduction by bioinformatics control, thereby achieving a more excellent noise reduction effect. Therefore, in some embodiments of the present application, the step 5) can further include the following steps:
[0124] Step 6) constructing a bioinformatics control and counting the classification units thereof;
[0125] Step 7) comparing the results of the test sample and the bioinformatics control, and calculating the classification unit correlation step;
[0126] Step 8) Comparing the results of the test sample and the bioinformatics control, identifying the hidden subgroup.
[0127] The term "bioinformatics control" can include multiple parts: in the bioinformatics analysis process, there are two major methods to obtain the composition of the classification units of the test sample: one is based on sequence alignment method (such as BLASTN), and the other is not based on sequence alignment method (such as kmer method). Regardless of which method, due to the similarity between classification units and the errors introduced by sequencing, sequencing read sequences from real components will be incorrectly grouped into other classification units, forming bioinformatics noise. Specifically, in the case of classification unit grouping based on sequence alignment, bioinformatics noise is often aligned to some conservative sequencing reads at the same time as real components, forming multiple alignments. At this time, multiple alignments can indicate the source of bioinformatics noise, so the alignment results are the basis for good bioinformatics noise signals, which can be used as bioinformatics controls; and in the case of not grouping classification units based on sequence alignment, the reference sequence of the real component can be used to simulate sequencing read data according to the error rate distribution of the sequencer, and the classification results of the simulated data can be used to construct the basis reference of the bioinformatics noise signal as the bioinformatics control.
[0128] Further, the bioinformatics control of step 6) is based on the alignment results of step 2), and the alignment results are used as the bioinformatics control according to the alignment of the sequencing read sequences.
[0129] In some embodiments, the bioinformatics control of step 6) is based on the non-alignment results of step 2), and data simulation is performed on the reference genome of each classification unit, the sequencing read sequences of the classification unit are simulated according to the error distribution of the sequencer, and the simulated sequencing read sequences are grouped. The grouping results of the simulated data can be used as the bioinformatics control of the classification unit.
[0130] In some embodiments, step 7) is a statistical quantity for the classification units of step 6), in which each two classification units in the test sample are paired, and it is calculated whether the proportion of the two classification units in the pair is higher or flat in the test sample than in the bioinformatics control;
[0131] Preferably, if it is higher or flat, it is considered that the classification units of the pair have a relationship with each other;
[0132] More preferably, the proportion of the two classification units is obtained by dividing the statistical quantities of the classification units; or is obtained by statistical testing of the unit statistical quantities.
[0133] In some embodiments, the step 7) is to pair each two classification units in the test sample according to the statistical quantity of the classification units in step 6) and calculate whether the proportion of the two classification units in the pair is higher or flat in the test sample than in the bioinformatics control;
[0134] Preferably, if higher or flat, it is considered that the classification units in the pair are associated with each other;
[0135] More preferably, the proportion of the two classification units is obtained by dividing the statistical quantities of the classification units; or is obtained by performing statistical test on the unit statistical quantities.
[0136] In some embodiments, the step 8) is to perform classification unit association processing and screening according to the association between the classification units of the test sample in step 2) and the classification units in step 7), and identify the association between the remaining classification units as the hidden subgroup.
[0137] In some embodiments, the step 8) is to perform hidden subgroup identification and / or analysis according to the association between the classification units of the test sample in step 2) and the classification units in step 7) after removing the classification units forming the hidden subgroup in step 5).
[0138] Further, the identification and / or analysis is performed by a hidden subgroup analysis without prior information; the identification and / or analysis is performed by using an analysis method for analyzing the association between two or more elements or the characteristics of the elements themselves;
[0139] Preferably, the identification and / or analysis is performed by using a graph method in computer science, i.e., each classification unit of the test sample in step 2) is taken as a node of a graph, and each pair of associated classification units in step 7) is taken as an edge of the graph to form a complete undirected graph; in the undirected graph, a complete subgraph is found by using a classical graph analysis method, and the complete subgraph can be used as the hidden subgroup;
[0140] More preferably, the number of nodes of the complete subgraph should be 2 or more; such as 2, 3, 4, 5, 6.
[0141] In some embodiments, the above method further comprises the following steps:
[0142] Step 9) removing background noise in the test sample.
[0143] Preferably, the step 9) rejection is to the classification units of the sample to be tested in step 2), all classification units forming hidden subgroups in step 5) (contamination introduced in the experiment) are rejected; and / or, to the classification units of step 6), all classification units forming hidden subgroups in step 8) and having low statistical quantities in the paired two classification units (i.e. false positives introduced by bioinformatics analysis) are rejected.
[0144] In some embodiments, further comprising the following steps:
[0145] Step 10) classification unit abundance calibration step;
[0146] Preferably, the abundance calibration step is to re-group the sequencing read sequences according to the classification units supported by the alignment results after the noise classification units are rejected, and count the number of sequences in each group and the proportion of the total read sequences.
[0147] In addition, based on the core idea of the present application, when only bioinformatics noise is considered without considering experimental pollution, the noise reduction method performed is also one of the core methods of the present application, although the noise reduction effect of this single method is slightly inferior to the comprehensive noise reduction, but it cannot be denied that it can also achieve the effect claimed by the present application, and also belongs to the protection content of the present application, such as Figure 5 The schematic diagram of noise reduction based on bioinformatics control alone.
[0148] Therefore, according to some aspects of the present application, the bioinformatics analysis method for distinguishing bioinformatics true signals from background signals of the present application can also include the following steps:
[0149] Step 1) sequencing step of the sample to be tested;
[0150] Step 2) grouping the sequencing data of the sample to be tested into different classification units step;
[0151] Step 3) counting the classification units of the sample to be tested step;
[0152] Step 4) constructing a bioinformatics control and counting its classification units step;
[0153] Step 5) comparing the results of the sample to be tested with the bioinformatics control, calculating the relationship between the classification units step;
[0154] Step 6) comparing the results of the sample to be tested with the bioinformatics control, identifying hidden subgroups step.
[0155] In some embodiments, the grouping in step 2) is performed by a grouping method based on alignment; preferably, the sequencing read sequences are grouped after sequence alignment using alignment software that retains non-single alignment results (such as BLASTN software);
[0156] In some embodiments, the grouping in step 2) is based on a non-alignment method; preferably, the grouping is performed by a method including but not limited to a kmer method, a hash table method or a string matching method.
[0157] In some embodiments, the statistics obtained in step 3) include but are not limited to the number of supporting sequencing reads of each classification unit, the relative proportion of each classification unit or the statistics of each classification unit after normalization.
[0158] In some embodiments, the bioinformatics control in step 4) is based on the alignment results of step 2);
[0159] Preferably, the bioinformatics control is used according to the alignment of the sequencing reads, wherein the alignment results are used as the bioinformatics control.
[0160] Alternatively, the bioinformatics control in step 4) is based on the non-alignment results of step 2);
[0161] Preferably, the reference genome of each classification unit is simulated according to the error distribution of the sequencer, the sequencing reads of the classification unit are simulated, and the simulated sequencing reads are grouped, and the grouping results of the simulated data are used as the bioinformatics control of the classification unit.
[0162] In some embodiments, step 5) is to pair each two classification units in the sample to be tested according to the statistics of the classification units in step 4), and to calculate whether the proportion of the two classification units in the pair is higher or flat in the sample to be tested than in the bioinformatics control; if higher or flat, it is considered that the classification units in the pair are related to each other;
[0163] Preferably, the proportion of the two classification units is obtained by dividing the statistics of the classification units; or is obtained by performing statistical test on the statistics of the units.
[0164] In some embodiments, step 6) is to identify and / or analyze the relationship between the classification units in the sample to be tested in step 3) and the classification units in step 5), and to identify the relationship between the remaining classification units as the hidden subgroup.
[0165] In some embodiments, the identification and / or analysis is performed by a hidden subgroup analysis without prior information.
[0166] Preferably, the identification is performed by using an analysis method for analyzing the relationship between two or more elements or the characteristics of the elements themselves.
[0167] More preferably, the identifying and / or analyzing is performed by using graph method of computer science, i.e. taking each classification unit in step 2) as a node of a graph, and each pair with connection in step 5) as an edge of the graph, to form a complete undirected graph; in the undirected graph, using classical graph analysis method to find a complete subgraph, which can be used as the hidden subgroup.
[0168] Further preferably, the complete subgraph has a node number of 2 or more.
[0169] In some embodiments, it further comprises the following steps:
[0170] Step 7) removing background noise in the sample to be tested.
[0171] Preferably, the removing in step 7) is removing all classification units in step 6) that form a hidden subgroup and have low statistical quantity in the two classification units in the pair, from the classification units in step 2).
[0172] It can be understood that the method of the present application is not limited to sequencing platforms, sequencing sample sources, etc., and therefore in any of the foregoing methods, the sequencing can be second-generation sequencing, third-generation sequencing or fourth-generation sequencing.
[0173] Preferably, the sequencing is from Illumina, BGI, ION TORRENT, PacBio, Roche, Helicos, ABI or nanopore sequencing.
[0174] More preferably, the sequencing is from Illumina, BGI or nanopore sequencing.
[0175] In some embodiments, the sequencing data is gene sequencing data.
[0176] Preferably, the sequencing data is genome sequencing data.
[0177] More preferably, the sequencing data is metagenome sequencing data.
[0178] Further preferably, the sequencing data is low-biomass metagenome sequencing data.
[0179] For better explanation of the present application, according to some specific embodiments of the present application, the method for distinguishing between true signals and background signals can be a specific method of the following steps, and this explanation does not limit the scope of the invention:
[0180] Step 1) sequencing step of the test sample and the negative control sample; Step 2) grouping step of the sequencing data of the test sample and the negative control sample into different classification units; Step 3) statistical step of the classification units of the test sample and the negative control sample; Step 4) step of constructing a bioinformatics control and counting the classification units thereof; Step 5) step of comparing the results of the test sample and the negative control and calculating the correlation between the classification units; Step 6) step of comparing the results of the test sample and the negative control and identifying hidden subgroups; and the method can further comprise: Step 7) step of comparing the results of the test sample and the bioinformatics control and calculating the correlation between the classification units; Step 8) step of comparing the results of the test sample and the bioinformatics control and identifying hidden subgroups; and Step 9) step of removing background noise in the test sample.
[0181] In some embodiments, the step 1) is parallel experimental operation of the test sample and the negative control sample, and a sequencing library is constructed. The parallel experimental operation requires that the experimental site, the operator, the reagent instrument, the operation process, the operation step, the operation method, the operator, the sequencing platform and other experimental conditions are as consistent as possible, so that the negative control can capture various background noise signals introduced in the experimental process systematically and comprehensively.
[0182] In some embodiments, the step 2) is grouping based on the sequencing data (generally base text files) of step 1) using an alignment-based method, using alignment software to perform sequence alignment on the sequencing readout sequence, and grouping according to the alignment results, for example, all sequencing readout sequences aligned to the same classification unit are grouped together. In some embodiments, the step 2) is grouping based on the sequencing data (generally base text files) of step 1) using a non-alignment-based method, the kmer method first divides the sequencing readout sequence into kmer, and divides the alignment database into kmer, and counts which reference sequence in the database is more consistent with the sequencing readout sequence, thereby grouping the sequencing readout sequence into the classification unit of the reference sequence. Preferably, for the alignment-based grouping method, the alignment software (such as BLASTN software) that retains non-single alignment results is used for sequence alignment and grouping of the sequencing readout sequence. Preferably, for the non-alignment-based grouping method, methods including but not limited to kmer method, hash table method, string matching method are used. This step processes the sample data and the negative control data in the same way.
[0183] In some embodiments, the step 3) calculates the statistics of each classification unit for the test sample and the negative control sample respectively based on the grouping results of the classification unit in step 2). For example, the number of sequencing read sequences supporting each classification unit, the relative proportion of each classification unit, and the statistics of each classification unit after certain normalization (such as reads per million (RPM), etc.) are calculated. In some embodiments, the step 3) divides the alignment results of the sequencing read sequences into two groups according to the alignment of the sequencing read sequences, i.e. single alignment (i.e. the classification unit to which each sequencing read sequence is aligned is the same) and multiple alignment (i.e. the classification unit to which each sequencing read sequence is aligned is different). When calculating the statistics of the classification unit, the error of the single alignment result is smaller, and only the results of the single alignment are used to calculate the number and percentage of the sequencing read sequences aligned to each classification unit. In some embodiments, the step 3) directly uses the grouping method (such as kmer method) that is not based on alignment to obtain the number and percentage of sequencing read sequences supporting each classification unit based on the non-alignment results of step 2). The same processing is performed on the test sample data and the negative control data to obtain the classification unit composition results of the test sample and the negative control.
[0184] In some embodiments, the step 4) is used to extract the biological control from the alignment results of the test sample or simulate the biological control from the classification unit results of the test sample for subsequent steps to eliminate the background noise signal introduced by the biological analysis. If the classification unit A' obtained in step 3) is a biological false detection of the classification unit A, then the number of A' must be lower than the number of A' in the multiple alignment results of the real classification unit A (or the number of A' in the classification unit of the simulation data of the real classification unit A). Based on this, the biological control is constructed for each classification unit obtained in step 3), and the pairing relationship in the biological control is limited to the classification unit with a high number as a potential real signal and the classification unit with a low number as a potential noise signal. Here, the number refers to the statistics in step 3). The high and low numbers are determined according to the meaning in the experimental design.
[0185] In some embodiments, the step 4) uses the alignment results as the biological control based on the alignment results of step 2) according to the alignment of the sequencing read sequences. In some embodiments, the step 4) simulates the reference genome of each classification unit based on the non-alignment results of step 2), simulates the sequencing read sequences of the classification unit according to the error distribution law of the current sequencer, and groups the simulated sequencing read sequences. The grouping results of the simulation data can be used as the biological control of the classification unit.
[0186] In some embodiments, the step 5) is to pair each two classification units in the test sample and calculate whether the ratio of the two classification units in the pair is stable in the test sample and the negative control. If it is stable, it is considered that the two classification units in the pair are associated with each other. In some specific embodiments, the ratio of the two classification units in each pair is obtained by dividing the classification unit statistics, and whether it is stable is obtained by calculating whether the ratio of the two (or more) conditions is close to 1. For example, the statistics of two classification units A and B in condition one are 10 and 30, respectively, and the ratio is 0.333. The statistics in condition two are 1 and 3, respectively, and the ratio is 0.333. The ratio of the pair A and B in condition one (0.333) and condition two (0.333) is (0.333:0.333) 1, which is stable. In some specific embodiments, the ratio of the two classification units in each pair is obtained by statistical test on the unit statistics, such as the p value of goodness-of-fit test. For example, the statistics of two classification units A and B in condition one are 10 and 30, respectively, and the ratio is 0.333. The statistics in condition two are 1 and 3, respectively, and the goodness-of-fit test is performed on the four numbers 10, 30, 1, and 3 (two-tailed), and the p value is calculated. If the p value is greater than or equal to a certain threshold (such as 0.05), it is considered to be stable.
[0187] In some specific embodiments, it is necessary to set a threshold to determine whether the ratio of a pair in the test sample and the negative control is close to 1, such as greater than or equal to 0.5 and less than or equal to 2. The setting of the threshold is related to the specific application scenario, and cannot be unified in different scenarios, and the specific threshold range is not protected. In some specific embodiments, the ratio of a pair in the test sample and the negative control is determined by statistical test, and the threshold is often 0.05, 0.01, etc. The setting of the threshold is related to the specific application scenario, and cannot be unified in different scenarios, and the specific threshold range is not protected.
[0188] In some embodiments, the step 6) is to identify and / or analyze the hidden subgroup based on the association between the classification units of the sample to be tested in step 2) and the classification units in step 5). In some embodiments, the hidden subgroup analysis is performed without prior information, for example, using a graph method in computer science, i.e., each classification unit in step 2) is taken as a node of a graph, and each pair of classification units with association in step 5) is taken as an edge of the graph, to form a complete undirected graph. In the undirected graph, a complete subgraph, also known as a "clique", is found using classical graph analysis methods. These "cliques" can be used as hidden subgroups. Preferably, the number of nodes of the complete subgraph should be 2 or more. More preferably, the number of nodes of the complete subgraph should be 2, 3, 4, 5, 6.
[0189] In some embodiments, the step 7) is to pair each two classification units in the sample to be tested and calculate whether the ratio of the two classification units (potential true component to potential noise component) in the sample to be tested is higher or flat than that in the bioinformatics control based on the classification unit statistics in step 4). In some embodiments, the step 7) is to pair each two classification units in the sample to be tested and calculate whether the ratio of the two classification units (potential true component to potential noise component) in the sample to be tested is higher or flat than that in the bioinformatics control after removing the classification units forming hidden subgroups in step 6) based on the classification unit statistics in step 4). If it is higher or flat, it is considered that the classification units in the pair have association with each other. In some specific embodiments, the ratio of the two classification units in each pair is obtained by dividing the classification unit statistics, and whether it is higher or flat in the sample to be tested than in the bioinformatics control is compared. In some specific embodiments, the ratio of the two classification units in each pair is obtained by statistical test on the unit statistics, such as the p-value of goodness-of-fit test. For example, the statistics of two classification units A and B in condition one are 10 and 30, respectively, and the ratio is 0.333. The statistics in condition two are 1 and 3, respectively. The goodness-of-fit test is performed on the four numbers 10, 30, 1, and 3 (one-tailed test), and the p-value is calculated. If the p-value is greater than or equal to a certain threshold (such as 0.05), it is considered that the ratio in the sample to be tested is higher or flat than that in the bioinformatics control.
[0190] In some embodiments, the step 8) performs the identification and / or analysis of the hidden subgroup based on the connection between the classification unit of the sample to be tested in step 2) and the classification unit in step 7). In some embodiments, the step 8) performs the identification and / or analysis of the hidden subgroup based on the connection between the classification unit of the sample to be tested in step 2) and the classification unit in step 7) after the classification unit forming the hidden subgroup in step 6) is removed. In some embodiments, the hidden subgroup analysis is performed in a non-priori manner, for example, using a graph method in computer science, i.e., each classification unit of the sample to be tested in step 2) is taken as a node of the graph, and each pair of connected classification units in step 7) is taken as an edge of the graph, to form a complete undirected graph. In the undirected graph, a complete subgraph, also called a "clique", is found using a classical graph analysis method. These "cliques" can be used as hidden subgroups. Preferably, the number of nodes of the complete subgraph should be 2 or more. More preferably, the number of nodes of the complete subgraph should be 2, 3, 4, 5, 6.
[0191] In some embodiments, the step 9) considers all the classification units forming the hidden subgroup in step 6) as contamination introduced by the experimental process, and all the classification units forming the hidden subgroup in step 8) and having a lower statistical quantity in the two classification units of the pair as false positives introduced by the bioinformatics analysis, and marks them as noise signals and removes them.
[0192] The present application also provides a computer system comprising a memory, a processor and a computer program stored on the memory, wherein the processor executes the computer program to implement the steps of any of the above-mentioned methods.
[0193] According to the core idea of the present application, in some embodiments, the present application can also relate to a computer-readable storage medium having a computer program stored thereon, characterized in that the computer program is executed by a processor to implement the steps of any of the above-mentioned methods.
[0194] In some embodiments, the present application can also relate to a computer program product comprising a computer program, characterized in that the computer program is executed by a processor to implement the steps of any of the above-mentioned methods.
[0195] In some embodiments, the present application can also relate to the application of a hidden subgroup in distinguishing true signals from background signals in bioinformatics analysis, wherein the hidden subgroup comes from the same source of noise signal introduction in the experimental process and / or the bioinformatics analysis process, and the proportion between the elements in the hidden subgroup remains stable under two or more conditions;
[0196] Preferably, the hidden subgroup is a hidden subgroup formed between the test sample and the negative control, and / or, the hidden subgroup is a hidden subgroup formed between the test sample and the bioinformatics control.
[0197] The present application will now be described in conjunction with specific embodiments.
[0198] Example 1: Invention Exploration Design
[0199] First, this application examines the entire process of genome (metagenomic) sequencing methods and, after careful analysis, breaks it down as follows. For example... Figure 1 As shown in steps 1-4, during sample preparation, a series of experimental contaminants are simultaneously introduced into both the sample and the negative control. Specifically, the sample originally contains only real nucleic acid molecules (solid dots), while the negative control sample contains nothing. Step 1 introduces a subgroup of contaminated nucleic acid molecules (squares), Step 2 introduces a new subgroup of contaminated nucleic acid molecules (hollow circles), and the proportion of the contaminated nucleic acid molecule subgroup represented by the squares changes (shrinks overall), as does the proportion of the real nucleic acid molecule subgroup represented by the solid dots. In Step 3, a new group of contaminated nucleic acid molecules (small solid dots A) is introduced; this nucleic acid molecule is also a real nucleic acid molecule (solid dot A), i.e., coexisting nucleic acid molecules. Furthermore, Step 3 also involves changes in the proportions of the original nucleic acid molecule subgroups (real nucleic acid molecule subgroup, square contaminated nucleic acid molecules, hollow circle contaminated nucleic acid molecules). In Step 4, cross-contamination between samples (such as aerosols) also occurs.
[0200] During sequencing, all nucleic acid molecules are randomly sampled using a Bernoulli process and their nucleic acid sequences are obtained by a sequencer (e.g., ...). Figure 1 (Sequencing step). After this, the nucleic acid sequences of all real nucleic acid molecules will undergo bioinformatics (such as...) Figure 1, bioinformatics) analysis, and some false alignments inevitably occur, forming some false detections, such as A' and A" in the figure, C' and C". That is, in the whole process, for the sample, from the original true signals A B C to the final detection results A B C D E F G H I A' A" C' C", only A B C are true signals, and D E F G H I A' A" C' C" are all noise signals, and D E F G H I are experimental pollution, and A' A" C' C" are bioinformatics false detections. Moreover, for the true signal A, as a coexisting nucleic acid molecule, the final detection result is the superposition of the true nucleic acid molecule A and the contaminated nucleic acid molecule A. For the negative control, it does not contain any true nucleic acid molecule, and the final detection results A B D E F G H I J A' A" are all noise signals, and among them, A B D E F G H I J are derived from experimental pollution, which are true nucleic acid molecules (D E F G H I J are introduced by steps 1 and 2, A is introduced by step 3, and B is introduced by step 4). And A' A" are bioinformatics false detections.
[0201] Secondly, the present application finds that if the present application looks at each detection separately to trace back how this detection is introduced into the sample, the present application will be lost in the above-mentioned overall transformation process. However, the present application finds that the proportional relationship between the detections is more important than the detection of the microbial sequence itself. A subgroup of noise signals introduced in the same step will maintain the group proportion during subsequent processing and transformation. As shown in Figure 2 , even if the proportion between different subgroups changes greatly, the core idea of the present application is that the proportion within the group remains unchanged. Moreover, the present application finds that, through sequencing analysis of samples with known composition, the proportional relationship between true signals and noise signals cannot be maintained stable in the sample to be tested and the negative control (the stable proportional relationship is a dark dot, and the unstable proportional relationship is a light dot), but, on the contrary, the proportional relationship between noise signals and noise signals is mostly maintained stable in the sample to be tested and the negative control (as shown in Figure 3 ).
[0202] Based on this feature, the present application establishes a denoising algorithm based on hidden subgroup analysis (Hugo-DNA) Figure 4). The sequencing read sequences of the test samples (SP) and the negative control (NC) are quality controlled and host filtered. The filtered data are data aligned, wherein the single alignment results (test samples and negative control) are the basis for the statistics of the classification units, and the multiplex alignment results of the test samples are retained and used for the construction of bioinformatics controls later. First, the test samples are compared with the negative control, and whether the proportional relationship of the pairwise pairs of the classification units composed of the test samples remains stable in the test samples and the negative control is calculated. All the classification units of the test samples are the vertices of the undirected graph, and the pairwise pairs whose proportional relationship remains stable are the edges of the undirected graph, to form an undirected graph. Hidden subgroups (i.e. complete subgraphs) are found in the undirected graph, and the classification units of all hidden subgroups are marked as noise signals. Second, for the classification units that do not form hidden subgroups (i.e. outliers), bioinformatics controls based on alignment results are constructed, and the test samples are compared with the bioinformatics controls. If the proportion in the test sample is not higher than the proportion in the bioinformatics control, it is considered that there is a true signal and an error detection similar to the true signal in the paired classification unit, and at this time, the classification unit with less supporting sequencing read sequence number is considered as an error detection, which is classified as a noise signal and is removed. Finally, all the signals left are marked as true signals.
[0203] In addition, based on the core idea of the present application, the method of only reducing noise for experimental pollution or bioinformatics noise is also the core method of the present application, although the noise reduction effect of this single method is slightly inferior to the above, but it can also achieve the effect claimed in the present application, and also belongs to the protection content of the present application, such as Figure 5 The schematic diagram of noise reduction based on bioinformatics control alone is shown.
[0204] Example 2 Establishment of the method of the present application
[0205] Based on Example 1, the specific steps of establishing the Hugo-DNA method of the present application are as follows:
[0206] 1) Constructing a test sample with known composition: the present application adds high-purity artificially synthesized nucleic acid molecules with known sequences to nucleic acid-free water. The true signal of this sample is the artificially synthesized nucleic acid molecules with known sequences. All other detections are considered as noise signals.
[0207] 2) Constructing a negative control sample: the present application uses nucleic acid-free water as a negative control, and processes the negative control completely according to the processing of the test sample, to ensure that the negative control can capture the noise signals of the whole process.
[0208] 3) Library construction: The VAHTS Universal Plus DNA Library Prep Kit for Illumina (Vazyme, ND617) was used for library construction, and the operation was performed according to the instructions. The Agilent 2100 bioanalyzer, Agilent 2100 DNA 1000 kit (Agilent, 5067-1504) was used for library quality inspection, and the Qubit 4.0 fluorometer was used for quantification. Each library was normalized to the same concentration and mixed, and finally completed double-end 150bp sequencing on the Nextseq 550 Dx sequencing platform.
[0209] 4) Microbial community composition analysis of off-line data: The genome database used for annotation was extracted from the nt database of NCBI, including archaea, bacteria, fungi and viruses. The BLASTN was used for sequence alignment, and the single alignment reads were used for subsequent sample detection. These detections constitute the list of taxonomic units detected by the sample. Both the test sample and the negative control can obtain a taxonomic unit composition to facilitate further removal of experimental introduced pollution. For the sample, the multiple alignment reads are also counted. These multiple alignment reads are used to construct the multiple alignment results, i.e. a taxonomic unit can be aligned to the number of reads of other taxonomic units at the same time, to facilitate further removal of false positives introduced by bioinformatics.
[0210] 5) The taxonomic unit composition of the test sample and the taxonomic unit composition of the negative control are compared to find the subset of taxonomic units commonly detected. These taxonomic units constitute the vertices in the undirected graph. For these commonly detected taxonomic units, the ratio between each other is calculated, and if the ratio of the pair is consistent in the sample and in the negative control, the application considers that there is a connection between the pair of taxonomic units, forming an edge of the undirected graph.
[0211] 6) Whether the ratio of a pair of taxonomic units in the test sample and the negative control is consistent is determined by the following rules:
[0212] a) If the difference between the ratios is less than 2 times;
[0213] b) If the p-value of the goodness-of-fit test is higher than 0.05.
[0214] 7) Search for a complete subgraph in the above undirected graph using a depth-first algorithm. The elements contained in the complete subgraph are marked as noise signals. Other taxonomic units are marked as candidate real signals.
[0215] 8) For each candidate true signal, compare it with the multiple alignment results. Similarly, check if the ratio between each pair of classification units is consistent with the ratio in the multiple alignment results. The rule of consistency is the same as 6) above. If the ratio of a pair of classification units remains consistent, the present application considers the classification unit with less reads as a false positive of the bioinformatics alignment, and marks it as a noise signal. The rest of the signals are marked as true signals.
[0216] Example 3: Hidden subgroups in the test sample
[0217] The present application further verifies the existence of the hidden subgroups in the core hypothesis of the Hugo-DNA algorithm. The specific steps are as follows. Based on the method of Example 2 (steps 1-7), the present application constructs an undirected graph of the test sample ( Figure 6 ). Among them, each circle point is a vertex in the undirected graph, that is, a classification unit of the test sample; each gray line is an edge of the undirected graph, indicating that the proportion relationship of the two classification units connected remains stable in the test sample and the negative control. As shown in Figure 6 , all vertices in the undirected graph except true signals (indicated by arrows) are connected, indicating that all classification units except true signals come from different hidden subgroups and are all noise.
[0218] Example 4: Comparison of the present application method with existing methods
[0219] Further, this embodiment compares the results of Hugo-DNA with other existing methods, and the specific steps are as follows:
[0220] 1) Based on the method of Example 2 (steps 1-4), this embodiment obtains the composition of classification units of the test sample and the negative control. These results will be used as input for subsequent different noise reduction methods. And also as the original data case is shown in Figure 6 (RAW). Among them, the true signals are indicated by arrows.
[0221] 2) For the Hugo-DNA method, this embodiment uses steps 5-8 of Example 2 to analyze and obtain the annotation of true signals and noise signals. Figure 7 , DNA)
[0222] 3) For the Z-score method of noise reduction, this embodiment normalizes the relative abundance of classification units of the test sample and the negative control by Z-score, and marks the classification units with a relative abundance in the test sample that is more than twice that in the negative control after normalization as true signals, and the others as noise signals. Figure 7 , ZSC)
[0223] 4) For Decontam method denoising, the present embodiment inputs the classification unit results of the test sample and the negative control into the software Decontam (version 1.10.0) for frequency-based denoising analysis. The selected threshold is 0.5, which additionally requires library concentration information when building a library, and the library concentration information uses Invitrogen TM Qubit TM Fluorometer before the machine is determined. Figure 7 , DCN
[0224] 5) For SourceTracker method denoising, the present embodiment inputs the classification unit results of the test sample and the negative control into the software (version 1.0) for analysis. The SourceTracker method outputs the probability of each classification unit being contaminated and the probability of being real, and the present embodiment marks the classification unit with a contamination probability greater than the real probability as contaminated, and marks the remaining classification units as real signals. Figure 7 , STR
[0225] The results are shown in Figure 7 The test sample includes about 2000 classification units, and among them, only one is a real signal (red dot) and the rest are noise signals (green dot). Compared with other methods, Hugo-DNA retains fewer noise signals while retaining real signals.
[0226] The above description of the specific embodiments of the present application does not limit the present application, and those skilled in the art can make various changes or modifications to the present application according to the present application, as long as they do not deviate from the spirit of the present application, and all should belong to the scope of the claims attached to the present application.
Claims
1. A bioinformatics analysis method for distinguishing between real bioinformatics signals and background signals, characterized in that, The method comprises the following steps: Step 1) sequencing the sample to be tested; Step 2) grouping the sequencing data of the sample to be tested into different classification units; the grouping is based on alignment or non-alignment method; Step 3) counting the classification units of the sample to be tested; Step 4) constructing a bioinformatics control and counting the classification units thereof; the bioinformatics control is based on the alignment or non-alignment result of step 2); Step 5) comparing the results of the sample to be tested and the bioinformatics control, and calculating the correlation between the classification units; for the count of the classification units of step 4), each two classification units in the sample to be tested are paired, and it is calculated whether the proportion of the two classification units in the pair is higher or flat in the sample to be tested than in the bioinformatics control; if it is higher or flat, it is considered that the classification units in the pair are correlated with each other; Step 6) comparing the results of the sample to be tested and the bioinformatics control, and identifying the hidden subgroup; for the classification units of the sample to be tested in step 3) and the correlation between the classification units in step 5), the correlation between the classification units is identified and / or analyzed, and the correlation between the remaining classification units is identified as a hidden subgroup; The hidden subgroup is a form in which signals of two or more conditions appear together, and is composed of two or more elements. The proportion between the elements in the hidden subgroup remains stable under two or more conditions. The hidden subgroup is introduced from the same source in the experimental process or bioinformatics analysis process. The hidden subgroup is selected from the hidden subgroup formed between the sample to be tested and the bioinformatics control.
2. The bioinformatics analysis method according to claim 1, characterized in that, In step 2), the grouping is based on alignment method; the sequencing read sequences are grouped after sequence alignment by using alignment software that retains non-single alignment results.
3. The bioinformatics analysis method of claim 1, wherein, In step 2), the grouping is based on non-alignment method; the grouping is performed by using kmer method, hash table method or string matching method.
4. The bioinformatics analysis method of claim 1, wherein, In step 3), the counting obtains the following statistical quantities: the number of supporting sequencing read sequences of each classification unit, the relative proportion of each classification unit, or the statistical quantity of each classification unit after certain normalization.
5. The bioinformatics analysis method according to claim 1, characterized in that, In step 4), the bioinformatics control is based on the alignment result of step 2), and the alignment result is used as the bioinformatics control according to the alignment of the sequencing read sequences.
6. The bioinformatics analysis method of claim 1, wherein, In step 4), the bioinformatics control is based on the non-alignment result of step 2). The reference genome of each classification unit is simulated, the sequencing read sequences of the classification unit are simulated according to the error distribution law of the sequencer, and the grouped simulated sequencing read sequences are used as the bioinformatics control of the classification unit.
7. The bioinformatics analysis method of claim 1, wherein, In step 5), the proportion of the two classification units is obtained by dividing the classification unit count; or is obtained by statistical test on the unit count.
8. The bioinformatics analysis method according to claim 7, characterized in that, The identification and / or analysis is a hidden subgroup analysis without prior information.
9. The bioinformatics analysis method of claim 8, wherein, The hidden subgroup analysis without prior information is performed by using an analysis method for analyzing the correlation between two or more elements or the characteristics of the elements.
10. The bioinformatics analysis method of claim 8, wherein, The identification and / or analysis is analyzed by using the graph theory method of computer science, that is, each classification unit in step 2) is taken as a vertex of a graph, and each pair with a connection in step 5) is taken as an edge of the graph to construct a complete undirected graph; in the undirected graph, a complete subgraph is found, and the complete subgraph is taken as a hidden subgroup; the number of vertices of the complete subgraph is 2 or more than 2.
11. The bioinformatics analysis method according to any one of claims 1-10, characterized in that, The sequencing is second-generation sequencing, third-generation sequencing or fourth-generation sequencing.
12. The bioinformatics analysis method of claim 11, wherein, The sequencing is from Illumina, BGI, ION TORRENT, PacBio, Roche, Helicos, ABI or nanopore sequencing.
13. The bioinformatics analysis method of claim 12, wherein, The sequencing is from Illumina, BGI or nanopore sequencing, and the sequencing data is gene sequencing data.
14. The bioinformatics analysis method of claim 13, wherein, The gene sequencing data is metagenomic sequencing data.
15. The bioinformatics analysis method of claim 14, wherein, The metagenomic sequencing data is low-biomass metagenomic sequencing data.
16. A bioinformatics analysis method for background noise removal of sequencing data, characterized in that, The method of any one of claims 1-10 further comprises the following steps: Step 7) eliminating background noise in the sample to be tested; the elimination of step 7) is for the classification units of step 2, and all classification units that form a hidden subgroup in step 6) and have a low statistical amount in the two classification units of the pair are eliminated.
17. A bioinformatics analysis system for distinguishing between real bioinformatics signals and background signals, characterized in that, The method comprises the following modules: Module 1) a sample to be tested sequencing module; Module 2) a sample to be tested sequencing data grouping into different classification units module; the grouping is based on alignment or non-alignment grouping; Module 3) a sample to be tested classification unit statistics module; Module 4) a bioinformatics control construction and classification unit statistics module; the bioinformatics control is based on the alignment or non-alignment results of step 2); Module 5) a sample to be tested and bioinformatics control result comparison and classification unit correlation calculation module; the calculation is for the statistical amount of the classification units of step 4, each two classification units in the sample to be tested are paired, and it is calculated whether the proportion of the two classification units in the pair is higher or flat in the sample to be tested than in the bioinformatics control; if it is higher or flat, it is considered that the classification units in the pair have a connection with each other; Module 6) a sample to be tested and bioinformatics control result comparison and hidden subgroup identification module; the identification is for the classification unit correlation of step 3) of the sample to be tested and step 5), the identification and / or analysis of the classification unit correlation, and the identification of the connection between the remaining classification units as a hidden subgroup; Module 7) a sample to be tested background noise elimination module; The hidden subgroup is a form in which signals commonly appear under two or more conditions, is composed of two or more elements, and the proportion between the elements in the hidden subgroup remains stable under two or more conditions; the hidden subgroup is a signal introduced from the same source in the experimental process or the bioinformatics analysis process; the hidden subgroup is selected from the hidden subgroup formed between the sample to be tested and the bioinformatics control.
18. The bioinformatics analysis system of claim 17, wherein, The sequencing is second-generation sequencing, third-generation sequencing or fourth-generation sequencing.
19. The bioinformatics analysis system of claim 18, wherein, The sequencing is from Illumina, BGI, ION TORRENT, PacBio, Roche, Helicos, ABI, or Nanopore sequencing.
20. The bioinformatics analysis system of claim 19, wherein, The sequencing is from Illumina, BGI, or Nanopore sequencing, and the sequencing data is gene sequencing data.
21. The bioinformatics analysis system of claim 20, wherein, The gene sequencing data is metagenomic sequencing data.
22. The bioinformatics analysis system of claim 21, wherein, The metagenomic sequencing data is low-biomass metagenomic sequencing data.
23. A computer system comprising a memory, a processor, and a computer program stored on the memory, wherein, The processor executes the computer program to implement the steps of the method of any one of claims 1-10.
24. A computer readable storage medium having stored thereon a computer program, characterized in that, The computer program, when executed by the processor, implements the steps of the method of any one of claims 1-10.
25. A computer program product comprising a computer program, characterised in that, The computer program, when executed by the processor, implements the steps of the method of any one of claims 1-10.
Citation Information
Patent Citations
Hidden subgroup-based raw signal noise reduction analysis method and system
CN115719614A