A bioinformatics analysis method and system based on hidden subgroups

By using a bioinformatics analysis method based on hidden subgroups and constructing an undirected graph using graph theory, the problem of random perturbation of single-element research objects in bioinformatics analysis is solved, achieving more efficient noise reduction and quantification, and is applicable to various sequencing platforms and samples.

CN115732032BActive Publication Date: 2025-11-25HUGOBIOTECH BEIJING CO LTD +1
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202111000286.X
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2021-08-27
Publication Date
2025-11-25
Estimated Expiration
2041-08-27

AI Technical Summary

Technical Problem

Existing bioinformatics analysis methods focus on a single element, making them susceptible to random disturbances and resulting in inaccurate results. Furthermore, existing noise reduction and quantification methods may lose true results or introduce erroneous results.

Method used

A bioinformatics analysis method based on hidden subgroups was adopted. By grouping sequencing readouts, statistically analyzing taxonomic units, calculating cross-relationships, and identifying hidden subgroups, an undirected graph was constructed using graph theory methods from computer science to find and analyze hidden subgroups.

Benefits of technology

It effectively reduces the impact of random perturbations, improves the sensitivity and specificity of bioinformatics analysis, enhances the accuracy of noise reduction and quantification, and is suitable for various sequencing platforms and sample sources.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN115732032B_ABST
    Figure CN115732032B_ABST
Patent Text Reader

Abstract

The application relates to the field of bioinformatics, and particularly relates to a sequencing data analysis method and system. The method is based on "hidden subgroup theory" to analyze components in a sequencing process. The method can effectively analyze bioinformatics without relying on prior knowledge of any to-be-measured signal.
Need to check novelty before this filing date? Find Prior Art

Description

TECHNICAL FIELD

[0001] The present application relates to the field of bioinformatics analysis, in particular to a bioinformatics analysis method and system based on hidden subgroups. TECHNICAL BACKGROUND

[0002] Existing bioinformatics analysis methods often focus on a single element and perform subsequent analysis based on the data characteristics of the single element. For example, in microbiome research, the existing process takes the data characteristics of a single microorganism (such as presence or absence, abundance) as the research object. Taking the most basic denoising analysis of microbial communities as an example, an existing denoising method calculates the relative abundance of each microorganism and compares it with the relative abundance of the negative control. Microorganisms with a relative abundance lower than that of the negative control in the sample are considered to be noise. Other existing denoising methods also take the data characteristics of a single microorganism or prior information as the research object. This type of denoising analysis based on the data characteristics of a single microorganism does not perform well, as it easily loses real species while being unable to efficiently eliminate background microorganisms. Another application of microbial community analysis is the absolute quantification of the content of microorganisms in the community. Existing methods generally add a small amount of internal reference sequences to the sample to be tested for quantification. However, the addition of internal references itself is a disturbance to the composition of the microbial community, not to mention that in the case where the amount of microorganisms in the sample to be tested is unknown, it is difficult to determine the amount of internal reference to be added. Too little internal reference is difficult to detect, thereby losing the basis for quantification; too much internal reference obscures the real microorganism species and loses the real results.

[0003] This kind of analysis idea taking a single element as the research object is often affected by random disturbance, leading to inaccurate results, which may lose part of the real results on the one hand, and may also introduce false results on the other hand. In order to reduce the influence of random disturbance, researchers often need to design multiple repetitions or multiple condition data to search for specific patterns, and to clarify the real results by comparing the patterns. However, due to the fact that the overall influence of random disturbance on samples under different conditions may be very inconsistent, the randomness between data is difficult to effectively offset, so the idea of offsetting random disturbance has some effect, but its effect is not ideal. Taking the correlation between two or more elements in the same sample (i.e., hidden subgroup) as the research object for analysis can effectively avoid the influence of random disturbance on the real signal. First, the same sample ensures the consistency of the overall level of random disturbance. Second, because the probability of random disturbance of two or more elements in the same direction with the same amplitude is much lower than that of a single element, the correlation between two or more elements is less affected by random disturbance. In summary, in order to solve the problem that a single element is affected by random disturbance and cannot obtain satisfactory analysis performance, it is necessary to develop a research idea framework taking hidden subgroup as the research object, and apply the framework to noise reduction and quantitative analysis scenarios.

[0004] In view of this, the present application is proposed. SUMMARY

[0005] The core problem to be solved by the present application is how to use the information provided by the current experiment to reduce the influence of random disturbance on experimental data and efficiently analyze the current experimental results.

[0006] Therefore, the first object of the present application is to provide a sequencing data bioinformatics analysis method, system or computer program;

[0007] The second object of the present application is to provide the application of hidden subgroup in sequencing data bioinformatics analysis, including but not limited to the application in noise reduction, quantification and the like.

[0008] In order to achieve the above objects, the present application provides the following technical solutions:

[0009] The present application first provides a sequencing data bioinformatics analysis method, comprising the following steps:

[0010] Step 1) grouping sequencing read sequences into different classification units;

[0011] Step 2) classification unit statistics step;

[0012] Step 3) classification unit correlation calculation step;

[0013] Step 4) hidden subgroup identification step.

[0014] Further, the hidden subgroup is in the form of co-occurrence of signals under two or more conditions, and is composed of two or more elements; the ratio between any two elements in the hidden subgroup remains stable under two or more conditions.

[0015] Further, the hidden subgroup is introduced from the same source in the experimental process or bioinformatics analysis process.

[0016] Further, the hidden subgroup is formed between the sample to be tested and the positive control, the hidden subgroup is formed between the sample to be tested and the negative control, the hidden subgroup is formed between the sample to be tested and the bioinformatics control, the hidden subgroup is formed between the sample to be tested and other controls, the hidden subgroup is formed between the sample to be tested and the sample to be tested, the hidden subgroup is formed between the negative control and the positive control, the hidden subgroup is formed between the bioinformatics control and the negative control, the hidden subgroup is formed between the other control and the positive control, the hidden subgroup is formed between the other control and the negative control, or the hidden subgroup is formed between the bioinformatics control and the other control.

[0017] Further, the grouping method in step 1) includes but is not limited to alignment-based grouping method or non-alignment-based grouping method.

[0018] In some embodiments, the alignment-based grouping method uses alignment software to perform sequence alignment on the sequencing read sequences, and all sequencing read sequences aligned to the same taxonomic unit are grouped into one group.

[0019] In some embodiments, the grouping method without alignment includes but is not limited to kmer method, hash table method or string matching method.

[0020] Further, step 2) is based on the grouping results of the taxonomic units in step 1), and the statistical quantity of each taxonomic unit is calculated.

[0021] In some embodiments, the statistical quantity includes but is not limited to: the number of supporting sequencing read sequences of each taxonomic unit, the relative proportion of each taxonomic unit, and the quantity of each taxonomic unit after certain normalization.

[0022] Further, step 3) is to pair each two taxonomic units based on the statistical quantity of the taxonomic units in step 2), and to calculate whether the proportion of the two taxonomic units in the pair remains stable under two or more conditions; in some embodiments, the stable proportion indicates that the paired taxonomic units are related to each other.

[0023] Further, the conditions in step 3) include, but are not limited to, under the condition of the sample to be tested, under the condition of the negative control, under the condition of the positive control, under the condition of the bioinformatics control, or under the condition of other controls.

[0024] In some embodiments, in step 3), the pairing can be random pairing; the ratio of the two classification units can be obtained by dividing the classification unit statistics, or by statistical test of the classification unit statistics.

[0025] Further, the identification in step 4) is the identification and / or analysis of the hidden subgroup for the correlation between the classification units in step 2) and the classification units in step 3).

[0026] Further, the identification of the hidden subgroup in step 4) is performed by a priori information-free method.

[0027] Further, the identification and / or analysis is performed by using an analysis method for analyzing the correlation between two or more elements or the characteristics of the elements themselves.

[0028] In some embodiments, the identification and / or analysis is performed by using a graph method in computer science, i.e., each classification unit in step 2) is taken as a node of a graph, and each pair with a connection in step 3) is taken as an edge of the graph, to form a complete undirected graph; in the undirected graph, a complete subgraph is found by using a classical graph analysis method, and the complete subgraph is taken as the hidden subgroup; preferably, the number of nodes of the complete subgraph is 2 or more than 2.

[0029] Further, the sequencing data described above comes from a second-generation sequencing platform, a third-generation sequencing platform, or a fourth-generation sequencing platform; preferably, from an Illumina, BGI, ION TORRENT, PacBio, Roche, Helicos, ABI, or nanopore sequencing platform; more preferably, from an Illumina, BGI, or nanopore sequencing platform.

[0030] Further, the sequencing data described above is gene sequencing data, preferably genomic sequencing data; more preferably, metagenomic sequencing data; further preferably, low-biomass metagenomic sequencing data.

[0031] Further, when the purpose is to remove the background noise of the sequencing data, the identification of the hidden subgroup can be performed between the sample to be tested and the negative control sample, and / or between the sample to be tested and the bioinformatics control sample in the above method.

[0032] Further, when quantitative purposes are to be achieved, the identification of the hidden subgroup can be selected between the above-mentioned negative control and positive control samples, and between the to-be-tested sample and the negative control sample.

[0033] The application also provides a sequencing data bioinformatics analysis system, characterized in that it comprises the following modules:

[0034] Module 1) a module for grouping sequencing read sequences into different classification units;

[0035] Module 2) a classification unit statistics module;

[0036] Module 3) a classification unit correlation calculation module;

[0037] Module 4) a hidden subgroup identification module.

[0038] These modules respectively perform the above-mentioned respective steps.

[0039] The application also provides a computer-readable storage medium having a computer program stored thereon, characterized in that the computer program, when executed by a processor, implements the steps of the above-mentioned method.

[0040] The application also provides a computer program product comprising a computer program, characterized in that the computer program, when executed by a processor, implements the steps of the above-mentioned method.

[0041] The application also provides an application of a hidden subgroup in sequencing data bioinformatics analysis, characterized in that the hidden subgroup is a signal introduced from the same source in the experimental process or the bioinformatics analysis process, and the proportion between elements in the hidden subgroup remains stable under two or more conditions.

[0042] In some embodiments, the hidden subgroup is formed between the to-be-tested sample and the positive control, the hidden subgroup is formed between the to-be-tested sample and the negative control, the hidden subgroup is formed between the to-be-tested sample and the bioinformatics control, the hidden subgroup is formed between the to-be-tested sample and other controls, the hidden subgroup is formed between the to-be-tested sample and the to-be-tested sample, the hidden subgroup is formed between the negative control and the positive control, the hidden subgroup is formed between the bioinformatics control and the negative control, the hidden subgroup is formed between the other controls and the positive control, the hidden subgroup is formed between the other controls and the negative control, or the hidden subgroup is formed between the bioinformatics control and the other controls.

[0043] The application has the following beneficial technical effects:

[0044] 1. The present application is a revolution in sequencing data bioinformatics algorithm, which creatively proposes the hypothesis that "hiding subgroups as research units can effectively reduce random disturbance", and conducts efficient analysis on genomic sequencing sample composition data.

[0045] 2. The present application first compares the proportion between each two compositions to determine whether various classification units are derived from the same subgroup, thereby solving various bioinformatics analysis practical problems. For example, by comparing the proportion between each two compositions, whether the classification units are derived from the same subgroup, i.e. noise, is determined by comparing whether the proportion between each two compositions is consistent in sample composition and negative control composition, multiple alignment composition. The problem that the existing noise reduction method cannot fully utilize the current experimental true information and needs a large amount of prior knowledge is solved, and the sensitivity and specificity of the true signal are effectively improved; for example, by comparing the proportion between each two compositions, whether the classification units are derived from the same subgroup is determined by comparing whether the proportion between each two compositions is consistent in positive composition and negative control composition, and the composition of the sample to be tested and the negative control composition, thereby realizing absolute quantification.

[0046] 3. The calculation framework of the present application is independent of the selection of specific sequencing platforms, and can be applied to sequencing data of various platforms such as second-generation sequencing technology and third-generation sequencing technology, and can be applied to detection samples of different sources or different species. BRIEF DESCRIPTION OF DRAWINGS

[0047] Figure 1 : Macro-genome noise signal introduction schematic diagram.

[0048] Figure 2 : The proportion relationship between each two classification units in the subgroup remains unchanged in the sample and the negative control.

[0049] Figure 3 : The proportion relationship between the true signal and the noise signal in the sample to be tested and the negative control, and the proportion relationship between the noise signal and the noise signal in the sample to be tested and the negative control.

[0050] Figure 4 : Hugo-DNA process schematic diagram;

[0051] Figure 5 : Schematic diagram of noise reduction based on bioinformatics multiple alignment noise alone.

[0052] Figure 6 : Undirected graph constructed by classification units of the sample to be tested.

[0053] Figure 7 : Comparison of the noise reduction effect of Hugo-DNA with other methods.

[0054] Figure 8: Schematic diagram of absolute quantification method based on hidden subgroup. A demonstrates the identification of hidden subgroup and its procedure as a hypothetical internal control. B demonstrates the theoretical experiment and calculation procedure of absolute quantification using the absolute quantification method of the present application and using internal control method, and proves the advantage of the absolute quantification method of the present application compared with the internal control method.

[0055] Figure 9 : Quantitative performance of the present application, demonstrating the results of three data. DETAILED DESCRIPTION

[0056] The embodiments of the present application will be described in detail below with examples, but those skilled in the art will understand that the following examples are only for illustration of the present application and should not be regarded as limiting the scope of the present application. The specific conditions are not specified in the examples, and the conventional conditions or the conditions recommended by the manufacturer are used. The reagents or instruments used are not specified by the manufacturer, and are all conventional products that can be purchased on the market.

[0057] Definitions of some terms

[0058] Unless otherwise defined in the following, all technical and scientific terms used in the present application are intended to have the same meaning as commonly understood by one of ordinary skill in the art. Although the following terms are believed to be well understood by one of ordinary skill in the art, the following definitions are set forth to better clarify the application.

[0059] As used in the present application, the terms "comprise", "contain", "have", "include" or "involve" are inclusive or open-ended and do not exclude additional, unrecited elements or method steps. The term "consist of" is considered to be a preferred embodiment of the term "comprise". If in the following a group is defined to comprise at least a certain number of embodiments, this is also to be understood as disclosing a group which preferably consists only of these embodiments.

[0060] The indefinite article "a" or "an" or the definite article "the" used in connection with a singular noun, e.g. "a" or "an", "the", includes a plurality of such nouns.

[0061] The term "about" in the present application means an interval of accuracy which the person skilled in the art is able to understand as still ensuring the technical effect of the feature in question. The term usually means ±10%, preferably ±5% deviation from the indicated numerical value.

[0062] Furthermore, the terms first, second, third, (a), (b), (c), and the like, as used in the specification, are used as identifiers to distinguish between like elements, and are not necessarily descriptive of an order or temporal sequence. It is to be understood that the terms so used are interchangeable under appropriate circumstances and that the embodiments of the application described herein are capable of practical implementation in other sequences than the examples given herein.

[0063] The following terms or definitions are provided only to aid in the understanding of the present application. These definitions should not be construed to limit the scope or

[0064] The term "read" or "reads" in the present application refers to one or a set of nucleic acid sequences read by a sequencing platform.

[0065] The term "alignment" in the present application refers to the correspondence between a sequencing read and a reference sequence. A sequencing read can have multiple alignments.

[0066] The term "taxon" in the present application refers to a classification concept at the same hierarchical level in biological terms. For example, when studying the evolutionary relationship of organisms, a taxon is a biological population concept at each of the levels of kingdom, phylum, class, order, family, genus, species, and strain. For example, the kingdom Protozoa, the order Primates, the species Staphylococcus aureus, and the subspecies Salmonella enterica subsp. enterica are each a taxon. In the study of systems biology, a taxon can be a unit of study of different genes, different genomic elements, different genomic positions, or different sequence characteristics.

[0067] The term "species" in the present application refers to a special taxon, which is a group of organisms that can interbreed and produce offspring.

[0068] The term "(a set of taxa) mutually exclusive" in the present application refers to any two taxa A and B selected from the set of taxa, satisfying the condition that taxon A neither contains taxon B nor is contained by taxon B. For example, the three taxa "Escherichia coli, Salmonella enterica, and Klebsiella pneumoniae" are mutually exclusive; while the two taxa "Klebsiella and Klebsiella pneumoniae" are not mutually exclusive.

[0069] The term "a sequencing read aligns to a taxon", "a sequencing read aligns onto a taxon", or "a sequencing read supports a taxon" in the present application refers to the alignment result of the sequencing read containing a reference sequence from the taxon.

[0070] The term "single alignment" in the present application refers to the alignment result of a sequencing read containing reference sequences from the same taxon.

[0071] The term "multiple alignment" in the present application refers to the alignment result of a sequencing read sequence which contains reference sequences from two or more mutually exclusive classification units.

[0072] The sequencing data bioinformatics analysis method described in the present application generally comprises the following steps:

[0073] Step 1) grouping sequencing read sequences into different classification unit steps;

[0074] Step 2) classification unit statistics step;

[0075] Step 3) classification unit interrelation calculation step;

[0076] Step 4) hidden subgroup identification step.

[0077] The term "hidden subgroup", "subgroup" or "sub-group" in the present application refers to a form of co-occurrence of signals that are assumed to exist, which can be composed of two or more elements. The hidden subgroup requires mutual correlation or connection between its elements. The hidden subgroup can be introduced by signals from the same source in the experimental process or in the bioinformatics analysis process, and the ratio between the elements of the hidden subgroup remains stable under two or more conditions. In some embodiments, the hidden subgroup is formed between the test sample and the positive control, the hidden subgroup is formed between the test sample and the negative control, the hidden subgroup is formed between the test sample and the bioinformatics control, the hidden subgroup is formed between the test sample and other controls, the hidden subgroup is formed between the test sample and the test sample, the hidden subgroup is formed between the negative control and the positive control, the hidden subgroup is formed between the bioinformatics control and the negative control, the hidden subgroup is formed between the other control and the positive control, the hidden subgroup is formed between the other control and the negative control, or the hidden subgroup is formed between the bioinformatics control and the other control.

[0078] For example, in the process of bioinformatics noise reduction, the hidden subgroup is the same source of noise signal introduced in the experimental process or analysis process. These noise signals of the same source undergo various treatments under different conditions, but the internal proportional relationship remains stable, so the proportion between the elements of the hidden subgroup remains stable under two or more conditions. In the process of absolute quantification, high-abundance background signals often appear in negative controls, positive controls and test samples at the same time. The proportional relationship between these high-abundance background signals may remain stable in negative controls, positive controls and test samples, forming a hidden subgroup. These hidden subgroups can be used as a pseudo-internal control (PIC). The PIC can be used to calculate the proportion between negative controls and positive controls, and also to calculate the proportion between negative controls and test samples. Through the connection relationship, the proportion between positive controls and test samples can be calculated, and the absolute quantification of the test sample can be calculated according to the absolute quantification (known) of the positive control.

[0079] In some embodiments, the grouping method of step 1) can be any logical classification method, including but not limited to alignment-based grouping method or non-alignment-based grouping method.

[0080] In some embodiments, the alignment-based grouping method uses alignment software to perform sequence alignment on sequencing read sequences, and all sequencing read sequences aligned to the same classification unit are grouped into a group.

[0081] In some embodiments, the grouping method without alignment includes but is not limited to kmer method, hash table method or string matching method.

[0082] In some embodiments, step 2) is based on the grouping results of the classification units in step 1), and the statistical quantity of each classification unit is calculated.

[0083] In some embodiments, the statistical quantity includes but is not limited to: the number of support sequencing read sequences of each classification unit, the relative proportion of each classification unit, and the quantity of each classification unit after certain normalization.

[0084] In some embodiments, step 3) is to pair each two classification units according to the statistical quantity of the classification units in step 2), and to calculate whether the proportion of the two classification units in the pair remains stable under two or more conditions; in some embodiments, the stable proportion indicates that the paired classification units are related to each other.

[0085] In some embodiments, the conditions in step 3) include, but are not limited to, conditions of the sample to be tested, conditions of the negative control, conditions of the positive control, conditions of the bioinformatics control, or conditions of other controls.

[0086] In some embodiments, in step 3), the pairing can be random pairing; the ratio of the two classification units can be obtained by dividing the classification unit statistics, or by statistical testing of the classification unit statistics, without limitation.

[0087] In some embodiments, the identification in step 4) is the identification and / or analysis of the hidden subgroup for the correlation between the classification units in step 2) and the classification units in step 3); the hidden subgroup is a form of common occurrence of signals under two or more conditions, composed of two or more elements; the ratio between each pair of elements in the hidden subgroup remains stable under two or more conditions.

[0088] Preferably, the hidden subgroup is introduced by signals from the same source in the experimental process or bioinformatics analysis process, and the ratio between each pair of elements in the hidden subgroup remains stable under two or more conditions.

[0089] More preferably, the hidden subgroup is formed between the sample to be tested and the positive control, the hidden subgroup is formed between the sample to be tested and the negative control, the hidden subgroup is formed between the sample to be tested and the bioinformatics control, the hidden subgroup is formed between the sample to be tested and other controls, the hidden subgroup is formed between the sample to be tested and the sample to be tested, the hidden subgroup is formed between the negative control and the positive control, the hidden subgroup is formed between the bioinformatics control and the negative control, the hidden subgroup is formed between other controls and the positive control, the hidden subgroup is formed between other controls and the negative control, or the hidden subgroup is formed between the bioinformatics control and other controls.

[0090] In some embodiments, the identification of the hidden subgroup in step 4) is performed by a method without prior information.

[0091] Preferably, the identification and / or analysis is performed by an analysis method for analyzing the correlation between two or more elements or the characteristics of the elements themselves.

[0092] The hidden subgroup in the present application is different from a single element, which only has the data characteristics of the element itself as the research object of the subsequent analysis method. The hidden subgroup includes not only the elements in the subgroup but also the associations of the elements in the subgroup. Therefore, the analysis method for the hidden subgroup can not only take the associations between the elements in the subgroup as the research object, but also take the sum (or other combination methods such as product, statistical quantity, etc.) of the same data characteristics of all elements in the subgroup as the research object. Such analysis methods are collectively referred to as "analysis methods for analyzing the associations between two or more elements or the characteristics of the elements themselves".

[0093] In some embodiments, the identification and / or analysis is performed using the graph method of computer science, i.e., each classification unit in step 2) is taken as a node of a graph, and each pair with a connection in step 3) is taken as an edge of the graph, to form a complete undirected graph; in the undirected graph, a complete subgraph is found using the classical graph analysis method, and the complete subgraph is taken as the hidden subgroup; preferably, the number of nodes of the complete subgraph is 2 or more than 2.

[0094] The terms "node" and "edge" in the present application refer to the constituent elements of a "graph" in computer data structure. A node is an element in a set, and an edge is a connection between two nodes. In the present application, a node is a classification unit, and an edge is a connection between classification units.

[0095] The term "subgraph" in the present application refers to a subset of elements in an undirected graph.

[0096] The term "complete subgraph" in the present application refers to a subset of elements in an undirected graph, which satisfies that there is an edge between any two nodes in the subset.

[0097] The term "outlier" in the present application refers to an element in an undirected graph, which has no connection with other elements.

[0098] The present application also provides a bioinformatics analysis system / device for sequencing data, characterized in that it comprises the following modules:

[0099] Module 1) module for grouping sequencing read sequences into different classification units;

[0100] Module 2) classification unit statistics module;

[0101] Module 3) classification unit interrelation calculation module;

[0102] Module 4) hidden subgroup identification module.

[0103] These modules respectively perform the corresponding steps described above.

[0104] The application further provides a computer readable storage medium, which stores a computer program, and the computer program is characterized in that the computer program is executed by a processor to realize the steps of the above method.

[0105] The application further provides a computer program product, which comprises a computer program, and the computer program is characterized in that the computer program is executed by a processor to realize the steps of the above method.

[0106] Based on the basic idea of the application, it can be understood that the above-mentioned method or system of the application is not limited to the sequencing platform, the source of the sequencing sample, and the like. Therefore, in any of the above-mentioned methods, the sequencing can be second-generation sequencing, third-generation sequencing or fourth-generation sequencing.

[0107] Preferably, the sequencing is from Illumina, BGI, ION TORRENT, PacBio, Roche, Helicos, ABI or nanopore sequencing.

[0108] More preferably, the sequencing is from Illumina, BGI or nanopore sequencing.

[0109] Similarly, the method of the application is also not limited to the source and type of sequencing data, and any sequencing data can meet the idea or rule of the application. Therefore, the sequencing data can be any gene sequencing data.

[0110] Preferably, the sequencing data is genomic sequencing data.

[0111] More preferably, the sequencing data is metagenomic sequencing data.

[0112] Further preferably, the sequencing data is low-biomass metagenomic sequencing data.

[0113] The application further provides an application of a hidden subgroup in bioinformatics analysis of sequencing data, and the hidden subgroup is a signal introduced from the same source in an experimental process or a bioinformatics analysis process, and the proportion between elements in the hidden subgroup remains stable under two or more conditions.

[0114] In some embodiments, the hidden subgroup is formed between the test sample and the positive control, the hidden subgroup is formed between the test sample and the negative control, the hidden subgroup is formed between the test sample and the bioinformatics control, the hidden subgroup is formed between the test sample and the other control, the hidden subgroup is formed between the test sample and the test sample, the hidden subgroup is formed between the negative control and the positive control, the hidden subgroup is formed between the bioinformatics control and the negative control, the hidden subgroup is formed between the other control and the positive control, the hidden subgroup is formed between the other control and the negative control, or the hidden subgroup is formed between the bioinformatics control and the other control.

[0115] Further, when the purpose is to remove the background noise of the sequencing data, the hidden subgroup can be identified between the test sample and the negative control sample, and / or between the test sample and the bioinformatics control sample in the above method.

[0116] Further, when the purpose is to remove the background noise of the sequencing data, the hidden subgroup can be identified between the test sample and the negative control sample, and / or between the test sample and the bioinformatics control sample in the above method.

[0117] In order to further understand the present application, one specific embodiment of the present application is exemplarily provided as follows: a sequencing data bioinformatics analysis method can specifically include the following steps: step 1) grouping sequencing reads into different classification units; step 2) classification unit statistics step; step 3) classification unit correlation calculation step; and step 4) hidden subgroup identification step.

[0118] In some embodiments, the step 1) can include, but is not limited to, the following two grouping methods: alignment-based grouping method and non-alignment-based grouping method. First, the alignment-based grouping method uses alignment software to align the sequencing reads, and groups according to the alignment results, for example, all sequencing reads aligned to the same classification unit are grouped together. Second, the non-alignment-based grouping method, for example, the kmer method, first divides the sequencing reads into kmers, and divides the alignment database into kmers, and then counts which reference sequence in the database is more consistent with a certain sequencing read, so as to group the sequencing read into the classification unit of the reference sequence.

[0119] In some embodiments, the step 2) calculates the statistics of each classification unit based on the grouping results of the classification unit in step 1). For example, the number of support sequencing reads of each classification unit, the relative proportion of each classification unit, the statistics of each classification unit after certain normalization (such as reads per million (RPM), etc.).

[0120] In some embodiments, the step 3) pairs each two classification units for the classification unit statistics in step 2), and calculates whether the proportion of the two classification units in the pair is stable under two or more conditions. If it is stable, it is considered that the classification units in the pair are associated with each other. In some specific embodiments, the proportion of the two classification units in each pair is obtained by dividing the classification unit statistics, and whether it is stable is obtained by calculating whether the proportion of the proportion of the pair under two (or more) conditions is close to 1. For example, the statistics of two classification units A and B under condition one are 10 and 30, respectively, and the proportion is 0.333, and the statistics under condition two are 1 and 3, respectively, and the proportion is 0.333, so the proportion of the pair A and B under condition one (0.333) and condition two (0.333) is (0.333:0.333)1, which is stable. In some specific embodiments, the proportion of the two classification units in each pair is obtained by statistical test on the unit statistics, such as the p value of goodness-of-fit test. For example, the statistics of two classification units A and B under condition one are 10 and 30, respectively, and the proportion is 0.333, and the statistics under condition two are 1 and 3, respectively, and the goodness-of-fit test is performed on the four numbers 10, 30, 1, and 3, and the p value is calculated. If the p value is greater than or equal to a certain threshold (such as 0.05), it is considered to be stable.

[0121] In some specific embodiments, in order to determine whether the proportion of a pair under two (or more) conditions is close to 1, a threshold value needs to be set, such as greater than or equal to 0.5 and less than or equal to 2. The setting of the threshold value is related to the specific application scenario, and cannot be unified in different scenarios, and the specific threshold range is not protected and limited. In some specific embodiments, whether the proportion of a pair under two (or more) conditions is stable is determined by statistical test, and the threshold value is often 0.05, 0.01, etc. The setting of the threshold value is related to the specific application scenario, and cannot be unified in different scenarios, and the specific threshold range is not protected and limited.

[0122] In some embodiments, the step 4) identifies and analyzes the hidden subgroups based on the association between the classification units in step 2) and the classification units in step 3).

[0123] In some embodiments, the hidden sub-group analysis is performed in a non-priori way, which is to analyze the mathematical properties of the hidden sub-group. We can use any analysis method that can be used to analyze the correlation between two or more elements and the characteristics of the elements themselves to deal with the connection relationship.

[0124] For example, the graph method of computer theory is used to efficiently deal with the connection relationship, such as the identification of core nodes, the calculation of various indicators of node correlation (centrality, betweeness, etc.), the calculation of connection paths, the calculation of the maximum and minimum range of connection, the calculation of the similarity of connection, etc. In the noise reduction application, we mainly use the graph method to efficiently find the completely connected subgraph, which is only one method of hidden sub-group analysis; in the quantitative application, we use the strength of the elements associated with the hidden sub-group to calculate the proportional relationship between two samples. These applications are all methods of hidden sub-group analysis. Therefore, in general, it can be understood that the graph method of computer science is only an optional method, and any analysis method that can be used to analyze the correlation between two or more elements and the characteristics of the elements themselves is applicable.

[0125] In some embodiments, the graph method of computer science is: each classification unit in step 2) is taken as a node of the graph, and each pair with a connection in step 3) is taken as an edge of the graph, to form a complete undirected graph. In the undirected graph, using the classical graph analysis method, find the complete subgraph in it, which is also called "clique". These "cliques" can be used as hidden sub-groups. In some embodiments, the hidden sub-group analysis is performed in a priori way, such as using known biological pathways to find out whether the genes in a certain pathway are enriched in the classification units with mutual correlation found in step 3).

[0126] When the method of the present application is used for noise reduction purposes, it is known in the art that in bioinformatics analysis, the composition of the sample to be observed is usually the superposition of true signals and noise signals. In order to remove noise signals, the present application needs to compare the composition of the sample to be observed with the background composition containing only noise signals, find the factors that remain constant, and construct a model according to these constant factors, so as to effectively remove noise.

[0127] First of all, an empirical observation is that the composition of noise signals in any kind of gene sequencing results (such as metagenome) is not single, but contains extremely diverse species combinations, which indicates that the introduction of noise is in the form of a sub-group, rather than a single noise signal introduced from a single source of contamination.

[0128] Secondly, by investigating each step of the sequencing method, the present application finds that the source of the introduction of noise signals can be mainly divided into two categories. The first category is the pollution introduced in the experimental process. These noise signals are all real nucleic acid molecules. The second category is the error signal introduced in the bioinformatics analysis process, such as multiple alignment from real nucleic acid molecules, which is not a real nucleic acid molecule. According to the definition of the negative control, the negative control itself does not include any real signal. The composition of the observed negative control is theoretically only the pollution introduced in the experimental process. Therefore, the negative control can be considered as the basic reference of the experimental noise signal in the experimental process.

[0129] In the bioinformatics analysis process, there are usually two categories of methods that can obtain the composition of the classification units of the sample to be tested: one is based on sequence alignment method (such as BLASTN), and the other is not based on sequence alignment method (such as kmer method). Regardless of which method, due to the similarity between classification units and sequencing errors, sequencing read sequences from real components are incorrectly grouped into other classification units, forming bioinformatics noise. Specifically, in the case of classification unit grouping based on sequence alignment, bioinformatics noise is often aligned to some conservative sequencing reads at the same time as real components, forming multiple alignment. At this time, multiple alignment can indicate the source of bioinformatics noise, so the alignment result is a good basic reference for bioinformatics noise signal, which can be used as a bioinformatics control. In the case of not grouping classification units based on sequence alignment, the reference sequence of the real component can be used to simulate sequencing read data according to the error rate distribution rule of the sequencer, and the classification result of the simulated data can be used to construct a basic reference for bioinformatics noise signal as a bioinformatics control. Therefore, the composition of the negative control and the composition of the bioinformatics control are important references for noise signals.

[0130] Thirdly, the removal of noise signals needs to consider the specific way of introducing noise signals. The present application finds that in the experimental process, the experimental pollution noise signals introduced by the same operation (such as components D, E, F, all 1 units) will be simultaneously deformed in subsequent operations (such as simultaneous amplification by 2 times, becoming D, E, F, all 2 units), but the proportion of noise signals in the same group will remain relatively constant (such as D:E:F=1:1:1 before and after deformation). Although there may be no stable relationship between the pollution noise signals introduced in different steps, the mutual proportion relationship of the same group of pollution signals introduced at the same time and in the same step remains constant. Therefore, for experimental noise, by calculating whether the proportion relationship between classification units remains consistent in the sample to be tested and the negative control, it can be judged whether it is experimental noise.

[0131] The present application finds that, during the experiment, if the classification unit A' is close to the sequence A, and A' is not a bioinformatics false detection of A, then the number of A' should not be lower than the number of A' obtained by simulating the error comparison with the current number A (or the number of A' obtained by multiple comparison with the current number A). As described above, the proportional relationship between the bioinformatics noise and the true component is similar in the test sample and the bioinformatics control. Therefore, for the bioinformatics noise, by calculating whether the proportional relationship between the classification units remains consistent in the test sample and the bioinformatics control, it can be judged whether it is a bioinformatics noise.

[0132] In summary, whether it is pollution introduced during the experiment or false detection introduced during bioinformatics analysis, it can be judged whether it is a noise signal by judging whether the proportion of the composition within the subgroup remains unchanged (comparison of the composition of the test sample with the composition of the negative control, and comparison of the composition of the test sample with the bioinformatics control result) such hidden subgroup.

[0133] In the present application, the present application independently develops a hidden subgroup-based de-noising algorithm (Hugo-DNA) to effectively reduce noise signals (pollution introduced during the experiment and false detection introduced during bioinformatics analysis). Hugo-DNA is based on the following assumptions: a pair of noise signals of the same source will maintain the same ratio during subsequent processing and transformation. Specifically, for pollution introduced during the experiment, the pollution signals introduced by the completed experimental steps will maintain the same ratio between each other during the subsequent experimental process; for false detection introduced during bioinformatics analysis, the proportion of a true signal incorrectly identified as a noise signal in the current data will not change. By comparing the composition data of the test sample with the composition data of the negative control, the pollution noise signal introduced during the experiment can be eliminated. By comparing the composition data of the test sample with the composition data of the bioinformatics control, the false detection noise signal introduced during bioinformatics analysis can be eliminated. Hugo-DNA does not require prior knowledge and does not need to set a threshold value, and can efficiently remove pollution.

[0134] Therefore, according to some aspects of the present application, the present application relates to a bioinformatics analysis method for distinguishing true signals from background signals, comprising the following steps:

[0135] Step 1) sequencing step of test sample and negative control sample;

[0136] Step 2) grouping step of test sample and negative control sample sequencing data according to classification units;

[0137] Step 3) classification unit statistics step of test sample and negative control sample;

[0138] Step 4) comparison of test sample and negative control results, calculation of classification unit relationship step;

[0139] Step 5) The results of the test sample are compared with the negative control to identify the hidden subgroup.

[0140] Some terms in this application are explained as follows when it comes to the application of noise reduction:

[0141] The term "negative control" in this application refers to a blank control set up during the sequencing experiment, which undergoes the same operation process as the test sample. This negative control has the meaning of traditional negative control. For example, during the sequencing experiment, it can be a nucleic acid-free water sample. In theory, the detection signal of this sample should be empty, i.e. nothing should be detected. However, due to experimental contamination and false positives introduced by bioinformatics analysis, this sample will become the best representative of noise signals in the same batch experiment. For example, during the sequencing experiment, the negative control is a blank system that does not contain the detection target compared to the detection target. In theory, as long as a detection target is considered to be contained in sample 1 but not in sample 2, sample 2 can be used as the negative control of sample 1, so as to find the element contained only in sample 1.

[0142] The term "false positive" in this application refers to the alignment result of the sequencing readout sequence actually from a certain taxon containing the reference sequence from another taxon that neither contains nor is contained by the former taxon. It should be noted that there is no error in the alignment result in calculation, and the base identity rate can be very high, but the classification unit in the alignment does not match the source classification unit of the sample. This false alignment generally occurs between classification units with close genetic relationship.

[0143] The term "experimental contamination" in this application refers to non-sample-derived nucleic acid molecules introduced during the processing of the sample. These signals are not nucleic acid molecules from the sample itself, but contaminated nucleic acid molecules introduced during the experiment. Possible sources include sampling, instruments, reagents, etc.

[0144] The term "detection signal" in this application refers to all taxons reported after analyzing the sequencing data of a sample. It includes true signals and noise signals.

[0145] The term "true signal" in this application refers to the nucleic acid molecules actually contained in the sample. In the experimental design, this application will incorporate specific known high-purity nucleic acid molecules into the nucleic acid-free water. These nucleic acid molecules are true signals. When the test sample is an unknown sample, the nucleic acid molecules without any treatment are true signals.

[0146] The term "noise signal" in the present application refers to non-real signals introduced during the processing of samples after processing. These signals are not nucleic acid molecules of the sample itself, but are contaminated nucleic acid molecules introduced during the experiment or false positives introduced during bioinformatics analysis.

[0147] The term "correlation" or "connection" in the present application refers to the proportional relationship between two paired classification units, which can be strictly quantified, and the proportional relationship remains stable under two or more conditions.

[0148] The term "co-detection" in the present application refers to the simultaneous detection of a classification unit under different conditions, such as the detection of a sample and a negative control, or the simultaneous detection of a sample and a positive control, or the simultaneous detection of a sample and a bioinformatics control, etc.

[0149] The term "single detection" in the present application refers to the detection of a classification unit under a single condition, such as the detection of a sample or a negative control. In the expression, it is often indicated that the sample is detected alone or the negative control is detected alone.

[0150] The term "coexistence" in the present application refers to a classification unit that is both a real signal and a noise signal. This often occurs when the real signal in the sample contains nucleic acid molecules of environmental microorganisms, which are a common source of contamination.

[0151] The term "single existence" in the present application refers to a classification unit that is only one of a real signal or a noise signal.

[0152] The "classification unit component calculation method of sequencing data" in the present application generally refers to a method for calculating whether various classification units appear in sequencing data and the proportion of each classification unit by bioinformatics analysis of sequencing data. This method can effectively remove false positive results in classification unit component calculation. It can be understood that the present application solves the noise removal problem that the original conventional classification unit component calculation method cannot solve by introducing the calculation of "hidden subgroups", effectively improving the specificity and accuracy of pathogen detection; at the same time, the calculation framework of the present application is independent of the selection of specific sequencing platforms and is not limited to sequencing platforms, and can be applied to sequencing data of various platforms such as second-generation sequencing technology and third-generation sequencing technology; moreover, the calculation framework is also independent of the specific sequencing data source. It can be understood in the art that the calculation framework of the present application is for any homologous sequence misalignment, therefore, the sequence source does not limit the application of the present application, in addition to the preferred macrogenomic data source, other genomic or genetic data sources are also suitable for the present application.

[0153] In some embodiments, the grouping in step 2) is based on alignment method, respectively grouping the test sample and the negative control sample; preferably, the alignment method is grouping after sequence alignment of the sequencing read sequences using alignment software that retains non-single alignment results (such as BLASTN software).

[0154] In some embodiments, the grouping in step 2) is based on non-alignment method, respectively grouping the test sample and the negative control sample; preferably, the non-alignment method includes but is not limited to kmer method, hash table method or string matching method.

[0155] The classification unit statistics described in the present application include but are not limited to the following statistical quantities: the number of supporting sequencing read sequences of each classification unit, the relative proportion of each classification unit, or the statistical quantity of each classification unit after certain normalization.

[0156] In some embodiments, the calculation of the correlation between classification units in step 4) is to pair each two classification units in the test sample according to the statistical quantity of the classification units in step 3), and calculate whether the proportion of the two classification units in the pair is stable in the test sample and the negative control; if it is stable, it is considered that the classification units in the pair are related to each other.

[0157] In some embodiments, the identification of the hidden subgroup in step 5) is based on the classification units in step 2) and the correlation between classification units in step 4), processing and screening the correlation between classification units, and identifying the remaining correlation between classification units as the hidden subgroup; preferably, the identification and / or analysis is a hidden subgroup analysis without prior information; the identification and / or analysis is performed using an analysis method for analyzing the correlation between two or more elements or the characteristics of an element itself.

[0158] The hidden subgroup in the present application is different from a single element whose data characteristics can be used as the research object of subsequent analysis methods. The hidden subgroup includes not only the elements in the subgroup, but also the correlation between the elements in the subgroup. Therefore, the analysis method for the hidden subgroup can not only use the correlation between the elements in the subgroup as the research object, but also use the sum (or other combination methods such as product, statistical quantity, etc.) of the same data characteristics of all elements in the subgroup as the research object. Such analysis methods are collectively referred to as "analysis methods for analyzing the correlation between two or more elements or the characteristics of an element itself".

[0159] More preferably, the identification and / or analysis is performed using a graph method of computer science, i.e., each classification unit in step 2) is taken as a node of a graph, and each pair with a connection in step 4) is taken as an edge of the graph, to form a complete undirected graph; in the undirected graph, a complete subgraph is found using a classical graph analysis method, and the complete subgraph is taken as the hidden subgroup.

[0160] Further preferably, the number of nodes of the complete subgraph is 2 or more; for example, 2, 3, 4, 5, or 6.

[0161] It can be understood that, based on the negative control, the present application can further be subjected to noise reduction by a bioinformatics control, so as to achieve a more excellent noise reduction effect. Therefore, in some embodiments of the present application, the step 5) can further include the following steps:

[0162] Step 6) constructing a bioinformatics control and counting the classification units thereof;

[0163] Step 7) comparing the results of the test sample and the bioinformatics control, and calculating the relationship between the classification units;

[0164] Step 8) comparing the results of the test sample and the bioinformatics control, and identifying the hidden subgroup.

[0165] The term "bioinformatics control" can include multiple parts: in the bioinformatics analysis process, there are two major methods to obtain the composition of the classification units of the test sample: one is based on sequence alignment (such as BLASTN), and the other is not based on sequence alignment (such as kmer method). Regardless of which method, due to the similarity between the classification units and the errors introduced by sequencing, the sequencing read sequences from the real components will be incorrectly grouped into other classification units, forming bioinformatics noise. Specifically, in the case of classification unit grouping based on sequence alignment, bioinformatics noise often aligns to some conservative sequencing reads at the same time as the real components, forming multiple alignments. At this time, multiple alignments can indicate the source of bioinformatics noise, so the alignment result is a good reference for bioinformatics noise signal, and can be used as a bioinformatics control; and in the case of classification unit grouping without sequence alignment, the reference sequence of the real component can be used to simulate sequencing read data according to the error rate distribution of the sequencer, and the classification result of the simulated data can be used to construct a reference for the bioinformatics noise signal, as a bioinformatics control. Further, the bioinformatics control of step 6) is based on the alignment result of step 2), and the alignment result thereof is used as a bioinformatics control according to the alignment of the sequencing read sequences.

[0166] In some embodiments, the bioinformatics control of step 6) is that based on the non-aligned results of step 2), the reference genome of each classification unit is simulated with data, the sequencing read sequences of the classification unit are simulated according to the error distribution law of the sequencer, and the grouping is performed with the simulated sequencing read sequences. The grouping results of the simulation data can be used as the bioinformatics control of the classification unit.

[0167] In some embodiments, step 7) is a statistical quantity of the classification unit of step 6), each two classification units in the test sample are paired, and it is calculated whether the proportion of the two classification units in the pair is higher or flat in the test sample than in the bioinformatics control; preferably, if higher or flat, it is considered that the classification units in the pair have a relationship with each other; more preferably, the proportion of the two classification units is obtained by dividing the classification unit statistics; or is obtained by statistical testing of the unit statistics.

[0168] In some embodiments, step 7) is that after removing the classification units forming the hidden subgroup in step 5), each two classification units in the test sample are paired, and it is calculated whether the proportion of the two classification units in the pair is higher or flat in the test sample than in the bioinformatics control; preferably, if higher or flat, it is considered that the classification units in the pair have a relationship with each other; more preferably, the proportion of the two classification units is obtained by dividing the classification unit statistics; or is obtained by statistical testing of the unit statistics.

[0169] In some embodiments, step 8) is a classification unit relationship between the classification units of the test sample in step 2) and the classification units of step 7), and the classification unit relationship is processed and screened, and the relationship between the remaining classification units is identified as a hidden subgroup.

[0170] In some embodiments, step 8) is that after removing the classification units forming the hidden subgroup in step 5), the relationship between the classification units of the test sample in step 2) and the classification units of step 7) is processed and screened, and the relationship between the remaining classification units is identified as a hidden subgroup.

[0171] Further, the identification and / or analysis is performed by a hidden subgroup analysis without prior information; the identification and / or analysis is performed by using an analysis method for analyzing the correlation between two or more elements or the characteristics of an element itself; preferably, the identification and / or analysis is performed by using a graph method in computer science, i.e., each classification unit of the sample to be tested in step 2) is taken as a node of the graph, and each pair with a connection in step 7) is taken as an edge of the graph, to form a complete undirected graph; in the undirected graph, a complete subgraph is found by using a classical graph analysis method, and the complete subgraph can be used as a hidden subgroup; more preferably, the number of nodes of the complete subgraph should be 2 or more; such as 2, 3, 4, 5, 6.

[0172] In some embodiments, the above method further comprises the following steps:

[0173] Step 9) background noise elimination step for the sample to be tested; preferably, the elimination in step 9) is for the classification units of the sample to be tested in step 2), eliminating all classification units (contamination introduced in the experiment) that form hidden subgroups in step 5); and / or, for the classification units in step 6), eliminating all classification units (i.e., false positives introduced by bioinformatics analysis) that form hidden subgroups and have low statistics in the two classification units in the pair in step 8).

[0174] In some embodiments, the above method further comprises the following steps:

[0175] Step 10) classification unit abundance calibration step; preferably, the abundance calibration step is for the comparison results after eliminating the noise classification units, re-grouping the sequencing read sequences according to the classification units supported by the comparison results in priority, and counting the number of sequences in each group and the proportion of the total read sequences.

[0176] In addition, based on the core idea of the present application, when only bioinformatics noise is considered without considering experimental pollution, the noise reduction method is also one of the core methods of the present application, although the noise reduction effect of this single method is slightly inferior to the comprehensive noise reduction, but it cannot be denied that it can also achieve the effect claimed in the present application, and also belongs to the protection content of the present application, such as Figure 5 The schematic diagram of noise reduction based on bioinformatics control alone.

[0177] Therefore, according to some aspects of the present application, the bioinformatics analysis method for distinguishing between bioinformatics true signals and background signals in the present application can also include the following steps:

[0178] Step 1) sequencing step for the sample to be tested;

[0179] Step 2) Grouping sequencing data of the sample to be tested into different classification units;

[0180] Step 3) Counting the classification units of the sample to be tested;

[0181] Step 4) Constructing a bioinformatics control and counting the classification units thereof;

[0182] Step 5) Comparing the results of the sample to be tested and the bioinformatics control, and calculating the correlation between the classification units;

[0183] Step 6) Comparing the results of the sample to be tested and the bioinformatics control, and identifying hidden subgroups.

[0184] In some embodiments, the grouping in step 2) is based on alignment; preferably, the sequencing reads are grouped after sequence alignment using alignment software that retains non-single alignment results, such as BLASTN software.

[0185] In some embodiments, the grouping in step 2) is based on non-alignment; preferably, the grouping is performed using methods including but not limited to kmer method, hash table method or string matching method.

[0186] In some embodiments, the counting in step 3) includes but is not limited to counting the number of supporting sequencing reads of each classification unit, the relative proportion of each classification unit or the statistical quantity of each classification unit after normalization.

[0187] In some embodiments, the bioinformatics control of step 4) is based on the alignment results of step 2).

[0188] Preferably, the alignment results are used as the bioinformatics control according to the alignment of the sequencing reads.

[0189] Alternatively, the bioinformatics control of step 4) is based on the non-alignment results of step 2).

[0190] Preferably, the reference genome of each classification unit is simulated according to the error distribution of the sequencer, the sequencing reads of the classification unit are simulated, and the simulated sequencing reads are grouped. The grouping results of the simulated data can be used as the bioinformatics control of the classification unit.

[0191] In some embodiments, step 5) is to pair each two classification units in the sample to be tested, and calculate whether the proportion of the two classification units in the pair is higher or flat in the sample to be tested than in the bioinformatics control; if it is higher or flat, it is considered that the classification units in the pair are related to each other.

[0192] Preferably, the ratio of the two classification units is obtained by dividing the classification unit statistics; or is obtained by statistical testing of the unit statistics.

[0193] In some embodiments, the step 6) is to perform identification and / or analysis of classification unit correlation for the classification units of the sample to be tested in step 3) and the classification units in step 5), and identify the correlation between the remaining classification units as a hidden subgroup.

[0194] In some embodiments, the identification and / or analysis is performed by a hidden subgroup analysis without prior information;

[0195] Preferably, the identification is performed by using an analysis method for analyzing the correlation between two or more elements or the characteristics of the elements themselves;

[0196] More preferably, the identification and / or analysis is performed by using a graph method in computer science, i.e., each classification unit in step 2) is taken as a node of a graph, and each pair with correlation in step 5) is taken as an edge of the graph, to form a complete undirected graph; in the undirected graph, a complete subgraph is found by using a classical graph analysis method, and the complete subgraph can be taken as a hidden subgroup.

[0197] Further preferably, the number of nodes of the complete subgraph is 2 or more.

[0198] In some embodiments, it further comprises the following steps:

[0199] Step 7) removing background noise in the sample to be tested.

[0200] Preferably, the removal in step 7) is for the classification units in step 2), and removes all classification units in step 6) that form a hidden subgroup and have low statistics in the pair of two classification units.

[0201] For better explanation of the present application, according to some specific embodiments of the present application, the bioinformatics analysis method for distinguishing true signals from background signals can be a specific method including the following steps (this explanation does not limit the scope of the invention):

[0202] Step 1) sequencing step of the test sample and the negative control sample; Step 2) grouping step of the sequencing data of the test sample and the negative control sample into different classification units; Step 3) statistical step of the classification units of the test sample and the negative control sample; Step 4) step of constructing a bioinformatics control and counting the classification units thereof; Step 5) step of comparing the results of the test sample and the negative control and calculating the correlation between the classification units; Step 6) step of comparing the results of the test sample and the negative control and identifying hidden subgroups; and the method can further comprise: Step 7) step of comparing the results of the test sample and the bioinformatics control and calculating the correlation between the classification units; Step 8) step of comparing the results of the test sample and the bioinformatics control and identifying hidden subgroups; and Step 9) step of removing background noise in the test sample.

[0203] In some embodiments, the step 1) is parallel experimental operation of the test sample and the negative control sample, and a sequencing library is constructed. The parallel experimental operation requires that the experimental site, the operator, the reagent instrument, the operation process, the operation step, the operation method, the operator, the sequencing platform and other experimental conditions are as consistent as possible, so that the negative control can capture various background noise signals introduced in the experimental process systematically and comprehensively.

[0204] In some embodiments, the step 2) is grouping based on the sequencing data (generally base text files) of step 1) using an alignment-based method, using alignment software to perform sequence alignment on the sequencing readout sequence, and grouping according to the alignment results, for example, all sequencing readout sequences aligned to the same classification unit are grouped together. In some embodiments, the step 2) is grouping based on the sequencing data (generally base text files) of step 1) using a non-alignment-based method, the kmer method first divides the sequencing readout sequence into kmer, and divides the alignment database into kmer, and counts which reference sequence in the database is more consistent with a certain sequencing readout sequence, thereby grouping the sequencing readout sequence into the classification unit of the reference sequence. Preferably, for the alignment-based grouping method, the alignment software (such as BLASTN software) that retains non-single alignment results is used for sequence alignment and grouping of the sequencing readout sequence. Preferably, for the non-alignment-based grouping method, methods including but not limited to kmer method, hash table method, string matching method are used. This step processes the sample data and the negative control data in the same way.

[0205] In some embodiments, the step 3) calculates the statistics of each classification unit for the test sample and the negative control sample respectively based on the grouping results of the classification unit in step 2). For example, the number of sequencing read sequences supporting each classification unit, the relative proportion of each classification unit, and the statistics of each classification unit after certain normalization (such as reads per million (RPM), etc.) are calculated. In some embodiments, the step 3) divides the alignment results of the sequencing read sequences into two groups according to the alignment of the sequencing read sequences, i.e. single alignment (i.e. the classification unit to which each sequencing read sequence is aligned is the same) and multiple alignment (i.e. the classification unit to which each sequencing read sequence is aligned is different). When calculating the statistics of the classification unit, the error of the single alignment result is smaller, and only the results of the single alignment are used to calculate the number and percentage of the sequencing read sequences aligned to each classification unit. In some embodiments, the step 3) directly uses the grouping method (such as kmer method) that is not based on alignment to obtain the number and percentage of sequencing read sequences supporting each classification unit based on the non-alignment results of step 2). The same processing is performed on the test sample data and the negative control data to obtain the classification unit composition results of the test sample and the negative control.

[0206] In some embodiments, the step 4) is used to extract the biological control from the alignment results of the test sample or simulate the biological control from the classification unit results of the test sample for subsequent steps to eliminate the background noise signal introduced by the biological analysis. If the classification unit A' obtained in step 3) is a biological false detection of the classification unit A, then the number of A' must be lower than the number of A' in the multiple alignment results of the real classification unit A (or the number of A' in the classification unit of the simulation data of the real classification unit A). Based on this, the biological control is constructed for each classification unit obtained in step 3), and the pairing relationship in the biological control is limited to the classification unit with a high number as a potential real signal and the classification unit with a low number as a potential noise signal. Here, the number refers to the statistics in step 3). The high and low numbers are determined according to the meaning in the experimental design.

[0207] In some embodiments, the step 4) uses the alignment results as the biological control based on the alignment results of step 2) according to the alignment of the sequencing read sequences. In some embodiments, the step 4) simulates the reference genome of each classification unit based on the non-alignment results of step 2), simulates the sequencing read sequences of the classification unit according to the error distribution law of the current sequencer, and groups the simulated sequencing read sequences. The grouping results of the simulation data can be used as the biological control of the classification unit.

[0208] In some embodiments, the step 5) is to pair each two classification units in the test sample and calculate whether the ratio of the two classification units in the pair is stable in the test sample and the negative control. If it is stable, it is considered that the two classification units in the pair are associated with each other. In some specific embodiments, the ratio of the two classification units in each pair is obtained by dividing the classification unit statistics, and whether it is stable is obtained by calculating whether the ratio of the two (or more) conditions is close to 1. For example, the statistics of two classification units A and B in condition one are 10 and 30, respectively, and the ratio is 0.333. The statistics in condition two are 1 and 3, respectively, and the ratio is 0.333. The ratio of the pair A and B in condition one (0.333) and condition two (0.333) is (0.333:0.333) 1, which is stable. In some specific embodiments, the ratio of the two classification units in each pair is obtained by statistical test on the unit statistics, such as the p value of goodness-of-fit test. For example, the statistics of two classification units A and B in condition one are 10 and 30, respectively, and the ratio is 0.333. The statistics in condition two are 1 and 3, respectively, and the goodness-of-fit test is performed on the four numbers 10, 30, 1, and 3 (two-tailed), and the p value is calculated. If the p value is greater than or equal to a certain threshold (such as 0.05), it is considered to be stable.

[0209] In some specific embodiments, it is necessary to set a threshold to determine whether the ratio of a pair in the test sample and the negative control is close to 1, such as greater than or equal to 0.5 and less than or equal to 2. The setting of the threshold is related to the specific application scenario, and cannot be unified in different scenarios, and the specific threshold range is not protected. In some specific embodiments, the ratio of a pair in the test sample and the negative control is determined by statistical test, and the threshold is often 0.05, 0.01, etc. The setting of the threshold is related to the specific application scenario, and cannot be unified in different scenarios, and the specific threshold range is not protected.

[0210] In some embodiments, the step 6) is to identify and / or analyze the hidden subgroup based on the association between the classification units of the sample to be tested in step 2) and the classification units in step 5). In some embodiments, the hidden subgroup analysis is performed without prior information, for example, using a graph method in computer science, i.e., each classification unit in step 2) is taken as a node of a graph, and each pair of classification units with association in step 5) is taken as an edge of the graph, to form a complete undirected graph. In the undirected graph, a complete subgraph, also known as a "clique", is found using classical graph analysis methods. These "cliques" can be used as hidden subgroups. Preferably, the number of nodes of the complete subgraph should be 2 or more. More preferably, the number of nodes of the complete subgraph should be 2, 3, 4, 5, 6.

[0211] In some embodiments, the step 7) is to pair each pair of classification units in the sample to be tested and calculate whether the ratio of the two classification units in the pair (potential true component ratio potential noise component) is higher or flat in the sample to be tested than in the bioinformatics control based on the classification unit statistics in step 4). In some embodiments, the step 7) is to pair each pair of classification units in the sample to be tested and calculate whether the ratio of the two classification units in the pair (potential true component ratio potential noise component) is higher or flat in the sample to be tested than in the bioinformatics control after removing the classification units forming the hidden subgroup in step 6) based on the classification unit statistics in step 4). If it is higher or flat, it is considered that the classification units in the pair have association with each other. In some specific embodiments, the ratio of the two classification units in each pair is obtained by dividing the classification unit statistics, and whether it is higher or flat in the sample to be tested than in the bioinformatics control is compared. In some specific embodiments, the ratio of the two classification units in each pair is obtained by statistical test on the unit statistics, such as the p value of goodness-of-fit test. For example, the statistics of two classification units A and B in condition one are 10 and 30, respectively, and the ratio is 0.333. The statistics in condition two are 1 and 3, respectively, and the goodness-of-fit test is performed on the four numbers 10, 30, 1, and 3 (one-tailed test), and the p value is calculated. If the p value is greater than or equal to a certain threshold (such as 0.05), it is considered that the ratio in the sample to be tested is higher or flat than in the bioinformatics control.

[0212] In some embodiments, the step 8) performs the identification and / or analysis of the hidden subgroup based on the connection between the classification unit of the sample to be tested in step 2) and the classification unit in step 7). In some embodiments, the step 8) performs the identification and / or analysis of the hidden subgroup based on the connection between the classification unit of the sample to be tested in step 2) and the classification unit in step 7) after the classification unit forming the hidden subgroup in step 6) is removed. In some embodiments, the hidden subgroup analysis is performed in a non-priori manner, for example, using a graph method in computer science, i.e., taking each classification unit of the sample to be tested in step 2) as a node of a graph, and each pair of classification units having a connection in step 7) as an edge of the graph, to form a complete undirected graph. In the undirected graph, a complete subgraph, also called a "clique", is found using a classical graph analysis method. These "cliques" can be used as hidden subgroups. Preferably, the number of nodes of the complete subgraph should be 2 or more. More preferably, the number of nodes of the complete subgraph should be 2, 3, 4, 5, 6.

[0213] In some embodiments, the step 9) considers all the classification units forming the hidden subgroup in step 6) as contamination introduced by the experimental process, and all the classification units forming the hidden subgroup in step 8) and having a lower statistical quantity in the two classification units of the pair as false positives introduced by the bioinformatics analysis, and marks them as noise signals and removes them.

[0214] The present application also provides a computer system comprising a memory, a processor and a computer program stored on the memory, wherein the processor executes the computer program to implement the steps of any of the above-mentioned methods.

[0215] According to the core idea of the present application, in some embodiments, the present application can also relate to a computer-readable storage medium having a computer program stored thereon, characterized in that the computer program is executed by a processor to implement the steps of any of the above-mentioned methods.

[0216] In some embodiments, the present application can also relate to a computer program product comprising a computer program, characterized in that the computer program is executed by a processor to implement the steps of any of the above-mentioned methods.

[0217] In some embodiments, the present application can also relate to the application of a hidden subgroup in distinguishing real signals from background signals in bioinformatics analysis, the hidden subgroup comes from the same source of noise signal introduction in the experimental process and / or bioinformatics analysis process, the proportion between the elements in the hidden subgroup is stable under two or more conditions; preferably, the hidden subgroup is the hidden subgroup formed between the test sample and the negative control, and / or the hidden subgroup is the hidden subgroup formed between the test sample and the bioinformatics control.

[0218] According to some aspects of the present application, when the present application is applied to quantitative analysis in the sequencing process, it relates to a bioinformatics analysis method for absolute quantification of sequencing data, comprising the following steps:

[0219] Step 1) grouping sequencing read sequences into different classification units;

[0220] Step 2) classification unit statistics step;

[0221] Step 3) classification unit interrelation calculation step;

[0222] Step 4) identifying hidden subgroup step;

[0223] Step 5) absolute quantification of negative control by hidden subgroup step;

[0224] Step 6) absolute quantification of test sample by hidden subgroup step.

[0225] The term "positive control" refers to a nucleic acid sample with a known sequence and concentration set during the sequencing experiment. In theory, the detection signal of this sample should only include the known sequence, but due to experimental contamination and false positives introduced by bioinformatics analysis, the sample will contain part of the noise detection. Since the amount of nucleic acid input is known, the overall nucleic acid amount of the sample can be calculated by dividing the proportion of known sequence nucleic acid in the sample by the amount of nucleic acid input, and the overall nucleic acid amount minus the known input can obtain the overall nucleic acid amount of noise nucleic acid molecules in the sample.

[0226] In the absolute quantification process of the present application, the term "hidden subgroup" refers to the hidden subgroup formed between the negative control and the positive control, and the hidden subgroup formed between the sample and the negative control. High abundance background signals often appear in the negative control, the positive control and the sample. The proportional relationship between these high abundance background signals can remain stable in the negative control, the positive control and the sample, forming a hidden subgroup. These hidden subgroups can be used as a pseudo-internal control (PIC), which can be used to calculate the ratio between the negative control and the positive control, and the ratio between the negative control and the sample. Through the chain equation, the ratio between the positive control and the sample can be calculated, and the absolute quantification of the sample can be calculated according to the absolute quantification of the positive control (known).

[0227] Therefore, in some embodiments, the hidden subgroup of step 5) is the hidden subgroup formed between the negative control and the positive control; and the hidden subgroup of step 6) is the hidden subgroup formed between the sample and the negative control.

[0228] As for other steps, in some embodiments, the grouping method of step 1) can be various grouping methods commonly used in the art, including but not limited to alignment-based grouping methods or non-alignment-based grouping methods.

[0229] In some embodiments, the alignment-based grouping method can use various alignment software to perform sequence alignment on the sequencing read sequences, and all sequencing read sequences aligned to the same classification unit are grouped into a group.

[0230] In some embodiments, the non-alignment-based grouping method includes but is not limited to kmer method, hash table method or string matching method, etc.

[0231] In some embodiments, the statistics of step 2) are based on the grouping results of the classification units of step 1), and various statistical quantities of each classification unit are calculated.

[0232] In some embodiments, the statistical quantity refers to a mathematical statistical quantity, including but not limited to: counting the number of supporting sequencing read sequences of each classification unit, counting the relative proportion of each classification unit, and counting the statistical quantity of each classification unit after certain normalization.

[0233] In some embodiments, step 3) is to pair each two classification units according to the statistical quantities of the classification units of step 2), and to calculate whether the proportion of the two classification units in the pair remains stable under two or more conditions; in some embodiments, the stable proportion indicates that the paired classification units are related to each other.

[0234] In some embodiments, the conditions in step 3) include: under the conditions of the sample to be tested, under the conditions of the negative control, under the conditions of the positive control.

[0235] In some embodiments, in step 3), the pairing can be random pairing; the ratio of the two classification units can be obtained by dividing the classification unit statistics, or by statistical test of the classification unit statistics.

[0236] In some embodiments, the identification of the hidden subgroup in step 4) is specifically the identification and analysis of the hidden subgroup based on the correlation between the classification units in step 2) and the classification units in step 3); the hidden subgroup is a form in which signals appear together under two or more conditions, and is composed of two or more elements; the ratio between each two elements in the hidden subgroup remains stable under two or more conditions.

[0237] In some embodiments, the identification of the hidden subgroup in step 4) can be performed by hidden subgroup analysis without prior information.

[0238] In some embodiments, the identification and / or analysis is performed using a mathematical analysis method for analyzing the correlation between two or more elements or the characteristics of the elements themselves.

[0239] In this application, the hidden subgroup is different from a single element, which only has the data characteristics of the element itself as the research object for subsequent analysis methods. The hidden subgroup includes not only the elements within the subgroup, but also the correlation between the elements within the subgroup. Therefore, the analysis method for the hidden subgroup can not only take the correlation between the elements within the subgroup as the research object, but also take the sum (or other combination methods such as product, statistical quantity, etc.) of the same data characteristics of all elements within the subgroup as the research object. Such analysis methods are collectively referred to as "analysis methods for analyzing the correlation between two or more elements or the characteristics of the elements themselves".

[0240] In some embodiments, the identification and analysis are performed using the graph method in computer science, i.e., each classification unit in step 2) is taken as a node of the graph, and each pair with a connection in step 3) is taken as an edge of the graph, to form a complete undirected graph; in the undirected graph, a complete subgraph is found using the classical graph analysis method, and the complete subgraph is taken as the hidden subgroup; preferably, the number of nodes of the complete subgraph is 2 or more.

[0241] In some embodiments, the step 5) is to identify the hidden subgroup from the subgroup stably existing in the negative control and the positive control in step 4), and select the subgroup with the highest proportion in the negative control as the pseudo-internal control (PIC). The proportion of the PIC in the negative control and the proportion of the PIC in the positive control are compared, and the proportion of the negative control relative to the whole positive control is obtained. Since the absolute number of the positive control is known, the absolute number of the negative control can be calculated.

[0242] In some embodiments, the step 6) is to identify the hidden subgroup from the subgroup stably existing in the negative control and the sample to be tested in step 4), and select the subgroup with the highest proportion in the negative control as the pseudo-internal control (PIC). The proportion of the PIC in the negative control and the proportion of the PIC in the sample to be tested are compared, and the proportion of the negative control relative to the whole sample to be tested is obtained. Since the absolute number of the negative control is known from step 5), the absolute number of the sample to be tested can be calculated.

[0243] To further illustrate the bioinformatics absolute quantification method based on hidden subgroup analysis described in the present application, the following is a specific embodiment of the present application, which is not limited, and which includes the following steps: step 1) grouping sequencing reads into different classification units; step 2) classification unit statistics step; step 3) classification unit correlation calculation step; step 4) identifying hidden subgroup step; step 5) absolute quantification of the negative control by the hidden subgroup step; step 6) absolute quantification of the sample to be tested by the hidden subgroup step.

[0244] In some embodiments, the step 1) has two grouping methods, namely the grouping method based on alignment and the grouping method not based on alignment. First, the grouping method based on alignment uses alignment software to perform sequence alignment on sequencing reads, and groups according to the alignment results, for example, all sequencing reads aligned to the same classification unit are grouped into a group. Second, the grouping method not based on alignment, for example, the kmer method first divides the sequencing reads into kmers, and divides the alignment database into kmers, and counts which reference sequence in the database is more consistent with a sequencing read, thereby grouping the sequencing read into the classification unit of the reference sequence. Preferably, for the grouping method based on alignment, the alignment software that retains non-single alignment results (such as BLASTN software) is used for sequence alignment and grouping of sequencing reads. Preferably, for the grouping method not based on alignment, the kmer method, hash table method, and string matching method are used.

[0245] In some embodiments, the step 2) calculates the statistical quantity of each classification unit based on the grouping result of the classification unit in step 1). For example, the number of support sequencing read sequences of each classification unit, the relative proportion of each classification unit, the statistical quantity of each classification unit after certain normalization (such as reads per million (RPM), etc.).

[0246] In some embodiments, the step 3) pairs each two classification units for the classification unit statistical quantity in step 2), and calculates whether the proportion of the two classification units in the pair is stable under two or more conditions. If it is stable, it is considered that the classification units in the pair are related to each other. In some specific embodiments, the proportion of the two classification units in each pair is obtained by dividing the classification unit statistical quantity, and whether it is stable is obtained by calculating whether the proportion of the proportion of the pair under two (or more) conditions is close to 1. For example, the statistical quantities of two classification units A and B under condition one are 10 and 30, respectively, and the proportion is 0.333. The statistical quantities under condition two are 1 and 3, respectively, and the proportion is 0.333. The proportion of the pair A and B under condition one (0.333) and condition two (0.333) is (0.333:0.333)1, which is stable. In some specific embodiments, the proportion of the two classification units in each pair is obtained by statistical test on the unit statistical quantity, such as the p value of goodness-of-fit test. For example, the statistical quantities of two classification units A and B under condition one are 10 and 30, respectively, and the proportion is 0.333. The statistical quantities under condition two are 1 and 3, respectively, and the goodness-of-fit test is performed on the four numbers 10, 30, 1, and 3, and the p value is calculated. If the p value is greater than or equal to a certain threshold (such as 0.05), it is considered to be stable.

[0247] In some specific embodiments, in order to determine whether the proportion of a pair under two (or more) conditions is close to 1, a threshold value needs to be set, such as greater than or equal to 0.5 and less than or equal to 2. The setting of the threshold value is related to the specific application scenario, and cannot be unified in different scenarios, and the specific threshold range is not protected and limited. In some specific embodiments, whether the proportion of a pair under two (or more) conditions is stable is determined by statistical test, and the threshold value is often 0.05, 0.01, etc. The setting of the threshold value is related to the specific application scenario, and cannot be unified in different scenarios, and the specific threshold range is not protected and limited.

[0248] In some embodiments, the step 4) identifies and analyzes the hidden subgroup based on the relationship between the classification units in step 2) and the classification units in step 3).

[0249] In some embodiments, the hidden subgroup analysis is performed without prior information,

[0250] In some embodiments, the identification and / or analysis is performed using an analysis method for analyzing the association between two or more elements or the characteristics of an element itself; for example, using a graph method in computer science, i.e., taking each classification unit in step 2) as a node of a graph and each pair with a connection in step 3) as an edge of the graph to form a complete undirected graph. In the undirected graph, using a classical graph analysis method, a complete subgraph is found, which is also called a "clique". These "cliques" can be used as hidden subgroups. Preferably, the number of nodes of the complete subgraph should be 2 or more. More preferably, the number of nodes of the complete subgraph should be 2, 3, 4, 5, 6.

[0251] Further, step 5) is to select the subgroup with the highest proportion in the negative control as a pseudo-internal control (PIC) for the hidden subgroup obtained in step 4) that stably exists in the negative control and the positive control. Comparing the proportion of the PIC in the negative control with the proportion of the PIC in the positive control can obtain the proportion of the negative control relative to the whole positive control. Since the absolute number of the positive control is known, the absolute number of the negative control can be calculated.

[0252] Further, step 6) is to select the subgroup with the highest proportion in the negative control as a pseudo-internal control (PIC) for the hidden subgroup obtained in step 4) that stably exists in the negative control and the sample to be tested. Comparing the proportion of the PIC in the negative control with the proportion of the PIC in the sample to be tested can obtain the proportion of the negative control relative to the whole sample to be tested. Since the absolute number of the negative control is known from step 5), the absolute number of the sample to be tested can be calculated.

[0253] It can be understood that the calculation framework of the present application is independent of the selection of a specific sequencing platform, and therefore in any of the above methods or systems, the sequencing data can come from a second-generation sequencing platform, a third-generation sequencing platform, or a fourth-generation sequencing platform, or even other types of sequencing methods or approaches; preferably, the sequencing data comes from an Illumina, BGI, ION TORRENT, PacBio, Roche, Helicos, ABI, or nanopore sequencing platform; more preferably, the sequencing data comes from an Illumina, BGI, or nanopore sequencing platform.

[0254] Furthermore, the computational framework of this application is independent of specific sample types or sample sources. Therefore, in any of the above methods or systems, the sequencing data is gene sequencing data, preferably, the sequencing data is genome sequencing data; more preferably, the sequencing data is metagenomic sequencing data; and even more preferably, the sequencing data is low-biomass metagenomic sequencing data.

[0255] The present application will now be described in conjunction with specific embodiments.

[0256] Example 1: Exploratory Design in Bioinformatics Noise Reduction

[0257] First, this application examines the entire process of genome (metagenomic) sequencing methods and, after careful analysis, breaks it down as follows. For example... Figure 1 As shown in steps 1-4, during sample preparation, a series of experimental contaminants are simultaneously introduced into both the sample and the negative control. Specifically, the sample originally contains only real nucleic acid molecules (solid dots), while the negative control sample contains nothing. Step 1 introduces a subgroup of contaminated nucleic acid molecules (squares), Step 2 introduces a new subgroup of contaminated nucleic acid molecules (hollow circles), and the proportion of the contaminated nucleic acid molecule subgroup represented by the squares changes (shrinks overall), as does the proportion of the real nucleic acid molecule subgroup represented by the solid dots. In Step 3, a new group of contaminated nucleic acid molecules (small solid dots A) is introduced; this nucleic acid molecule is also a real nucleic acid molecule (solid dot A), i.e., coexisting nucleic acid molecules. Furthermore, Step 3 also involves changes in the proportions of the original nucleic acid molecule subgroups (real nucleic acid molecule subgroup, square contaminated nucleic acid molecules, hollow circle contaminated nucleic acid molecules). In Step 4, cross-contamination between samples (such as aerosols) also occurs.

[0258] During sequencing, all nucleic acid molecules are randomly sampled using a Bernoulli process and their nucleic acid sequences are obtained by a sequencer (e.g., ...). Figure 1 (Sequencing step). After this, the nucleic acid sequences of all real nucleic acid molecules will undergo bioinformatics (such as...) Figure 1, bioinformatics) analysis, and some false alignments inevitably occur, forming some false detections, such as A', A", C' and C" in the figure. That is, in the whole process, for the sample, from the original true signal ABC to the final detection result ABC DEF GHIA' A" C' C", only ABC is the true signal, and DEF GHIA' A" C' C" are all noise signals, and DEF GH I are the experimental pollution introduced, and A' A" C' C" are the bioinformatics introduced false detections. Moreover, for the true signal A, as a coexisting nucleic acid molecule, its final detection result is the superposition of the true nucleic acid molecule A and the contaminated nucleic acid molecule A. For the negative control, it does not contain any true nucleic acid molecule, and the final detection result ABC DEF GH IJ A' A" are all noise signals, and among them ABC DEF GH IJ are derived from experimental pollution, which are true nucleic acid molecules (DEF GH IJ are introduced by steps 1 and 2, A is introduced by step 3, and B is introduced by step 4). And A' A" are bioinformatics introduced false detections.

[0259] Secondly, the present application finds that if the present application looks at each detection separately to trace back how this detection is introduced into the sample, the present application will be lost in the above-mentioned overall transformation process. However, the present application finds that the proportional relationship between each detection is more important than the detection of the microbial sequence itself. A subgroup of noise signals introduced in the same step will maintain the group proportion during subsequent processing and transformation. As shown in Figure 2 , even if the proportion between different subgroups changes greatly, the core idea of the present application is that the proportion within the group remains unchanged. Moreover, the present application finds that through sequencing analysis of samples with known composition, the proportional relationship between true signals and noise signals cannot be maintained stable in the sample to be tested and the negative control (the stable proportional relationship is a dark dot, and the unstable proportional relationship is a light dot), but, on the contrary, the proportional relationship between noise signals and noise signals is mostly maintained stable in the sample to be tested and the negative control (as shown in Figure 3 ).

[0260] Based on this feature, the present application establishes a denoising algorithm based on hidden subgroup analysis (Hugo-DNA) Figure 4). The sequencing read sequences of the test samples (SP) and the negative control (NC) are quality controlled and host filtered. The filtered data is data aligned, wherein the single alignment results (test samples and negative control) are the basis for the statistics of the classification units, and the multiplex alignment results of the test samples are retained and used for the construction of bioinformatics controls later. First, the test samples are compared with the negative control, and whether the proportional relationship of the pairwise of the classification units composed of the test samples remains stable in the test samples and the negative control is calculated. All the classification units of the test samples are the vertices of the undirected graph, and the pairwise with stable proportional relationship is the edge of the undirected graph, which constitutes the undirected graph. Hidden subgroups (i.e. complete subgraphs) are found in the undirected graph, and the classification units of all hidden subgroups are marked as noise signals. Second, for the classification units that do not form hidden subgroups (i.e. outliers), bioinformatics controls based on multiplex alignment are constructed, and the test samples are compared with the bioinformatics controls. If the proportion in the test sample is not higher than the proportion in the bioinformatics control, it is considered that there is a true signal and an error detection similar to the true signal in the paired classification unit, and at this time, the classification unit with less supporting sequencing read sequence number is considered to be an error detection, which is classified as a noise signal and is removed. Finally, all the signals left are marked as true signals.

[0261] In addition, based on the core idea of the present application, the method of only reducing noise for experimental pollution or bioinformatics noise is also the core method of the present application, although the noise reduction effect of this single method is slightly inferior to the above, but it can also achieve the effect claimed in the present application, and also belongs to the protection content of the present application, such as Figure 5 The schematic diagram of noise reduction based on bioinformatics multiplex alignment alone is shown.

[0262] Example 2: Bioinformatics noise reduction method

[0263] Based on example 1, the specific steps of establishing Hugo-DNA method of the present application are as follows:

[0264] 1) Constructing test samples with known composition: the present application adds high-purity artificially synthesized nucleic acid molecules with known sequences to nucleic acid-free water. The true signal of this sample is the artificially synthesized nucleic acid molecules with known sequences. All other detections are considered as noise signals.

[0265] 2) Constructing negative control sample: the present application uses nucleic acid-free water as negative control, and processes the negative control according to the processing of the test sample to ensure that the negative control can capture the noise signals in the whole process.

[0266] 3) Library construction: The VAHTS Universal Plus DNA Library Prep Kit for Illumina (Vazyme, ND617) kit was used for library construction, and the operation was performed according to the instructions. The Agilent 2100 bioanalyzer, Agilent 2100 DNA 1000 kit (Agilent, 5067-1504) was used for library quality inspection, and the Qubit 4.0 fluorometer was used for quantification. Each library was normalized to the same concentration and mixed, and finally sequenced on the Nextseq 550 Dx sequencing platform for paired-end 150 bp sequencing.

[0267] 550Dx sequencing platform for paired-end 150 bp sequencing.

[0268] 4) Microbial community composition analysis of off-line data: The genome database used for annotation was extracted from the nt database of NCBI, including archaea, bacteria, fungi and viruses. The BLASTN was used for sequence alignment, and the single alignment reads were used for subsequent sample detection. These detections constitute the classification unit list of sample detection. Both the test sample and the negative control can obtain a classification unit composition to facilitate further removal of experimental introduced pollution. For the sample, the multiple alignment reads are also counted. These multiple alignment reads are used to construct a multiple alignment result, i.e. a classification unit can be aligned to the number of reads of other classification units, to facilitate further removal of bioinformatics introduced false positives.

[0269] 5) The classification unit composition of the test sample and the classification unit composition of the negative control are compared to find the subset of common detected classification units. These classification units constitute the vertices in the undirected graph. For these commonly detected classification units, the proportion between each other is calculated, and if the proportion of this pair is consistent in the sample and in the negative control, it is considered that there is a connection between the pair of classification units, forming an edge of the undirected graph.

[0270] 6) Whether the proportion of a pair of classification units in the test sample and the negative control is consistent is determined by the following rules:

[0271] a) If the difference between the proportions is less than 2 times;

[0272] b) If the p-value of the goodness-of-fit test is higher than 0.05.

[0273] 7) Search for a complete subgraph in the above undirected graph using a depth-first algorithm. The elements contained in the complete subgraph are marked as noise signals. Other classification units are marked as candidate real signals.

[0274] 8) For each candidate true signal, compare it with the multiple alignment results. Similarly, calculate whether the proportion of each pair of classification units is consistent with their proportion in the multiple alignment results. The rule of whether the proportion is consistent is the same as 6) above. If the proportion of a pair of classification units remains consistent, the present application considers that the classification unit with fewer reads is a false positive of bioinformatics error alignment, and marks it as a noise signal, and the remaining all signals are marked as true signals.

[0275] Example 3 Hidden subgroup display in the test sample in bioinformatics noise reduction

[0276] The present application further verifies the existence of the hidden subgroup in the core hypothesis of the Hugo-DNA algorithm. The specific steps are as follows. Based on the method of Example 2 (steps 1-7), the present application constructs an undirected graph of the test sample ( Figure 6 ). Among them, each circle point is a vertex in the undirected graph, that is, a classification unit of the test sample; each gray line is an edge of the undirected graph, indicating that the proportion relationship of the two classification units connected thereto remains stable in the test sample and the negative control. As shown in Figure 6 , all vertices in the undirected graph except the true signal (indicated by the arrow) are connected, indicating that all classification units except the true signal come from different hidden subgroups and are all noise.

[0277] Example 4 Comparison of the bioinformatics noise reduction method of the present application with existing methods

[0278] Further, this embodiment compares the results of Hugo-DNA with other existing methods, and the specific steps are as follows:

[0279] 1) Based on the method of Example 2 (steps 1-4), this embodiment obtains the composition of classification units of the test sample and the negative control. These results will be used as input for subsequent different noise reduction methods. And also as the original data case is shown in Figure 6 (RAW). Among them, the true signal is indicated by an arrow.

[0280] 2) For the Hugo-DNA method, this embodiment uses steps 5-8 of Example 2 to analyze and obtain the annotation of true signals and noise signals. Figure 7 , DNA)

[0281] 3) For Z-score method noise reduction, this embodiment normalizes the relative abundance of classification units of the test sample and the negative control by Z-score, and marks the classification units with a relative abundance in the test sample after normalization higher than 2 times in the negative control as true signals, and the others as noise signals. Figure 7 , ZSC)

[0282] 4) For Decontam method denoising, in this example, the classified unit results of the test sample and the negative control were input into the software Decontam (version 1.10.0) for frequency-based denoising analysis. The selected threshold was 0.5, which additionally required library concentration information when building the library, and the library concentration information was determined using the Invitrogen TM Qubit TM Fluorometer before the machine. Figure 7 , DCN

[0283] 5) For SourceTracker method denoising, in this example, the classified unit results of the test sample and the negative control were input into the software (version 1.0) for analysis. The SourceTracker method outputs the probability of each classified unit being contaminated and the probability of being real, and in this example, the classified unit with a probability of contamination greater than the probability of being real was marked as contaminated, and the rest of the classified units were marked as real signals. Figure 7 , STR

[0284] The results are shown in Figure 7 , the test sample includes about 2000 classified units, and among them, only one is a real signal (red dot) and the rest are noise signals (green dot). Compared with other methods, Hugo-DNA retains not only the real signal but also fewer noise signals.

[0285] Example 5 Method establishment of the present application in absolute quantification

[0286] Based on the similar ideas in the above denoising aspect, in this example, the application of the quantification aspect was explored. Due to the extremely low detection limit and extremely high sensitivity of the metagenomic method, any metagenomic detection result of a sample will contain nucleic acid molecule signals from the sample source ( Figure 8 A left one, gray column, A B C D) and nucleic acid molecule signals from the background source ( Figure 8 A left one, black column, E F G). These nucleic acid molecule signals from the background source can be represented by a negative control sample ( Figure 8 A left one, black column, E F G H). Although the relative proportions of these nucleic acid molecule signals from the background source may be different in the sample and in the negative control ( Figure 8 A left two), their sources are the same, and the proportional relationship between them remains unchanged ( Figure 8 A left three), forming a hidden subgroup (E F G). These hidden subgroups can be used as a hypothetical internal control (PIC), and the relative proportion of the hypothetical internal control in the sample and the negative control can be obtained by simple linear regression. Figure 8A right one). So far, by hiding the subgroups (PIC), we can get the relative quantitative results of the sample to be tested and the negative control Figure 8 A right one, the slope term a1 in the fitting equation, while also calculating the proportion of hidden subgroups in the sample to be tested (%PICsp).

[0287] The application also compares the quantitative method using the quantitative method proposed in this patent Figure 8 B black solid line upper half) and the quantitative method using internal reference Figure 8 B black solid line lower half). First, for the quantitative method proposed in this patent, in order to obtain absolute quantitative results, the metagenomic sequencing of the sample to be tested (sample, SP), the negative control (negative control, NC) and the positive control (positive control, PC) is performed, and the microbial composition results are obtained. Assuming that the microbial content in the sample to be tested is χ, the microbial content in the background (i.e. negative control) is λ, and the microbial content in the positive control is δ, and the microbial content δ in the positive control is known. For the sample to be tested, the ratio of the microbial content χ in the sample to be tested to the microbial content λ in the negative control can be calculated by the proportion of hidden subgroups in the sample to be tested (%PICsp) using the formula; similarly, the ratio of the positive control microbial content δ to the negative control microbial content λ can be calculated by the proportion of hidden subgroups in the sample to be tested (%PICpc) using the formula; by combining the formula, and since the microbial content δ in the positive control is known, we can calculate the microbial content χ in the sample to be tested by the formula. Second, for the quantitative method using internal reference, the microbial content δ of the internal reference is known, and its proportion in the sample to be tested can also be obtained (%IC). In order to obtain absolute quantitative results, the sum of the microbial content in the sample to be tested and the microbial content from the background (i.e. negative control) χ+λ can be calculated by the formula. But the method using internal reference cannot distinguish the microbial content χ in the sample to be tested and the microbial content λ from the background. This will lead to the fact that the quantitative method using internal reference cannot achieve true absolute quantification in theory.

[0288] Example 6 Bioinformatics absolute quantitative effect verification

[0289] The application uses zymo standard as real sample nucleic acid for sequencing, and designs a gradient of input nucleic acid amount from 10 2 ng / ml to 10 -6 ng / ml, with a decrease of 10 times each time. Each gradient is repeated 3 times in parallel.

[0290] As Figure 9As shown, with the concentration gradient of the input nucleic acid molecules decreasing, the ratio of the real sample to the negative control calculated by the hidden subgroup also decreases in gradient, and there is a good linear relationship in three parallel experiments, which shows the effectiveness of the quantitative method.

[0291] The above description of specific embodiments of the application does not limit the application, and various changes or modifications can be made to the application by those skilled in the art without departing from the spirit of the application, and all such changes or modifications shall fall within the scope of the appended claims of the application.

Claims

1. A method for bioinformatics analysis of sequencing data, characterized in that, The method comprises the following steps: Step 1) grouping sequencing read sequences into different classification units; Step 2) classification unit statistics step: based on the grouping results of step 1), calculate the statistics of each classification unit; Step 3) classification unit correlation calculation step: pairing each two classification units according to the classification unit statistics of step 2), and calculating whether the proportion of the two classification units in the pair is stable under two or more conditions; the stability indicates that the classification units in the pair have a relationship with each other; Step 4) hidden subgroup identification step: based on the classification units in step 2) and the classification unit correlations in step 3), the hidden subgroup is identified and / or analyzed. The hidden subgroup is a form in which signals coexist under two or more conditions, and is composed of two or more elements; the proportion between any two elements in the hidden subgroup remains stable under two or more conditions; the hidden subgroup comes from signals introduced from the same source in the experimental process or bioinformatics analysis process.

2. The method of claim 1, wherein, The hidden subgroup includes: a hidden subgroup formed between the test sample and the positive control, a hidden subgroup formed between the test sample and the negative control, a hidden subgroup formed between the test sample and the bioinformatics control, a hidden subgroup formed between the test sample and other controls, a hidden subgroup formed between the test sample and the test sample, a hidden subgroup formed between the negative control and the positive control, a hidden subgroup formed between the bioinformatics control and the negative control, a hidden subgroup formed between the other control and the positive control, a hidden subgroup formed between the other control and the negative control, or a hidden subgroup formed between the bioinformatics control and the other control.

3. The method of claim 1 or 2, wherein the method further comprises: The grouping method of step 1) includes an alignment-based grouping method or an alignment-free grouping method. 4.The method of claim 3, wherein, The alignment-based grouping method uses alignment software to perform sequence alignment on sequencing read sequences, and all sequencing read sequences aligned to the same classification unit are grouped into a group.

5. The method of claim 3, wherein, The grouping method without alignment includes a kmer method, a hash table method, or a string matching method.

6. The method of claim 1-2, wherein In step 2), the statistics include: counting the number of support sequencing read sequences of each classification unit, counting the relative proportion of each classification unit, and counting the amount of each classification unit after a certain normalization.

7. The method of claim 1-2, wherein, In step 3), the conditions include: under the condition of the test sample, under the condition of the negative control, under the condition of the positive control, under the condition of the bioinformatics control, or under the condition of the other control.

8. The method of claim 1-2, wherein, In step 3), the pairing is random pairing; the proportion of the two classification units is obtained by dividing the classification unit statistics, or by statistical test of the classification unit statistics.

9. The method of claim 1-2, wherein, The hidden subgroup identification of step 4) is performed in a priori-free manner.

10. The method of claim 9, wherein, In step 4), the identification and / or analysis is performed by using an analysis method for analyzing the relationship between two or more elements or the characteristics of an element itself.

11. The method of claim 10, wherein, In step 4), the analysis method is graph theory method using computer science, i.e. taking each classification unit in step 2) as a vertex of a graph, and taking each pair with connection in step 3) as an edge of the graph, to construct a complete undirected graph; in the undirected graph, find a complete subgraph, which is the hidden subgroup; the number of vertices of the complete subgraph is 2 or more.

12. The method of claim 1-2, wherein, The sequencing is from a second-generation sequencing platform, a third-generation sequencing platform, or a fourth-generation sequencing platform.

13. The method of claim 1-2, wherein, The sequencing is from a second-generation sequencing platform, a third-generation sequencing platform, or a fourth-generation sequencing platform. Illumina, BGI, ION TORRENT, PacBio, Roche, Helicos, ABI, or a nanopore sequencing platform. 14.The method of claim 13, wherein, The sequencing is from an Illumina, BGI, or nanopore sequencing platform.

15. The method of claim 1-2, wherein, The sequencing data is gene sequencing data.

16. The method of claim 15, wherein, The gene sequencing data is metagenomic sequencing data.

17. The method of claim 16, wherein, The metagenomic sequencing data is low-biomass metagenomic sequencing data.

18. A method of sequencing data denoising signal analysis, comprising: The method comprises any one of claims 1-17, and the hidden subgroup is a hidden subgroup formed between the to-be-tested sample and a negative control, and / or a hidden subgroup formed between the to-be-tested sample and a bioinformatics control; Label the classification units of all hidden subgroups as noise signals.

19. A computer readable storage medium having stored thereon a computer program, characterized in that, The computer program, when executed by a processor, implements the steps of the method of any one of claims 1-17.

20. A computer program product comprising a computer program, characterized in that, The computer program, when executed by a processor, implements the steps of the method of any one of claims 1-17.

Citation Information

Patent Citations

  • Method and device for detecting copy number variation based on amplicon next generation sequencing

    CN106372459A

  • Full-process quality control flora high-throughput sequencing detection method and application thereof

    CN113265453A