Low-frequency mutation detection method based on mixed sequencing
Through hybrid sequencing technology, the coding matrix and decoding algorithm are used to reduce costs and improve accuracy, solving the problems of high computational complexity and high cost in hybrid sequencing, and realizing efficient detection of low-frequency mutations in large populations.
Patent Information
- Application Number
- CN202510576240.4
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-05-06
- Publication Date
- 2025-09-05
AI Technical Summary
Existing hybrid sequencing technology has problems with high computational complexity, high cost and insufficient accuracy in low-frequency mutation detection, making it difficult to apply to low-frequency mutation screening in large populations.
A low-frequency mutation detection method based on hybrid sequencing is adopted. By designing a coding matrix to construct a hybrid pool, combined with a threshold adaptive variation detection algorithm and a decoding algorithm, sequencing costs are reduced and detection accuracy is improved.
While ensuring the accuracy of single nucleotide variation detection is not less than 95%, the detection cost is reduced by more than 60%, making it suitable for low-frequency mutation detection in large populations.
Smart Images

Figure CN120600102A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to a mutation detection method, in particular to a low-frequency mutation detection method based on hybrid sequencing, and belongs to the field of bioinformatics. Background Art
[0002] Low-frequency mutations (SNVs) are single nucleotide variants (SNVs) with a minor allele frequency (MAF) greater than 0.5% and less than 5%. Despite their low incidence in the human population, numerous studies have shown that SNVs play a significant role in the development and progression of several major diseases, including cancer, rare inherited metabolic disorders, autoimmune diseases, and cardiovascular diseases. Detection of SNVs is crucial for early disease detection, diagnosis, personalized treatment, and prognosis assessment.
[0003] Next Generation Sequencing (NGS) technology has become an important method for detecting low-frequency mutations due to its high detection throughput and high sensitivity. However, due to its high cost, NGS is currently not suitable for large-scale clinical screening of low-frequency mutations. With the continuous development of NGS technology, the price of sequencing has gradually decreased, and the main cost of sequencing comes from pre-sequencing library construction. Hybrid sequencing technology reduces sequencing costs by mixing multiple samples into the same sequencing pool in a certain encoding method and using mathematical models to restore individual variations.
[0004] Hybrid sequencing mixes each sample into multiple mixed pools, using the sample mixing pattern as the encoding method. After sequencing is completed, a specific decoding method is used to reconstruct the variation information of each sample based on the variation information of the mixed pool. Hybrid sequencing can reduce sequencing costs by reducing the number of sequencing times, but the mixing of samples will cause the proportion of DNA fragments containing SNVs to be significantly reduced, which poses a huge challenge to the accurate detection of SNV variations from mixed sequencing data. Although the current decoding algorithm has a high accuracy rate, it has a high computational complexity and requires a lot of computing resources. It is not suitable for restoring the variation of low-frequency mutations in samples from the variation information of the mixed pool. Therefore, it is necessary to develop variation detection tools suitable for mixed sequencing data and algorithms that take both accuracy and computational complexity into account. Summary of the Invention
[0005] Purpose of the invention: The technical problem to be solved by the present invention is to provide a low-frequency variation detection method based on hybrid sequencing to reduce the cost of low-frequency mutation detection in large-scale populations.
[0006] Technical solution: In order to solve the above technical problems, the present invention provides a low-frequency mutation detection method based on hybrid sequencing, and its operating steps are as follows: design a coding matrix according to the total number of samples to be tested and the target mutation frequency, and construct one or more mixed pools containing part of the genetic material of the samples through the coding matrix, and obtain sequencing data of each mixed pool, align the sequencing data to the reference genome, and use a threshold adaptive variation detection algorithm based on the distribution characteristics of the variation frequency of each base site to obtain the single nucleotide variation of the mixed pool, and use a decoding algorithm to decode the variation information of the mixed pool to obtain the variation information of each sample.
[0007] Furthermore, the threshold adaptive mutation detection algorithm specifically performs threshold judgment based on the mutation detection algorithm threshold T. When the ratio of the number of minor alleles to the number of major alleles at a site in the mixed sample exceeds the threshold, the site is determined to have mutated.
[0008] Furthermore, the threshold calculation method is:
[0009]
[0010] Where N is the number of samples in the mixing pool.
[0011] Furthermore, the decoding algorithm goes through two stages: deterministic filling and gradual elimination inference.
[0012] Furthermore, the deterministic filling will traverse the coding strategy to find samples marked as participating in sequencing in the mixed pool for the positive mutation sites in the mixed pool. The positive sites are possible mutation sites of the samples participating in the mixed pool. For the negative mutation sites in the mixed pool, the coding strategy will be traversed to find samples marked as participating in sequencing in the mixed pool. All samples have not mutated at this site. The specific formula is as follows:
[0013]
[0014] Where M represents the encoding matrix, x represents the reconstruction matrix of the sample to be tested, and y represents the mutation matrix of the mixing pool.
[0015] Furthermore, the stepwise elimination inference checks the variation of all mixed pools containing the same sample at possible sites. If all mixed pools mutate, the sample is considered to have mutated at the possible mutation site; otherwise, it is considered that no mutation has occurred. The specific formula is as follows:
[0016] x i =1if{y i |i∈P j}={1}
[0017] Among them, x represents the reconstruction matrix of the sample to be tested, y represents the variation matrix of the mixed pool, and P j Indicates the numbers of all samples participating in the j-th mixing pool.
[0018] Furthermore, the deterministic filling strategy is used to quickly obtain variation information of 80% of the samples, gradually eliminates inference to obtain variation information of the remaining 20% of the samples, and ensures the accuracy of all sample variation through iteration.
[0019] Furthermore, the target mutation frequency can be any frequency lower than 5%, such as 1%, 2%, 3%, 4%, etc.
[0020] Furthermore, the number of the mixed pool is 5% to 50% of the total number of samples, and the number of samples in the mixed pool is less than 50.
[0021] Furthermore, the total sequencing depth of each mixed pool was 10× to 10,000×, and the sequencing depth of a single sample was 5× to 200×.
[0022] Furthermore, the sequencing method can be whole genome sequencing, whole exome sequencing, or targeted sequencing.
[0023] Furthermore, the species of the sample to be tested may be one or more of humans, animals, plants, and microorganisms.
[0024] Beneficial effects: Compared with the existing technology, the present invention has the following advantages: The low-frequency mutation detection method based on hybrid sequencing designed by the present invention can reduce the detection cost by more than 60% while ensuring the accuracy of single nucleotide variation detection is not less than 95%, compared with the traditional single-sample detection method. This method uses a hybrid sequencing strategy and performs sequencing in units of mixed pools, which effectively reduces the number of reactions for library construction and the amount of reagents required, and significantly reduces the sequencing cost from the source. The present invention effectively solves the problem of high cost of low-frequency mutation detection in large populations with a lower sequencing cost, which is conducive to promoting the popularity of low-frequency mutation detection. BRIEF DESCRIPTION OF THE DRAWINGS
[0025] Figure 1 It is an operational flow chart of the present invention.
[0026] Figure 2 Schematic diagram of hybrid sequencing in the embodiment.
[0027] Figure 3 This encoding matrix was designed using the STD (Shifted Transversal Design) algorithm, based on a total sample size of 20 and a target mutation frequency of 1%. Each row in the matrix represents a pool, and each column represents a sample. A total of 12 pools were generated, each containing 5 samples.
[0028] Figure 4 This is the test result of Example 1 based on hybrid sequencing of 100 neonatal genetic metabolic disease samples.
[0029] Figure 5 Example 1 is based on the differences in accuracy of hybrid sequencing mutation detection for seven types of inherited metabolic diseases. AG are: amino acid metabolism defects, neuronal metabolism defects, cholesterol metabolism, glycogen and carbohydrate metabolism disorders, lysosomal storage diseases, peroxisomal storage diseases, and free fatty acid oxidation disorders. DETAILED DESCRIPTION
[0030] The specific technical solutions of the present invention are further described in detail below in conjunction with specific embodiments.
[0031] The present invention provides a method for detecting low-frequency mutations based on hybrid sequencing, and its operating steps are as follows: designing a coding matrix based on the total number of samples to be tested and the target mutation frequency, and constructing multiple mixed pools containing part of the genetic material of the samples through the coding matrix, and obtaining sequencing data for each mixed pool, aligning the sequencing data to the reference genome, and using a threshold adaptive variation detection algorithm based on the distribution characteristics of the variation frequency of each base site to obtain single nucleotide variations in the mixed pool, and using a decoding algorithm to decode the variation information of the mixed pool to obtain the variation information of each sample.
[0032] The specific steps are as follows:
[0033] Step 1: Use the total number of samples to be tested and the target mutation frequency to generate a coding matrix M for mixed pool construction based on the STD algorithm. The algorithm uses the number of samples and the target mutation frequency as input parameters, outputs the required number of mixed pools and the number of samples contained in each mixed pool, and generates a coding matrix M that meets the coverage requirements. Each row of the coding matrix M represents a mixed pool, each column represents a sample, and each element M(i,j) indicates whether the mixed pool in the i-th row contains the genetic material of the j-th sample. If M(i,j) is 0, it contains it, and if M(i,j) is 1, it does not contain it. Using the STD algorithm, for example, when the total number of samples is 20 and the mutation frequency is 1%, the generated coding matrix is as follows Figure 3 shown.
[0034] Step 2: Library construction and sequencing are performed on the genetic material in the mixed pool to obtain sequencing data for each pool. A threshold-adaptive variant detection algorithm is used to detect single nucleotide variants in the pool based on the distribution characteristics of the variant frequency at each base site. This variant detection algorithm first preprocesses the sequencing data of each pool to remove low-quality, low-depth sequence fragments. The filtering criteria used are a Phred quality value below 30 (Q < 30) or a sequencing depth less than 1 / 10 of the total sequencing depth of the mixed sample. Subsequently, the preprocessed sequencing data are aligned to the human reference genome and sorted according to chromosomal order. The base information for each site in the mixed pool is then extracted, and the ratio of the major allele support to the minor allele support is calculated as the variant frequency for each site. Finally, a threshold calculation method is used to determine the variant site: when the variant frequency of a site exceeds the set threshold, the site is determined to have undergone a mutation.
[0035] The threshold calculation method is:
[0036]
[0037] N is the number of samples in the mixing pool.
[0038] Step 3: The decoding algorithm combines the encoding matrix to decode the single nucleotide variants in the mixed pool, thereby obtaining the variation information of each sample. The decoding algorithm goes through two stages: deterministic filling and step-by-step elimination inference.
[0039] Deterministic filling will traverse the mutation detection information of the mixed pool. When a mutation is detected at site i in a certain mixed pool j, all samples marked as participating in sequencing in the mixed pool will be recorded based on the mapping relationship between samples and mixed pools in the encoding matrix M, and the value of the corresponding site i in the mutation matrix x of the sample will be directly set to 1, indicating that these samples have all mutated. The specific formula is as follows:
[0040]
[0041] Where M represents the encoding matrix, x represents the reconstruction matrix of the sample to be tested, and y represents the mutation matrix of the mixing pool.
[0042] The step-by-step elimination inference targets the sample sites whose states have not yet been determined in the mutation matrix x. First, the set P of all mixed pools where the sample site is located is searched based on the encoding matrix M; then the test results of these mixed pools at the same site in the mutation detection information y of the mixed pool are retrieved. If all the test results in the set are positive, the site of the corresponding sample in the mutation matrix x is directly marked as "1" to determine that it has mutated; if there is any mixed pool with a negative test result at this site, no value is assigned for the time being and it is left for subsequent iterations to judge. By continuously iterating the above rules, the impossible situations are gradually eliminated against the real-time information of y until the complete mutation state of the remaining samples is restored. The specific formula is as follows:
[0043] x i =1if{y i |i∈P j}={1}
[0044] Among them, x represents the reconstruction matrix of the sample to be tested, y represents the variation matrix of the mixed pool, and P j Indicates the numbers of all samples participating in the j-th mixing pool.
[0045] Example 1 Detection of 100 Neonatal Genetic Metabolic Disease Samples Based on Hybrid Sequencing
[0046] This example involves targeted sequencing data of 100 neonatal genetic metabolic diseases, where the mutation frequency of the target mutation site to be detected does not exceed 2%. Peripheral blood of 100 neonatal genetic metabolic diseases obtained from a hospital in Jiangsu Province was first collected and DNA was extracted using the QIAamp DNA Mini Kit (250). The library was constructed using the BerryRare TMInborn Error of Metabolism 174 Target Panel (Berry Genomics, catalog number BR-IEM174). This library construction kit covers 174 common genes associated with neonatal genetic metabolic diseases, including amino acid metabolism disorders, neurometabolism defects, and cholesterol metabolism disorders. Based on the number of samples to be tested and the target mutation frequency, the STD algorithm was selected for encoding design. The STD algorithm uses 100 samples and a target mutation frequency of 2% as input parameters, outputs and constructs an encoding matrix M with 20 mixed pools, each containing 20 samples. The average sequencing depth of each sample is set to 20×, and the sequencing depth of each mixed pool is set to 400×. Among them, the encoding matrix M is a binary matrix. Each row of the matrix represents a mixed pool; each column represents a sample of a neonatal genetic metabolic disease; each element M(i,j) indicates whether the mixed pool in the i-th row contains the genetic material of the j-th sample. If it does, it is 1, and if it does not, it is 0. Libraries were constructed for these 20 pools, following the library construction kit instructions for DNA fragmentation (approximately 300 ± 20 bp), end repair, adapter ligation, library amplification, and target region capture. Finally, a pooled sequencing library with a peak fragment length of approximately 350 bp was prepared and sequenced on an Illumina NovaSeq 6000PE150 platform. High-throughput sequencing was then performed on the pooled sequencing library to obtain sequencing data for the corresponding 20 pools. For the sequencing data of the mixed pool, the threshold adaptive algorithm of the present invention is used to perform variation detection. The threshold adaptive algorithm will first pre-process the sequencing data of each mixed pool to remove low-quality, low-depth sequence fragments; Subsequently, the pre-processed sequencing data is aligned to the human reference genome and sorted according to the chromosome order; Then the variation information of each site in the mixed pool is extracted, and according to the threshold calculation method, the number of samples in the mixed pool at this time is 20. According to the threshold calculation formula, the variation detection threshold is calculated to be 0.009, and it is set: When the variation frequency of a certain site exceeds 0.009, the site is determined to be a site where a variation has occurred. Therefore, single nucleotide variations of 20 mixed pools are obtained. For the single nucleotide variations of these 20 mixed pools, a variation detection result matrix Y for neonatal genetic metabolic diseases is constructed, where each row of the result matrix represents a mixed pool, and each column represents a single nucleotide variation. According to the encoding matrix M and the result matrix Y, the decoding algorithm is used to decode the variation information of each sample. The decoding algorithm will first use deterministic filling to quickly obtain the variation information of about 80% of the samples, and then use step-by-step elimination inference to obtain the variation information of the remaining approximately 20% of the samples, and then obtain the variation information of each sample.
[0047] The results showed that compared with single sample variation detection, the detection accuracy of SNV was 98.9% and the false positive rate was 1.1%; the detection accuracy of mutation sites in more than half of the samples was 100%, and the accuracy of mutation sites in the remaining samples was above 96%, the false positive rate did not exceed 4%, and the false negative rate was 0. Figure 4 The detection accuracy is calculated as follows: the number of mutation sites correctly identified by hybrid sequencing divided by the total number of mutation sites detected by a single sample; the false positive rate is calculated as follows: the number of sites misidentified as mutations by hybrid sequencing divided by the total number of mutation sites detected by a single sample; the false negative rate is calculated as follows: the number of mutation sites missed by hybrid sequencing divided by the total number of mutation sites detected by a single sample; the results are as follows: Figure 3 shown.
[0048] In terms of cost, this invention significantly reduces sequencing and library construction costs while ensuring detection accuracy. Traditional testing requires building libraries and sequencing independently for 100 samples, requiring 100 library constructions, a total sequencing depth of 2000×, and a total cost of 50,000 yuan. However, by constructing 20 mixed pools, this method only requires 20 library constructions, with each mixed pool sequencing at a depth of 400×, a total sequencing depth of 8000×, and a total cost of 17,600 yuan. The cost of sequencing mixed samples accounts for only 35.2% of the cost of single-sample variant detection, a savings of up to 64.8%.
[0049] Furthermore, the present invention also analyzes the detection accuracy based on different classifications of neonatal genetic metabolic diseases. Figure 5 As shown in the figure, the median detection accuracy of four categories of diseases, including neurometabolism defects, lysosomal storage diseases, peroxisomal storage diseases, and free fatty acid oxidation disorders, is close to 99.5% to 100%, and the distribution range is highly concentrated with minimal fluctuation; while the median accuracy of three categories of diseases, including amino acid metabolism disorders, cholesterol metabolism disorders, and glycogen and carbohydrate metabolism disorders, is about 97.5%, with an overall distribution range of 94% to 100%. Figure 5 shown.
[0050] Example 2 Detection of 376 samples suspected of having rare diseases based on hybrid sequencing
[0051] This example involves samples from 376 patients suspected of having rare genetic diseases. Peripheral blood samples from 376 patients suspected of having rare genetic diseases were collected from a hospital in Jiangsu Province using BerryRare TMDNA was collected and extracted using the Rare Disease Targeted Sequencing Kit (Berry Genomics, Cat. No. BR600). This kit covers 600 rare disease-associated genes, including mitochondrial diseases and neurodevelopmental disorders. The target mutation frequency for the target site to be detected was no more than 1%. Based on a total sample size of 376 and an expected mutation frequency of 1%, the STD algorithm was selected. Using the 376 samples and the target mutation frequency of 1% as input parameters, the algorithm output a matrix consisting of 57 pools, each containing 25 samples. The average sequencing depth per sample was set to 30×, and the sequencing depth per pool was set to 750×. Library construction was performed for these 57 pools according to the kit instructions, including DNA fragmentation (approximately 300 ± 20 bp), end-repair, adapter ligation, library amplification, and target region capture. A pooled sequencing library with a peak fragment length of approximately 380 bp was generated and sequenced on an Illumina NovaSeq 6000 PE150 platform. The mixed library was then subjected to high-throughput sequencing to obtain sequencing data for the corresponding 57 mixed pools. The threshold adaptive variation detection algorithm designed by the present invention was used for the sequencing data of the mixed pool. First, the original sequencing data was quality controlled and preprocessed to remove low-quality and low-depth sequences; the data was then aligned to the human reference genome, sorted by chromosome order, and the variation information of each site was extracted. Based on the variation detection threshold adaptive algorithm, the number of samples N in the mixed pool of the current experiment was 25. According to the threshold calculation formula, the variation detection threshold set at this time was 0.007. When the variation frequency of a certain site exceeded the threshold, it was determined to have mutated. Thus, single nucleotide variations of 57 mixed pools were obtained, and the result matrix Y was constructed. Wherein, each row of the result matrix Y represents a mixed pool, each column represents a single nucleotide variation, and each element Y(i, j) represents whether the j-th column site in the i-th row mixed pool has mutated. If a mutation occurs, Y(i, j) is 1, and if no mutation occurs, it is 0. Combining the encoding matrix M and the result matrix Y, a decoding algorithm is used to decode the variation information of each sample. The decoding algorithm first uses deterministic filling to quickly decode and obtain the variation information of about 90% of the samples, and then gradually eliminates and infers the variation information of the remaining samples. The algorithm then restores the variation information of each sample.
[0052] Results showed that compared to single-sample variant detection, SNV detection accuracy was 99.8% with a false-positive rate of 0.2%. The accuracy of reconstructing the mutation sites was 100% for 97% of samples, and higher than 98% for the remaining 3%. The false-positive rate for 97% of the samples was 0%, 1% for 2%, and 2% for only 1%. The false-negative rate was 0% for all samples, resulting in a cost savings of 73.6%.
[0053] Example 3 Detection of 608 cystic fibrosis samples based on hybrid sequencing
[0054] This embodiment involves whole exome sequencing data of 608 cystic fibrosis samples. Peripheral blood of 608 suspected cystic fibrosis patients obtained from a hospital in Jiangsu Province was collected and DNA was extracted using the QIAamp DNA Mini Kit (250) kit. Whole exome sequencing was performed using the exon capture kit developed by Twist Bioscience and VCGS Laboratory (product name: Human Core Exome Kit, item number: 102017). Based on the total number of samples and the expected mutation frequency, the STD algorithm was selected, and the number of samples 608 and the target mutation frequency 1% were input as input parameters. The output was a coding matrix M with a number of mixed pools of 92 and a number of samples N in the mixed pool of 15. The average sequencing depth of each sample was set to 20×, and the sequencing depth of each mixed pool was set to 300×. The coding matrix M is a binary matrix that represents the correspondence between the mixed pool and the sample. Libraries were constructed for 92 mixed pools. DNA fragmentation, end repair, adapter ligation, library amplification, and target region capture were completed in sequence according to the standard process of the kit. After obtaining the sequencing library, high-throughput sequencing was performed to obtain sequencing data for the 92 mixed pools. After preprocessing, reference genome alignment, and sorting, the sequencing data was preprocessed to extract the variant site information. Based on the threshold calculation formula, the number of samples N in the mixed pool was 15, and the mutation detection threshold of this experiment was calculated to be 0.012. Based on this, the presence of mutations at each site was identified. The extracted mutation results were used to construct the matrix Y, where each row of the result matrix Y represents a mixed pool, each column represents a single nucleotide variation, and each element Y(i, j) indicates whether a mutation occurred at the j-th column site in the i-th row mixed pool. If a mutation occurred, Y(i, j) is 1, and if no mutation occurred, it is 0. Combining the encoding matrix M and the result matrix Y, a decoding algorithm is used to decode the sample variation information. The decoding algorithm first uses deterministic filling to obtain 90% of the sample variation information, and then restores the remaining sample variation information by gradually eliminating inference, thereby obtaining the variation information of each sample.
[0055] Results showed that compared to single-sample variant detection, SNV detection accuracy was close to 100%, with a false-positive rate of less than 0.1%. The accuracy of reconstructing the mutation sites was 100% for 99% of the samples, and greater than 98% for the remaining 1%. The false-positive rate for 99% of the sample mutation sites was zero, with only 1% experiencing a false-positive rate of less than 2%. The false-negative rate for all samples was less than 0.1%.
Claims
1. A method for detecting low-frequency mutations based on hybrid sequencing, characterized in that: The method includes the following steps: constructing one or more mixed pools containing genetic material of part of the samples according to the total number of samples to be tested and the target mutation frequency, obtaining sequencing data of each mixed pool, using a threshold adaptive variation detection algorithm based on the variation frequency distribution characteristics of each base site to obtain single nucleotide variations in the mixed pool, and then using a decoding algorithm to decode the variation information of the mixed pool to obtain the variation information of each sample.
2. The method for detecting low-frequency mutations based on hybrid sequencing according to claim 1, characterized in that: The threshold adaptive variation detection algorithm described herein performs threshold judgment based on the variation detection algorithm threshold T. When the ratio of the number of minor alleles to the number of major alleles at a site in a mixed sample exceeds the threshold, the site is determined to have mutated; otherwise, it indicates that no mutation has occurred.
3. The method for detecting low-frequency mutations based on hybrid sequencing according to claim 2, characterized in that: The calculation formula of the mutation detection algorithm threshold T is: Where N is the number of samples in a mixing pool.
4. The method for detecting low-frequency mutations based on hybrid sequencing according to claim 1, wherein: The decoding algorithm includes two steps: deterministic filling and step-by-step elimination inference. The deterministic filling specifically includes the following steps: for the positive variant sites in the mixed pool, the encoding strategy is traversed to find the samples marked as participating in sequencing in the mixed pool. The positive variant sites are possible variant sites of the samples participating in the mixed pool. For the negative variant sites in the mixed pool, the encoding strategy is traversed to find the samples marked as participating in sequencing in the mixed pool. All samples have not mutated at this site. The specific formula is as follows: Where M represents the encoding matrix, x represents the reconstruction matrix of the sample to be tested, and y represents the mutation matrix of the mixing pool.
5. The method for detecting low-frequency mutations based on hybrid sequencing according to claim 4, characterized in that: The stepwise elimination inference includes the following steps: checking the variation of all mixed pools containing the same sample at possible sites. If all mixed pools mutate, it is considered that the sample has mutated at the possible variation site; otherwise, it is considered that no variation has occurred. The specific formula is as follows: x i =1if{y i |i∈P j }={1} Among them, x represents the reconstruction matrix of the sample to be tested, y represents the variation matrix of the mixed pool, and P j Indicates the numbers of all samples participating in the j-th mixing pool.
6. The method for detecting low-frequency mutations based on hybrid sequencing according to claim 1, characterized in that: The target mutation frequency is any frequency below 5%.
7. The method for detecting low-frequency mutations based on hybrid sequencing according to claim 1, characterized in that: The number of the mixing pool is 5% to 50% of the total number of samples, and the number of samples in the mixing pool does not exceed 50.
8. The method for detecting low-frequency mutations based on hybrid sequencing according to claim 1, characterized in that: The total sequencing depth of each mixed pool was 10× to 10,000×, and the sequencing depth of a single sample was 5× to 200×.
9. The method for detecting low-frequency mutations based on hybrid sequencing according to claim 1, characterized in that: Sequencing methods include whole-genome sequencing, whole-exome sequencing, or targeted sequencing.
10. The method for detecting low-frequency mutations based on hybrid sequencing according to claim 1, characterized in that: The species of the sample to be tested includes one or more of humans, animals, plants, and microorganisms.