tNGS detection pollution judgment model, method for establishing same and application

By establishing a random forest model based on primer amplification spectrum, the problem of contamination judgment in tNGS technology was solved, achieving high-precision sample contamination judgment. This model is applicable to the detection of pathogens such as Mycobacterium tuberculosis complex, improving the accuracy and efficiency of detection.

CN117352055BActive Publication Date: 2026-04-14GZ VISION GENE TECH CO LTD +5
View PDF 4 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2023-08-30
Publication Date
2026-04-14

AI Technical Summary

Technical Problem

Existing tNGS technology lacks an efficient and reliable method for judging contamination in pathogen detection, especially since nucleic acid extraction of pathogens such as Mycobacterium tuberculosis complex is difficult and PCR contamination is hard to distinguish, resulting in low detection accuracy.

Method used

A random forest model based on primer amplification spectrum was established. By statistically analyzing the distribution characteristics of reliable fragments amplified by each primer in the sample, and combining quantitative analysis and the proportion of strong positives within the batch, the model automatically interprets whether the sample is contaminated and uses the random forest model for judgment.

Benefits of technology

It achieves high-precision contamination detection with an AUC value of up to 0.95, is low-cost, highly versatile, requires no additional reagents or experimental procedures, and significantly improves the sample contamination discrimination.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN117352055B_ABST
    Figure CN117352055B_ABST
Patent Text Reader

Abstract

The present application relates to a tNGS detection pollution judgment model and its establishment method and application, belong to the macro gene group detection technical field. The method is based on the characteristics of tNGS detection, and the tNGS detection pollution judgment model for judging whether the sample is contaminated by strong positive sample is established according to the primer amplification spectrum characteristics, which has the advantages of low cost and high universality without additional special primer reagent requirements. And because the characteristics of historical strong positive samples, batch strong positive samples and strain differences are considered comprehensively, the method has the advantages of more accurate judgment and higher discrimination.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to the field of metagenomic detection technology, and in particular to a tNGS detection contamination judgment model, its establishment method, and its application. Background Technology

[0002] Currently, targeted next-generation sequencing (tNGS) technology is gaining increasing attention and application in the field of pathogen detection. Compared to metagenomic next-generation sequencing (mNGS), tNGS is unaffected by host ratio and exhibits high sensitivity. Furthermore, it possesses high specificity due to the use of specific primers to selectively amplify DNA or RNA fragments of interest from pathogens. In addition, because it only measures the fragments of interest, sequencing efficiency is high; only about 0.5M-1M of sequence is needed per sample to meet detection requirements, and a single chip can simultaneously detect over 100 samples, resulting in high throughput. By designing different primers, different types of microbial sequences can be selectively amplified, enabling diverse detection. Moreover, sequence alignment and analysis can determine the species and subtype information of microorganisms, thereby facilitating the tracing of pathogens and epidemiological studies.

[0003] For example, CN114974427A discloses a multi-target detection primer set and its design method. In order to improve detection sensitivity (obtain a lower detection limit), the primer panel is designed by selecting multi-copy regions as much as possible. That is, the prior frequencies obtained by statistical analysis of a large number of clinical samples are used. Each target species includes conserved multi-copy amplification primers of clinical strains (the primer design region sequence of different copies is the same, but the middle part sequence of the target may be different), as well as some non-multi-copy primers and drug resistance point mutation primers. This ensures that the amplification fragment of each target species retains some strain-specific characteristics (rich target fragment types) as much as possible while ensuring primer conservatism and high efficiency.

[0004] However, despite the high flexibility in target selection, detection sensitivity, and throughput of tNGS technology, the ability to identify contamination from nucleic acid extractions from non-sample sources within the same batch or from cross-batch PCR aerosols remains a significant challenge for the large-scale application of this technology.

[0005] Especially for important pathogens such as Mycobacterium tuberculosis complex, the thick cell walls of these microorganisms make nucleic acid extraction difficult. Furthermore, clinical requirements necessitate the detection of drug resistance point mutations, necessitating the design of multiple primer pairs to improve detection sensitivity and simultaneously meet the need for drug resistance point mutation detection. Simultaneously, the relatively slow growth rate of Mycobacterium tuberculosis and varying infection levels lead to significant fluctuations in viral load across different samples. This results in high-load samples being highly susceptible to PCR contamination, while low-load samples are difficult to distinguish from PCR contamination. Using sequence number or quantitative threshold methods during analysis can result in a large gray area and low accuracy.

[0006] Based on the above, various solutions have been developed in the existing technology, such as the one disclosed in CN115058490A, which uses a one-round PCR technique to avoid generating barcode-free PCR contamination, aiming to reduce cross-batch aerosol contamination. However, with this one-round PCR technique, the generated aerosol contamination can replace the barcode in the system after two PCR cycles, and the sensitivity is low, so it is not very effective. More commonly, experimental protocols use methods such as removing residual nucleic acid contamination from the environment with nucleic acid-degrading enzymes, improving ventilation equipment, and using separate operating rooms to prevent contamination. However, while experimental cleanup and prevention methods are effective for small sample volumes, their effectiveness is very limited for high-throughput, frequent applications.

[0007] Therefore, there is an urgent need for an efficient and reliable method for quality control of tNGS detection and for assessing sample contamination. Summary of the Invention

[0008] Therefore, it is necessary to provide a contamination judgment model for tNGS detection, which addresses the lack of efficient and reliable contamination monitoring methods. This tNGS contamination judgment model is based on features such as amplification spectrum and can automatically interpret and judge the contamination status of samples without the need for additional special primers, reagents, or detection requirements. It can not only accurately judge the contamination status of samples through data analysis, but also has the advantages of low cost and high versatility.

[0009] A method for establishing a tNGS-based contamination detection model includes the following steps:

[0010] Sample preparation: Take a number of samples that have detected the predetermined pathogen in the batch tNGS detection, and perform a retest on the samples. Samples with a positive retest result are defined as true positive samples, and samples with a negative retest result are defined as contaminated samples.

[0011] Feature acquisition: The distribution characteristics of reliable fragments without sequencing errors amplified by each primer for the target pathogen in each sample detected by the chip used in the above-mentioned tNGS detection and in the laboratory within a predetermined time period were statistically analyzed. The distribution characteristics include smrn and quant; smrn is the total number of specific sequences of the target pathogen detected in the sample and aligned to the target pathogen, and quant is the quantitative copy number of the target pathogen detected in the sample, obtained by correcting smrn with a fixed amount of internal reference nucleic acid.

[0012] Model establishment: Using the distribution features as the feature matrix X and the verification results as the model result Y, a random forest model is established and trained to obtain the tNGS detection contamination judgment model used to distinguish between positive samples and contaminated samples.

[0013] Based on thorough preliminary research and investigation, and taking into account the characteristics of tNGS detection, the inventors proposed a method for establishing an automated interpretation model for determining whether a sample is contaminated by a strongly positive sample based on primer amplification spectrum. In principle, since each sample differs in factors such as the strain of the predetermined pathogen (e.g., Mycobacterium tuberculosis complex), gene expression, gene mutation, and extraction fluctuations, the primer amplification spectrum of a contaminated sample should be consistent with that of a strongly positive sample, while that of a genuine pathogen sample may differ to some extent.

[0014] Furthermore, based on experience, PCR contamination in the laboratory, or contamination from extraction on the same day, will largely degrade within a certain timeframe (e.g., about three days). Therefore, by statistically analyzing the distribution of reliable fragments (fragments without sequencing errors) amplified by each primer for the predetermined pathogen in each sample from the laboratory within a certain timeframe, and combining this with quantification, the proportion of strong positives within the batch, and other characteristics, a random forest model can be used to determine whether the predetermined antigen sequence detected in the sample is contamination or from the sample's true origin.

[0015] In one embodiment, the distribution feature further includes a primer, which is the number of primer pairs that are effectively amplified in the sample.

[0016] In one embodiment, the distribution features also include s2max, q2max and pimps, where s2max is the proportion of the sample’s smrn to the highest smrn in the same detection chip;

[0017] The q2max is the ratio of the quant of the sample to the highest quant in the same detection chip.

[0018] The pimps refers to the number of non-contaminated primer amplification fragment types compared to all samples within a predetermined time period.

[0019] In one embodiment, the uncontaminated amplification primer sequence is obtained by comparing the sample with the sequences amplified by the same primers in all other samples within a predetermined time period. If the number of sequences amplified by the sample is not less than 5% of the number of sequences amplified by the same primers in other samples, then the sequence amplified by the amplification primers in the sample is considered an uncontaminated amplification primer sequence. Further, if the number of sequences amplified by the sample is not less than 20% of the number of sequences amplified by the same primers in other samples, then the sequence amplified by the amplification primers in the sample is considered an uncontaminated amplification primer sequence.

[0020] In one embodiment, the scheduled time is 2-5 days. Further, the scheduled time is 3-4 days.

[0021] In one embodiment, the reliable fragment is obtained by comparing the amplified sequence fragment with the highest-read fragment amplified by the same primers in the sample. If the read count of the amplified sequence fragment is not less than 1% of the read count of the highest-read fragment, then the sequence fragment is considered a reliable fragment without sequencing errors. Further, if the read count of the amplified sequence fragment is not less than 5% of the read count of the highest-read fragment, then the sequence fragment is considered a reliable fragment without sequencing errors.

[0022] In one embodiment, the predetermined pathogen is Mycobacterium tuberculosis complex. Because the cell walls of Mycobacterium tuberculosis complex are relatively thick, nucleic acid extraction is difficult. Furthermore, clinical requirements necessitate the detection of drug resistance point mutations, generally requiring the design of multiple primer pairs to improve the detection sensitivity of this pathogen and simultaneously meet the detection requirements for drug resistance point mutations. Simultaneously, due to the relatively slow growth rate of Mycobacterium tuberculosis and varying degrees of infection, the viral load of different samples fluctuates significantly, leading to situations where high-load samples are highly susceptible to PCR contamination, while low-load samples are difficult to distinguish from PCR contamination. The contamination monitoring and judgment method of this invention has significant advantages in analyzing detected Mycobacterium tuberculosis complex. However, it is understood that the tNGS detection contamination judgment model obtained above is not limited to Mycobacterium tuberculosis complex; this model strategy is also applicable to other microorganisms.

[0023] The present invention also discloses a tNGS detection contamination judgment model, including a computer-readable storage medium storing the above-mentioned tNGS detection contamination judgment model data.

[0024] The present invention also discloses a tNGS detection contamination judgment method, comprising: obtaining the distribution characteristics of the sample to be judged, substituting them into the above-mentioned tNGS detection contamination judgment model for analysis and judgment, and obtaining the model judgment result.

[0025] This invention also discloses a tNGS contamination detection and judgment device, comprising:

[0026] The data storage module is used to store the data information of the above-mentioned tNGS detection contamination judgment model;

[0027] The data analysis module is used to perform analysis according to the methods described above; and

[0028] The data display module is used to output and display the judgment results of the model.

[0029] The present invention also discloses a computer device, including a memory and a processor, wherein the memory stores a computer program, and the processor executes the computer program to implement the above-described method.

[0030] Compared with the prior art, the present invention has the following beneficial effects:

[0031] This invention discloses a method for establishing a tNGS detection contamination judgment model. Based on the characteristics of tNGS detection, it establishes an automatic interpretation model for determining whether a sample is contaminated by a strongly positive sample using primer amplification spectrum. The obtained tNGS detection contamination judgment model has an AUC value of 0.83 on the validation set, and under optimized conditions, its AUC can reach 0.90, or even 0.95.

[0032] Furthermore, the tNGS contamination detection model established using the above method not only requires no additional special primers or reagents, resulting in lower costs and higher versatility, but also, by comprehensively considering characteristics such as historical strongly positive samples, batch-specific strongly positive samples, and strain differences, offers advantages in more accurate judgment and higher discrimination. Attached Figure Description

[0033] Figure 1 This is a schematic diagram of the ROC curves obtained from different models in Example 2. Detailed Implementation

[0034] To facilitate understanding of the present invention, a more complete description will be given below with reference to the accompanying drawings. Preferred embodiments of the invention are shown in the drawings. However, the invention can be implemented in many different forms and is not limited to the embodiments described herein. Rather, these embodiments are provided to provide a thorough and complete understanding of the disclosure of the invention.

[0035] Unless otherwise defined, all technical and scientific terms used herein have the same meaning as commonly understood by one of ordinary skill in the art to which this invention pertains. The terminology used herein in the specification of this invention is for the purpose of describing particular embodiments only and is not intended to be limiting of the invention. The term "and / or" as used herein includes any and all combinations of one or more of the associated listed items.

[0036] Unless otherwise specified, all reagents used in the following examples are commercially available; and all methods used in the following examples are conventional methods unless otherwise specified.

[0037] Example 1

[0038] A tNGS detection contamination judgment model is described in this embodiment, taking the judgment model of Mycobacterium tuberculosis complex contamination as an example. It is specifically established through the following method.

[0039] 1. Sample preparation.

[0040] A total of 483 samples with Mycobacterium tuberculosis complex-specific sequences were obtained through a period of sample collection. These included 366 bronchoalveolar lavage fluid samples, 38 blood samples, 15 sputum samples, 11 cerebrospinal fluid samples, and 53 other sample types, including pleural effusion, pericardial effusion, lymph node puncture fluid, drainage fluid, pus, and lung tissue.

[0041] Sequencing was performed using tNGS on 75 different Illumina sequencing chips, employing a single-end 40bp sequencing mode. Samples showing specific detection of Mycobacterium tuberculosis complex were either re-extracted using capture technology for nucleic acid verification, or re-extracted and re-verified using tNGS. A positive result from either technique indicates that the tuberculosis in the sample originated from its true source, rather than being contaminated (this can be independently verified and defined as true), thus constituting a true positive sample. Samples that could not be verified were classified as contaminated samples (this could also be due to randomness near the detection limit, uniformly defined as false positives).

[0042] The final results showed 354 positive samples and 129 contaminated samples.

[0043] 2. Feature acquisition.

[0044] The distribution characteristics of reliable fragments without sequencing errors obtained by each primer amplification of Mycobacterium tuberculosis complex in each of the 483 samples tested using the chip and in each sample tested in the laboratory on the day of and the three days prior to the chip were statistically analyzed.

[0045] Since sequencing errors are generally below 1%, a fragment is considered error-free if it accounts for at least 5% of the highest-read fragment amplified by the same primers within the sample. Therefore, reliable fragments are obtained by comparing the amplified sequence fragment with the highest-read fragment amplified by the same primers within the sample. If the read count of the amplified fragment is at least 5% of the read count of the highest-read fragment, then the fragment is considered error-free.

[0046] Based on the above distribution characteristics, calculate the following 6 characteristic parameters for future reference.

[0047] (1) smrn (number of specific sequences aligned): The total number of specific sequences aligned to the Mycobacterium tuberculosis complex in each sample.

[0048] (2) quant (copy number): The quantitative copy number of Mycobacterium tuberculosis complex in each sample is obtained by correcting smrn with a fixed amount of internal reference nucleic acid as a reference.

[0049] (3) Primers: Within a sample, some primer pairs can amplify multiple reliable fragments. These different reliable fragments represent multiple different copies, pathogens carrying point mutations and pathogens not carrying point mutations within the same sample. Primers represent the number of primer pairs that can be effectively amplified in each sample. For example, if a pair of primers effectively amplifies both mutated and non-mutated pathogens, even if the number of effectively amplified fragments is ≥2, the primers here refer to the number of primer pairs that can be effectively amplified, which is still counted as 1.

[0050] (4) s2max (same chip strong anode ratio): the ratio of the sample’s smrn to the highest smrn in the same chip.

[0051] (5) q2max (proportion of strong positives in the same chip): the proportion of the sample’s quant to the highest quant in the same chip.

[0052] (6) pimps (lowest number of uncontaminated primer sequences): If the number of sequences amplified by a primer in sample A is not less than 20% of the number of sequences amplified by the same primer in sample B, then the source of the amplification by that primer in sample A is unlikely to be contaminated by sample B (the contamination rate is generally less than 5%). Therefore, compared to B, this sequence in sample A is an uncontaminated primer sequence. For example, if one primer pair amplifies two reliable fragments, both of which are uncontaminated sequences, then the number of uncontaminated primer amplification fragments is 2. Count the number of uncontaminated primer amplification fragments in sample A, and use the lowest number of uncontaminated primer amplification fragments as the pimps of sample A.

[0053] For each sample, the amplification spectra of each primer for the Mycobacterium tuberculosis complex in each sample of the laboratory were statistically compared on the day the chip was used and in the three days prior. Based on this, the statistical characteristics of each sample were calculated: the number of specific aligned sequences (smrn), the number of copies (quant), the number of primer pairs (primers), the proportion of strong positives on the same chip (s2max), the proportion of quantitative strong positives on the same chip (q2max), and the number of the lowest non-contamination amplification primer sequence types (pimps).

[0054] Each sample was only included in the statistical features calculated by the sequencing chip it belonged to on that day, meaning that a total of 6 features were obtained for each of the 483 samples. Some of the results are shown in the table below.

[0055] Table 1. Example of statistical data for a single chip

[0056]

[0057]

[0058] Table 1 is a schematic representation of the features of a chip. The primer names are the names of the primers used for amplification. The information corresponding to each primer pair represents the target site and its number in the panel. For example, Mycobacterium_tuberculosis|632 means that the target site of this primer is Mycobacterium_tuberculosis and its number in the panel is 632.

[0059] The sequencing sequences were those used during SE 40bp sequencing. 5M31905, 5M31881, 5M31700, 5M31722, 5M31770, 5M31776, 5M31790, 5M31850, 5M31855, and 5M31857 are sample numbers. The sample number corresponding to each sequencing sequence is the number of reads obtained from amplification sequencing.

[0060] Table 2. Distribution characteristics obtained from statistics

[0061]

[0062] Table 2 above shows the distribution characteristics of samples 5M31905, 5M31881, 5M31700, 5M31722, 5M31770, 5M31776, 5M31790, 5M31850, 5M31855, and 5M31857.

[0063] 3. Model establishment.

[0064] The distribution characteristics are used as the input feature matrix X, and the verification results are used as the model result Y. A random forest model is built using the sklearn package for training, thus obtaining the tNGS detection and judgment model for distinguishing between positive and contaminated samples.

[0065] In this embodiment, 338 samples, representing 70% of the 483 samples mentioned above, were randomly selected as the training set. Among these, 247 samples were positive (able to repeatedly detect positive results), 91 samples were negative (unable to repeatedly detect positive results), and the remaining 30% of the samples were used as the validation set for subsequent testing.

[0066] The following three strategies were adopted to train and establish a tNGS model for detecting contamination.

[0067] 1) Establish a model using smrn and quant as feature matrices X.

[0068] 2) Establish a model using smrn, quant, and primer as feature matrices X.

[0069] 3) Establish a model using smrn, quant, primer, s2max, q2max and pimps as feature matrices X.

[0070] Example 2

[0071] A method for judging contamination by tNGS detection is proposed, which uses three models established in Example 1 for judgment.

[0072] 1. Sample source.

[0073] The remaining 30% of the 483 samples in Example 1 were used as the samples for this experiment, including 107 positive samples and 38 negative samples. The ROC curve and AUC value were calculated based on the verification results.

[0074] 2. Results.

[0075] The results are as follows Figure 1 As shown, the results indicate that when the model is built using smrn and quant as feature matrices X, the AUC of the resulting random forest model is only 0.83. When the model is built using smrn, quant, and primer as feature matrices X, the AUC of the resulting random forest model is 0.90. However, when the model is built using smrn, quant, primer, s2max, q2max, and pimps as feature matrices X, the AUC of the resulting random forest model can reach 0.95 (sensitivity 98.6%, specificity 91.5%).

[0076] The above experiments demonstrate that the tNGS contamination detection model based on amplified spectral features constructed in this invention has higher accuracy than commonly used methods such as sequence number or quantitative threshold, and has the advantage of low cost without the need for additional reagents or new experimental procedures.

[0077] The technical features of the above embodiments can be combined in any way. For the sake of brevity, not all possible combinations of the technical features in the above embodiments are described. However, as long as there is no contradiction in the combination of these technical features, they should be considered to be within the scope of this specification.

[0078] The embodiments described above are merely illustrative of several implementations of the present invention, and while the descriptions are relatively specific and detailed, they should not be construed as limiting the scope of the invention patent. It should be noted that those skilled in the art can make various modifications and improvements without departing from the concept of the present invention, and these all fall within the protection scope of the present invention. Therefore, the protection scope of this invention patent should be determined by the appended claims.

Claims

1. A method for establishing a tNGS-based contamination detection and judgment model, characterized in that, Includes the following steps: Sample preparation: Take several samples that have been detected by the predetermined pathogen in the laboratory batch tNGS test, and perform a retest on the samples. Samples with a positive retest result are defined as true positive samples, and samples with a negative retest result are defined as contaminated samples. Feature acquisition: The distribution characteristics of reliable fragments without sequencing errors amplified by each primer for the target pathogen in each sample detected by the chip used in the above-mentioned tNGS detection and in the laboratory within a predetermined time period were statistically analyzed. The distribution characteristics include smrn and quant; smrn is the total number of specific sequences of the target pathogen detected in the sample and aligned to the target pathogen, and quant is the quantitative copy number of the target pathogen detected in the sample, obtained by correcting smrn with a fixed amount of internal reference nucleic acid. Model establishment: Using the distribution features as the feature matrix X and the verification results as the model result Y, a random forest model is established and trained to obtain the tNGS detection contamination judgment model used to distinguish between positive samples and contaminated samples.

2. The method for establishing a tNGS-based contamination detection model according to claim 1, characterized in that, The distribution characteristics also include a primer, which is the number of primer pairs that are effectively amplified in the sample.

3. The method for establishing a tNGS-based contamination detection model according to claim 2, characterized in that, The distribution characteristics also include s2max, q2max and pimps, where s2max is the proportion of the sample’s smrn to the highest smrn in the same detection chip. The q2max is the ratio of the quant of the sample to the highest quant in the same detection chip. The pimps is the lowest number of non-contaminated primer amplification fragment types compared to all samples within a predetermined time period.

4. The method for establishing a tNGS-based contamination detection model according to claim 3, characterized in that, The sequence of the non-contamination primer amplification fragment is obtained by comparing the sample with the sequences amplified by the same primers in all other samples within a predetermined time. If the number of sequences amplified by the sample is not less than 5% of the number of sequences amplified by the same primers in other samples, then the sequence amplified by the amplification primers in the sample is considered to be the sequence of the non-contamination primer amplification fragment.

5. The method for establishing a tNGS-based contamination detection model according to claim 1, characterized in that, The scheduled time is 2-5 days.

6. The method for establishing a tNGS-based contamination detection model according to claim 1, characterized in that, The reliable fragment is obtained by comparing the amplified sequence fragment with the highest sequence fragment amplified by the same amplification primer in the sample. If the number of reads of the amplified sequence fragment is not less than 1% of the number of reads of the highest sequence fragment, then the sequence fragment is considered to be a reliable fragment without sequencing errors.

7. The method for establishing a tNGS-based contamination detection model according to claim 1, characterized in that, The predetermined pathogen is the Mycobacterium tuberculosis complex.

8. A tNGS method for determining contamination, characterized in that, include: The distribution characteristics of the sample to be judged are obtained, and then substituted into the tNGS detection pollution judgment model established by the method described in any one of claims 1-7 for analysis and judgment, so as to obtain the model judgment result.

9. A tNGS contamination detection and judgment device, characterized in that, include: Data storage module, used to store data information of the tNGS detection pollution judgment model established by the establishment method according to any one of claims 1-7; The data analysis module is used to perform analysis according to the tNGS detection contamination judgment method described in claim 8; as well as The data display module is used to output and display the judgment results of the model.

Citation Information

Patent Citations

  • Multi-target detection primer group and design method thereof

    CN114974427A

  • Primer combination for constructing microorganism targeted sequencing library and application thereof

    CN115058490A

  • Method for judging background introduced microorganism sequence and application thereof

    CN113270145A

  • Digital Counting of Individual Molecules by Stochastic Attachment of Diverse Labels

    US20110160078A1