TNGS pollution monitoring method based on sample specific molecular internal reference
By designing sample-specific internal reference nucleic acids, the problem of false positives caused by aerosol contamination in tNGS experiments was solved, and high-sensitivity contamination monitoring and pathogen detection at low contamination levels were achieved. Parallel optimization was carried out to overcome the influence of amplification bias and maintain detection stability.
Patent Information
- Application Number
- CN202510789711.X
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-06-13
- Publication Date
- 2025-09-19
AI Technical Summary
False positives caused by aerosol contamination in tNGS experiments cannot be effectively monitored and optimized by existing methods for parallel contamination and pathogen detection. In addition, traditional internal controls are susceptible to amplification bias in multiplex PCR, resulting in decreased detection stability and sensitivity.
A sample-specific internal reference nucleic acid containing a complementary chain was designed. The length was 80-120 bp and included a universal primer amplification region, a sample tag sequence region, and a backbone sequence region. Multiple PCR reactions and sequencing analysis were performed. Contamination was determined by the proportion of the internal reference nucleic acid sequence, with a threshold of 0.1% and an absolute read number ≥10. Amplification bias and synthesis end contamination were taken into account, and primer concentrations were optimized to balance detection.
It achieves contamination monitoring at the one thousandth level, is compatible with tNGS technology, overcomes the influence of amplification bias, maintains pathogen detection sensitivity, and is superior to traditional methods.
Smart Images

Figure CN120666010A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to a tNGS contamination monitoring method based on sample-specific molecular internal references, belonging to the technical field of molecular biology detection technology. Background Art
[0002] Targeted metagenomic sequencing (tNGS) uses multiplex PCR to enrich specific genetic regions of pathogens in samples, combined with high-throughput sequencing to achieve highly sensitive detection of low-abundance pathogens. Due to its high efficiency and specificity, this technology has been widely used in the diagnosis of clinical infectious diseases (such as bloodstream infections and respiratory infections) and in environmental microbial analysis. However, the accuracy of tNGS detection is highly dependent on the standardization of experimental procedures, especially aerosol contamination generated during the multiplex PCR amplification process, which can lead to cross-contamination of nucleic acid fragments between different samples and cause false-positive results.
[0003] In the tNGS experimental process, aerosolization of amplification products is a major source of contamination. Trace amounts of nucleic acid fragments (as low as one thousandth) can be transferred to other samples through laboratory procedures (such as opening the cap and pipetting). Due to the extremely high sensitivity of sequencing, such contamination can easily be misinterpreted as a true pathogen signal. tNGS relies on multiple PCR primers to enrich target regions. Differences in amplification efficiency (i.e., amplification bias) between primers can lead to the selective amplification of contaminating fragments in specific samples, further increasing the risk of false positives.
[0004] Existing contamination control methods include physical isolation and standardized operations. While compartmentalized operations (e.g., separating pre-amplification and post-amplification areas) can reduce the probability of contamination, they cannot completely eliminate aerosol contamination and rely on the rigor of personnel. Furthermore, negative control (NTC) methods only reflect background contamination from reagents or the environment and cannot pinpoint the source of cross-contamination between specific samples.
[0005] For example, Karius' patent (US9976181B2) adds sample-specific internal controls (UICs or ID spikes) to the sample, predicting cross-contamination by detecting unusual distributions of UIC tags in sequencing results. However, mNGS is an unbiased sequencing technology that doesn't rely on PCR amplification, so UIC detection is unaffected by amplification bias. However, because tNGS involves multiplex PCR, UIC amplification efficiency can be affected by factors such as primer competition and template concentration. This can lead to traditional UIC designs being unable to reliably detect or even interfering with pathogen signals in tNGS. Furthermore, existing molecular internal controls often use chemically synthesized double-stranded DNA (dsDNA), but residual single-stranded DNA (ssDNA) or unpurified fragments from the synthesis process can become a source of contamination, further complicating the experiment. Some protocols attempt to balance detection signals by increasing or decreasing the amount of internal controls added, but excessive addition can increase sequencing data volume and reduce pathogen detection sensitivity, while insufficient addition can make contamination monitoring difficult.
[0006] Therefore, how to design a molecular internal reference compatible with the tNGS amplification process that can resist the influence of amplification bias and avoid contamination at the synthesis end, and how to achieve parallel optimization of contamination monitoring and pathogen detection in the multiplex PCR process without relying on the adjustment of the amount of internal reference added, so as to maintain detection stability at extremely low contamination levels, are technical problems that need to be solved. Summary of the Invention
[0007] The technical problem to be solved by the present invention is to solve the false positive problem caused by aerosol contamination in tNGS experiments and to achieve contamination monitoring at the thousandth level without affecting the sensitivity of pathogen detection.
[0008] An internal reference nucleic acid for monitoring sample contamination in tNGS amplification comprises a complementary chain with universal primer amplification regions at both ends and a sample tag sequence region and a backbone sequence region in the middle.
[0009] The nucleic acid is 80-120 bp in length.
[0010] The length of the sample tag sequence region is 20-30 bp, the length of the backbone sequence region is 30-40 bp, and the length of the primer amplification region is 18-22 bp.
[0011] A tNGS contamination monitoring method based on sample-specific molecular internal references comprises the following steps: Prepare at least two sample solutions for tNGS, add different internal reference nucleic acids to each solution, and then perform multiplex PCR reactions; After the reaction products are purified, they are subjected to library construction and NGS sequencing. The data is then analyzed to determine the content of different internal reference nucleic acids in the sequence results of each sample. If a sample contains sequences of internal reference nucleic acids that were not added during sample preparation, it is considered that the sample is contaminated.
[0012] When the proportion of the sequence of the unadded internal reference nucleic acid in the number of reads obtained by sequencing reaches a certain threshold, it is considered that the sample is contaminated.
[0013] The threshold is: the proportion of internal reference sequences of the contaminated sample in the contaminated sample is ≥0.1%, and the absolute number of reads is ≥10.
[0014] In the multiplex PCR reaction, the concentration of the amplification primers for the internal reference nucleic acid is 5-16.7 nM, and the concentration of the pathogen detection primers is 50 nM.
[0015] The amount of the internal reference nucleic acid added is 1-5 μL of 50 nM working solution per sample, and the total amount of the internal reference nucleic acid accounts for 0.01%-0.1% of the total DNA in the sample.
[0016] The present invention is compatible with tNGS technology, overcoming amplification bias and primer competition issues. By regulating primer concentration, it enables concurrent optimization of contamination monitoring and pathogen detection. Its sensitivity reaches 1 / 1000th the contamination level, surpassing traditional methods. BRIEF DESCRIPTION OF THE DRAWINGS
[0017] Figure 1 : Schematic diagram of UIC sequence structure.
[0018] Figure 2 : Schematic diagram of sample addition UIC steps.
[0019] Figure 3 : Schematic diagram of UIC test results. The left figure is a schematic diagram of theoretical test results, and the right figure is a schematic diagram of actual test results.
[0020] Figure 4 : Results of simulated pollution identification experiments. DETAILED DESCRIPTION
[0021] The method of the present invention is suitable for the detection of respiratory samples by tNGS. The molecular internal reference designed by this method has been verified and can have the sensitivity to specifically mark sample contamination in the detection of respiratory samples containing common bacteria, viruses, fungi, etc.
[0022] Example 1 (Design and Preparation of UIC): UIC sequence design principles: Figure 1The 99-bp sequence contains universal primers (for mPCR amplification), differentially differentiated sequences (as sample tags), and a backbone sequence (to maintain stability). Double-stranded DNA (dsDNA) is formed by annealing two completely reverse-complementary single-stranded DNAs (ssDNA) to avoid contamination at the synthetic ends.
[0023] The specific design of universal primer sequences (such as length, GC content, and design principles to avoid cross-reaction with non-target sequences) includes: the primer length should be controlled within 18-22bp to ensure a balance between amplification efficiency and specificity; the GC content should be adjusted to 40%-60% to prevent abnormal annealing efficiency caused by high GC regions (>65%) or low GC regions (<30%). The Tm value difference between primer sets should be ≤2°C, and the adjacent thermodynamic model (Breslauer algorithm) is used for accurate calculation, and the formula is T m =64.9 plus 41 multiplied by (number of G+C bases minus 16.4) divided by the primer length. Prefer G / C bases at the 3' end of the primer to enhance binding stability (e.g., designing the last two nucleotides as a GC clamp). Avoid four consecutive identical bases or reverse complementary sequences to prevent primer dimers. Use the OligoAnalyzer tool to assess and eliminate hairpin structures (ΔG value ≥ -3 kcal / mol). BLAST alignments were used to exclude homology to human genome and pathogen databases (consecutive base matches ≤ 12 bp, E-value ≤ 1e-5).
[0024] Possible design principles for generating differentially differentiated sequences (e.g., random sequence libraries or pre-defined encoding systems) include: 1. Based on random sequence library design, synthesize an oligonucleotide pool containing 20-30 bp of random bases to ensure the uniqueness of each sample tag; 2. Use a pre-defined encoding system, such as converting the sample ID into a binary code and mapping it to base pairs (e.g., 00=A, 01=T, 10=C, 11=G), and adding a check bit (e.g., Hamming code) to enable single-base error correction. Both strategies must meet the following requirements: The random or coded sequence must be aligned to exclude homology with the human genome, common pathogens, and experimental primers (consecutive matches ≤ 12 bp, E-value ≤ 1e-5); avoid four consecutive identical bases or reverse complementary structures (e.g., palindromes) within the sequence, and experimentally verify the consistency of amplification efficiency (Ct value difference ≤ 1.5 between different tags). In addition, the preset code needs to reserve redundant bits (such as increasing the total length by 20%) to cope with label recognition failures caused by sequencing errors or amplification deviations, ensuring accurate traceability at a contamination level of one thousandth.
[0025] Table 1 shows 5 example sequences:
[0026] The designed UIC sequence was subjected to Blast analysis, which confirmed that it could not be aligned to any species in the database (analysis date 2025.05.20).
[0027] UIC preparation steps: ssDNA synthesis, annealing conditions (temperature gradient, time).
[0028] The designed UIC sequence and its corresponding reverse complementary sequence were delivered to a biosynthesis company for synthesis. The synthesis amount was 5 nmol, the purification method was HPLC purification, and the delivery form was dry powder delivery. In a clean bench, add 50 μL of IDT's Duplex buffer to dilute the single-chain UIC stock solution to 100 μM. The two single-stranded UIC stock solutions were added to 0.2 mL PCR tubes in sequence and placed in a PCR instrument for the following reaction to obtain 100 μL of 50 μM double-stranded UIC stock solution.
[0029] Table 2 Double-strand UIC annealing reaction program
[0030] Pipette 1 μL of double-chain UIC stock solution into a 1.5 mL centrifuge tube containing 999 μL of Duplex buffer and vortex to mix to obtain 50 nM UIC working solution.
[0031] Example 2 (Verification of pollution monitoring): To simulate different contamination scenarios (such as cross-contamination between samples and high-concentration sample contamination) and demonstrate the correlation between UIC detection rate and contamination ratio, the following experiments were conducted: Sample preparation: Three clinically collected bronchoalveolar lavage fluid samples were selected and verified by previous mNGS results, including one high-pathogen content sample (sample 1), one medium-pathogen content sample (sample 2), and one low-pathogen content sample (sample 3).
[0032] Nucleic acid extraction: according to the attached Figure 2 In the example, 1 μL of different UIC working solutions were added to the original sample, and the sample nucleic acid was extracted using the magnetic bead method.
[0033] Multiplex PCR: Prepare the tNGS reaction system according to Table 3, where the universal primer sequences for UIC amplification are SEQ ID NO. 5-6: F- ACACTCTTTCCCTACACGACCGCTTCCGATCTGCTTATAGACCGCCATAGCCCTCTAG; R-GACTGGAGTTCAGACGTGTGCTCTTCCGATCTTACTCGACCGTCACCGATCC.
[0034] Place the reaction tube in a PCR instrument and run the reaction program in Table 4.
[0035] Table 3 tNGS multiplex PCR reaction system
[0036] Table 4 tNGS multiplex PCR reaction procedure
[0037] Purification: The mPCR product was purified using 0.9× magnetic beads and eluted with 50 μL of enzyme-free water. 1 μL of the purified product was then aspirated and the concentration was determined using Qubit.
[0038] Simulated contamination: To verify the purpose of UIC monitoring cross-contamination, a cross-contamination simulation was performed. 1‰ volume of mPCR product was taken to contaminate adjacent samples. The specific operation steps are as follows: 1 μL of mPCR product from each of the three samples was placed in 19 μL of enzyme-free water (diluted 20 times). After mixing, 1 μL of the diluted solution of sample 1 was placed in the original solution of sample 2, the diluted solution of sample 2 was placed in the original solution of sample 3, and the diluted solution of sample 3 was placed in the original solution of sample 1.
[0039] Library amplification: Add sequencing adapters according to the reaction system in Table 5 and the reaction procedure in Table 6 to obtain the library to be sequenced.
[0040] Table 5 tNGS second round amplification reaction system
[0041] Table 6 tNGS second round amplification reaction system
[0042] Library purification: Sequencing libraries were purified using 0.7× magnetic beads. After quantification and pooling, SE50 sequencing was performed on a DIFSEQ-200 instrument (Difei Medical Devices Co., Ltd.) with a preset data volume of 1 million reads.
[0043] Off-line quality control analysis: Table 7 shows basic quality control information for the sequenced samples. It can be seen that for sample 3, which has a low pathogen content, the UIC sequence proportion reaches 93.2%, which may have some impact on the test results. The sequence proportion refers to the proportion of reads.
[0044] Table 7 Quality control information of UIC contamination simulation test samples
[0045] UIC and Pathogen Detection Analysis: Table 8 compares the results of characteristic pathogen detection and UIC detection for three samples. After simulated contamination, each sample contained contaminated reads, including those for characteristic pathogens and UIC contamination. Ratio analysis revealed that the UIC contamination ratio was similar to that for pathogen contamination, thus validating the UIC's ability to monitor cross-contamination between samples.
[0046] Table 8 Analysis of UIC pollution simulation detection results
[0047] Example 3 (UIC primer concentration optimization): In Example 2, the UIC primer concentration was 50 nM. The proportion of UIC detection reads was too high in samples with low pathogen content, potentially leading to a decrease in pathogen detection reads or even the risk of missed detection. To further balance the proportion of pathogen sequences and UIC sequences in different samples, the amount of UIC primers used in the multiplex PCR process was optimized.
[0048] (1) Sample preparation: The experiment selected two samples with high pathogen content (samples 3 and 5), one sample with medium pathogen content (sample 4), two samples with low pathogen content (samples 1 and 2), and one blank control sample (NTC).
[0049] (2) Nucleic acid extraction: Add 1 μL of different UIC working solutions to the sample tubes, and use the magnetic bead method to extract the sample nucleic acid.
[0050] (3) Multiplex PCR: The multiplex PCR reaction solution was prepared according to Table 3. The difference was that the amount of UIC primers added during this optimization process was reduced to 16.7 nM, 5 nM, and 1.67 nM, respectively, thereby reducing the proportion of UIC in the total detected reads. The multiplex PCR reaction was then performed according to the procedure shown in Table 4.
[0051] (4) Subsequent experiments: Perform multiplex PCR product purification, library amplification, library purification, sequencing, and other operations as in Example 2.
[0052] (5) Off-line quality control data: As shown in Table 9, the proportion of UIC sequences in the sample off-line results increases with the increase in the amount of UIC primers used, while the proportion of pathogenic sequences decreases. Therefore, adjusting the amount of UIC primers used can effectively change the proportion of UIC sequences in the sequencing results.
[0053] Table 9 Quality control data of samples with optimized UIC primer concentration
[0054] (6) UIC and pathogen detection data: Appendix Figure 4 The detection results for all samples in this experiment are shown. The circle size in the figure represents the number of detected reads, blue represents the original detection, and green represents the result of simulated contamination. The results show that the pathogen contamination ratio and UIC contamination ratio are basically consistent under different UIC primer dosages. At the same time, in highly pathogenic samples, a UIC primer dosage of 5 nM can effectively detect a sufficient number of UIC sequences to distinguish a contamination rate of 1‰, while a primer dosage of 1.67 nM carries the risk of missing contaminated UICs. Therefore, to balance the sensitivity of tNGS for pathogen detection and the performance of UIC in predicting contamination, a UIC primer dosage of 5 nM was selected.
Claims
1. An internal reference nucleic acid for monitoring sample contamination during tNGS amplification, comprising a complementary strand with universal primer amplification regions at both ends and a sample tag sequence region and a backbone sequence region in the middle.
2. The internal reference nucleic acid according to claim 1, characterized in that The nucleic acid is 80-120 bp in length.
3. The internal reference nucleic acid according to claim 1, characterized in that The length of the sample tag sequence region is 20-30 bp, the length of the backbone sequence region is 30-40 bp, and the length of the primer amplification region is 18-22 bp.
4. A tNGS contamination monitoring method based on sample-specific molecular internal references, characterized in that: The steps include: Prepare at least two sample solutions for tNGS, add different internal reference nucleic acids to each solution, and then perform multiplex PCR reactions; After the reaction products are purified, they are subjected to library construction and NGS sequencing. The data is then analyzed to determine the content of different internal reference nucleic acids in the sequence results of each sample. If a sample contains sequences of internal reference nucleic acids that were not added during sample preparation, it is considered that the sample is contaminated.
5. The tNGS contamination monitoring method based on sample-specific molecular internal reference according to claim 3, characterized in that: When the proportion of the sequence of the unadded internal reference nucleic acid in the number of reads obtained by sequencing reaches a certain threshold, it is considered that the sample is contaminated.
6. The tNGS contamination monitoring method based on sample-specific molecular internal reference according to claim 3, characterized in that: The threshold is: the proportion of internal reference sequences of the contaminated sample in the contaminated sample is ≥0.1%, and the absolute number of reads is ≥10.
7. The tNGS contamination monitoring method based on sample-specific molecular internal reference according to claim 3, characterized in that: In the multiplex PCR reaction, the concentration of the amplification primers for the internal reference nucleic acid is 5-16.7 nM, and the concentration of the pathogen detection primers is 50 nM.
8. The tNGS contamination monitoring method based on sample-specific molecular internal reference according to claim 3, characterized in that: The amount of the internal reference nucleic acid added is 1-5 μL of 50 nM working solution per sample, and the total amount of the internal reference nucleic acid accounts for 0.01%-0.1% of the total DNA in the sample.
Citation Information
Patent Citations
Synthetic nucleic acid spike-ins
US9976181B2