Molecular identification tag for identifying cross contamination of biological samples in tNGS, reagent containing molecular identification tag and application of molecular identification tag

By introducing molecular recognition tags and paraffin oil physical isolation technology into the tNGS experiment, the problem of cross-contamination in the nucleic acid extraction stage is solved, the false positive rate is reduced, and the reliability and efficiency of detection are improved, making it suitable for clinical application.

CN121629028APending Publication Date: 2026-03-10SANSURE BIOTECH INC
View PDF 0 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-11-25
Publication Date
2026-03-10

AI Technical Summary

Technical Problem

Cross-contamination of biological samples in tNGS experiments is difficult to control effectively, especially in the nucleic acid extraction stage, leading to false positive results and diagnostic reliability issues. Existing technical measures suffer from drawbacks such as high cost, high complexity, or unstable effectiveness.

Method used

A molecular identification tag containing the index complementary sequence, sequencing primer complementary sequence, and Arabidopsis sequence was designed and combined with paraffin oil physical isolation for use in sample pretreatment and nucleic acid extraction stages, forming a closed-loop anti-contamination scheme. The unique molecular tag and physical barrier work together to reduce the risk of contamination.

Benefits of technology

It significantly reduces the false positive rate, improves the reliability and efficiency of detection, is suitable for large-scale clinical applications, is low in cost, is compatible with mainstream extraction kits, and achieves a 90% contamination control rate.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure SMS_1
    Figure SMS_1
  • Figure SMS_2
    Figure SMS_2
  • Figure SMS_3
    Figure SMS_3
Patent Text Reader

Abstract

The invention relates to the technical field of gene sequencing, and discloses a molecular recognition tag for recognizing cross contamination of biological samples in tNGS, a reagent containing the molecular recognition tag and application of the molecular recognition tag. The molecular identification tag comprises a sequence complementary to Index, a sequence complementary to a sequencing primer and an arabidopsis thaliana sequence. A unique sequence is designed and mainly comprises three parts, the first part is a sequence complementarily paired with second-round amplification index, the second part is a sequencing primer combination sequence, and the third part is an arabidopsis thaliana sequence with the length of about 200bp. Wherein an arabidopsis thaliana sequence is used as a unique molecular identification tag, and a proper amount of the sequence is added in a sample pretreatment process and is used as a unique molecular identification tag of each sample.
Need to check novelty before this filing date? Find Prior Art

Description

TECHNICAL FIELD

[0001] The present application relates to the field of gene sequencing technology, in particular to a molecular recognition tag for identifying cross-contamination of biological samples in tNGS, a reagent comprising the same and use thereof. BACKGROUND

[0002] Targeted next-generation sequencing (tNGS) is a precise detection method based on high-throughput sequencing technology. By enriching specific genomic regions before sequencing, it significantly improves the coverage depth and analysis sensitivity of target regions. The workflow of this technology mainly includes key steps such as nucleic acid extraction and quality control, fragmentation and library construction (such as adapter ligation, PCR amplification), targeted enrichment (using hybrid capture or amplicon sequencing strategy), high-throughput sequencing (commonly using Illumina, MGI or Ion Torrent platform) and bioinformatics analysis (such as reference genome alignment, variant detection and functional annotation). Compared with whole genome sequencing (WGS) or whole exome sequencing (WES), tNGS can achieve higher throughput of target region deep sequencing at lower cost, so it has wide application value in tumor driver gene detection, pathogenic mutation screening of genetic diseases, pathogen identification and drug resistance gene analysis, etc.

[0003] However, the sensitivity advantage of tNGS technology also makes the problem of nucleic acid contamination in the experimental process particularly prominent. Common sources of contamination can be divided into three stages: first, during sample processing and library construction, improper operation may lead to cross-contamination between samples, such as using the same pipette or centrifuge tube to handle different samples, liquid residue or sample confusion, etc.; second, during PCR amplification, high-concentration amplification products are prone to aerosol contamination, which in turn affects the accuracy of subsequent experiments; in addition, there may be low-level nucleic acid contamination in the laboratory environment or commercial reagents, such as residual DNA fragments from previous experiments or sequencing adapter contamination. If these contaminations are not controlled, they may lead to the generation of false positive results, seriously affecting the reliability of diagnostic or research data.

[0004] To effectively reduce the risk of contamination in tNGS experiments, strict preventive measures should be taken. Currently, the prevention and control measures for nucleic acid contamination in tNGS experiments reduce the risk of false positives, but also bring certain technical and management challenges. The introduction of unique molecular tags (UMI) can effectively distinguish between real variations and amplification or sequencing errors, significantly improving detection specificity, but its shortcomings are the increase in sequencing data volume and bioinformatics analysis complexity, and the capture ability of low-abundance targets may be limited. Negative controls (NTC) help identify reagent or environmental contamination, but their sensitivity is limited and cannot completely rule out the effects of low-level contamination. Strict experimental partitioning operations can effectively avoid cross-contamination between samples, but require additional space resources and personnel training, increasing the operating costs of the laboratory. UV and nuclease decontamination, although simple to operate, may interfere with certain experimental materials (such as enzyme activity), and the removal effect of adsorptive contaminants (such as aerosol DNA) is unstable. In addition, comprehensive quality control measures (such as multiple batch repeated detection) can improve the reliability of the results, but will prolong the detection period and increase the cost.

[0005] Therefore, in practical applications, the strictness and operability of the prevention and control measures need to be balanced according to different experimental needs, to optimize the experimental efficiency and economic benefits while ensuring data quality. SUMMARY

[0006] Based on this, the present application at least provides a molecular recognition tag for identifying cross-contamination of biological samples in tNGS, a reagent comprising the same and the use thereof.

[0007] In the first aspect of the present application, a molecular recognition tag for identifying cross-contamination of biological samples in tNGS is provided, comprising: a sequence complementary to Index, a sequence complementary to a sequencing primer, and an Arabidopsis thaliana sequence.

[0008] In the second aspect of the present application, the use of the molecular recognition tag for identifying cross-contamination of biological samples in tNGS in identifying cross-contamination of biological samples in tNGS is provided.

[0009] In the third aspect of the present application, a reagent for multi-sample tNGS detection is provided, comprising the molecular recognition tag for identifying cross-contamination of biological samples in tNGS as described in the first aspect.

[0010] The application designs a unique sequence, which mainly includes three parts, the first part is a sequence complementary to the second round of amplification index, the second part is a sequencing primer binding sequence, and the third part is an Arabidopsis sequence of about 200 bp in length. The Arabidopsis sequence is used as a unique molecular identification tag, and an appropriate amount of the sequence is added during sample pretreatment as a unique molecular identification tag for each sample, effectively indicating the cross contamination of biological samples in tNGS. DETAILED DESCRIPTION

[0011] In order to facilitate the understanding of the present application, the present application will be described more fully below. The preferred embodiments of the present application are given in the embodiments. However, the present application can be realized in many different forms and is not limited to the embodiments described herein. On the contrary, the purpose of providing these embodiments is to make the disclosure of the present application more thorough and comprehensive.

[0012] Unless otherwise defined, all technical and scientific terms used herein have the same meaning as commonly understood by one of ordinary skill in the art to which the present application belongs. The terms used in the specification of the present application are only for the purpose of describing the specific embodiments of the present application and are not intended to limit the present application. The term "and / or" used herein includes any and all combinations of one or more related listed items.

[0013] In the present application, "one or more" means any one of the listed items or any combination of the listed items, unless otherwise specified. Similarly, "one or more" and the like otherwise indicate "one or more" are also understood in the same way, unless otherwise specified.

[0014] As used in the present application, "combination thereof", "any combination thereof", "any combination thereof" and the like include all suitable combinations of any two or more of the listed items.

[0015] In the present application, "suitable combination", "suitable manner", "any suitable manner" and the like, "suitable" means that the technical solutions of the present application can be implemented, the technical problems of the present application can be solved, and the expected technical effects of the present application can be achieved.

[0016] In the present application, "further", "further", "in particular", "for example", "as", "example", "for example" and the like are used for description purposes, indicating that the different technical solutions before and after are related in terms of coverage, but should not be understood as a limitation of the previous technical solution, nor should it be understood as a limitation of the protection scope of the present application. In the present application, unless otherwise specified, A (such as B) means that B is one non-limiting example of A, and it can be understood that A is not limited to B.

[0017] In the present application, "optionally", "optional", "option" means optional, that is, selected from "have" or "no" two parallel schemes. If there are multiple "optional" in a technical solution, if there is no special description, and there is no contradiction or mutual restriction relationship, each "optional" is independent. If there is no other description, the present application is described as "optionally includes", "optionally contains" and the like. For example, "optionally includes" means "may include or not include".

[0018] The terms "contain", "include" and "comprise" used in the present application are synonymous terms, which are inclusive or open, and do not exclude additional, unmentioned members or features. Members or features, such as materials or components, structures, elements, instruments, etc.; Non-limiting examples of members or features also include actions, conditions, timing, states, etc.

[0019] In the present application, the technical features or technical solutions described in open language include closed technical features or technical solutions composed of listed contents, and also include open technical features or technical solutions containing listed contents.

[0020] In the present application, the exemplary description involving "in some embodiments (or examples)", "in one embodiment (or example)" and the like can cover but is not limited to the following meanings: these schemes can be combined with other schemes in a suitable manner to form new technical solutions.

[0021] In the present application, in the "first aspect", "second aspect", "third aspect", "fourth aspect", etc., the terms "first", "second", "third", "fourth" and the like are only for description purposes, and cannot be understood as indicating or implying relative importance or quantity, nor can it be understood as implicitly indicating the importance or quantity of the indicated technical features. Moreover, "first", "second", "third", "fourth" and the like only serve the purpose of non-exhaustive enumeration description, and it should be understood that they do not constitute a closed limitation on the quantity.

[0022] In this application, when a numerical interval (i.e., a numerical range) is involved, the distribution of the optional numbers in the numerical interval is considered to be continuous and includes both numerical endpoints (i.e., the minimum value and the maximum value) of the numerical interval and each number between the two numerical endpoints, unless otherwise specified. When a numerical interval refers only to integers in the numerical interval, including both endpoint integers and each integer between the two endpoints, it is equivalent to directly listing each integer, unless otherwise specified. When multiple numerical ranges are provided to describe a feature or a characteristic, these numerical ranges can be combined. In other words, unless otherwise indicated, the numerical ranges disclosed herein are to be understood as including any and all sub-ranges therein. The "numbers" in the numerical interval can be any quantitative value, such as a number, a percentage, a ratio, etc. The "numerical interval" is intended to broadly include numerical interval types such as percentage intervals, ratio intervals, and the like.

[0023] In this application, unless otherwise specified, the execution of the steps involved in the method flow does not have strict order restrictions, and these steps can be executed in other orders than described. Moreover, any step can include multiple sub-steps or multiple stages, which do not necessarily be executed at the same time, but can be executed at different times, and the execution order is not necessarily sequential, but can be alternated or simultaneously executed with other steps or sub-steps or stages of other steps.

[0024] Recent experiments have found that nucleic acid contamination in tNGS experiments mainly comes from the nucleic acid extraction stage, rather than the PCR amplification or sequencing link as previously thought. Studies have shown that more than 60% of contamination events can be traced back to cross-contamination between samples or reagent contamination during the nucleic acid extraction process, such as pipette contamination, centrifuge tube residue, automated extraction instrument pipeline residue, or nucleic acid impurities carried by reagents. This discovery has overturned the traditional contamination prevention and control approach, which is to excessively focus on aerosol contamination in the PCR amplification zone, while the contamination control in the extraction stage is relatively insufficient. The sources of these contaminations can include residual DNA from previous samples, laboratory environmental DNA (such as operator skin flakes or microorganisms), and nucleic acid fragments mixed in commercial extraction reagents (such as proteinase K or magnetic beads). Since the extraction step contamination has a cumulative effect, and the contaminated fragments can be amplified and enriched during the subsequent PCR or hybrid capture process, leading to false positive results, even misleading clinical diagnosis.

[0025] For this new discovery, the optimization of the experimental process should focus on the pollution prevention and control of the extraction link, including: (1) using disposable consumables (such as filter head, pre-inactivated centrifuge tube); (2) using automatic closed extraction system to reduce manual operation error; (3) introducing extraction reagent negative control (EBC) to monitor reagent pollution; (4) optimizing laboratory cleaning procedures, such as using DNase regularly to treat the table top and equipment; (5) adding pollution screening steps when quantifying and testing the quality of DNA after extraction. The company has upgraded and optimized the following two aspects for the extraction link.

[0026] In some embodiments, during the extraction process, paraffin oil is added to the grinding tube during the pre-treatment grinding process, and the oil seal interface needs to be preserved during the subsequent operation process until the dissolution.

[0027] Through the combination of the above two steps, the oil seal can effectively avoid cross contamination caused by liquid splashing and other reasons, and cooperate with the unique analysis tag to effectively monitor the pollution, achieving a 90% pollution control rate.

[0028] In one aspect of the present application, a molecular recognition tag for identifying biological sample cross contamination in tNGS is provided, which comprises: a sequence complementary to Index, a sequence complementary to sequencing primer, and an Arabidopsis thaliana sequence.

[0029] In some embodiments, the molecular recognition tag for identifying biological sample cross contamination in tNGS has a length of 200 nt to 300 nt; wherein:

[0030] The length of the sequence complementary to Index is 40 nt to 60 nt;

[0031] The length of the sequence complementary to sequencing primer is 100 nt to 110 nt;

[0032] The length of the Arabidopsis thaliana sequence is 320 nt to 400 nt.

[0033] In some embodiments, the molecular recognition tag for identifying biological sample cross contamination in tNGS has a length of 320 nt, 330 nt, 340 nt, 350 nt, 360 nt, 370 nt, 380 nt, 390 nt, 400 nt, and any range or value between any two values.

[0034] In some embodiments, the sequence complementary to the sequencing primer in the molecular recognition tag for identifying biological sample cross-contamination in tNGS has a length of 100 nt, 101 nt, 102 nt, 103 nt, 104 nt, 105 nt, 106 nt, 107 nt, 108 nt, 109 nt, or 110 nt.

[0035] In some embodiments, the sequence complementary to the sequencing primer in the molecular recognition tag for identifying biological sample cross-contamination in tNGS has a length of 100 nt, 101 nt, 102 nt, 103 nt, 104 nt, 105 nt, 106 nt, 107 nt, 108 nt, 109 nt, or 110 nt.

[0036] In some embodiments, the Arabidopsis thaliana sequence in the molecular recognition tag for identifying biological sample cross-contamination in tNGS has a length of 200 nt.

[0037] In some embodiments, the Arabidopsis thaliana sequence is as shown in any one of SEQ ID NOs: 1~98.

[0038] The inventors have creatively designed the molecular recognition tag for identifying biological sample cross-contamination in tNGS, in which the index complementary pairing sequence and the sequencing primer binding sequence are fixed sequences, the Arabidopsis thaliana sequence is a preferred sequence, has a unique alignment, and does not introduce amplification bias in the entire reaction system without specific amplification.

[0039] In some embodiments, the molecular recognition tag for identifying biological sample cross-contamination in tNGS comprises, from 5' end to 3' end, in order: a sequence complementary to the sequencing primer, a sequence complementary to the Index, and an Arabidopsis thaliana sequence.

[0040] In another aspect of the present application, there is provided use of the molecular recognition tag for identifying biological sample cross-contamination in tNGS as described above in identifying biological sample cross-contamination in tNGS.

[0041] In some embodiments, the method for identifying biological sample cross-contamination in tNGS comprises the following steps:

[0042] S100. Sample processing: adding the molecular recognition tag for identifying biological sample cross-contamination in tNGS to the biological sample to form a sample processing system;

[0043] S200. Nucleic acid extraction: extracting nucleic acid from the biological sample;

[0044] S300. Nucleic acid fragmentation and PCR amplification; and,

[0045] S400. Ligation of sequencing adapters.

[0046] Before nucleic acid extraction, adding molecular recognition tags containing specially designed unique alignment Arabidopsis sequences for identifying biological sample cross-contamination in tNGS can ensure that each molecule has a unique identification.

[0047] In some embodiments, in step S100, the concentration of the molecular recognition tag for identifying biological sample cross-contamination in tNGS in the sample processing system is 0.005 ng / μL~0.015 ng / μL. Exemplarily, for example, 0.005 ng / μL, 0.006 ng / μL, 0.007 ng / μL, 0.008 ng / μL, 0.009 ng / μL, 0.01 ng / μL, 0.011 ng / μL, 0.012 ng / μL, 0.013 ng / μL, 0.014 ng / μL, 0.015 ng / μL, or a range or value between any two values.

[0048] In some embodiments, before step S200, paraffin oil is added to the sample processing system for oil sealing.

[0049] In the nucleic acid extraction step (such as lysis, centrifugation), paraffin oil is added to cover the liquid surface, blocking the aerosol diffusion path, and ensuring no contamination of other nucleic acids and pathogenic microorganisms.

[0050] In some embodiments, compatibility optimization is performed, and low-viscosity paraffin oil (such as density 0.83 g / cm 3 ~0.86 g / cm 3 ) is used for screening. The advantage of this operation is at least that it does not affect the subsequent purification efficiency, and it is suitable for mainstream extraction kits, and can be applied to fully automated extraction kits.

[0051] For the first time, unique molecular tag design is combined with paraffin oil physical isolation in the extraction link, and the aerosol pollution problem in tNGS is solved from the molecular level (UMI monitoring) and the operation level (paraffin oil blocking) to significantly reduce the false positive rate. From sample pretreatment (UMI labeling) to nucleic acid extraction (paraffin oil isolation), a closed-loop anti-pollution solution is formed to make up for the shortcomings of existing technologies that only focus on a single link. In addition, paraffin oil is low in cost, and UMI labeling can be realized through conventional reagents, which is suitable for large-scale clinical application.

[0052] The method in the present application can also be matched with the development of high-precision UMI clustering algorithm to effectively distinguish between real mutations and amplification / sequencing errors.

[0053] In some embodiments, the sequences of the sequencing adapters are set forth in SEQ ID NO: 99 and SEQ ID NO: 100, respectively.

[0054] In another aspect of the present application, there is provided a reagent for multi-sample tNGS detection, comprising the molecular recognition tag for recognizing cross-contamination of biological samples in tNGS as described above.

[0055] In some embodiments, the reagent for multi-sample tNGS detection further comprises paraffin oil.

[0056] Some examples are provided below.

[0057] The embodiments of the present application will be described in detail below with reference to the examples. It should be understood that these examples are only used to illustrate the present application and not to limit the scope of the present application. The experimental methods not specified in the following examples are preferably referred to the guidelines given in the present application, and can also be performed according to the experimental manuals or conventional conditions in the art, or according to the conditions suggested by the manufacturers, or according to the known experimental methods in the art.

[0058] Example 1

[0059] Experimental procedure:

[0060] I. Sample processing

[0061] 1. Pretreatment

[0062] Pulmonary alveolar lavage fluid and cerebrospinal fluid samples Pleural effusion samples: If the sample state is clear (similar to water), sample enrichment is required: take 2 mL of pulmonary alveolar lavage fluid and centrifuge at 12000 rpm for 10 min;

[0063] Whole blood samples: Take 2 mL of whole blood sample, centrifuge at 1600 g at 4°C for 10 min, take the supernatant in a new centrifuge tube, centrifuge at 16000 g at 4°C for 10 min, and take the supernatant in a new centrifuge tube. Take 600 μL of sample for subsequent extraction steps.

[0064] 2. Mechanical disruption (blood plasma samples, ophthalmic samples need to skip this step): Add 5 μL of Error barcode corresponding to each sample, 50 μL of lysis solution 1, 150 μL of lysis solution 2, 30 μL of nucleic acid protection agent and 20 μL of proteinase K to the grinding tube after the pretreatment of the above sample, and add paraffin oil, shake and mix for 2 min. Use a tissue grinding homogenizer for mechanical disruption: if subsequent RNA library construction is required, mechanical disruption is not recommended.

[0065] 3. If DNA library is to be prepared later, the recommended procedure for breaking the cell wall is: 4°C, 6m / s, 45s, interval 45s, 2 cycles; after the grinding is completed, take all the liquid including the oil phase to a new 1.5mL centrifuge tube.

[0066] 4. Take 20ul proteinase K and add to the centrifuge tube for digestion: incubate at room temperature for 10 min.

[0067] 5. Magnetic bead binding: add 350ul anhydrous ethanol (or isopropanol) and 15ul magnetic bead suspension, vortex to mix, and stand at room temperature for 10 min, vortex to mix every 5 min for 30 sec.

[0068] 6. Place the centrifuge tube on the magnetic stand for 2 min, and carefully remove the lower aqueous phase with a pipette when the magnetic beads are completely adsorbed.

[0069] 7. Add 750ul of wash solution 1 (check if 36mL of anhydrous ethanol has been added before use), vortex to mix for 2 min to fully suspend the magnetic beads.

[0070] 8. Place the centrifuge tube on the magnetic stand for 1 min, and carefully remove the lower aqueous phase with a pipette when the magnetic beads are completely adsorbed.

[0071] 9. Repeat steps 7 and 8.

[0072] 10. Add 750ul of wash solution 2 (check if 80mL of anhydrous ethanol has been added before use), vortex to mix for 2 min to fully suspend the magnetic beads.

[0073] 11. Place the centrifuge tube on the magnetic stand for 1 min, and carefully remove all the liquid with a pipette when the magnetic beads are completely adsorbed.

[0074] 12. Repeat steps 10 and 11.

[0075] 13. Place the centrifuge tube on the magnetic stand and air dry at room temperature for 10 min.

[0076] Note: Ethanol residue will inhibit subsequent enzyme reactions, so make sure the ethanol is completely volatilized when air drying. However, do not dry for too long, as it may be difficult to elute the nucleic acids.

[0077] 14. Add 50-100ul of RNase-Free ddH2O (60ul is recommended), resuspend the magnetic beads with a pipette, and incubate at room temperature for 5 min (note: for RNA, incubate at room temperature), and gently shake every 2 min to fully elute the nucleic acids.

[0078] 15. Place the centrifuge tube on the magnetic stand for 2 min, and when the magnetic beads are completely absorbed, carefully transfer the nucleic acid solution to a new centrifuge tube and store at -20°C.

[0079] Note: The magnetic adsorption time can be appropriately extended, or the nucleic acid solution can be transferred to a new centrifuge tube after high-speed centrifugation for 2 min to ensure that there are no magnetic beads remaining in the solution to avoid affecting subsequent experiments.

[0080] II. Library construction

[0081] 1. Prepare the reverse transcription reaction system (reagent preparation)

[0082] Table 1

[0083]

[0084] 2. Reverse transcription reaction (template preparation)

[0085] Place the PCR reaction tube in the sample slot of the amplification instrument, and start the PCR reaction. The cycle parameters are set as shown in Table 2.

[0086] Table 2

[0087]

[0088] Place the reaction product cDNA solution on ice for subsequent experiments; or immediately store at -20°C for later use.

[0089] 3. Prepare the multiplex PCR amplification system

[0090] Table 3

[0091]

[0092] 4. Multiplex amplification

[0093] Place the reaction tube prepared in the first round of PCR in the sample slot of the amplification instrument, and start the PCR reaction. The cycle parameters are set as shown in Table 4.

[0094] Table 4

[0095]

[0096] 5. Purification of multiplex PCR products

[0097] Preparation: This kit does not contain anhydrous ethanol, please use commercially available anhydrous ethanol that has been performance verified. Take the DNA clean beads out of the refrigerator, and equilibrate at room temperature for at least 30 min. Vortex or thoroughly invert the DNA clean beads to fully

[0098] Resuspend and prepare 80% ethanol.

[0099] (1) Add 15 uL DNA clean beads to the first round PCR amplification product; vortex mix and incubate at room temperature for 5 min.

[0100] (2) Centrifuge the tube briefly and place it in a magnetic stand. When the solution is clear (about 2 min), carefully transfer the supernatant to a clean centrifuge tube.

[0101] (3) Add 15 uL DNA clean beads to the supernatant; vortex mix and incubate at room temperature for 5 min.

[0102] (4) Centrifuge the tube briefly and place it in a magnetic stand. When the solution is clear (about 2 min), carefully remove the supernatant.

[0103] (5) Add 30 uL Nuclease-Free Water to the DNA clean beads; vortex mix, then add another 30 uL DNA clean beads to the system, vortex mix and incubate at room temperature for 5 min.

[0104] (6) Centrifuge the tube briefly and place it in a magnetic stand. When the solution is clear (about 2 min), carefully remove the supernatant.

[0105] (7) Keep the centrifuge tube in the magnetic stand at all times, add 100 L freshly prepared 80% ethanol, and repeatedly adsorb the DNA clean beads on different two sides of the magnetic stand to fully suspend the DNA clean beads for washing. Carefully remove the supernatant.

[0106] (8) Repeat step (7).

[0107] (9) Keep the centrifuge tube in the magnetic stand at all times, open the cap and dry the DNA clean beads (about 5 min). Note: The drying time should not be too long, otherwise the DNA clean beads will be dried too much, which will affect the purification effect.

[0108] 6. Second round amplification

[0109] Transfer the prepared 30 uL reaction system to the product analysis area, directly add to the purified DNA clean beads (combined with the first round PCR product) of the corresponding sample, resuspend the DNA clean beads. After the reaction solution is centrifuged to the bottom of the tube, the following reaction is carried out.

[0110] Table 5

[0111]

[0112] Preparation:

[0113] Take the DNA clean beads out of the freezer and equilibrate at room temperature for at least 30 min. Vortex or invert the DNA clean beads repeatedly to ensure full resuspension and prepare 80% ethanol.

[0114] (1) Add 24 uL DNA clean beads to the two-cycle PCR amplification product; vortex well and incubate at room temperature for 5 min.

[0115] (2) Centrifuge the tube briefly and place it in the magnetic stand. After the solution is clear (about 2 min), carefully remove the supernatant.

[0116] (3) Keep the centrifuge tube in the magnetic stand at all times, add 100 uL of freshly prepared 80% ethanol, and use the magnetic stand to repeatedly adsorb the DNA clean beads on different two sides to fully suspend the DNA clean beads for washing. Carefully remove the supernatant.

[0117] (4) Repeat step (3).

[0118] (5) Keep the centrifuge tube in the magnetic stand at all times, and dry the DNA clean beads (about 5 min). Note: The drying time should not be too long, otherwise the DNA clean beads will be dried too much, which will affect the purification effect.

[0119] (6) Take the centrifuge tube out of the magnetic stand, add 25 uL of Nuclease-Free Water, vortex well, and incubate at room temperature for 2 min.

[0120] (7) Centrifuge the tube briefly and place it in the magnetic stand to separate the DNA clean beads and the liquid. After the solution is clear (about 2 min), carefully pipette the supernatant into a clean tube, and the library construction is complete. Label and store at -20±5°C for short-term storage or at -70±5°C for long-term storage.

[0121] Different molecular tags were added to the samples, with an addition amount of 0.005 ng, and the results are shown in Table 6.

[0122] Table 6

[0123]

[0124]

[0125] The molecular tag sequences between different samples can be clearly seen, and the molecular tag sequences account for 1% of all sequences. There is a risk of cross contamination of adjacent samples in individual wells, which proves that the molecular tag function is consistent with the expected results.

[0126] Using blank water samples for detection, one group of adding paraffin oil, one group without adding paraffin oil, it can be seen that the extraction of paraffin oil is very effective for pollution prevention and control (see Table 7).

[0127] Table 7

[0128]

[0129] The technical features of the above-described embodiments can be combined in any manner. In order to make the description concise, not all possible combinations of the technical features in the above-described embodiments are described, however, as long as the combinations of the technical features do not exist contradictions, they should be considered as the scope of the present disclosure.

[0130] The above-described embodiments only express several implementation manners of the present application, and the description is relatively specific and detailed, but it should not be understood as a limitation on the scope of the patent. It should be pointed out that for ordinary skilled in the art, without departing from the concept of the present application, a number of modifications and improvements can be made, which are all within the scope of the present application. Therefore, the scope of the patent of the present application should be subject to the appended claims, and the description can be used to explain the content of the claims.

Claims

1. A molecular identification tag for identifying cross-contamination of biological samples in tNGS, characterized in that, complementary to Index, a sequence complementary to sequencing primer, and Arabidopsis thaliana sequence.

2. The molecular identification tag for identifying cross contamination of biological samples in tNGS of claim 1, wherein, The length is 320 nt-400 nt. The length of the sequence complementary to Index is 40 nt-60 nt. The length of the sequence complementary to sequencing primer is 100 nt-110 nt. The length of the Arabidopsis thaliana sequence is 180 nt-220 nt.

3. The molecular identification tag for identifying cross-contamination of biological samples in tNGS of claim 1 or 2, wherein, The Arabidopsis thaliana sequence is shown in any one of SEQ ID NO: 1-98.

4. The molecular identification tag for identifying cross contamination of biological samples in tNGS according to any one of claims 1-3, wherein, From 5' end to 3' end, it comprises, in sequence, a sequence complementary to sequencing primer, a sequence complementary to Index, and Arabidopsis thaliana sequence.

5. Use of the molecular identification tag for identifying biological sample cross-contamination in tNGS according to any one of claims 1-4 in identifying biological sample cross-contamination in tNGS.

6. Use according to claim 5, characterized in that, The method for identifying biological sample cross-contamination in tNGS comprises the following steps: Sample processing: adding the molecular identification tag for identifying biological sample cross-contamination in tNGS into the biological sample to form a sample processing system; Nucleic acid extraction: extracting nucleic acid from the biological sample; Nucleic acid fragmentation and PCR amplification; Connecting sequencing adaptor.

7. Use according to claim 6, characterized in that, In the step of sample processing, the use concentration of the molecular identification tag for identifying biological sample cross-contamination in tNGS in the sample processing system is 0.005 ng / μL-0.015 ng / μL.

8. The use according to claim 6, characterized in that, Before the step of nucleic acid extraction, paraffin oil is added to the sample processing system for oil sealing.

9. Use according to any one of claims 5 to 8, characterized in that, The sequences of the sequencing adaptors are shown in SEQ ID NO: 99 and SEQ ID NO: 100 respectively.

10. A reagent for use in a multi-sample tNGS assay, characterized in that, It comprises the molecular identification tag for identifying biological sample cross-contamination in tNGS according to any one of claims 1-4; Optionally, the reagent for multi-sample tNGS detection further comprises paraffin oil.