Tag, library construction method and application thereof for pathogenic microorganism detection

By using nucleotide tags composed of 6-8 bases as adapters in pathogen detection, the contamination problem in the library construction process is solved, achieving efficient pathogen detection, improving detection accuracy and efficiency, and reducing costs.

CN122104959APending Publication Date: 2026-05-29TIANJIN NUOHE MEDICAL LAB CO LTD

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
TIANJIN NUOHE MEDICAL LAB CO LTD
Filing Date
2026-02-09
Publication Date
2026-05-29

Smart Images

  • Figure CN122104959A_ABST
    Figure CN122104959A_ABST
Patent Text Reader

Abstract

The application provides a tag for pathogenic microorganism detection, a library construction method and application thereof. The library construction method comprises the following steps: S1, breaking and performing end repair on DNA fragments of a pathogenic microorganism sample, and then connecting adapters to obtain adapter-connected gene fragments; and S2, purifying the adapter-connected gene fragments, and then performing PCR amplification to obtain a sample library; the adapters are connected with first tags and second tags; the first tags are nucleotides randomly composed of 6-8 bases. The application can solve the problem that it is difficult to exclude contamination in the library construction process of the pathogenic microorganism sample in the prior art, and is suitable for the field of pathogenic microorganism detection.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to the field of pathogen detection, and more specifically, to a label, a library construction method, and its application for pathogen detection. Background Technology

[0002] In the field of pathogen detection, especially in metagenomics (mNGS) and targeted pathogen sequencing (tNGS) using next-generation sequencing technologies, high sample purity is a prerequisite for improving detection accuracy. However, the sources of contamination in pathogen detection are diverse and complex. First, the sample extraction stage may be affected by microbial DNA from the environment; this type of contamination can be monitored and controlled to some extent by adding specific internal standards (i.e., internal standards). During library construction and sequencing, the risk of contamination increases significantly, and this type of contamination is often difficult to detect and handle.

[0003] Cross-contamination becomes more common during library preparation because multiple samples are processed in the same reaction environment. Even with strict quality control measures, it is difficult to completely avoid crosstalk between samples, especially when processing large numbers of samples. This contamination can only be indirectly inferred after sequencing data is generated, using bioinformatics methods such as comparing single nucleotide polymorphisms (SNPs) between different samples. Current analytical techniques increase the time and cost of testing and analysis, leading to secondary experiments, which poses an obstacle to the delivery and analysis of clinical samples.

[0004] Index hopping can also occur during sequencing. Index hopping refers to the mismatch between molecules with specific index sequences that originally belonged to different samples, due to physical or chemical factors of the sequencer, leading to disordered reads. Despite continuous improvements in the performance of modern sequencing platforms, index hopping persists, especially on high-density sequencing chips, further increasing the complexity of data analysis.

[0005] In summary, the existing technology lacks a method for screening and removing contaminants during the detection of pathogenic microorganisms, which increases the time and cost of sample detection and analysis, hindering the implementation of clinical treatment. Summary of the Invention

[0006] The main objective of this invention is to provide a label, a library construction method, and its application for the detection of pathogenic microorganisms, in order to solve the problem of difficulty in eliminating contamination during the library construction process of pathogenic microorganism samples in the prior art.

[0007] To achieve the above objectives, according to a first aspect of the present invention, a method for constructing a library of a pathogenic microorganism sample is provided. The method includes: S1) breaking down the DNA fragment of the pathogenic microorganism sample and performing end repair, and then ligating it with a adapter for pathogen detection to obtain a gene fragment with the adapter attached; S2) purifying the gene fragment with the adapter attached and performing PCR amplification to obtain a library of the sample; the adapter is attached with a first tag and a second tag; the first tag is a nucleotide composed of 6 to 8 bases randomly.

[0008] Furthermore, the pathogenic microorganism sample is derived from a biological sample of the infectious pathogenic microorganism; preferably, the biological sample includes one or more of the following: oral sample, sputum, bronchoalveolar lavage fluid, cerebrospinal fluid, blood, or pleural or peritoneal fluid.

[0009] Furthermore, the adapter includes a first strand and a second strand, wherein the first strand includes, from 5' to 3', a P5 adapter, a second tag, a sequencing primer binding region, and a first tag;

[0010] The second strand, from 5' to 3', includes a first tag, a sequencing primer binding region, a second tag, and a P7 adapter; preferably, the second tag is an index tag.

[0011] Furthermore, the first and second chains are complementary and paired, with the joint having a Y-shaped structure.

[0012] Further, the first tag consists of any of the following nucleotide sequences: TCGGTA, GACGTT, AGTACA, CTAGGT, AATTG, TCTGCT, CTGGAC, GACAGC, GCTCAT, TGAGCA, ATACCT, CATGCG, ATCATG, TATGAC, CCGCTT, GTGAGT, ATAAGC, TAATCG, GCTTCG, CGTCTG, ATGCAA, TCCTAT, GGATTA, CAAGTA, ATTGAG, TATCCA, GGCAAG, CATTGT, ACAGAT, TTGTCA, GGAACC, CTCCTC, AACCAG, TACTTC, GTAGAA, CTGATA, AGATGT, TGTATC, GCACCA, CCATAG, ATCTAC, TGTTAG, GATGGA, CAACAT, ACAACG, TCGACC, GTTAAC, or GCAATT; preferably, the number of first tags is 8 to 10; preferably, the nucleotide sequence of each first tag is different.

[0013] To achieve the above objectives, according to a second aspect of the present invention, a label for detecting pathogenic microorganisms is provided, the label comprising nucleotides randomly composed of 6 to 8 bases.

[0014] Furthermore, the tag consists of any of the following nucleotide sequences: TCGGTA, GACGTT, AGTACA, CTAGGT, AATTG, TCTGCT, CTGGAC, GACAGC, GCTCAT, TGAGCA, ATACCT, CATGCG, ATCATG, TATGAC, CCGCTT, GTGAGT, ATAAGC, TAATCG, GCTTCG, CGTCTG, ATGCAA, TCCTAT, GGATTA, CAAGTA, ATTGAG, TATCCA, GGCAAG, CATTGT, ACAGAT, TTGTCA, GGAACC, CTCCTC, AACCAG, TACTTC, GTAGAA, CTGATA, AGATGT, TGTATC, GCACCA, CCATAG, ATCTAC, TGTTAG, GATGGA, CAACAT, ACAACG, TCGACC, GTTAAC, or GCAATT.

[0015] To achieve the above objectives, according to a third aspect of the present invention, a sequencing method for pathogenic microorganism samples is provided, wherein a library is constructed using the aforementioned tag for pathogenic microorganism detection or the aforementioned library construction method for pathogenic microorganism samples, and then sequencing is performed to obtain a sequencing data file of the sample.

[0016] To achieve the above objectives, according to a fourth aspect of the present invention, an application is provided of the above-described method for constructing a library of pathogenic microorganisms, or the above-described tag for detecting pathogenic microorganisms, or the above-described method for sequencing samples of pathogenic microorganisms in the detection of pathogenic microorganisms; the application includes: identifying gene fragments linked to tags in sequencing data files; deleting the gene fragments linked to tags when sequencing crosstalk or library construction contamination occurs, thereby obtaining the final analysis file of the pathogenic microorganism sample.

[0017] To achieve the above objectives, according to a fifth aspect of the present invention, a kit for detecting pathogenic microorganisms is provided, the kit comprising a connector; a first label and a second label are attached to the connector; the first label is the aforementioned label for detecting pathogenic microorganisms.

[0018] Further, the adapter comprises a first strand and a second strand. The first strand, from 5' to 3', sequentially comprises a P5 adapter, a second tag, a sequencing primer binding region, and a first tag; the second strand, from 5' to 3', sequentially comprises a first tag, a sequencing primer binding region, a second tag, and a P7 adapter. Preferably, the second tag is an index tag. Preferably, the first and second strands are complementary, and the adapter has a Y-shaped structure. Preferably, the kit further comprises one or more of the following: end repair reaction reagent, ligase, ATP, index PCR amplification primers, dNTPs, or DNA polymerase.

[0019] By applying the technical solution of this invention, the short base nucleotide tags in this application are attached to the adapters of pathogenic microorganism sample library construction. After sequencing, these tag sequences can be accurately identified. If crosstalk sequences appear in the data file, they can be removed to obtain a pure data file for clinical delivery. This avoids secondary experiments after SNP site analysis, reduces the time and experimental cost of pathogenic microorganism detection, and improves the efficiency of detection. Attached Figure Description

[0020] The accompanying drawings, which form part of this application, are used to provide a further understanding of the invention. The illustrative embodiments of the invention and their descriptions are used to explain the invention and do not constitute an undue limitation of the invention. In the drawings:

[0021] Figure 1 A schematic diagram of the connector and label structure of Embodiment 1 of this application is shown.

[0022] Figure 2 A schematic diagram of the structure of a gene fragment in the library of Example 1 of this application is shown. Detailed Implementation

[0023] It should be noted that, unless otherwise specified, the embodiments and features described in this application can be combined with each other. The present invention will now be described in detail with reference to the embodiments.

[0024] As mentioned in the background section, existing pathogen detection technologies are prone to contamination during extraction, library construction, and index hopping during sequencing. Contamination during library construction often requires SNP analysis after data is processed to determine if cross-contamination exists, and only a second experiment is possible, without subsequent remedial measures. This increases the time, analysis, and experimental costs of pathogen detection. Therefore, in this application, the inventors attempt to provide a new label for pathogen detection, and thus propose a series of protection schemes.

[0025] In a first typical embodiment of this application, a method for constructing a library of pathogenic microorganism samples is provided. The method includes: S1) breaking the DNA fragment of the pathogenic microorganism sample and performing end repair, and then ligating it with a adapter for pathogen detection to obtain a gene fragment with the adapter; S2) purifying the gene fragment with the adapter and performing PCR amplification to obtain a library; wherein, the adapter is attached with a first tag and a second tag; the first tag is a nucleotide composed of 6 to 8 bases (including but not limited to 6, 7 or 8) randomly arranged.

[0026] In existing technologies, after library construction for pathogenic microorganisms is completed and sequencing begins, a paired-end indexing mechanism is often used to further refine the classification of reads. Index hopping refers to the mismatch between samples caused by mislabeling DNA fragments with adapters during high-throughput sequencing. When multiple samples are sequenced together, if there is a possibility of index crossover, some reads may be incorrectly assigned to other samples, causing data contamination. The PCR amplification described above is index PCR amplification.

[0027] This application applies adapters containing the aforementioned tags to library construction of pathogenic microorganism samples. First, the DNA fragments from the samples are mechanically or enzymatically fragmented to generate short fragments suitable for sequencing. These short fragments undergo end-repair processing, resulting in each DNA fragment having an A-tail, facilitating subsequent adapter ligation. These are standard techniques known to those skilled in the art during library construction. Subsequently, the adapters used for pathogen detection are ligated to both ends of the DNA fragments. This process assigns a tag to each fragment, facilitating the analysis and attribution of subsequent sequencing data.

[0028] This application combines next-generation sequencing technology and uses randomly assigned 6-8 base nucleotides as short adapter tags to achieve unique identification of sample nucleic acids. By introducing short adapters with short nucleotide sequences, the nucleic acids of each sample can be accurately labeled. Even if sample mixing occurs in subsequent operations, this tag can be relied upon for accurate sample separation, ensuring the complete traceability of original information. Furthermore, the introduction of 6-8 base tags in the early stages of this application will not lead to difficulties in determining the modulus due to excessively long sequences, thus affecting the sequencing quality. Those skilled in the art can flexibly select and set the base composition and length of the first tag sequence of this application in actual operation.

[0029] Furthermore, this application purifies the adapter-ligated gene fragments to remove unligated adapters and other impurities, ensuring the efficiency and purity of subsequent PCR amplification. Subsequently, index PCR amplification is performed. This process uses primers with specific index sequences to amplify the adapter-ligated DNA fragments, thereby constructing the initial library. The index sequence serves as a tag to distinguish different samples, allowing multiple libraries to be sequenced in the same sequencing stream, and enabling subsequent data splitting based on the index sequence.

[0030] After obtaining the library and performing sequencing, this application further processes and splits the data by identifying gene fragments linked to the aforementioned tags in the sequencing data file. If crosstalk, i.e., tag confusion, or index hopping is found in the tagged gene fragments during this process, these abnormal gene fragments will be specifically identified and deleted from the sequencing data. By identifying short adapters with nucleotide tags consisting of the aforementioned 6-8 bases, this application can effectively distinguish and exclude unexpected sequences, thereby significantly reducing index hopping and the interference of possible sample mixing during library construction on the detection results, ensuring data purity, and maintaining high-standard data delivery even under contamination conditions. The library construction method of this application significantly improves the purity of the final analysis file and the reliability of sequencing results, avoids "secondary experiments," enables timely delivery to clinical samples for analysis, improves the detection efficiency of pathogenic microorganisms, and reduces data analysis and experimental costs.

[0031] In a preferred embodiment, the pathogenic microorganism sample is derived from a biological sample of the infectious pathogenic microorganism; preferably, the biological sample includes, but is not limited to, one or more of the following: oral sample, sputum, bronchoalveolar lavage fluid, cerebrospinal fluid, blood, or pleural or peritoneal fluid.

[0032] In a preferred embodiment, the adapter includes a first strand and a second strand. The first strand includes, from 5' to 3', a P5 adapter, a second tag (i.e., a specific tag commonly used in sequencing (index1)), a sequencing primer binding region, and a first tag. The second strand includes, from 5' to 3', a first tag, a sequencing primer binding region, a second tag (a specific tag (index2)), and a P7 adapter.

[0033] The first linker header contains three key regions from the 5' end to the 3' end, along with the first tag of this application: the P5 adapter sequence, the second tag (i.e., the specific tag index), the sequencing primer binding region, and the first tag. The P5 adapter sequence is a universal adapter compatible with sequencing platforms, used to connect to the sequencing chip and provide the necessary environment for sequencing. The aforementioned second tag, the specific tag (index1), is a specific tag used for sequence splitting after sequencing and is a commonly used specific tag in the sequencing field. The sequencing primer binding region is the initiation point for subsequent sequencing reactions, ensuring that the sequencing primers can bind and guide the sequencing process. The first tag, as described above, marks the nucleic acid fragments of each sample to facilitate accurate sample differentiation in subsequent processing.

[0034] The layout of the second linker, from the 5' end to the 3' end, consists of the first tag, the sequencing primer binding region, the second tag, and the P7 adapter sequence. The P7 adapter sequence is a universal adapter compatible with the sequencing platform, used to connect to the sequencing chip and provide the necessary environment for sequencing. The functions of the remaining regions and tags are the same as those of the first linker.

[0035] In a preferred embodiment, the first chain and the second chain are complementary and paired, and the connector has a Y-shaped structure.

[0036] In addition to the first label mentioned above, the connectors in this application, including the other areas, the specific label (i.e., the second label), the P5 connector, and the P7 connector, are all common settings and commonly used component connectors in the pathogen detection library construction process. Those skilled in the art can flexibly select them according to actual needs.

[0037] In a preferred embodiment, the nucleotide sequence of the first tag includes any one of the following: TCGGTA, GACGTT, AGTACA, CTAGGT, AATTG, TCTGCT, CTGGAC, GACAGC, GCTCAT, TGAGCA, ATACCT, CATGCG, ATCATG, TATGAC, CCGCTT, GTGAGT, ATAAGC, TAATCG, GCTTCG, CGTCTG, ATGCAA, TCCTAT, GGATTA, CAAGTA, ATTGAG, TATCCA, GGCAAG, CATTGT, ACAGAT, TTGTCA, GGAACC, CTCCTC, AACCAG, TACTTC, GTAGAA, CTGATA, AGATGT, TGTATC, GCACCA, CCATAG, ATCTAC, TGTTAG, GATGGA, CAACAT, ACAACG, TCGACC, GTTAAC, or GCAATT. Preferably, the number of first tags is 8 to 10, more preferably 8; preferably, the nucleotide sequence of each first tag is different. Preferably, the first label includes any group in Table 1.

[0038] Table 1

[0039]

[0040] The preferred design principle for the first tag in this application is that each tag in each group (one group is sufficient for each library construction) should have a difference of at least two bases, and each group preferably has eight specific tags (8 to 10 are acceptable, and those skilled in the art can select and set them according to actual needs). The eight tags in each group should have A, T, C, and G distributed as evenly as possible at each base position from the first to the sixth base. Those skilled in the art can design and select tags according to the design principles of this application.

[0041] In a second typical embodiment of this application, a label for detecting pathogenic microorganisms is provided, the label comprising nucleotides randomly composed of 6 to 8 bases.

[0042] In a preferred embodiment, the tag consists of a tag with any of the following nucleotide sequences: TCGGTA, GACGTT, AGTACA, CTAGGT, AATTG, TCTGCT, CTGGAC, GACAGC, GCTCAT, TGAGCA, ATACCT, CATGCG, ATCATG, TATGAC, CCGCTT, GTGAGT, ATAAGC, TAATCG, GCTTCG, CGTCTG, ATGCAA, TCCTAT, GGATTA, CAAGTA, ATTGAG, TATCCA, GGCAAG, CATTGT, ACAGAT, TTGTCA, GGAACC, CTCCTC, AACCAG, TACTTC, GTAGAA, CTGATA, AGATGT, TGTATC, GCACCA, CCATAG, ATCTAC, TGTTAG, GATGGA, CAACAT, ACAACG, TCGACC, GTTAAC, or GCAATT.

[0043] In a third typical embodiment of this application, a sequencing method for samples of pathogenic microorganisms is provided, wherein the library constructed by the above-mentioned library construction method for pathogenic microorganism detection is sequenced to obtain a sequencing data file of the sample.

[0044] In a preferred embodiment, sequencing includes mNGS or tNGS.

[0045] In a fourth typical embodiment of this application, the above-mentioned label for pathogen detection, or the above-mentioned connector for pathogen detection, or the above-mentioned library construction method for pathogen detection is provided for application in pathogen detection.

[0046] In a preferred embodiment, the above application includes: identifying tagged gene fragments in sequencing data files; deleting tagged gene fragments when crosstalk or library construction contamination occurs, thereby obtaining the final analysis file of the pathogenic microorganism sample.

[0047] Using the adapters described in this application for pathogen detection can improve contamination control at each stage of the pathogen detection process. Specifically, it can prevent cross-contamination by pathogens and contamination during sequencing from affecting the accuracy of test results, ensuring normal data delivery and improving the efficiency of pathogen detection and analysis. This application utilizes short adapters with 6-8 base tags during adapter ligation, combined with independent paired-end indexes. This allows the sequenced data to be initially split based on the paired-end indexes, followed by secondary screening using the 6-8 base tags on the short adapters. This ensures that each read is correctly assigned to its corresponding sample, and even if tag crosstalk occurs, it can be accurately removed, significantly improving the reliability and clinical application value of sequencing data.

[0048] Through the steps described above, the library construction method proposed in this application not only optimizes the nucleic acid fragmentation and labeling process of pathogenic microorganism samples, but also introduces advanced quality control strategies to effectively address common contamination problems in pathogen detection, especially index hopping, ensuring the accuracy and clinical application value of high-throughput sequencing data. This method is not only applicable to the detection of pathogenic microorganisms, but also provides a solid technical foundation for handling complex samples and ensuring the consistency of high-throughput sequencing results.

[0049] In a fifth typical embodiment of this application, a kit for detecting pathogenic microorganisms is provided, the kit including the adapter for detecting pathogenic microorganisms as described above.

[0050] In a preferred embodiment, the kit further includes, but is not limited to, one or more of the following: end-repair reaction reagents, ligases, ATP, index PCR amplification primers, dNTPs, or DNA polymerases. The reagents in the above kit are used in the library construction process for pathogen detection. Those skilled in the art can flexibly select conventional pathogen library construction reagents, all of which will achieve the effects described in this application.

[0051] The beneficial effects of this application will be explained in more detail below with reference to specific embodiments.

[0052] Example 1

[0053] The pathogen detection process includes: extracting sample DNA → constructing a library → pooling (preparing to mix libraries from multiple samples for sequencing) → sequencing.

[0054] In this embodiment, a library was built for clinical respiratory infection samples (i.e., positive samples, numbered S12), and a library was also built for control group samples (i.e., negative samples, numbered S1-S11).

[0055] 1. Extract DNA from the sample

[0056] ① Sample cell disruption; ② Lysis with lysis buffer to lyse pathogens and release nucleic acids; ③ Adsorption of nucleic acids by magnetic beads under the action of binding solution; ④ Cleaning of magnetic beads to remove various salts and impurities; ⑤ Eluting of nucleic acids with TE buffer.

[0057] 2. Perform DNA fragmentation and end repair on the sample.

[0058] ① Take a 20ng sample based on the qubit quantification result; ② Add the cleavage enzyme and the terminal repair A mixture to the PCR reaction to complete the cleavage and terminal A addition.

[0059] 3. Connector connection (This step involves adding a short connector with a 6-base label)

[0060] ① Add the 6-base short linker of the present invention using a linking reagent; ② Perform the linking reaction.

[0061] A schematic diagram of the connector and label structure is shown below. Figure 1 As shown, the red part is the tag sequence, and the blue "Y" part is the short connector composed of the first chain and the second chain.

[0062] The base T after the tag on the first chain forms a complementary base pair with the A tail of the inserted fragment during the linker ligation process, thereby improving ligation efficiency. This is a common processing method used by those skilled in the art in the library construction process.

[0063] In this embodiment, the label for the short connector is selected from group 1 in Table 1.

[0064] 4. Connector purification

[0065] The ligation product was purified using purifying magnetic beads, with the volume of magnetic beads added at a 1:1 ratio to the sample volume.

[0066] ① Add magnetic beads and mix well; ② Incubate at room temperature for 10 min; ③ Adsorb using a magnetic rack for 5 min; ④ Discard the supernatant; ⑤ Wash with 80% ethanol; ⑥ Air dry the magnetic beads; ⑦ Elute with TE.

[0067] 5. Index PCR amplification

[0068] ① Add the eluted product to the PCR amplification reagent; ② Add the index amplification primers; ③ Perform the PCR amplification reaction.

[0069] 6. Purification after amplification

[0070] The PCR products were purified using magnetic beads, with the volume of magnetic beads added at a 1:1 ratio to the sample volume.

[0071] ① Add magnetic beads and mix well; ② Incubate at room temperature for 10 min; ③ Adsorb using a magnetic rack for 5 min; ④ Discard the supernatant; ⑤ Wash with 80% ethanol; ⑥ Air dry the magnetic beads; ⑦ Elute with TE.

[0072] The fragment structure in the library after library construction using the above connectors is as follows: Figure 2 As shown, during insert sequencing, the corresponding adapter tag is identified. This tag can be used to determine the attribution of subsequent reads. Figure 2 In this context, "Insert" refers to the insertion of a gene fragment.

[0073] 7. After pooling the library constructed using the above method, the sample is tested. After sequencing, the sequencing results of the first 6 bases of each read are compared and analyzed. If the sequence matches the expected sequence, it is considered to be the sample. If the sequence does not match, the read is deleted directly. (Note: ① The actual reads are homogenized to 30M reads; ② Negative and positive samples are prepared and tested in the same batch.)

[0074] The bioinformatics analysis results after sewage discharge in this embodiment are shown in Table 2.

[0075] In this embodiment, samples S1-11 were all negative samples, and S12 was a positive sample, indicating resistance to *Haloxysporum tobira* (Haloxysporum tobira). Imtechella halotolerans ) and salt-resistant heterologous bacteria ( Allobacillus halotolerans ) to conduct testing.

[0076] The numbers in Table 2 refer to the number of sequencing reads for the two pathogenic microorganisms. The number before the " / " is the actual number of reads detected, and the number after the " / " is the number of reads after normalization.

[0077] Table 2

[0078]

[0079] It is evident that the wastewater treatment process described in this application can accurately distinguish between negative and positive samples, preventing misjudgments and further improving the accuracy of clinical detection.

[0080] Comparative Example 1

[0081] The difference from Example 1 is that the first label from Example 1 is not added.

[0082] The bioinformatics analysis results of this embodiment are shown in Table 3.

[0083] Table 3

[0084]

[0085] It is evident that by using only the index tag and not the first tag of this application for library construction, even after decontamination treatment, hundreds of reads were detected in negative samples S1-S11. According to the reporting rules, this may lead to false positive results in the test report, affecting the accuracy of clinical detection.

[0086] Comparative Example 2

[0087] The difference between this comparative example and Example 1 is that the connector of this application is replaced with a UMI connector.

[0088] Results: The library constructed using the UMI connector in this comparative example was difficult to distinguish between samples and difficult to clean. The specific results are shown in Table 4.

[0089] Table 4

[0090]

[0091] It is evident that using the existing UMI connector for library construction has resulted in the detection of excessive reads in negative samples, which may lead to false positive results and thus affect the accuracy of clinical detection.

[0092] As can be seen from the above description, the embodiments of the present invention achieve the following technical effects: By introducing a specific 6-8 base tag and an optimized adapter design, this application enables accurate determination of sample attribution in a high-throughput sequencing environment, avoiding secondary experiments and significantly reducing resource consumption and time costs. Technically, this promotes the accuracy and efficiency of pathogen detection; moreover, it enhances the value of detection technology in clinical applications, providing strong data support for the rapid diagnosis and treatment of infectious diseases.

[0093] The above description is merely a preferred embodiment of the present invention and is not intended to limit the invention. Various modifications and variations can be made to the present invention by those skilled in the art. Any modifications, equivalent substitutions, improvements, etc., made within the spirit and principles of the present invention should be included within the scope of protection of the present invention.

Claims

1. A method for constructing a library of pathogenic microorganism samples, characterized in that, The database construction method includes: S1) The DNA fragments of the pathogenic microorganism sample are broken and the ends are repaired, and then the adapters are connected to obtain the gene fragments with the adapters; S2) After purifying the gene fragment of the ligator, perform PCR amplification to obtain the library of the sample; The connector is attached to a first label and a second label; The first tag is a nucleotide composed of 6 to 8 bases randomly.

2. The database construction method according to claim 1, characterized in that, The pathogenic microorganism samples were derived from biological samples infected with pathogenic microorganisms; Preferably, the biological sample includes one or more of the following: oral sample, sputum, bronchoalveolar lavage fluid, cerebrospinal fluid, blood, or pleural or peritoneal fluid.

3. The database construction method according to claim 1, characterized in that, The adapter includes a first strand and a second strand, wherein the first strand comprises, from 5' to 3', a P5 adapter, a second tag, a sequencing primer binding region, and the first tag; The second strand, from 5' to 3', includes the first tag, the sequencing primer binding region, the second tag, and the P7 adapter. Preferably, the second tag is an index tag.

4. The database construction method according to claim 3, characterized in that, The first chain and the second chain are complementary and paired, and the connector has a Y-shaped structure.

5. The database construction method according to claim 1, characterized in that, The first tag consists of a tag with any of the following nucleotide sequences: TCGGTA, GACGTT, AGTACA, CTAGGT, AATTG, TCTGCT, CTGGAC, GACAGC, GCTCAT, TGAGCA, ATACCT, CATGCG, ATCATG, TATGAC, CCGCTT, GTGAGT, ATAAGC, TAATCG, GCTTCG, CGTCTG, ATGCAA, TCCTAT, GGATTA, CAAGTA, ATTGAG, TATCCA, GGCAAG, CATTGT, ACAGAT, TTGTCA, GGAACC, CTCCTC, AACCAG, TACTTC, GTAGAA, CTGATA, AGATGT, TGTATC, GCACCA, CCATAG, ATCTAC, TGTTAG, GATGGA, CAACAT, ACAACG, TCGACC, GTTAAC, or GCAATT; Preferably, the number of the first tags is 8 to 10; Preferably, the nucleotide sequence of each of the first tags is different.

6. A label for detecting pathogenic microorganisms, characterized in that, The tag consists of nucleotides composed of 6 to 8 bases randomly.

7. The label according to claim 6, characterized in that, The tag consists of a tag with any of the following nucleotide sequences: TCGGTA, GACGTT, AGTACA, CTAGGT, AATTG, TCTGCT, CTGGAC, GACAGC, GCTCAT, TGAGCA, ATACCT, CATGCG, ATCATG, TATGAC, CCGCTT, GTGAGT, ATAAGC, TAATCG, GCTTCG, CGTCTG, ATGCAA, TCCTAT, GGATTA, CAAGTA, ATTGAG, TATCCA, GGCAAG, CATTGT, ACAGAT, TTGTCA, GGAACC, CTCCTC, AACCAG, TACTTC, GTAGAA, CTGATA, AGATGT, TGTATC, GCACCA, CCATAG, ATCTAC, TGTTAG, GATGGA, CAACAT, ACAACG, TCGACC, GTTAAC, or GCAATT.

8. A sequencing method for pathogenic microorganism samples, characterized in that, After constructing a library using the library construction method for pathogenic microorganism samples according to any one of claims 1 to 5, or the tag for pathogenic microorganism detection according to any one of claims 6 or 7, sequencing is performed to obtain the sequencing data file of the sample.

9. The method for constructing a library of pathogenic microorganism samples according to any one of claims 1 to 5, or the label for pathogenic microorganism detection according to any one of claims 6 or 7, or the application of the sequencing method for pathogenic microorganism samples according to claim 8 in the detection of pathogenic microorganisms; The applications include: The gene fragments linked to the tag in the sequencing data file are identified. When sequencing crosstalk or library construction contamination occurs in the gene fragments linked to the tag, the gene fragments linked to the tag are deleted to obtain the final analysis file of the pathogenic microorganism sample.

10. A kit for detecting pathogenic microorganisms, characterized in that, The kit includes a connector; The connector is attached to a first label and a second label; The first tag is a nucleotide sequence for the detection of pathogenic microorganisms as described in any one of claims 6 or 7.

11. The reagent kit according to claim 10, characterized in that, The adapter includes a first strand and a second strand, wherein the first strand comprises, from 5' to 3', a P5 adapter, a second tag, a sequencing primer binding region, and the first tag; The second strand, from 5' to 3', includes the first tag, the sequencing primer binding region, the second tag, and the P7 adapter. Preferably, the second tag is an index tag; Preferably, the first chain and the second chain are complementary pairs, and the connector has a Y-shaped structure; Preferably, the kit further includes one or more of the following: end repair reaction reagent, ligase, ATP, index PCR amplification primers, dNTPs, or DNA polymerase.