A Hybrid Primer and Its Design Method and Use

Customized hybridization probes and double-strand specific nucleases address host DNA interference in sequencing, enhancing pathogen detection and sequencing accuracy by specifically targeting and removing host DNA.

CN118813608BActive Publication Date: 2025-07-15SHENZHEN GENEPLUS CLINICAL LAB +3
View PDF 4 Cites 0 Cited by

Patent Information

Application Number
CN202310440677.6
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2023-04-20
Publication Date
2025-07-15
Estimated Expiration
2043-04-20

AI Technical Summary

Technical Problem

In the metagenomic second-generation high-throughput sequencing technology, existing hybrid primers cannot effectively remove host nucleic acids, resulting in low detection rate of pathogenic microorganisms and affecting infection judgments. Incomplete information on the repeat sequence of human genome, which can easily affect the identification region of the nucleic acid of the species to be tested.

Method used

A hybrid primer was designed to obtain the target biological reference genome sequence, filter similar sequences, build backup hybrid primers, and use double-strand specific nuclease to remove the host genome, and combine probes for targeted capture to ensure that the pathogenic microbial nucleic acid is not affected.

Benefits of technology

It improves the detection rate of pathogenic microorganisms, reduces the interference of host nucleic acids, and ensures the accuracy of sequencing results and the integrity of pathogenic microorganisms.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN118813608B_ABST
    Figure CN118813608B_ABST
Patent Text Reader

Abstract

A hybrid primer, its design method and use, wherein the hybrid primer comprises at least one of the nucleotide sequences shown in SEQ ID No. 1 to 500. The nucleic acid sequence of the hybrid primer provided by the present invention is clear, highly effective, has a wide coverage range of the host genome, and can be effectively applied to the binding and removal of host nucleic acids.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the field of sequencing technology, and particularly relates to a hybridization primer, a design method thereof, and uses thereof. Background Art

[0002] In recent years, high-throughput sequencing technology has been widely used in fields such as research and clinical practice. During the construction of a sequencing library, the use of hybridization primers has been widely applied in the field of molecular biology. Hybridization primers or probes are mainly used to bind to specific regions of specific species in a sample nucleic acid, enrich or block specific regions of specific species.

[0003] Currently, metagenomic next-generation high-throughput sequencing technology (mNGS) and pathogen-targeted capture sequencing technology (tNGS) are widely used for the auxiliary identification of clinical pathogenic microorganism infections by sequencing the nucleic acids of pathogenic microorganisms and comparing them with a database. The mNGS technology has the characteristic of a wide database, which is extremely helpful for the detection of rare pathogens, etc. However, due to the excessive content of the host in metagenomic sequencing, it is easy to result in few or no detections of pathogenic microorganisms, thus affecting the judgment of the true infection situation of the sample.

[0004] Removing host nucleic acids in the wet experiment of mNGS can not only reduce the amount of sequencing data, but also increase the detection ratio of pathogenic microorganisms. While removing host nucleic acids, it is necessary to ensure that the nucleic acids of pathogenic microorganisms are not affected so as to avoid the loss of nucleic acids of pathogenic microorganisms.

[0005] The tNGS technology solves the problem of few detections in the mNGS technology by reducing the cost of the detected species. By designing a targeted region for specific capture and enrichment for common target microorganisms, the detection of pathogenic microorganisms is improved, and the accuracy of clinical interpretation is increased. For the tNGS technology based on hybridization capture, in the hybridization step, host genome blocking primers are often used to block the nucleic acids of the host genome, thereby reducing non-specific hybridization and reducing the proportion of host sequences in the final library.

[0006] Currently, the commonly used hybridization method for removing or blocking host nucleic acid sequences generally uses human Cot-1 sequences or human genomic fragments. These nucleic acid fragments containing repetitive regions of the human genome or the complete nucleic acid sequence of the human genome can bind to the human genomic nucleic acids in the nucleic acids of the sample to be processed during hybridization to form a double-stranded structure. For forward enrichment, this double-stranded structure formed can effectively block this region from being bound by hybridization probes, reducing the risk of non-specific hybridization; for reverse enrichment (removing host nucleic acids), this double-stranded structure can be recognized by double-strand specific enzymes, thereby removing this nucleic acid sequence.

[0007] However, the sequence information in the human Cot-1 sequence or human genomic fragments is extremely large, and it is impossible to fully grasp the information of each fragment of these nucleic acids. Although it has been widely used at present, there are potential risks in its application for the enrichment of microbial nucleic acids in metagenomics. There will be partially homologous sequences between the human genomic sequence and the genome of the species to be detected, which is likely to cause the loss of important recognition regions near the homologous sequences of the nucleic acids of the species to be detected during the process of removal or blocking. Summary of the Invention

[0008] According to a first aspect, in one embodiment, a hybridization primer is provided. In one embodiment, the hybridization primer comprises at least one of the nucleotide sequences shown in SEQ ID No.1 to 500.

[0009] According to a second aspect, in one embodiment, a method for designing a hybridization primer is provided, comprising:

[0010] A step of constructing a backup hybridization primer, comprising fragmenting the host genome to obtain the backup hybridization primer;

[0011] A step of obtaining a reference genome of a target organism, comprising obtaining a reference genome sequence of the target organism;

[0012] A filtering step, comprising aligning the backup hybridization primer to the reference genome of the target organism, removing the sequences in the backup hybridization primer that are similar to the reference genome sequence of the target organism, and obtaining the hybridization primer.

[0013] According to a third aspect, in one embodiment, an application of the hybridization primer according to any one of the first aspect in blocking a genomic sequence and / or removing a genomic sequence is provided.

[0014] According to a fourth aspect, in one embodiment, an application of the hybridization primer designed by the method according to any one of the second aspect in blocking a genome or removing a genome is provided.

[0015] According to a fifth aspect, in one embodiment, a method for constructing a pathogen-targeted capture sequencing library is provided, comprising:

[0016] A hybridization step, comprising providing the hybridization primer according to any one of the first aspect, or the hybridization primer designed by the method according to any one of the second aspect, hybridizing the hybridization primer with a library to be tested, and obtaining a hybridized library;

[0017] A targeted capture step, comprising using a probe to perform targeted capture on the hybridized library to obtain a library after targeted capture.

[0018] According to a sixth aspect, in one embodiment, a method for removing a genome in a host is provided, comprising:

[0019] Hybridization step, including providing a sample to be tested, using the hybridization primers described in any one of the first aspect, or the hybridization primers designed by the method described in any one of the second aspect to hybridize with the sample to be tested to obtain a hybridization library;

[0020] Double-stranded nucleic acid molecule removal step, including using a double-stranded specific nuclease to remove double-stranded nucleic acid molecules in the hybridization library to obtain a library after removing the host genome.

[0021] In one embodiment, the hybridization primer provided by the present invention has a clear nucleic acid sequence, high effectiveness, and a wide coverage range of the host genome, and can be effectively applied to the binding and removal of host nucleic acids.

[0022] In one embodiment, the hybridization primer provided by the present invention can be used to block repetitive sequences of the host genome during forward capture without affecting the binding of target regions of other microorganisms.

[0023] In one embodiment, the hybridization primer is not likely to cause the loss of microbial sequences during the host removal process of the mNGS DSN method. Brief Description of the Drawings

[0024] Figure 1 Schematic diagram of the application of the host nucleic acid hybridization primer in forward capture and reverse capture (human host genome removal) in microbial metagenomic sequencing in one embodiment of the present invention;

[0025] Figure 2 Coverage of host hybridization primers of different lengths on the human genome;

[0026] Figure 3 For Example 1, the capture and blocking effect of the first 384 hybridization primers for tNGS (compared with the Cot-1 fragment of the human genome);

[0027] Figure 4 For Example 2, the capture and blocking effects of the first 384 and the first 500 hybridization primers for tNGS (compared with the Cot-1 fragment of the human genome);

[0028] Figure 5 For Example 3, the effect of the first 384 hybridization primers in DSN host removal (compared with the Cot-1 fragment of the human genome). Detailed Description of the Invention

[0029] The present invention will be further described in detail below in conjunction with the specific embodiments and the accompanying drawings. In the following embodiments, many details are described to enable a better understanding of the present application. However, those skilled in the art can easily recognize that some of the features can be omitted in different situations, or can be replaced by other materials or methods. In some cases, some operations related to the present application are not shown or described in the specification in order to avoid the core part of the present application being overwhelmed by excessive description. For those skilled in the art, it is not necessary to describe these related operations in detail, and they can fully understand the related operations based on the description in the specification and the general technical knowledge in the art.

[0030] In addition, the features, operations or characteristics described in the specification can be combined in any appropriate manner to form various embodiments. At the same time, the steps or actions in the method description can also be reordered or adjusted in a manner obvious to those skilled in the art. Therefore, the various sequences in the specification and the drawings are only for clearly describing a certain embodiment and do not mean that they are the necessary sequences, unless it is stated that a certain sequence must be followed.

[0031] The serial numbers assigned to the components in this text, such as "first", "second", etc., are only used to distinguish the described objects and do not have any sequential or technical meaning.

[0032] In this text, "Kmer" refers to a DNA fragment with a length of k, which is obtained by cutting a part of the sequencing reads. k is odd or even. For example: the length of the sequencing reads is 100 bp, and these 100 bp are broken into short fragments of 17 bp. The broken 17-bp fragment is called 17mer, and (100 - 17 + 1) k-mer sequences can be obtained.

[0033] In this text, "DSN" is short for duplex-specific nuclease, and its Chinese name is duplex-specific nuclease, which is a heat-stable nuclease. This enzyme can selectively degrade double-stranded DNA and DNA in DNA-RNA hybrids, but has little effect on single-stranded nucleic acid molecules.

[0034] Since it is impossible to determine whether each nucleic acid fragment of Cot-1 can bind to a non-host or non-target closed region, resulting in the loss of non-host sequences or a decrease in the binding efficiency of the capture probe to the target region. Therefore, the use of hybridization primers for known human genomic repeats in microbial metagenomic sequencing can effectively avoid the occurrence of this problem.

[0035] The existing hybridization primers mainly have the following disadvantages:

[0036] (1) The number of hybridization primers is small, and the effect of the closed region is not ideal enough.

[0037] (2) The length of the hybridization primer is long, and some regions are prone to bind to non-host nucleic acid sequences, which is not conducive to the removal of the host by the DSN method of mNGS;

[0038] (3) The length of the human genome Cot-1 fragment is between 50-300 bp, in a double-stranded form and the specific sequence ratio is not clear, and the closing effect is unstable.

[0039] (4) The human genome Cot-1 fragment or hybridization primer has not been screened after comparison with the microbial database, which is not conducive to the removal of host nucleic acid by mNGS and the closure of host nucleic acid by t NGS.

[0040] According to the first aspect, in one embodiment, a hybridization primer is provided. In one embodiment, the hybridization primer comprises at least one of the nucleotide sequences shown in SEQ ID No. 1 to 500.

[0041] In one embodiment, the hybridization primer comprises at least one of the groups consisting of the nucleotide sequences shown in SEQ ID No. 1 to 500.

[0042] In one embodiment, the hybridization primer comprises all of the groups consisting of the nucleotide sequences shown in SEQ ID No. 1 to 500.

[0043] In one embodiment, the hybridization primer comprises at least one of the nucleotide sequences shown in SEQ ID No. 1 to 384.

[0044] In one embodiment, the hybridization primer comprises at least one of the groups consisting of the nucleotide sequences shown in SEQ ID No. 1 to 384.

[0045] In one embodiment, the hybridization primer comprises all of the groups consisting of the nucleotide sequences shown in SEQ ID No. 1 to 384.

[0046] In one embodiment, the hybridization primer may comprise any number of the nucleotide sequences shown in SEQ ID No. 1 to 500, for example, including but not limited to the first 500, first 400, first 300, first 200, first 100, first 90, first 80, first 70, first 60, first 50, first 40, first 30, first 20, first 10 nucleotide sequences.

[0047] In one embodiment, when the hybridization primers can be all of the group consisting of the nucleotide sequences shown in SEQ ID No. 1 to 384, any number (the any number can be 0, 1, 2... 384) of nucleotides can be selectively selected from the group consisting of the nucleotide sequences shown in SEQ ID No. 385 to 500 and added to the first 384 nucleotides, and effects similar to those of the first 384 nucleotides can be obtained, such as the blocking effect or the de-host effect.

[0048] In one embodiment, the hybridization primers can include at least any number of nucleotides in the group consisting of the nucleotide sequences shown in SEQ ID No. 1 to 500. For example, it can be at least 500, 400, 300, 200, 100, 90, 80, 70, 60, 50, 40, 30, 20 or 10 nucleotide sequences. Any number of hybridization primer combinations among the 500 nucleotides that can achieve effects similar to those of the present invention are within the protection scope of the present invention.

[0049] In one embodiment, the hybridization primers can include any number of nucleotides in the group consisting of the nucleotide sequences shown in SEQ ID No. 1 to 500. For example, it includes but is not limited to any 500, 400, 300, 200, 100, 90, 80, 70, 60, 50, 40, 30, 20, 10 nucleotide sequences. Any number of hybridization primer combinations among the 500 nucleotides that can achieve effects similar to those of the present invention are within the protection scope of the present invention.

[0050] In one embodiment, the hybridization primers are used to capture the nucleic acid sequences of the host.

[0051] In one embodiment, the hybridization primers are used to capture the sequences of the genomic repetitive regions of the host.

[0052] In one embodiment, the host includes but is not limited to the host of microorganisms.

[0053] In one embodiment, the host includes but is not limited to animals or plants.

[0054] In one embodiment, the animals include but are not limited to mammals, such as humans, mice, non-human primates, rabbits or other mammals, or non-mammals, such as zebrafish samples, etc.

[0055] According to the second aspect, in one embodiment, a method for designing hybridization primers is provided, including:

[0056] A step of constructing backup hybridization primers, including fragmenting the host genome to obtain backup hybridization primers;

[0057] Steps for obtaining the reference genome of the target organism, including obtaining the reference genome sequence of the target organism;

[0058] Filtering step, including aligning the backup hybridization primers to the reference genome of the target organism, removing the sequences in the backup hybridization primers that are similar to the reference genome sequence of the target organism, and obtaining the hybridization primers.

[0059] In one embodiment, in the step of constructing the backup hybridization primers, the host genome is fragmented into Kmers of 40 bp to 120 bp. Including but not limited to Kmers of 40 bp, 50 bp, 60 bp, 70 bp, 80 bp, 90 bp, 100 bp, 110 bp, 120 bp.

[0060] In one embodiment, in the step of constructing the backup hybridization primers, the host genome is fragmented into Kmers of 60 bp to 80 bp.

[0061] In one embodiment, in the step of constructing the backup hybridization primers, the host genome is fragmented into Kmers of 60 bp and / or 80 bp. The length here is related to the target length of the hybridization primers, and the length of the hybridization primers will affect the hybridization temperature. In one embodiment, we hope the target length of the hybridization primers is 60 bp or 80 bp, so it is set to 60 and 80. In one embodiment, 40 - 120 bp is a relatively reasonable range.

[0062] In one embodiment, in the step of constructing the backup hybridization primers, Kmer fragments with a GC content of 30% to 70% are screened from the fragmented host genome sequences, sorted from largest to smallest according to the frequency, and the Kmer fragments with the highest ranking are the highly repetitive regions suitable for designing hybridization primers. The frequency represents the number of times the sequence appears in the genome, and the more times, the better, and the larger the covered region.

[0063] There is genome coverage in the sequence list. The more the frequency, the larger the genome coverage. It gradually decreases from top to bottom. The more forward, the larger the genome coverage of the hybridization primers and the better the effect.

[0064] In one embodiment, in the step of constructing the backup hybridization primers, after obtaining the highly repetitive regions, the screened hybridization primers are clustered according to homology, and each representative sequence generated after clustering is the backup hybridization primer for removing host repetitive sequences.

[0065] In one embodiment, in the step of constructing the backup hybridization primers, when clustering the screened hybridization primers according to homology, representative sequences are selected from the hybridization primers with homology ≥ threshold to obtain the screened backup hybridization primers.

[0066] In one embodiment, the threshold of homology can be 2 / 3.

[0067] The homology threshold is used to cluster and screen the Kmer fragments obtained by breaking the host genome into fragments through bioinformatics methods. For example, when the homology screening threshold is 2 / 3, there should be no more than 2 / 3 consecutive identical sequences between fragment sequences. Otherwise, duplicates are removed and only one of them is retained. The homology here is used to remove the homologous fragments between the Kmers of the host genome to obtain the backup hybridization primers.

[0068] In one embodiment, the target biological reference genome includes, but is not limited to, the genomes of at least one of viruses, fungi, bacteria, archaea, and parasites.

[0069] In one embodiment, in the filtering step, the backup hybridization primers (i.e., host hybridization primers) with an alignment length greater than the first percentage of the length of the backup hybridization primers and a similarity greater than the second percentage to the target biological reference genome sequence are removed to obtain the hybridization primers. These hybridization primers can be used for forward capture or reverse capture experiments.

[0070] In one embodiment, the first percentage is 85% - 95%, including but not limited to 85%, 90%, 95%, etc.

[0071] In one embodiment, the second percentage is 85% - 95%, including but not limited to 85%, 90%, 95%, etc.

[0072] In one embodiment, the first percentage and the second percentage can be the same or different.

[0073] In one embodiment, in the filtering step, the backup hybridization primers with an alignment length greater than 90% of the length of the backup hybridization primers (i.e., host hybridization primers) and a similarity greater than 90% to the target biological reference genome sequence are removed to obtain the hybridization primers.

[0074] The backup hybridization primers are aligned with the microbial database, and according to the similarity, the backup hybridization primers homologous to the microbial database are removed.

[0075] In one embodiment, in the filtering step, after removing the backup hybridization primers with an alignment length greater than 90% of the length of the backup hybridization primers (i.e., host hybridization primers) and a similarity greater than 90%, the hybridization primer sequences are finally output according to the frequency ranking. For example, the backup hybridization primers aligned with the microbial genome are removed, and they are re-ranked according to the frequency of occurrence on the host genome, and the top 500 or top 384 sequences are selected.

[0076] In one embodiment, the hybridization primers according to any one of the first aspects are designed by the method according to any one of the second aspects.

[0077] According to a third aspect, in one embodiment, there is provided an application of the hybridization primer described in any one of the first aspects in blocking genomic sequences and / or removing genomic sequences.

[0078] In one embodiment, the hybridization primer is used to block the genome.

[0079] In one embodiment, the hybridization primer is used to block genomic sequences in targeted pathogen capture sequencing (tNGS).

[0080] In one embodiment, the hybridization primer is used to bind to DSN to remove host nucleic acids in metagenomic next-generation sequencing (mNGS).

[0081] According to a fourth aspect, in one embodiment, there is provided an application of the hybridization primer designed by the method described in any one of the second aspects in blocking or removing the genome.

[0082] According to a fifth aspect, in one embodiment, there is provided a method for constructing a targeted pathogen capture sequencing library, including:

[0083] A hybridization step, including providing the hybridization primer described in any one of the first aspects, or the hybridization primer designed by the method described in any one of the second aspects, hybridizing the hybridization primer with a library to be tested to obtain a hybridized library, and at least part of the host gene sequences in the hybridized library bind to the hybridization primer and thus are blocked;

[0084] A targeted capture step, including using a probe to perform targeted capture on the hybridized library to obtain a library after targeted capture. The probe is used to perform targeted capture on nucleic acid molecules of a pathogenic organism (such as a pathogenic microorganism).

[0085] In one embodiment, in the targeted capture step, after using a probe to perform targeted capture on the hybridized library, PCR amplification is performed using a sequencing primer to obtain a sequencing library, which is a library that can be used for on-machine sequencing.

[0086] In one embodiment, in the targeted capture step, after using a probe to perform targeted capture on the hybridized library, magnetic bead purification is first performed, and then PCR amplification is performed using a sequencing primer.

[0087] In one embodiment, the library to be tested is a library after end repair, A-tailing, adapter ligation, and PCR amplification.

[0088] In one embodiment, the library to be tested is a library constructed from nucleic acids extracted from a host. The nucleic acids of the host may contain a pathogenic organism, such as a pathogenic microorganism. After blocking the host genomic nucleic acids with a hybridization primer, and then using a probe to perform targeted capture on the nucleic acid molecules of the possible pathogenic organism, the pathogenic organism can be detected.

[0089] According to a sixth aspect, in one embodiment, a method for removing a genome from a host is provided, including:

[0090] A hybridization step, including providing a sample to be tested, and hybridizing the sample to be tested with the hybridization primers described in any one of the first aspect or the hybridization primers designed by the method described in any one of the second aspect to obtain a hybridization library;

[0091] A double-stranded nucleic acid molecule removal step, including using a double-stranded specific nuclease to remove double-stranded nucleic acid molecules in the hybridization library to obtain a library after removing the host genome.

[0092] In one embodiment, in the hybridization step, the sample to be tested is a nucleic acid sample extracted from a host. The sample may contain nucleic acids of at least one pathogenic organism (such as a pathogenic microorganism).

[0093] In one embodiment, in the hybridization step, the sample to be tested contains DNA and DNA reverse-transcribed from RNA, and these nucleic acid molecules are mainly derived from the host.

[0094] In one embodiment, in the hybridization step, the nucleic acid molecules in the sample to be tested are nucleic acid molecules after fragmentation, end repair, and A-addition reactions.

[0095] In one embodiment, in the hybridization step, the preparation method of the sample to be tested includes: extracting total DNA and total RNA from a biological sample, reverse-transcribing the RNA to obtain total DNA and DNA reverse-transcribed from RNA, and then performing fragmentation, end repair, and A-addition reactions to obtain the sample to be tested.

[0096] In one embodiment, in the double-stranded nucleic acid molecule removal step, the library after removing the host genome is subjected to PCR amplification to obtain a library that can be used for next-generation sequencing.

[0097] In one embodiment, using the primers and methods of the present invention for target region enrichment can solve the following technical problems:

[0098] 1) The nucleic acid sequences of the hybridization primers are clear, with high effectiveness. The lengths are 60 nt and 80 nt, and the coverage of the host genome is wide, and they can be effectively applied to the binding and removal of host nucleic acids. As Figure 2 can be seen, both 60 nt and 80 nt have relatively high coverage, and 60 nt is preferably used.

[0099] 2) Screening using a microbial database can be used to block host genome repetitive sequences during forward capture without affecting the binding of the target regions of other microorganisms.

[0100] 3) Using a microbial database for screening is less likely to result in the loss of microbial sequences during the host removal process of the mNGS DSN method.

[0101] In one embodiment, if Figure 1 As shown, the hybridization primers provided by the present invention can be used for both forward capture (the hybridization primers are used to block invalid nucleic acids) and reverse capture (for example, for removing host nucleic acids).

[0102] In one embodiment, the design process of hybridization primers for human genome repetitive sequences in the present invention is as follows:

[0103] 1. Preparation of hybridization primer pool for host genome: Use jellyfish software to break the human Grch37 genome into 60bp and 80bp kmers, screen kmer fragments with GC content between 30% and 70%, and sort them from high to low in frequency. The top ranked regions are highly repetitive regions suitable for designing hybridization primers (such as Figure 2 ). Then, the screened hybridization primers were clustered according to 2 / 3 homology using vsearch software, and each representative sequence generated after clustering was a backup hybridization primer for human repetitive sequence removal.

[0104] 2. Preparation of microbial reference genome files:

[0105] Download the nt database nt.gz from the ncbi website (download address: https: / / ftp.ncbi.nih.gov / blast / db / FASTA / nt.gz). Use the software taxonkit to process the database and obtain the taxid of all species under viruses, fungi, bacteria, archaea and parasites. Then, according to the nucleic acid sequence ID (Accession) and the taxid corresponding file "nucl_gb.accession2taxid.gz", the nucleic acid sequence ID of all viruses, fungi, bacteria, archaea and parasites is obtained. Finally, the software seqtk is used to obtain the nucleic acid reference genome sequence microbe.fa of all viruses, fungi, bacteria, archaea and parasites from the file nt.gz.

[0106] 3. Filter and screen the hybridization primers of the host genome

[0107] The microbe.fa file obtained in step 2 was used to construct a blastn index using the software makeblastdb. The human host hybrid primers constructed in step 1 were aligned to microbe.fa using the software blastn. Human host hybrid primers with alignment length greater than 90% of the length of the human host hybrid primer and similarity greater than 90% were removed. Finally, a list of sequences was output according to ranking.

[0108] In one embodiment, compared with the commonly used human genome Cot-1 fragments, the sequences of the hybridization primers of the present invention are known, which can avoid hybridization or blocking errors.

[0109] In one embodiment, the hybridization primer sequences are aligned with and removed from a microbial database, and can be used for host removal or blocking in mNGS or tNGS.

[0110] In one embodiment, since the human genome is about 3G and the repetitive sequences account for 30-50% of the total genome, it is very important to select the most effective regions and design primers with appropriate lengths and quantities. Secondly, the workload of parallel alignment and removal with the microbial database is extremely large.

[0111] In one embodiment, the beneficial effects that can be brought by enriching nucleic acid target regions using the primers and / or methods of the present invention are as follows:

[0112] 1) It can avoid hybridization or blocking errors.

[0113] 2) It can be used for host removal or blocking in mNGS or tNGS.

[0114] 3) This solution is convenient for targeted applications for different hybridization requirements. The specific regions to be removed can be removed using the hybridization primer design method of the present invention. For example, if there is interference from nucleic acids of other species in the sample and they need to be removed together, the screening method of the present invention can be used to screen the genomic sequences of that species, and then the obtained hybridization primers can be synthesized and used together.

[0115] Example 1

[0116] This example tests the capture and blocking effect of hybridization primers for tNGS, and compares them with human genome Cot-1 fragments.

[0117] Taking the construction of a hybridization capture library and then detecting the hybridization effect and detection of pathogenic microorganisms in the sample by sequencing as an example, the performance of the hybridization primers of the human genomic sequences in the present invention is illustrated. The specific steps are as follows:

[0118] 1. Construction of the Pathogenic Microorganism Standard Library: Add 49 ng of fragmented A549 DNA to TE buffer to prepare a human-derived nucleic acid background solution with a concentration of approximately 5×10^4 copies / mL. Then, add Klebsiella oxytoca at a final concentration of 50 copies / mL, Staphylococcus aureus at 50 copies / mL, Candida albicans at 50 copies / mL, Legionella pneumophila at 100 copies / mL, Listeria monocytogenes at 100 copies / mL, Fusarium oxysporum at 100 copies / mL, adenovirus at 100 copies / mL, and respiratory syncytial virus at 100 copies / mL to the prepared human-derived nucleic acid background solution to prepare a pathogenic microorganism standard.

[0119] Use the Yisheng cDNA&gDNA Library Prep kit LOT:H2123021 to construct a library for the prepared standard nucleic acid mixture. The steps include end repair, A-tailing, adapter ligation, and PCR amplification.

[0120] The library was quantified using Qubit HS dsDNA quantification reagent, and the fragment size was quality controlled using Qsep.

[0121] 2. Hybridization capture of the library using Boke hybridization reagents and Geneplus pathogen capture probes (product number: KB0032048). The human genome in the library was blocked using the host nucleic acid binding primers of the present invention (see the sequences shown in SEQ ID No.1 - 384 in Table 8), with conventional Cot-1 DNA as a control. The specific steps are as follows:

[0122] 2.1 Preparation of the hybridization system:

[0123] Take 1000 ng of the library to be hybridized and prepare a nucleic acid hybridization system in a 1.5 mL centrifuge tube according to the following table:

[0124] Table 1

[0125] Component Volume (μL) Library to be hybridized 8 Host nucleic acid hybridization primers shown as SEQ ID No. 1 - 384 in Table 8 of the present invention (1 μg / μL) 5 Adapter sequence blocking primer (3 μg / μL) 2 Total volume 15

[0126] Use a vacuum concentrator to evaporate the above-prepared nucleic acid to be hybridized at 60°C.

[0127] 2.2 Hybridization reaction

[0128] Prepare the hybridization buffer according to the following table:

[0129] Table 2

[0130] Component Volume (μL) Boke 2X Hybridization Buffer 8.5 Boke Hybridization Enhancer 2.7 Nuclease-Free Water 1.8 Total volume 13

[0131] After mixing the prepared hybridization buffer, add it to the nucleic acid to be hybridized that has been dried in step 2.1. After incubating at room temperature for 10 min, transfer the liquid to a 0.2 mL low-binding centrifuge tube, and add 4 μL of Geneplus pathogen hybridization probe. After mixing, perform the hybridization reaction according to the following procedure: 95 °C, 10 min; 65 °C, 16 h.

[0132] 2.3 Magnetic bead binding after hybridization

[0133] 2.3.1. Take 10 μL of M270 Streptavidin magnetic beads, wash them with 1X Beads Wash Buffer, and discard the supernatant.

[0134] 2.3.2. Transfer the hybridization reaction solution to the magnetic beads and incubate at 65 °C for 45 min using a PCR instrument.

[0135] 2.3.3. Vortex and mix for 3 sec every 9 min to ensure that the magnetic beads are in a suspended state.

[0136] 2.4 Washing after hybridization

[0137] 2.4.1 Heat washing (65 °C)

[0138] Take 100 μL of 1X Wash Buffer 1 preheated at 65 °C and add it to the product after hybridization. After mixing, place it on a magnetic stand to remove the supernatant.

[0139] Subsequently, add 200 μL of preheated 1X Wash Buffer S, vortex to resuspend, incubate with shaking at 1200 r at 65 °C for 5 min, and then place it on a magnetic stand to remove the supernatant. Repeat the washing once with 200 μL of preheated 1X Wash Buffer S, and then remove the supernatant.

[0140] 2.4.2 Room temperature washing

[0141] Add 180 μL of 1X Wash Buffer I to the product after heat washing, vortex to resuspend, and remove the supernatant after instantaneous centrifugation. Subsequently, add 180 μL of 1X Wash Buffer II, vortex to resuspend, and remove the supernatant after instantaneous centrifugation. Then add 180 μL of 1X Wash Buffer III, vortex to resuspend, and remove the supernatant after instantaneous centrifugation. Finally, resuspend the washed magnetic bead product with 20 μL of Nuclease-free water.

[0142] 2.5 Amplification after hybridization:

[0143] Add 25 μL of 2×Ultima HF Amplification Mix and 4 μL of F / R Index primer (20 μM, a universal primer provided by the sequencing platform) to the system after the completion of the washing in Step 2.4. After thorough mixing, perform cyclic amplification: 98°C for 1 min; 98°C for 10 s, 60°C for 30 s, 72°C for 30 s (8 cycles); 72°C for 1 min; hold at 4°C. After the amplification is completed, add 50 μL of magnetic beads to the amplification product for magnetic bead purification, and dissolve the purified product in 30 μL of TE.

[0144] 2.5 Sequencing and Data Analysis

[0145] Use the MGI single-stranded circularization kit and the MGISEQ-200RS high-throughput sequencing reagent kit (SE50) to prepare DNBs for the library amplified in Step 2.4, and use the MGISEQ200 sequencer to sequence it. The sequencing results are compared and analyzed with the pathogen database using BWA.

[0146] Example 2

[0147] This example tests the difference in the capture and blocking effects of the hybrid primers output by the method of the present invention (384 and 500) for tNGS.

[0148] The specific steps of this example are the same as those of Example 1, except that the pathogen standards are different from those of Example 1 and are Klebsiella oxytoca at 20 copies / mL, Legionella pneumophila at 50 copies / mL, and Listeria monocytogenes at 50 copies / mL.

[0149] Example 3

[0150] This example tests the effect of the hybrid primers in DSN dehosting: compare with the human genomic Cot-1 fragment

[0151] This part of the experiment is carried out according to the Chinese patent "A Method and Kit for Constructing a Microbial Sequencing Library Based on DSN" with the patent application number 202211139551.7. The primers used for hybridizing to form double strands are the hybrid primers of the present invention, and Cot-1 is used as a control for the hybrid primers. The specific steps are as follows:

[0152] 1. Construction of pathogen microorganism standards: Add 49 ng of fragmented A549 DNA to TE buffer to prepare a human-derived nucleic acid background solution with a concentration of approximately 5×10^4 copies / mL. Then, add fragmented Zymo standards (8 bacteria and 2 fungi) with a final concentration of 1×10^(-3) ng / μL to the prepared human-derived nucleic acid background solution, and mix well to prepare a standard nucleic acid mixture.

[0153] 2. Use Yisheng cDNA&gDNA Library Prep kit LOT:H2123021 library construction kit to perform end repair, add A tail and connect adapters to the prepared standard nucleic acid mixture:

[0154] 2.1 Reverse transcription and cDNA duplex synthesis of RNA in the sample were performed according to the following reaction system:

[0155] Table 3

[0156] Component Volume (μL) Total DNA and RNA after extraction 14 Random primer (1 μM) 1 Total volume 15

[0157] After mixing thoroughly according to the formula in Table 3 above, place it at 70℃ for heat denaturation for 5 minutes, and place it on ice for 3 minutes for primer binding. Then add 8μL 1st t Reaction Buffer and 2μL 1st Stand Enzyme Mix and mix thoroughly, and perform reverse transcription reaction on RNA at 25℃ for 5min, 42℃ for 30min, and 85℃ for 5min. After the reverse transcription is completed, add 7μL 2nd Reaction Buffer and 3μL 2nd Stand Enzyme Mix to the system and mix thoroughly, and place it at 16℃ for 5min for second-strand synthesis of cDNA.

[0158] 2.2 After the second-strand synthesis is completed, place the reaction product on ice, add 10μL Smearase Buffer and 5μL Smearase Enzyme Mix, and add 10μL NF H2O to adjust the reaction system to 60μL. After thorough mixing, place it at 30℃ for 25min, and then at 72℃ for 20min. This step can fragment, end-repair and add "A" tail to the total DNA and RNA in the reaction system after reverse transcription and second-strand synthesis.

[0159] 2.3 Add the corresponding reagents to the above reaction products in order according to the following table for connector connection:

[0160] Table 4

[0161] Component Volume (μL) Product in Step 2.2 60 Adapter solution (3 μM) 5 Ligation enhancer 30 T4 DNA ligase 5 Total volume 100

[0162] After the reaction mixture is fully mixed, it is placed at 20°C for 15 minutes for connector connection. After the reaction is completed, 60 μL of magnetic beads are added for magnetic bead purification, and the purified product is re-dissolved in 16 μL NF H2O (Water Nuclease-Free, nuclease-free water).

[0163] 3. Hybridization primer binding:

[0164] Add the hybridization primers of the present invention to the product after joint connection according to the following table, using the human genomic Cot-1 fragment as a control:

[0165] Table 5

[0166] Component Volume (μL) Adapter ligation product in Step 5 15 Host nucleic acid hybridization primers shown as SEQ ID No. 1 - 384 in Table 8 of the present invention (1 μg / μL) 3 DSN buffer 2 Total volume 20

[0167] After thoroughly mixing the above reaction mixture, perform the hybridization binding of the hybridization primers at the following reaction temperatures:

[0168] Table 6

[0169] Temperature Time 95℃ 5 min 65℃ 20 min

[0170] After the hybridization is completed, keep the hybridization system at 65 °C.

[0171] 4. DSN enzyme digestion treatment:

[0172] Add 1 μL of DSN nuclease to the system after the hybridization of the hybridization primers in step 3. After thoroughly mixing, continue to remove the double-stranded nucleic acid molecules in the system at the following reaction temperatures.

[0173] Table 7

[0174] Temperature Time 65℃ 10 min 95℃ 5 min

[0175] 5. PCR amplification:

[0176] Place the system after the digestion in step 4 on ice. Add 25 μL of 2×KAPA HiFi hotstart ready Mix and 4 μL of F / R Index primer (20 μM) in sequence. After thoroughly mixing, perform cycle amplification: 98 °C for 1 min; 98 °C for 10 s, 60 °C for 30 s, 72 °C for 30 s (8 cycles); 72 °C for 1 min; hold at 4 °C. After the amplification is completed, add 45 μL of magnetic beads to the amplification product for magnetic bead purification, and the purified product is redissolved in 30 μL of TE.

[0177] 6. Sequencing and data analysis

[0178] Use the MGI single-stranded circularization kit and the MGISEQ-200RS high-throughput sequencing reagent kit (SE50) to prepare DNB (DNA nanoballs) for the library amplified in step 5, and use the MGISEQ200 sequencer to sequence it. The sequencing results are compared and analyzed with the pathogen database using BWA.

[0179] The experimental results of Example 1 are as follows:

[0180] Using the hybridization primers in Table 8 for blocking, the obtained results are as Figure 3 .

[0181] Figure 3 Among them, the vertical coordinate "RPM" refers to the number of reads detected for the microorganism divided by the total amount of nucleic acid sequencing data (M) in the sample, that is, the commonly used index RPM (Reads per million) for mNGS microorganism detection.

[0182] The results show that using the hybridization step of the hybridization primers of the present invention for metagenomic tNGS targeted sequencing can effectively detect pathogenic standards such as Klebsiella oxytoca, Staphylococcus aureus, Candida albicans, Legionella pneumophila, Listeria monocytogenes, Fusarium oxysporum, adenovirus, and respiratory syncytial virus in the background of normal human nucleic acid, indicating that the hybridization primers of the present invention can be effectively used for metagenomic tNGS sequencing and have good performance. Compared with the hybridization step using conventional human genomic Cot-1 fragments, the detection effect of known pathogens of the hybridization primers of the present invention is the same, but the closed sequence information of the present invention is known, which is more conducive to monitoring the loss of nucleic acids of more unknown pathogenic microorganisms after being closed.

[0183] The experimental results of Example 2 are as follows:

[0184] Using the first 384 hybridization primers (SEQ ID No.1 - 384) output by the method of the present invention in Table 8 and the first 500 hybridization primers (SEQ ID No.1 - 500) output by the method of the present invention for blocking, the obtained results are as Figure 4 .

[0185] Figure 4 Among them, the vertical coordinate "RPM" refers to the number of reads detected for the microorganism divided by the total amount of nucleic acid sequencing data (M) in the sample, that is, the commonly used index RPM (Reads per million) for mNGS microorganism detection.

[0186] The results show that the blocking effect of the first 384 hybridization primers selected by the present invention is the same as that of the human genomic Cot-1 fragment and the first 500 hybridization primers output by the method of the present invention, and their pathogenic detection effects on the standards are basically the same. Since the synthesizer can complete one-round synthesis and has lower cost when using a 384-well plate for hybridization primer synthesis, the first 384 hybridization primers are selected as the number of hybridization primers of the present invention.

[0187] The experimental results of Example 3:

[0188] After hybridizing with the hybridization primers shown in SEQ ID No.1 - 384 in Table 8 of the present invention and removing host nucleic acids, the obtained results are as Figure 5 .

[0189] Figure 5Among them, the vertical coordinate "RPM" refers to the number of reads detected for the microorganism divided by the total amount of nucleic acid sequencing data (M) in the sample, that is, the commonly used mNGS microorganism detection index RPM (Reads per million).

[0190] In the DSN host removal step for metagenomic mNGS sequencing using the hybridization primers of the present invention, compared with the control group without host nucleic acid removal, the hybridization primers of the present invention in combination with the DSN host removal method can effectively increase the detection of pathogenic microorganisms in the background of normal human nucleic acid by about 3 to 4 times. Compared with using human genomic Cot-1 fragments for double-stranded hybridization, the double-stranded binding effect of the hybridization primers of the present invention to host nucleic acids is better, and the nucleic acid sequences of the hybridization primers of the present invention are known, which is more conducive to monitoring the loss of more unknown pathogenic microorganism nucleic acids after double-stranded binding.

[0191] In one embodiment, the hybridization primers provided by the present invention can not only be used for hybridization capture, host nucleic acid removal and other purposes, but also can be used in other human or other organism genome closure schemes, such as locked nucleic acid inhibiting the reverse transcription of RNA expressed in human regions.

[0192] Locked nucleic acid (LNA) is a modified RNA in which the 2' and 4' carbons on a part of the ribose are linked together. It is generally found in A-DNA or RNA.

[0193] In one embodiment, the hybridization primer design method provided by the present invention is not only applicable to humans, but also to other species, such as plants, other animals, etc.

[0194] Table 8 Partial nucleic acid sequences of the hybridization primers of the present invention and their coverage of the host genome

[0195]

[0196]

[0197]

[0198]

[0199]

[0200]

[0201]

[0202]

[0203]

[0204]

[0205]

[0206]

[0207]

[0208]

[0209]

[0210]

[0211]

[0212]

[0213] The above uses specific examples to elaborate on the present invention, which is only used to help understand the present invention and is not intended to limit the present invention. For those skilled in the art of the present invention, based on the idea of the present invention, several simple deductions, deformations or substitutions can also be made.

Claims

1. A hybrid primer combination, characterized in that, Composed of hybridization primers shown by nucleotide sequences such as SEQ ID No.1 to 384.

2. A hybrid primer combination, characterized in that, Composed of all of the group composed of hybridization primers shown by nucleotide sequences such as SEQ ID No.1 to 384 and at least one of the group composed of hybridization primers shown by nucleotide sequences such as SEQ ID No.385 to 500.

3. The hybridization primer combination according to claim 2, wherein The hybridization primer combination is composed of all of the hybridization primers shown by SEQ ID No.1 to 500.

4. Use of the hybridization primer combination according to any one of claims 1 to 3 in blocking or removing human genomic sequences.

5. A method for constructing a pathogen-targeted capture sequencing library, characterized in that, Including: A hybridization step, including providing the hybridization primer combination according to any one of claims 1 to 3, hybridizing the hybridization primer combination with a library to be tested, and obtaining a hybridized library. A target capture step, including using a probe to perform target capture on the hybridized library to obtain a library after target capture.

6. A method for removing a host genome, characterized in that, Including: A hybridization step, including providing a sample to be tested, using the hybridization primer combination according to any one of claims 1 to 3 to hybridize with the sample to be tested, and obtaining a hybridized library. A double-stranded nucleic acid molecule removal step, including using a double-stranded specific nuclease to remove double-stranded nucleic acid molecules in the hybridized library to obtain a library after removing the host genome. Wherein, the host genome is a human genome.

7. The method according to claim 6, characterized in that, In the hybridization step, the sample to be tested is a nucleic acid sample extracted from a host. Or, in the hybridization step, the sample to be tested contains DNA and DNA reverse transcribed from RNA. Or, in the hybridization step, the nucleic acid molecules in the sample to be tested are nucleic acid molecules after fragmentation, end repair, and "A" addition reaction. Or, in the double-stranded nucleic acid molecule removal step, the library after removing the host genome is subjected to PCR amplification to obtain a library that can be used for on-machine sequencing.

Citation Information

Patent Citations

  • Method and kit for constructing microbial sequencing library based on DSN

    CN117721178A

  • Method for removing non-target RNA (Ribonucleic Acid) in RNA sample

    CN107119043A

  • Method and kit for detecting infection lines based on high-throughput sequencing

    CN110093409A

  • Closing primer for blocking repetitive sequence on human genome, as well as composition and application of closing primer

    CN113969277A