Pathogen detection method based on targeted high-throughput sequencing
By adopting targeted high-throughput sequencing technology and innovative bioinformatics algorithms in pathogen detection, the sample preparation and data analysis process is optimized, and the problems of slow detection speed, high cost and insufficient sensitivity in the existing technology are solved, and fast and accurate pathogen detection is achieved, which is suitable for a variety of application scenarios.
Patent Information
- Application Number
- CN202510111401.2
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-01-23
- Publication Date
- 2025-05-13
AI Technical Summary
The existing pathogen detection technology has problems such as long sample preparation time, complex data analysis, high cost, and difficulty in distinguishing the nucleic acids of the host from the pathogen, resulting in slow detection speed and insufficient sensitivity and specificity.
Using a pathogen detection method based on targeted high-throughput sequencing, we can reduce detection costs, improve sensitivity and specificity, and simplify the data analysis process by optimizing sample preparation process, introducing innovative bioinformatics algorithms and computing strategies. Specific steps include data acquisition, data quality control, comparison of pathogen genomic data, pathogen signal recognition, NT database verification and pathogen grading.
It realizes the rapid and accurate identification and quantitative analysis of various pathogens in clinical samples, reduces false positives and background signal interference in the detection results, improves the accuracy and efficiency of detection, and is suitable for disease diagnosis, epidemiological investigation and public health prevention and control.
Smart Images

Figure CN119993263A_ABST
Abstract
Description
Technical Field
[0001] The present invention belongs to the technical field of gene detection and analysis, and specifically relates to a pathogen detection method based on targeted high-throughput sequencing. Background Art
[0002] In the fields of medicine and biology, rapid and accurate detection of pathogens is crucial for disease diagnosis, epidemiological investigations, and public health prevention and control. Traditional pathogen detection methods include culture, serological testing, PCR (polymerase chain reaction), etc., but these methods often have problems such as low sensitivity, insufficient specificity, long time consumption, or can only target specific pathogens. With the development of genomics, especially the advancement of high-throughput sequencing (HTS) technology, we can analyze a large number of DNA or RNA sequences in a short period of time, thereby greatly improving the speed and accuracy of pathogen detection.
[0003] High-throughput sequencing, also known as massively parallel sequencing, is a technology that can perform parallel sequencing on a large number of nucleic acid molecules at a time. Usually, a sequencing reaction can produce no less than 100Mb of sequencing data and is widely used in genomic research. It allows all nucleic acids in a complex sample mixture to be sequenced at a relatively low cost and high efficiency. However, in pathogen detection applications, high-throughput sequencing methods still have problems such as long sample preparation time, complex data analysis, high cost, and difficulty in distinguishing host and pathogen nucleic acids.
[0004] Based on this, in order to overcome some of the limitations of existing pathogen detection technologies, increase detection speed, reduce costs, improve sensitivity and specificity, and simplify data analysis processes, it is urgent to improve or develop new algorithms, optimize analysis steps, or introduce innovative bioinformatics tools so that high-throughput sequencing technology can be used more effectively for pathogen detection. Summary of the invention
[0005] In view of this, the present invention provides a pathogen detection method based on targeted high-throughput sequencing. The method can quickly and accurately identify and quantitatively analyze a variety of pathogens in clinical samples, including but not limited to bacteria, viruses, fungi and parasites. This method reduces detection costs, improves sensitivity and specificity, and simplifies the data analysis process by optimizing the sample preparation process and introducing innovative bioinformatics algorithms and computing strategies.
[0006] In order to achieve the above object, the present invention adopts the following technical solutions: A pathogen detection method based on targeted high-throughput sequencing comprises the following steps: S1. Data acquisition: Obtain the NGS raw sequencing data of the pathogen samples to be tested and the negative control samples; S2. Data quality control: First, remove low-quality data from the original sequencing data, then filter the host sequence, and finally filter the plasmid sequence to obtain high-quality sequencing reads; S3. Align pathogen genome data: Align high-quality sequencing reads with pathogen reference genomes, retain read data based on alignment identity and InDel length, and obtain specific read data; S4. Pathogen signal identification: Classify the specific read data according to pathogens, calculate the number of specific reads and their RPM value for each pathogen, and retain the sequence information of all specific reads; Among them, specific reads are defined as: if a read can only be matched to one pathogen or it can be matched to multiple pathogens but these pathogens belong to different subspecies of the same species; RPM is defined as: the number of specific reads for the pathogen per million reads; The specific algorithm for the number of specific reads and RPM is as follows: Each read may be mapped to ≥1 pathogen, each pathogen belongs to a species, and each species belongs to a genus. According to the unique classification number of the pathogen, species, and genus, they are recorded as TaxID, S_TaxID, and G_TaxID respectively; Count the TaxID list, S_TaxID list, and G_TaxonID list on the read alignment one by one; if the number of S_TaxonID is 1, the number of specific reads for each pathogen in the TaxID list will increase by 1; if the number of G_TaxonID is 1, the number of specific reads for each species in the S_TaxID list will increase by 1, and the number of specific reads for each genus in the G_TaxID will increase by 1; According to the formula , you can get the RPM, S_RPM, and G_RPM of each pathogen, species, and genus; where the total number of sequencing reads in the formula is the original sequencing data in step S1; S5. NT database verification: construct a microbial reference genome database based on the nucleic acid sequence database, compare the sequence data of the specific Read in step S4 with the constructed microbial reference genome database, and verify whether each Read still has species specificity in the microbial reference genome based on the alignment identity and InDel length, so as to calculate the S_RPM value of each pathogen in the NT database according to the RPM formula in step S4 and record it as NT_RPM; S6. Pathogen classification: define the signal value Signal = RPM-0.001×the maximum RPM of the same batch of samples, and classify each pathogen detection signal into strong, medium and weak intensity levels according to the signal value. The specific levels are as follows: RPM≥1000 is a strong signal, 1000>RPM≥200 is a medium intensity, and RPM<200 is a weak signal.
[0007] Furthermore, the negative control sample in step S1 is enzyme-free water.
[0008] In some specific embodiments, preferably, the indicators for removing low-quality data in step S2 include: the base quality score should be Q30 ≥ 95%, the GC content should be in the range of 40% to 60%, reads with a quality lower than 30 are removed and the proportion of low-quality reads should be lower than 10%, and reads with a complexity lower than 60% are removed.
[0009] Furthermore, the host sequence filtering in step S2 is specifically as follows: comparing the sequencing data after removing low-quality data with the human reference genome, and extracting and retaining those read sequences that fail to be effectively aligned with the human genome; Plasmid sequence filtering is specifically as follows: the sequencing data filtered by the host sequence is compared with the plasmid reference genome, and the read sequences that cannot be effectively compared with the plasmid genome are extracted and retained.
[0010] In some specific embodiments, preferably, the degree of identity of the alignment in step S3 is ≥90%, and the InDel length is ≤3 bp.
[0011] In some specific embodiments, preferably, the pathogen genome data in step S3 and the pathogen classification information in step S4 are both derived from the species information in the "Catalogue of Pathogenic Microorganisms Transmitted Among Humans".
[0012] Furthermore, the microbial reference genome database in step S5 includes nucleotide sequences of archaea, bacteria, fungi, viruses and protists.
[0013] In some specific embodiments, preferably, the degree of identity of the comparison in step S5 is ≥90%, and the InDel length is ≤3 bp.
[0014] Furthermore, the level division in step S6 also includes: For strong or medium signals, if , then the signal strength is degraded; For strong or medium signals, define the negative control sample RPM as the maximum RPM value of the pathogen in all negative control samples. It is set as background signal; For strong or medium signals, the number of signal samples in the same batch is defined as the number of samples of the pathogen in the same batch whose signal strength is not lower than that of the current sample. , it is set as background signal.
[0015] The above detection method is used for detecting pathogens.
[0016] Compared with the prior art, the present invention has the following beneficial effects: Based on the high-throughput sequencing results, the present invention first removes impurities from the data, removes host data and plasmid data, and then sequentially compares the pathogen genome and the microbial NT library to obtain the pathogen-specific Read and its RPM value to be tested, and finally classifies the detected signal into four intensity levels of strong, medium, weak, and background, thereby determining the pathogen information in the sample to be tested. The detection method of the present invention can effectively reduce the false positive of the test results and the interference of the background signal, and the test results are accurate, and can be well applied to disease diagnosis, epidemiological investigation and public health prevention and control. BRIEF DESCRIPTION OF THE DRAWINGS
[0017] Figure 1 It is a schematic diagram of the detection method of the present invention. DETAILED DESCRIPTION
[0018] The present invention is further described in detail below in conjunction with specific embodiments so that those skilled in the art can understand the present invention more clearly.
[0019] Sources and physical and chemical parameters of key test materials: Sample CB5002G101KD was provided by the laboratory of Wuhan Kangsheng Zhenyuan Medical Laboratory Co., Ltd. as a test sample, and mainly contained pathogens such as Aspergillus and Veillonella parvum.
[0020] Example 1 This example performs pathogen detection on sample CB5002G101KD, and the specific steps are as follows: S1. Take 10 samples of CB5002G101KD for detection and 3 negative control samples (enzyme-free water) as samples from the same batch for high-throughput sequencing and obtain NGS raw sequencing data.
[0021] S2. First, the quality of the NGS raw sequencing data (referred to as Reads) was evaluated using quality control software (Fastp). The base quality score should be Q30 ≥ 95%, the GC content should be in the range of 40% to 60%, and the Reads with a quality lower than 30 should be removed. The proportion of low-quality Reads should be lower than 10%, and the Reads with a complexity lower than 60% should be removed to remove low-quality data.
[0022] Then, the low-quality data were compared with the human reference genome using the alignment software (BBTools), and the read sequences that failed to be effectively aligned with the human genome were extracted and retained. The purpose of this step is to remove the host-derived data introduced during the experiment and sequencing process.
[0023] Finally, the host sequence-filtered sequencing data was aligned with the plasmid reference genome using alignment software (BBTools), and the read sequences that failed to be effectively aligned with the plasmid genome were extracted and retained. There is a lot of homologous sequence interference in the plasmid data, and removing the plasmid data in this step helps to improve the specificity of the detection results and optimize the use of computing resources.
[0024] High-quality sequencing reads are obtained through the above processing.
[0025] S3. Prepare pathogenic microorganism reference genomes according to the species information of the "Catalogue of Pathogenic Microorganisms Transmitted in Humans" issued by the National Health Commission. Use the alignment software (BBTools) to align the obtained high-quality sequencing reads with the pathogenic microorganism reference genomes and extract the read data with a retention alignment rate of ≥90% and an insertion and deletion (InDel) length of ≤3bp, that is, obtain specific read data.
[0026] S4. Classify the pathogens according to the alignment results, calculate the number of specific reads and their RPM values for each pathogen, and retain the sequence information of all specific reads.
[0027] in; A specific read is defined as a read that can only be matched to one pathogen or that can be matched to multiple pathogens but these pathogens belong to different subspecies of the same species.
[0028] RPM is defined as the number of pathogen-specific reads per million reads.
[0029] The specific algorithm for the number of specific reads and RPM is as follows: Each read may be mapped to ≥1 pathogen, each pathogen belongs to a species, and each species belongs to a genus. According to the unique taxonomy ID of the pathogen, species, and genus, they are recorded as TaxID, S_TaxID, and G_TaxID respectively; Count the TaxID list, S_TaxID list, and G_TaxonID list on the read alignment one by one; if the number of S_TaxonID is 1, the number of specific reads for each pathogen in the TaxID list will increase by 1; if the number of G_TaxonID is 1, the number of specific reads for each species in the S_TaxID list will increase by 1, and the number of specific reads for each genus in the G_TaxID will increase by 1; According to the formula , the RPM, S_RPM, and G_RPM of each pathogen, species, and genus can be obtained; where the total number of sequencing reads in the formula is the original sequencing data in step S1.
[0030] In this step, RPM is a data normalization algorithm that can effectively solve the problem of inconsistent specific read quantity standards caused by inconsistent sequencing quantities of different samples.
[0031] S5. Verify the above results with the NT database: extract the archaea in the NCBI nucleic acid sequence database (NT) archa ),bacteria( bacteria ), fungi ( mushrooms ),Virus( viral ) and protists ( protozoa ) nucleic acid sequence to construct a microbial reference genome database, compare the sequence data of the specific Read in step S4 with the constructed microbial reference genome database, and verify whether each Read still has species specificity in the microbial reference genome one by one based on the Read data with a retained alignment rate ≥ 90% and an insertion and deletion (InDel) length ≤ 3bp, thereby calculating the S_RPM value of each pathogen in the NT database according to the RPM formula in step S4, recorded as NT_RPM.
[0032] S6. Finally, the detection signal of each pathogen detected in the sample to be tested is graded into four intensity levels: strong, medium, weak, and background. The specific algorithm for each pathogen is as follows: 1) Define the signal value Signal = RPM-0.001×maximum RPM of samples in the same batch; 2) Define RPM ≥ 1000 as a strong signal, 1000> RPM ≥ 200 as a medium strength, and RPM < 200 as a weak signal; 3) For strong or medium signals, if , then the signal strength is degraded; 4) For strong or medium signals, define the negative control sample RPM as the maximum RPM value of the pathogen in all negative control samples. It is set as background signal; 5) For strong or medium signals, the number of signal samples in the same batch is defined as the number of samples of the pathogen in the same batch whose signal strength is not lower than that of the current sample. , it is set as background signal.
[0033] The final test results are shown in Table 1.
[0034] Table 1 Results of sample CB5002G101KD
[0035]
[0036] From the results in Table 1, we can see that the sample mainly detected several pathogens of the genus Aspergillus and Campylobacterconcisus, etc. Among them, Clostridium perfringens and Human gammaherpesvirus 4 had higher signals in negative control samples or batch samples and were determined to be background bacteria; while Aspergillus flavus, Veillonella parvula, etc. were downgraded after NT verification. The downgrade of the signal indicates that there is a false positive in the detection of the pathogen, and once it is downgraded to a weak signal, it is considered a false signal.
[0037] Comparative Example 1 This comparative example performs pathogen detection on sample CB5002G101KD, and its steps are basically the same as those of Example 1, except that the verification of the result with the NT database is omitted, and the rest remain unchanged. The specific detection results are shown in Table 2.
[0038] Table 2 Partial test results of sample CB5002G101KD
[0039] As can be seen from Table 2: Although there are a large number of reads for Aspergillus parasiticus and Aspergillus tubingensis, most of the reads are sequences homologous to other pathogens, that is, the pathogen source of the reads cannot be determined. Therefore, comparing the sequence data of the obtained specific reads with the constructed microbial reference genome database can effectively reduce the false positives of the test results.
[0040] Comparative Example 2 This comparative example performs pathogen detection on sample CB5002G101KD, and its steps are basically the same as those in Example 1, except that the standard for signal grading in step S6 is not adopted (i.e., the negative control sample RPM is not used for grading), and the rest remains unchanged. The specific detection results are shown in Table 3.
[0041] Table 3 Partial test results of sample CB5002G101KD
[0042] As can be seen from Table 3, Clostridium perfringens and Human gammaherpesvirus 4 both had RPM values that met the medium intensity before the negative control algorithm, but because Clostridium perfringen negative control RPM also had a high RPM, its signal intensity was adjusted to the background. At the same time, Human gammaherpesvirus 4 had a signal intensity above medium in 6 of the 10 samples in the same batch, so its signal intensity was adjusted to the background. While Aspergillus flavus and Campylobacter concisus also met the medium intensity signal, their negative control RPM and the same batch of samples did not affect the signal intensity. This approach can effectively prompt report interpreters to pay attention in a timely manner.
[0043] The specific raw materials not described in the present invention are all existing materials and can be directly purchased from the market.
[0044] The above are only preferred embodiments of the present invention and are not intended to limit the protection scope of the present invention. Any modifications, equivalent substitutions, improvements, etc. made within the spirit and principles of the present invention should be included in the protection scope of the present invention.
Claims
1. A pathogen detection method based on targeted high-throughput sequencing, characterized in that: The following steps are involved: S1. Data acquisition: Obtain the NGS raw sequencing data of the pathogen samples to be tested and the negative control samples; S2. Data quality control: First, remove low-quality data from the original sequencing data, then filter the host sequence, and finally filter the plasmid sequence to obtain high-quality sequencing reads; S3. Align pathogen genome data: Align high-quality sequencing reads with pathogen reference genomes, retain read data based on alignment identity and InDel length, and obtain specific read data; S4. Pathogen signal identification: Classify the specific read data according to pathogens, calculate the number of specific reads and their RPM value for each pathogen, and retain the sequence information of all specific reads; Among them, specific reads are defined as: if a read can only be matched to one pathogen or it can be matched to multiple pathogens but these pathogens belong to different subspecies of the same species; RPM is defined as: the number of specific reads for the pathogen per million reads; The specific algorithm for the number of specific reads and RPM is as follows: Each read may be mapped to ≥1 pathogen, each pathogen belongs to a species, and each species belongs to a genus. According to the unique classification number of the pathogen, species, and genus, they are recorded as TaxID, S_TaxID, and G_TaxID respectively; Count the TaxID list, S_TaxID list, and G_TaxonID list on the read alignment one by one; if the number of S_TaxonID is 1, the number of specific reads for each pathogen in the TaxID list will increase by 1; if the number of G_TaxonID is 1, the number of specific reads for each species in the S_TaxID list will increase by 1, and the number of specific reads for each genus in the G_TaxID will increase by 1; According to the formula , you can get the RPM, S_RPM, and G_RPM of each pathogen, species, and genus; where the total number of sequencing reads in the formula is the original sequencing data in step S1; S5. NT database verification: construct a microbial reference genome database based on the nucleic acid sequence database, compare the sequence data of the specific Read in step S4 with the constructed microbial reference genome database, and verify whether each Read still has species specificity in the microbial reference genome based on the alignment identity and InDel length, so as to calculate the S_RPM value of each pathogen in the NT database according to the RPM formula in step S4 and record it as NT_RPM; S6. Pathogen classification: define the signal value Signal = RPM-0.001×the maximum RPM of the same batch of samples, and classify each pathogen detection signal into strong, medium and weak intensity levels according to the signal value. The specific levels are as follows: RPM≥1000 is a strong signal, 1000>RPM≥200 is a medium intensity, and RPM<200 is a weak signal.
2. The detection method according to claim 1, characterized in that: The negative control sample in step S1 is enzyme-free water.
3. The detection method according to claim 1, characterized in that: Indicators for removing low-quality data in step S2 include: base quality score should be Q30 ≥ 95%, GC content should be in the range of 40% to 60%, reads with quality lower than 30 should be removed and the proportion of low-quality reads should be lower than 10%, and reads with complexity lower than 60% should be removed.
4. The detection method according to claim 1, characterized in that: The host sequence filtering in step S2 is specifically as follows: the sequencing data after removing low-quality data are compared with the human reference genome, and the read sequences that cannot be effectively compared with the human genome are extracted and retained; Plasmid sequence filtering is specifically as follows: the sequencing data filtered by the host sequence is compared with the plasmid reference genome, and the read sequences that cannot be effectively compared with the plasmid genome are extracted and retained.
5. The detection method according to claim 1, characterized in that: The alignment identity in step S3 is ≥90%, and the InDel length is ≤3 bp.
6. The detection method according to claim 1, characterized in that: The pathogen genome data in step S3 and the pathogen classification information in step S4 are both derived from the species information in the Catalogue of Pathogenic Microorganisms Transmitted Among Humans.
7. The detection method according to claim 1, characterized in that: The microbial reference genome database in step S5 includes nucleotide sequences of archaea, bacteria, fungi, viruses and protists.
8. The detection method according to claim 1, characterized in that: The alignment in step S5 has a degree of identity of ≥90% and an InDel length of ≤3 bp.
9. The detection method according to claim 1, characterized in that: The level division in step S6 also includes: For strong or medium signals, if , then the signal strength is degraded; For strong or medium signals, define the negative control sample RPM as the maximum RPM value of the pathogen in all negative control samples. It is set as background signal; For strong or medium signals, the number of signal samples in the same batch is defined as the number of samples of the pathogen in the same batch whose signal strength is not lower than that of the current sample. , it is set as background signal.
10. Use of the detection method according to any one of claims 1 to 9 in detecting pathogens.
Citation Information
Cited By
Clinical pathogen infection interpretation method based on machine learning
CN121122395A