High-throughput sequencing detection method and system for pathogenic microorganisms based on multiple PCR (Polymerase Chain Reaction) detection
By quantitatively adding internal reference nucleic acid markers and performing multiplex PCR amplification, combined with quality control judgment and host genome comparison, the problem of insufficient amplification specificity and sensitivity in high-throughput sequencing of pathogenic microorganisms was solved, achieving efficient and accurate detection of pathogenic microorganisms.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-12-05
- Publication Date
- 2026-03-27
AI Technical Summary
Existing multiplex PCR detection methods suffer from problems such as insufficient amplification specificity and sensitivity, difficulty in guaranteeing data quality and accuracy, and insufficient integration of intelligent algorithm analysis with clinical information in high-throughput sequencing of pathogenic microorganisms.
A standardized DNA template library was generated by quantitatively adding internal reference nucleic acid markers, followed by multiplex PCR amplification and quality control assessment. A PCR quality control result table was generated, a sequencing library was constructed, and the host reference genome was compared with the genome. The host sequence was removed, and finally, a pathogen species comparison analysis was performed.
It improved the comparability and accuracy of experimental results, ensured the detection of low-abundance pathogens, saved time and costs, increased pathogen species coverage, optimized the library construction process, reduced inter-sample variability, improved detection sensitivity and specificity, and ensured the accuracy and reproducibility of sequencing data.
Smart Images

Figure CN121737281A_ABST
Abstract
Description
Technical Field
[0001] This invention relates to the field of DNA sequencing, and more particularly to a high-throughput sequencing detection method and system for pathogenic microorganisms based on multiplex PCR detection. Background Technology
[0002] In recent years, with the rapid development of next-generation sequencing (NGS) technology, its application prospects in the field of pathogen detection have gradually attracted widespread attention. High-throughput sequencing technology has advantages such as simultaneous detection of large numbers of pathogens, no need for culturing, and no need for pre-hypothesizing of pathogens. It can rapidly and comprehensively identify various types of bacteria, viruses, fungi, and other microorganisms in pathogen communities, greatly improving the detection sensitivity and accuracy of pathogens. However, due to the wide variety and complexity of pathogens, the sequencing depth, data volume, and analysis difficulty of high-throughput sequencing still pose certain technical challenges to the efficient detection of pathogens. Existing multiplex PCR detection methods still face some problems in high-throughput sequencing of pathogens. First, due to the differences in gene characteristics among different pathogens, designing efficient multiplex PCR primers to improve the specificity and sensitivity of amplification is a key issue. Second, when dealing with complex pathogen populations, how to remove background noise and improve data quality and accuracy remains a technical challenge for existing high-throughput sequencing technologies. Finally, further optimization and improvement are needed to determine how to quickly and accurately analyze high-throughput sequencing data using intelligent algorithms and combine it with clinical information to provide effective diagnostic support. Summary of the Invention
[0003] To address the aforementioned technical problems, this invention proposes a high-throughput sequencing detection method and system for pathogenic microorganisms based on multiplex PCR detection, thereby solving at least one of the aforementioned technical problems.
[0004] To achieve the above objectives, this invention provides a high-throughput sequencing detection method for pathogenic microorganisms based on multiplex PCR detection, comprising the following steps: Step S1: Quantitatively add internal reference nucleic acid markers to the first microbial sample to be tested to generate a standardized DNA template library; Step S2: Perform multiplex PCR amplification on the standardized DNA template library and conduct quality control assessment to generate a PCR quality control result table; Step S3: Construct sequencing libraries based on the PCR quality control result table to generate raw sequencing library data; Step S4: Align the raw sequencing library data with the host reference genome and remove any missing data, then extract the dataset to be compared. Step S5: Perform pathogen species comparison analysis based on the dataset to be compared, and output the final pathogen detection list.
[0005] This specification provides a high-throughput sequencing detection system for pathogenic microorganisms based on multiplex PCR detection, used to perform the high-throughput sequencing detection method for pathogenic microorganisms based on multiplex PCR detection as described above, including: The internal control addition unit is used to quantitatively add internal control nucleic acid markers to the first microbial sample to be tested, and generate a standardized DNA template library. The PCR amplification unit is used to perform multiplex PCR amplification on a standardized DNA template library and to evaluate and determine the quality control results, generating a PCR quality control result table. The sequencing library unit is used to construct sequencing libraries based on PCR quality control results and generate raw sequencing library data. The host genome comparison unit is used to align the raw sequencing library data with the host reference genome and remove it, and extract the dataset to be compared. The pathogen species comparison unit is used to perform pathogen species comparison analysis based on the dataset to be compared, and output the final list of pathogenic microorganisms to be detected.
[0006] The specific benefits of this invention are as follows: By adding an internal reference nucleic acid marker, the detection conditions between different samples can be ensured to be uniform, thereby improving the comparability of experimental results. This helps overcome differences caused by factors such as sample quality and concentration, ensuring the accuracy of high-throughput sequencing. Quantitative addition of the internal reference marker helps detect trace amounts of pathogen DNA, ensuring accurate detection results even in cases of low pathogen abundance. Generating a standardized DNA template library provides a uniform basis for subsequent PCR amplification, reducing the impact of sample differences on downstream detection processes. Multiplex PCR can simultaneously amplify multiple target pathogen DNA sequences, saving time and costs and improving pathogen coverage. This step can detect a wider range of pathogens, especially in cases of complex or mixed infections. Quality control assessments (such as the specificity of amplified products and amplification efficiency) ensure the reliability of PCR amplification. If problems occur in PCR (such as non-specific amplification or low amplification efficiency), they can be identified and adjusted in a timely manner, avoiding distortion of downstream sequencing data. PCR quality control result tables help comprehensively evaluate the success of each PCR reaction, providing necessary information for subsequent steps (such as sequencing library construction) and ensuring the accuracy of the next operation. Quality control result sheets guide researchers to eliminate failed PCR reactions, ensuring high-quality DNA libraries for sequencing. High-quality sequencing libraries ensure the accuracy and reproducibility of sequencing data. Detailed checks on the PCR amplification of each sample optimize the library construction process, minimizing bias and improving the depth and coverage of sequencing data. Quality control during library construction prevents low-quality samples and PCR products from entering the sequencing process, reducing erroneous results due to sample contamination or other adverse factors. The presence of the host genome can interfere with the accurate detection of pathogenic microorganisms. By aligning and removing the host genome, most non-pathogenic DNA can be effectively removed, leaving more pathogenic information. This step significantly improves the sensitivity and specificity of microbial detection. After host DNA removal, the remaining dataset for comparison is more representative of the true pathogenic information, contributing to the accuracy of subsequent analyses. Removing the host genome sequence further optimizes data storage and analysis efficiency, reducing the consumption of computational resources and time by useless data. By comparing with known pathogen information in the database, the types of pathogens in the sample can be identified, and their abundance can be determined. This species comparison analysis can provide highly accurate identification of the infectious pathogens. High-throughput sequencing technology allows for the simultaneous detection of multiple pathogens, eliminating the need for individual detection of each pathogen. This step significantly improves the efficiency of the testing protocol, especially important when dealing with complex or mixed infections. Attached Figure Description
[0007] Fig. 1This is a schematic diagram of the steps of a high-throughput sequencing detection method for pathogenic microorganisms based on multiplex PCR detection according to the present invention. Fig. 2 This is a detailed flowchart illustrating the implementation steps of step S1. Fig. 3 This is a flowchart illustrating the detailed implementation steps of step S2. Detailed Implementation
[0008] It should be understood that the specific embodiments described herein are for illustrative purposes only and are not intended to limit the scope of the invention.
[0009] This application provides a method and system for high-throughput sequencing detection of pathogenic microorganisms based on multiplex PCR detection. The execution entities of the method and system include, but are not limited to, mechanical equipment, data processing platforms, cloud server nodes, and network upload devices mounted on the system, which can be considered as general computing nodes in this application. The data processing platform includes, but is not limited to, at least one of an audio / image management system, an information management system, and a cloud data management system.
[0010] Please see Figs. 1 to 3 This invention provides a high-throughput sequencing detection method for pathogenic microorganisms based on multiplex PCR detection, comprising the following steps: Step S1: Quantitatively add internal reference nucleic acid markers to the first microbial sample to be tested to generate a standardized DNA template library; Step S2: Perform multiplex PCR amplification on the standardized DNA template library and conduct quality control assessment to generate a PCR quality control result table; Step S3: Construct sequencing libraries based on the PCR quality control result table to generate raw sequencing library data; Step S4: Align the raw sequencing library data with the host reference genome and remove any missing data, then extract the dataset to be compared. Step S5: Perform pathogen species comparison analysis based on the dataset to be compared, and output the final pathogen detection list.
[0011] In the embodiments of the present invention, see Fig. 1 This is a schematic flowchart of a high-throughput sequencing detection method for pathogenic microorganisms based on multiplex PCR detection according to the present invention. In this example, the steps of the high-throughput sequencing detection method for pathogenic microorganisms based on multiplex PCR detection include: Step S1: Quantitatively add internal reference nucleic acid markers to the first microbial sample to be tested to generate a standardized DNA template library; In this embodiment, the microbial samples to be tested need to be processed to ensure the consistency and standardization of nucleic acid levels in the samples, laying the foundation for subsequent PCR amplification and high-throughput sequencing. First, DNA is extracted from the microbial samples to be tested, and then pre-designed internal control nucleic acid markers are added. Internal control markers are generally standardized sequences unrelated to the microorganisms to be tested, used to quantify and calibrate the efficiency and results of PCR amplification. By quantitatively adding these internal control markers, it can be ensured that the amount of different target DNA sequences in the sample is controllable and consistent during the amplification process. The selection of internal control markers usually considers their non-cross-reactivity with the microorganisms to be tested and should have stable amplification efficiency. Then, the DNA in the sample is mixed with the internal control markers and standardized to obtain a standardized DNA template library. The characteristic of this library is that the DNA concentration and the amount of internal control marker added for each sample are consistent across multiple samples, thereby reducing deviations caused by differences in sample concentrations during the experiment and providing an accurate starting point for subsequent PCR amplification and high-throughput sequencing.
[0012] Step S2: Perform multiplex PCR amplification on the standardized DNA template library and conduct quality control assessment to generate a PCR quality control result table; In this embodiment, after preparing the standardized DNA template library, the library is then subjected to multiplex PCR amplification. Multiplex PCR refers to the simultaneous amplification of multiple target regions in a single reaction, which is highly effective in pathogen detection because it can identify multiple pathogen species simultaneously. The PCR reaction system includes multiple target primers that specifically amplify different fragments of pathogen DNA. The amplification process is typically performed in a thermal cycler, where reaction conditions, such as annealing temperature, number of amplification cycles, and enzyme selection, need to be optimized to ensure balanced amplification efficiency for different pathogens.
[0013] To ensure the quality of PCR reactions, quality control is essential. Quality control includes checking the specificity and efficiency of PCR amplification, as well as the presence of non-specific amplification products. Fluorescence signal detection is typically used to monitor the PCR amplification process in real time, analyzing the amplification curve for each reaction, and assessing the sensitivity and specificity of amplification using the Ct value (cycle threshold). The PCR quality control results table will record key parameters including Ct value, amplification efficiency, plateau intensity (intensity during the signal stabilization phase), background fluorescence level, and curve slope. This data helps determine the success of the PCR reaction and provides quality assurance for subsequent analysis. If the quality control parameters do not meet the standards, it may be necessary to adjust the PCR reaction system or primer design to improve the specificity and accuracy of amplification.
[0014] Step S3: Construct sequencing libraries based on the PCR quality control result table to generate raw sequencing library data; In this embodiment, after passing the PCR quality control assessment, the next step is the construction of the sequencing library. Based on the PCR amplification results, qualified amplification products are selected for library construction. The library construction process includes end repair of the PCR-amplified DNA fragments, adapter addition (ligation of sequencing aptamers), and then enrichment through PCR amplification. This process converts the PCR amplification products into a library format suitable for high-throughput sequencing. Special attention must be paid to adapter selection and ligation reactions to ensure the correct library structure and compatibility of the aptamers with the sequencing platform. The addition of adapters not only helps guide sequencing but also provides crucial identifying information for subsequent data analysis.
[0015] After the constructed sequencing library passes quality checks, its concentration and purity are calculated to ensure it meets sequencing standards. Concentration is typically measured using devices such as Qubit or NanoDrop, while fragment size distribution is detected using an Agilent Bioanalyzer or similar device. If the library's concentration and fragment distribution meet the requirements, it can proceed to the high-throughput sequencing stage, generating raw sequencing library data. This data will provide raw information for subsequent pathogen comparison and analysis.
[0016] Step S4: Align the raw sequencing library data with the host reference genome and remove any missing data, then extract the dataset to be compared. In this embodiment, after sequencing, the raw data needs to undergo a series of preprocessing steps to remove contamination from the host genome and extract the DNA information of the pathogenic microorganism. First, the raw sequencing data needs to be compared with the host genome. The host genome generally refers to the common host DNA sequence of the experimental subject, such as the genomes of humans, animals, or plants. The comparison process typically uses alignment tools such as BWA and Bowtie2 to remove the host genome sequence from the sequencing data based on sequence similarity.
[0017] After alignment with the host reference genome, all sequences aligned to the host genome are filtered out, retaining only those sequences that do not match the host genome. This yields a clean dataset for comparison, containing the DNA sequences of pathogenic microorganisms that may be present in the sample. This dataset undergoes a host removal step, eliminating interference from the host genome and ensuring the accuracy of subsequent pathogen identification and analysis.
[0018] Step S5: Perform pathogen species comparison analysis based on the dataset to be compared, and output the final pathogen detection list.
[0019] In this embodiment, after obtaining the dataset to be compared, the next step is to perform pathogen species comparison analysis. This process mainly involves comparing the dataset to be compared with a pathogen species-level reference database to identify the pathogens that may be present in the sample. Commonly used comparison databases include NCBI, PATRIC, and MycoBank, which contain the genome sequences of various pathogens and are classified according to species and genus. During the comparison process, each sequence in the sequencing data is compared with the pathogen genome in the reference database, and the matching degree and similarity of sequence fragments are calculated to determine which pathogens may be present in the sample.
[0020] The alignment results will display information such as sequence abundance and similarity for each pathogen, helping to further screen for pathogen sequences with high similarity. By setting appropriate alignment parameters (such as similarity and abundance thresholds), it can be ensured that the finally screened pathogens are clinically relevant and highly reliable. Finally, these results are compiled into a pathogen detection checklist, and a final detection report is generated. This checklist lists all pathogen species detected in the sample and provides detailed information for each pathogen through alignment and evaluation, including possible abundance, clinical manifestations, etc., supporting disease diagnosis.
[0021] In this embodiment, see Fig. 2 The diagram below illustrates the detailed implementation steps of step S1. In this embodiment, the detailed implementation steps of step S1 include: The first microbial sample to be tested was quantitatively added with an internal control nucleic acid marker to generate the second sample to be tested. Obtain the pre-defined positive control samples and blank control samples; Nucleic acid extraction was performed on the second sample to be tested, the preset positive control sample, and the blank control sample to generate the original DNA sample library; The original DNA sample library is standardized and adjusted to generate a standardized DNA template library.
[0022] In this embodiment, In this embodiment, The specific steps for standardizing and adjusting the original DNA sample library to generate a standardized DNA template library are as follows: Based on the original DNA sample library, the DNA concentration, purity, and integrity are calculated to generate a DNA quality assessment table; The recovery rate of internal reference nucleic acids for each sample was calculated using the DNA quality assessment form to obtain the extraction efficiency coefficient for multiple samples. The original DNA sample library was standardized in concentration and adjusted based on the extraction efficiency coefficient of multiple samples to generate a standardized DNA template library.
[0023] In this embodiment, an internal control nucleic acid marker is first added to the first microbial sample to be tested. The internal control marker is a standardized tool to ensure accurate detection and quantification of the target pathogenic microorganism in multiplex PCR reactions; it is typically a known synthetic DNA or RNA fragment. To ensure an appropriate amount of internal control marker is added, its concentration needs to be quantified (usually within the range of 10^5 to 10^8 copies / μL). The marker can be added precisely by mixing the solution, allowing it to act as a quantitative standard in the PCR reaction. The purpose of this step is to ensure a consistent internal control nucleic acid concentration in each sample to be tested, serving as a reference in subsequent data analysis, helping to correct differences between samples, and improving experimental accuracy. Pre-set positive control samples and blank control samples play a role in verification and control during the experiment. Positive control samples are standard samples known to contain the nucleic acid of the target pathogenic microorganism, used to ensure that the PCR reaction system can normally amplify the target sequence and provide the expected results. Blank control samples do not contain any DNA of the target pathogenic microorganism and are mainly used to check for contamination in the experiment, ensuring that the signal in the reaction comes only from the sample to be tested and the internal control. Positive control samples typically use known pathogen DNA standards or synthetic target gene fragments, while blank control samples are pure extraction solutions or reaction systems, ensuring that all samples are subject to the same controlled conditions during extraction and reaction. These two types of control samples effectively eliminate systematic errors and external contamination during the experimental process.
[0024] Nucleic acid extraction is a crucial step in obtaining microbial DNA. During this process, the second sample to be tested, the positive control sample, and the blank control sample all need to undergo nucleic acid extraction to isolate pure DNA from the samples. Common extraction methods include the phenol / chloroform method, the silica gel membrane method, and the magnetic bead method. These methods can effectively isolate high-quality DNA from complex samples. During extraction, it is essential to ensure standardized procedures to avoid sample contamination or inconsistent extraction efficiency. After extraction, the DNA concentration and purity of the samples are measured using a UV spectrophotometer, typically requiring a 260 / 280 ratio between 1.8 and 2.0. The extracted DNA will form the basis for subsequent multiplex PCR detection, generating the original DNA sample library. The key to this step is ensuring the integrity and purity of the DNA to guarantee the reliability of subsequent experimental results. The DNA concentration of each sample in the original DNA sample library is measured using a spectrophotometer or quantitative real-time PCR device. Based on the measurement results, the samples are diluted to a uniform concentration (e.g., 10-50 ng / μL) to ensure that the DNA input volume for each sample is equal. Standardized DNA samples will be used for multiplex PCR reactions. This process effectively avoids variations in amplification efficiency caused by differences in DNA concentration, ensuring the stability and accuracy of the PCR reaction. After standardization, all DNA samples will be processed at a uniform concentration, laying the foundation for subsequent multiplex PCR analysis and high-throughput sequencing.
[0025] In this embodiment, see Fig. 3 The diagram below illustrates the detailed implementation steps of step S2. In this embodiment, the detailed implementation steps of step S2 include: Multiplex PCR amplification was performed on a standardized DNA template library, fluorescence signals were collected during the amplification process, and PCR amplification curves were generated. PCR amplification curves were processed to extract reaction parameters and generate a PCR reaction parameter table. The PCR reaction parameter table includes Ct value, amplification efficiency, plateau intensity, background fluorescence level, and curve slope; Based on the PCR reaction parameter table, a comparative analysis was conducted on three types of samples: the sample to be tested, the positive control, and the blank control, to obtain the difference parameters among the three types of samples. Quality control assessment was conducted based on three types of sample difference parameters, and positive, negative and suspicious samples were marked. Statistical analysis of positive, negative, and suspicious samples was performed to generate a PCR quality control result table.
[0026] In this embodiment, during multiplex PCR amplification, the standardized DNA template library is placed into the reaction system along with primers, fluorescently labeled probes, and PCR reagents for multiplex PCR amplification. Multiplex PCR amplification uses multiple specific primer combinations to simultaneously amplify the DNA sequences of various pathogenic microorganisms. Each target pathogenic microorganism is equipped with a different fluorescently labeled probe, and real-time data is obtained by monitoring changes in the fluorescence signal. During PCR amplification, the instrument records and plots amplification curves in real time by reading changes in the fluorescence signal after each cycle. The intensity of the fluorescence signal is directly proportional to the amount of amplified product; therefore, by monitoring changes in the fluorescence signal, the amplification progress of each target can be observed intuitively. Typically, the PCR instrument performs amplification according to a preset reaction program, collects fluorescence signals, and generates amplification curves. These curves reflect the relative amount of PCR amplification products at different time points. As the reaction proceeds, the fluorescence signal of the curve gradually increases, eventually reaching a plateau and saturation. The changes in the fluorescence signal of each target are crucial for analyzing the efficiency and quality of the PCR reaction.
[0027] After the PCR amplification curve is generated, the next step is to extract key information from the curve to generate a reaction parameter table. The main parameters include Ct value, amplification efficiency, plateau intensity, background fluorescence level, and curve slope. The Ct value (threshold cycle number) reflects the initial concentration of the target DNA; a lower Ct value indicates a higher target concentration. Amplification efficiency is calculated using the curve slope; ideally, it should be between 90% and 110%. Inefficient amplification may indicate primer problems or poor sample DNA quality. Plateau intensity refers to the intensity of the fluorescence signal when it reaches its maximum value and tends to stabilize during PCR amplification, representing the total amount of amplified products. Background fluorescence level represents stray signals in the reaction system; high background fluorescence may affect the accuracy of detection. The curve slope reflects the amplification speed; ideally, a steeper slope indicates high amplification efficiency. Using professional real-time PCR analysis software, these parameters are automatically extracted and presented in tabular form, providing a comprehensive assessment of the reaction process. By comparing and analyzing the PCR reaction parameters of the sample to be tested, positive control samples, and blank control samples, the reaction quality and results of each sample can be evaluated. First, the Ct value and amplification efficiency of the sample to be tested are compared with those of the positive control to determine whether the sample contains the DNA of the target pathogen. If the Ct value of the sample to be tested is close to that of the positive control and the amplification curve is clear, the sample is considered positive. Next, the sample to be tested is compared with the blank control sample to ensure there is no contamination or background interference. If the signal of the sample to be tested is significantly higher than that of the blank control and conforms to the expected amplification curve, the amplification signal can be considered to originate from the sample itself, rather than from reagent contamination. By comparing and analyzing these reaction parameters (such as Ct value, plateau intensity, etc.), differential parameters can be further obtained to help identify and exclude abnormal samples and ensure the accuracy of the experiment.
[0028] After obtaining the differential analysis results of PCR reaction parameters, quality control judgment is required based on these differences. Positive samples typically have low Ct values (fewer cycle numbers) and strong plateau intensity, indicating abundant amplification products and good amplification efficiency. Negative samples should show no amplification signal, or have very high Ct values and extremely low plateau intensity, indicating the absence of target DNA. Suspicious samples may show moderate amplification signals, or have low amplification efficiency but still have some background fluorescence signal; these samples require further confirmation. The purpose of quality control judgment is to identify positive, negative, and suspicious samples through comprehensive analysis of sample reaction parameters, providing a basis for subsequent further verification and data analysis. The quality control judgment results for positive, negative, and suspicious samples will be statistically integrated to generate a PCR quality control result table. This result table will include a unique identifier for each sample, key PCR reaction parameters (such as Ct value, amplification efficiency, etc.), and the final quality control result (positive, negative, suspicious). This information comprehensively reflects the quality of the experiment and helps researchers understand the detection status of each sample. The statistical results table allows for the evaluation of the overall experimental performance, including the positive detection rate, negative exclusion rate, and proportion of suspicious samples, thus providing reliable data support for subsequent high-throughput sequencing analysis. The quality control results table, as part of the experimental report, will be used for final analysis, evaluation, and result interpretation.
[0029] In this embodiment, step S3 includes the following steps: Positive samples were extracted based on the PCR quality control results table; The positive samples were subjected to electrophoretic analysis of the corresponding PCR products to obtain the electrophoretic patterns of the products; Based on the electrophoretic patterns of the products, the molecular weight, number of bands, and band clarity are calculated to generate the physical characteristics of the PCR products. The expected number and location of specific bands in the PCR products were verified, and products that passed the verification were extracted. Based on the validated products, sequencing libraries are constructed to generate raw sequencing library data.
[0030] In this embodiment, after generating the PCR quality control result table, positive samples need to be screened first. Based on the samples marked "positive" in the quality control result table (usually samples with low Ct values and high amplification efficiency), these samples are extracted from all samples to be tested. The criteria for a positive sample are typically that the sample produced a clear amplification signal during PCR amplification, its Ct value was below a set threshold, its plateau intensity was strong, and its amplification efficiency met expectations. Positive samples indicate that they contain the DNA of the target pathogen; therefore, in subsequent experiments, positive samples will become the focus of further analysis and verification. The purpose of this step is to ensure that only samples that clearly show the target amplification signal are subjected to subsequent electrophoresis analysis, physical characteristic verification, and sequencing, avoiding erroneous negative results or contaminated samples interfering with the analysis. The PCR reaction products are mixed with DNA loading buffer and loaded into an agarose gel electrophoresis tank. By setting an appropriate voltage (usually 80-120V), the PCR products are separated according to their molecular size under the influence of an electric field. During electrophoresis, smaller DNA fragments will run faster than larger fragments. After electrophoresis, the electrophoretic pattern is observed under ultraviolet light. Typically, staining with dyes such as ethidium bromide or SYBR Green causes the DNA fragments to fluoresce under ultraviolet light. By comparing the bands of the PCR products with known molecular weight standard control samples, the electrophoretic pattern of each positive sample can be clearly obtained. This pattern visually displays the molecular weight, number of bands, and band distribution of the amplified products in different samples, serving as an important basis for subsequent analysis.
[0031] After obtaining the electrophoresis pattern, the next step is to perform quantitative and qualitative analysis to extract the physical characteristics of the PCR products. First, the molecular weight of the PCR amplification products is calculated by comparing the migration speed of each band in the electrophoresis pattern with known molecular weight standards. Different target DNA fragments migrate at different rates in the gel, allowing for estimation of their size. Next, the number of clearly visible bands in the pattern is counted. Ideally, the amplification product of each target pathogen should form a specific band. The presence of multiple bands may indicate non-specific amplification or primer dimers. Band clarity is a crucial indicator of product purity; ideally, bands should be clear and free from tailing, bending, or other defects. All these physical characteristics provide necessary quality information for subsequent validation and analysis. Describing the physical characteristics of the products ensures the specificity and reliability of the PCR reaction, allowing for the selection of PCR products that meet expectations for further validation. After obtaining the physical characteristics of the PCR products, they need to be compared with the expected number and location of specific bands to ensure the specificity of the PCR reaction. Based on the primer information in the experimental design, each target pathogen should produce a specific band. If multiple bands or misaligned bands appear in the amplification product of a certain target, it may indicate non-specific amplification in the PCR reaction or a primer design problem. By comparing with the preset band positions, it is determined whether each amplification product meets the expected molecular weight range and number of bands. If the bands are clear and the band positions are within the expected size, the PCR reaction is successful and has high specificity, and these products will be extracted for subsequent sequencing analysis. If the products do not meet expectations, these unqualified samples need to be excluded, and PCR amplification should be repeated or experimental conditions optimized. After confirming that the PCR products meet the expected specificity requirements, the next step is to construct a sequencing library based on the validated PCR products. This step usually involves purifying the amplification products to remove impurities and primers from the reaction system. Common purification methods include using DNA purification kits (such as column purification or magnetic bead purification). The purified PCR products will be used to construct a high-throughput sequencing library. The library construction process includes adapter ligation, PCR amplification, and selection of DNA fragments of appropriate size for library preparation. Adapters are pre-designed short DNA sequences that provide recognition sites for sequencing, ensuring that the sequencing platform can correctly identify and read the target region in the sample. After library construction, the generated library undergoes quality control to ensure that its complexity and uniformity meet sequencing requirements. Finally, the quality-controlled library is used as input for high-throughput sequencing to generate raw data, providing fundamental data for subsequent data analysis and pathogen identification.
[0032] In this embodiment, the specific steps for constructing sequencing libraries based on the validation products and generating raw sequencing library data are as follows: The validated products were purified to generate a purified PCR product sample set. The concentration and purity of purified PCR product sample sets are calculated to generate product quality assessment data; Downstream sequencing compatibility screening was performed based on product quality assessment data, and a list of qualified samples was extracted. Sequencing libraries were constructed from the list of qualified samples to generate raw sequencing library data.
[0033] In this embodiment, after verification through electrophoresis and physical characteristics, PCR products meeting the expected criteria were selected. These products require further purification to remove impurities from the reaction system, such as primers that did not participate in amplification, dNTPs, enzymes, and other stray components. Common purification methods include column purification and magnetic bead purification. Column purification involves loading the PCR products into a DNA purification column and using the affinity of the silica membrane for DNA to separate DNA from impurities. After multiple washes and centrifugations, the purified PCR products are finally collected. Magnetic bead purification involves binding DNA to functionalized magnetic beads and using magnetic force to separate and purify the target DNA. The purified PCR products are collected under DNA-free conditions and their concentration and purity are further evaluated. This step ensures the purity and quality of the products, preparing a high-quality DNA sample set for subsequent sequencing and ensuring the accuracy and reliability of the sequencing data. The purified PCR product sample set will be subjected to concentration and purity determination, which is a crucial step in evaluating product quality. Concentration is typically determined using a UV-Vis spectrophotometer via absorbance (A260). DNA has a specific absorption peak at 260 nm, so the concentration of the sample can be calculated by measuring its absorbance and combining it with the dilution factor. Purity assessment is performed using the A260 / A280 ratio. An ideal ratio for DNA samples should be 1.8-2.0, indicating that the sample does not contain excessive protein or other contaminants. A low A260 / A280 ratio may indicate the presence of protein or other impurities in the sample, requiring further purification. In addition, purity can be further confirmed using the A260 / A230 ratio. This step, by determining the concentration and purity of each sample, generates product quality assessment data, providing a basis for the next step of suitability screening.
[0034] Based on the product quality assessment data generated in the previous step, qualified samples that meet the sequencing requirements are screened. Key criteria for sequencing suitability screening include sample concentration, purity, and quality. First, the sample concentration needs to meet a certain range; generally, it should be between 10 ng / µL and 100 ng / µL. Concentrations that are too high or too low may affect the efficiency and uniformity of downstream library construction. Second, high purity is required; the A260 / A280 ratio should be close to 2.0, and the A260 / A230 ratio should be greater than 1.8 to ensure that there are no significant proteins or other contaminants affecting the sequencing reaction. For unqualified samples (e.g., samples with too low concentration or poor purity), PCR amplification or purification should be considered again. Samples that meet the above criteria are listed as qualified samples, and a list of qualified samples is extracted. These qualified samples will enter the downstream sequencing library construction stage, ensuring that only quality-qualified samples proceed to subsequent processes to improve the quality and accuracy of high-throughput sequencing. After the screening of qualified samples is completed, the sequencing library construction stage begins. The purpose of library construction is to convert purified DNA products into a library format suitable for high-throughput sequencing platforms (such as the Illumina platform). Key steps in library construction include adapter ligation, PCR amplification, library cleaning, and fragment selection. First, specific adapter sequences are ligated to both ends of the DNA fragments. These adapters contain primer binding sites and barcode sequences required by the sequencing platform. Adapter ligation is typically performed using T4 DNA ligase. Next, appropriate primers are added for PCR amplification to increase the number of adapter-ligated DNA fragments and ensure the homogeneity of each library sample. After PCR amplification, the library is purified to remove primers, unligated adapters, and other impurities. Finally, DNA fragments meeting the required length (typically in the 300-600 bp range) are selected using methods such as gel electrophoresis or magnetic bead selection. After quality checks (e.g., using instruments like Qubit or Bioanalyzer to assess library concentration and fragment distribution), the generated sequencing library is used as input data for sequencing in a high-throughput sequencer. Through this series of steps, the raw data from the generated sequencing library provides the foundation for subsequent sequence analysis and identification of pathogenic microorganisms.
[0035] In this embodiment, step S4 includes the following steps: Next-generation sequencing (NGS) is performed on the raw sequencing library data to generate NGS data. The second-generation sequencing data was filtered to generate the first-layer quality control dataset; The filtration process includes quality control (QC) screening, connector contamination sequence removal, and low-complexity sequence filtration. Align the first-level quality control dataset with the host reference genome and remove the host genome sequence to obtain the non-host sequence dataset; A second-level quality control report is generated by performing in-depth analysis on non-host sequence datasets; the in-depth analysis includes sequence length distribution, GC content distribution analysis, and sequencing depth calculation. Quality assessment is performed based on the second-level quality control report, and the dataset to be compared is extracted.
[0036] In this embodiment, after library construction, the next step is to feed the sequencing library into a next-generation sequencing platform for high-throughput sequencing. Common next-generation sequencing platforms include Illumina and Ion Torrent, which can generate large amounts of sequencing data quickly and efficiently. The sequencing process uses short-read sequencing, amplifying and synthesizing DNA fragments based on adapter sequences in the library to generate short DNA sequences (typically 75-300 bp). On the Illumina platform, four fluorescently labeled dNTPs are used for the reaction, reading one base at a time, and monitoring the changes in fluorescence signals after the reaction. The sequencer detects these signals and converts them into digitized base sequences, thereby generating raw sequencing data. In this step, the generated data typically includes a series of sequence files (such as FASTQ format), each sequence recording the base sequence of the DNA fragment and its corresponding quality value (Phred value). This data constitutes the raw output of the sequencing data, serving as the basis for subsequent data analysis. Next-generation sequencing data typically contains a large number of raw sequences, some of which may be of poor quality, contain adapter sequence contamination, or have low complexity. Therefore, filtering is necessary to ensure data quality. The first step in filtering is quality control (QC) screening, which removes sequences with a QC value below a set threshold. The QC value is an indicator of sequencing data quality, with common thresholds set at Q20 or Q30, representing mismatch rates of 1% and 0.1%, respectively. Sequences below this threshold are usually filtered out due to their large errors. Secondly, adapter contamination sequences need to be removed. During sequencing, adapter sequences may not be completely removed, affecting subsequent analysis. Therefore, specialized algorithms are needed to identify and remove adapter sequences. Finally, low-complexity sequences (such as sequences with many repetitive sequences or high AT / GC content) are also filtered. These sequences generally contain limited information and are unlikely to provide effective data for subsequent pathogen analysis. Through these steps, high-quality and reliable sequences can be selected, generating the first-level quality control dataset, providing high-quality input data for subsequent alignment and analysis.
[0037] Before performing host reference genome alignment, it is necessary to determine the host species in the experiment, which typically refers to the human genome in the patient sample or other potentially existing host genomes. The purpose of alignment is to remove sequences from the host genome from the sequencing data. These sequences do not provide information about the pathogenic microorganisms and may even affect the accuracy of the data and the validity of the analysis results. Commonly used alignment tools include BWA (Burrows-Wheeler Aligner) and Bowtie2. These tools can identify and remove sequences that are highly similar to the host genome by aligning the reference host genome with the sequencing data. During the alignment process, a certain alignment threshold is usually set (e.g., sequences with an alignment rate of 95% or higher are considered host genome sequences). The alignment results will show which sequences belong to the host genome. After alignment, by removing these host genome sequences, the remaining non-host sequences become potential sequence data for the target pathogenic microorganisms. The resulting non-host sequence dataset is the foundational data for further pathogenic microorganism analysis and identification.
[0038] In-depth analysis typically includes the following aspects: Sequence length distribution: This involves statistically analyzing the distribution of the lengths of all valid sequences. Ideally, the vast majority of sequences should fall within a relatively consistent length range, typically 300-400 bp. Abnormal length distributions or the presence of a large number of extremely short or extremely long sequences may indicate PCR amplification bias or library construction problems.
[0039] GC content distribution analysis: GC content analysis helps identify potential biases in sequencing data. The GC content distribution should roughly correspond to the GC ratio of the target genome. An abnormal GC content distribution may indicate contamination in the sample or problems with data quality. By statistically analyzing the GC content distribution, the reliability of the data can be further assessed.
[0040] Sequencing depth calculation: Sequencing depth refers to the number of times each site is sequenced. Higher sequencing depth provides better data coverage and more accurate variant detection. Typically, sequencing depth should reach a certain standard (e.g., 30x or higher) to ensure sufficient coverage of each target region, especially in multiplex PCR assays, where some low-abundance pathogens may require deeper sequencing depths to improve detection sensitivity.
[0041] After generating the second-level quality control report, a comprehensive quality assessment of the data in the report is required. Based on indicators such as sequence length distribution, GC content distribution, and sequencing depth, it is determined whether the data meets the analytical requirements. For example, if the sequencing depth is insufficient, it may be necessary to repeat sequencing or increase the sequencing depth. If the GC content deviation is too large, it may be necessary to re-evaluate the experimental procedures to ensure that the sample quality is sound. Through these quality assessments, sequence data that meet the criteria are selected to form a comparison dataset. This comparison dataset contains high-quality, host-interference-free pathogenic microorganism sequence data that meets quality standards, serving as the foundation for subsequent pathogen identification and further analysis. This process ensures high accuracy and reliability for downstream high-throughput sequencing data analysis.
[0042] In this embodiment, the specific steps of step S5 are as follows: Obtain pathogen species-level reference database; The dataset to be compared is processed based on a pathogen species-level reference database to generate preliminary comparison results. Based on the primary alignment results, the number of sequence fragments, alignment coverage, sequence similarity, and abundance percentage are calculated to generate a species abundance statistical matrix. Set thresholds for pathogen sequence fragments; A detailed comparison of the species abundance statistical matrix was performed based on pathogen sequence fragment thresholds, and pathogen sequence fragments exceeding the thresholds were marked. Pathogen sequence fragments exceeding the threshold are screened based on their confidence threshold, and the final list of pathogenic microorganisms is output.
[0043] In this embodiment, before performing high-throughput sequencing analysis of pathogenic microorganisms, it is necessary to prepare pathogenic species-level reference databases. These databases contain the genome sequences of known pathogenic microorganisms, classified according to different species and genera, and provide standard reference data for alignment analysis. Commonly used pathogen databases include public databases such as NCBI, GenBank, and RefSeq, as well as databases specifically for pathogenic microorganisms, such as PATRIC and MycoBank. When acquiring these databases, it is necessary to ensure that they cover a variety of pathogenic microorganisms, including bacteria, fungi, viruses, and other pathogenic species. In this step, the reference genome data of pathogenic microorganisms is first downloaded from these databases, ensuring that their versions and contents are up-to-date. In addition, depending on the needs of the experiment, it may be necessary to preprocess the database, such as screening for specific species and removing redundant sequences, to improve the efficiency and accuracy of the alignment. After obtaining the reference database, preparations can be made for subsequent alignment analysis. The dataset to be compared is compared with these reference data. The purpose of the comparison is to identify sequences in the test sample that are similar to the genomes of known pathogenic microorganisms. Common alignment tools include BLAST (Basic Local Alignment Search Tool), Bowtie2, and BWA. During alignment, each sequence in the dataset to be compared is compared with all sequences in the reference database, and the degree of match is calculated. The alignment results will include multiple alignment parameters, such as the number of aligned sequence fragments, alignment position, and alignment similarity (usually expressed as an E-value or alignment score). By setting appropriate alignment thresholds (e.g., an E-value less than 1e-5, or an alignment similarity greater than 90%), it is possible to determine which sequences are similar to pathogen sequences in the reference database and generate preliminary alignment results. These preliminary alignment results will contain alignment information for each pathogen species and the sequence to be tested, including which sequence fragments match the pathogen genome and the quality of the match, providing basic data for subsequent analysis.
[0044] After obtaining the initial alignment results, it is necessary to further calculate the abundance of each pathogen and generate a species abundance statistical matrix. First, the number of sequence fragments refers to the number of sequence fragments that match the genome of a specific pathogen, reflecting the frequency of the pathogen's occurrence in the sample. Next, alignment coverage indicates the degree of coverage of the aligned sequence relative to the pathogen genome, usually expressed as a percentage. High alignment coverage indicates that the pathogen genome is more representative of the sample. Sequence similarity refers to the similarity between the tested sequence and the reference pathogen genome, usually calculated as a percentage; higher similarity means more reliable alignment results. Finally, abundance percentage indicates the proportion of each pathogen in all aligned sequences, reflecting the relative abundance of that pathogen. For example, if a pathogen accounts for 20% of the total aligned sequences, its abundance percentage in the sample can be considered to be 20%. These data are ultimately summarized into a species abundance statistical matrix. This matrix provides a clear understanding of the relative abundance of each pathogen in the sample, laying the foundation for further analysis and judgment. To screen for reliable sequences related to pathogens, a pathogen sequence fragment threshold needs to be set. Threshold settings are typically based on indicators such as alignment coverage, alignment similarity, and abundance percentage of pathogen sequence fragments. Generally, if the alignment coverage of a pathogen sequence reaches 50% or more, the similarity reaches 90% or more, and the abundance percentage is high (e.g., exceeding 1%), the pathogen sequence fragment can be considered highly reliable and worthy of further analysis. Threshold settings need to be optimized based on experimental design and sample background information. For example, in clinical samples, screening for certain pathogens may require a lower abundance percentage threshold, while in environmental samples, a higher abundance threshold may be needed to reduce interference from contaminants. The rationality of the threshold setting directly affects the accuracy and sensitivity of subsequent analytical results.
[0045] After setting the threshold for pathogen sequence fragments, the next step is to perform a fine-grained comparison of the species abundance statistical matrix. Based on the thresholds set in step four, the abundance proportion of each species is screened, and pathogen sequence fragments exceeding the threshold are marked. This process is essentially a quality control and filtering step in the comparison results, designed to exclude sequence fragments with low abundance or that do not meet the requirements. Specifically, pathogen sequences with an abundance proportion exceeding the set threshold should be considered reliable and worthy of further analysis. These marked pathogen sequence fragments represent the pathogenic microorganisms that may be present in the sample and have high detection reliability. Sequence fragments that do not reach the threshold can be considered noise or background data and are generally not considered. This step ensures the rigor and accuracy of data screening, providing a reliable basis for subsequent pathogen identification. After marking the pathogen sequence fragments exceeding the threshold, a reliability threshold screening is finally required for these sequence fragments to further ensure the reliability of the detection results. The criteria for reliability screening typically include alignment similarity, sequence coverage, and the detection background in the experiment. Common screening methods involve setting stricter comparison quality standards (e.g., similarity of 95% or higher) and coverage thresholds (e.g., coverage exceeding 70%). Furthermore, priority screening criteria can be set for specific pathogens based on clinical experience or experimental needs. After this screening, the resulting pathogen sequence fragments will form the final pathogen detection list. This list includes all reliable and significant pathogens, and their abundance and presence will provide strong support for subsequent clinical diagnosis, pathogen identification, and surveillance. The final output detection list will be a comprehensive and highly reliable result, providing high-quality data support for disease diagnosis and further pathogen analysis.
[0046] In this embodiment, the specific steps for performing confidence threshold screening on pathogen sequence fragments exceeding the threshold and outputting the final pathogen detection list are as follows: For pathogen sequence fragments exceeding the threshold, a clinical reference database association query is performed to generate clinically relevant annotation data; Quantitative calibration calculations were performed based on clinically relevant annotation data, and a reliability score was generated to produce a pathogen reliability score table. The reliability threshold is screened based on the pathogen reliability scoring table, and the final list of pathogenic microorganisms to be detected is output.
[0047] In this embodiment, after filtering out pathogen sequence fragments exceeding a threshold, the next step is to correlate these fragments with clinical reference databases to generate clinically relevant annotation data. Clinical reference databases typically contain information on disease-related pathogens, such as the pathogenicity, prevalence, and clinical symptoms of different pathogens. Common clinical databases include CLINVAR, Disease Mutation, and PubMed. The purpose of using these databases is to verify the association between pathogen sequence fragments and clinically known diseases, thereby providing diagnostic information. Pathogen sequence fragments exceeding the threshold are compared with pathogen genomes in the clinical reference database using alignment tools (such as BLAST) to obtain the matching status of each sequence fragment with known pathogens. The alignment results not only provide similarity to sequences in the reference database but also incorporate clinical information, such as whether the pathogen is clearly associated with certain diseases and the clinical manifestations of the pathogen. This information is of significant value for further pathogen identification. Finally, the generated clinically relevant annotation data includes the correlation between each pathogen sequence fragment and known diseases, helping to confirm whether these pathogens may cause specific clinical symptoms or diseases, thus supporting the final diagnosis.
[0048] After generating clinically relevant annotation data, the next step is to perform quantitative calibration calculations based on this data and assign a confidence score to each pathogen based on the calculation results. Quantitative calibration calculations aim to accurately assess the confidence of each pathogen sequence fragment based on factors such as clinical data, alignment similarity, sample abundance, and the known pathogenicity of the pathogen. The core of this step is to calculate a confidence score for each pathogen based on parameters such as the similarity of each sequence fragment, alignment coverage, sequencing depth, and relevant information from clinical reference databases. The confidence score is typically calculated based on the following factors: Similarity comparison: The higher the similarity comparison (e.g., greater than 95%), the stronger the similarity between the sequence fragment and the known pathogen genome, and the higher the confidence level.
[0049] Abundance percentage: Pathogen sequence fragments with a higher abundance percentage are more likely to be the main pathogens in the sample and have higher credibility.
[0050] Clinical relevance: If a pathogen sequence is strongly associated with known clinical symptoms or diseases, then the reliability of the pathogen should also be high.
[0051] Sequence coverage: Pathogen sequences with high coverage are more likely to represent the actual pathogens in the sample, and therefore their reliability score is correspondingly higher.
[0052] A comprehensive evaluation of the above factors generates a reliability score sheet for each pathogenic microorganism. The score sheet not only lists the reliability score for each pathogen but also categorizes them according to established criteria, such as high reliability, moderate reliability, and low reliability. The key to this step is combining clinical information and experimental data to ensure that the reliability score fully reflects the authenticity and clinical relevance of each pathogen sequence fragment.
[0053] After obtaining the pathogen reliability score sheet, the final step is to perform reliability threshold screening to output the final list of pathogen detection microorganisms. The reliability threshold is set based on the experimental requirements, the sensitivity of pathogen detection, and the actual application scenario. A common approach is to consider a pathogen reliable and worthy of inclusion in the final detection list if its reliability score exceeds a preset threshold (e.g., a reliability score greater than 0.8 or 0.9). If the reliability score is low (e.g., below 0.5), the pathogen can be excluded from the final detection list, as it may be a contaminant or a false positive.
[0054] In practice, the credibility threshold can be adjusted based on the following factors: Experimental sensitivity requirements: Some high-risk pathogens or emerging pathogens may require a lower confidence threshold to improve sensitivity and avoid missed detection.
[0055] Pathogenicity of the disease: For pathogens known to be highly pathogenic, a higher confidence threshold may be required to ensure the accuracy and specificity of the test results.
[0056] Background noise control: For environmental or low-abundance samples, the confidence threshold may be increased to reduce the interference of background noise and thus reduce false positive results.
[0057] After screening, the final pathogen detection list will list the pathogens that have undergone rigorous reliability screening, typically including the name of each pathogen, abundance percentage, reliability score, and possible clinical associations. This list will serve as the final experimental results and will be provided to clinicians or researchers to assist them in disease diagnosis, pathogen detection, or disease surveillance.
[0058] In this embodiment, a high-throughput sequencing detection system for pathogenic microorganisms based on multiplex PCR detection is provided, for performing the high-throughput sequencing detection method for pathogenic microorganisms based on multiplex PCR detection as described above, including: The internal control addition unit is used to quantitatively add internal control nucleic acid markers to the first microbial sample to be tested, and generate a standardized DNA template library. The PCR amplification unit is used to perform multiplex PCR amplification on a standardized DNA template library and to evaluate and determine the quality control results, generating a PCR quality control result table. The sequencing library unit is used to construct sequencing libraries based on PCR quality control results and generate raw sequencing library data. The host genome comparison unit is used to align the raw sequencing library data with the host reference genome and remove it, and extract the dataset to be compared. The pathogen species comparison unit is used to perform pathogen species comparison analysis based on the dataset to be compared, and output the final list of pathogenic microorganisms to be detected.
[0059] Therefore, the embodiments should be considered as exemplary and non-limiting in all respects, and the scope of the invention is defined by the appended claims rather than the foregoing description. Thus, all variations falling within the meaning and scope of the equivalents of the application are intended to be included within the invention.
[0060] The above description is merely a specific embodiment of the present invention, enabling those skilled in the art to understand or implement it. Various modifications to these embodiments will be readily apparent to those skilled in the art, and the general principles defined herein are implemented in other embodiments without departing from the spirit or scope of the invention. Therefore, the present invention is not to be limited to the embodiments shown herein, but is to be accorded the widest scope consistent with the principles and novel features of the invention herein.
Claims
1. A high-throughput sequencing detection method for pathogenic microorganisms based on multiplex PCR detection, characterized in that, Includes the following steps: Step S1: Quantitatively add internal reference nucleic acid markers to the first microbial sample to be tested to generate a standardized DNA template library; Step S2: Perform multiplex PCR amplification on the standardized DNA template library and conduct quality control assessment to generate a PCR quality control result table; Step S3: Construct sequencing libraries based on the PCR quality control result table to generate raw sequencing library data; Step S4: Align the raw sequencing library data with the host reference genome and remove any missing data, then extract the dataset to be compared. Step S5: Perform pathogen species comparison analysis based on the dataset to be compared, and output the final pathogen detection list.
2. The method for high-throughput sequencing detection of pathogenic microorganisms based on multiplex PCR detection according to claim 1, characterized in that, The specific steps of step S1 are as follows: The first microbial sample to be tested was quantitatively added with an internal control nucleic acid marker to generate the second sample to be tested. Obtain the pre-defined positive control samples and blank control samples; Nucleic acid extraction was performed on the second sample to be tested, the preset positive control sample, and the blank control sample to generate the original DNA sample library; The original DNA sample library is standardized and adjusted to generate a standardized DNA template library.
3. The method for high-throughput sequencing detection of pathogenic microorganisms based on multiplex PCR detection according to claim 2, characterized in that, The specific steps for standardizing and adjusting the original DNA sample library to generate a standardized DNA template library are as follows: Based on the original DNA sample library, the DNA concentration, purity, and integrity are calculated to generate a DNA quality assessment table; The recovery rate of internal reference nucleic acids for each sample was calculated using the DNA quality assessment form to obtain the extraction efficiency coefficient for multiple samples. The original DNA sample library was standardized in concentration and adjusted based on the extraction efficiency coefficient of multiple samples to generate a standardized DNA template library.
4. The high-throughput sequencing detection method for pathogenic microorganisms based on multiplex PCR detection according to claim 1, characterized in that, The specific steps of step S2 are as follows: Multiplex PCR amplification was performed on a standardized DNA template library, fluorescence signals were collected during the amplification process, and PCR amplification curves were generated. PCR amplification curves were processed to extract reaction parameters and generate a PCR reaction parameter table. The PCR reaction parameter table includes Ct value, amplification efficiency, plateau intensity, background fluorescence level, and curve slope; Based on the PCR reaction parameter table, a comparative analysis was conducted on three types of samples: the sample to be tested, the positive control, and the blank control, to obtain the difference parameters among the three types of samples. Quality control assessment was conducted based on three types of sample difference parameters, and positive, negative and suspicious samples were marked. Statistical analysis of positive, negative, and suspicious samples was performed to generate a PCR quality control result table.
5. The high-throughput sequencing detection method for pathogenic microorganisms based on multiplex PCR detection according to claim 1, characterized in that, Step S3 is as follows: Positive samples were extracted based on the PCR quality control results table; The positive samples were subjected to electrophoretic analysis of the corresponding PCR products to obtain the electrophoretic patterns of the products; Based on the electrophoretic patterns of the products, the molecular weight, number of bands, and band clarity are calculated to generate the physical characteristics of the PCR products. The expected number and location of specific bands in the PCR products were verified, and products that passed the verification were extracted. Based on the validated products, sequencing libraries are constructed to generate raw sequencing library data.
6. The high-throughput sequencing detection method for pathogenic microorganisms based on multiplex PCR detection according to claim 5, characterized in that, The specific steps for constructing sequencing libraries based on the validation products and generating raw sequencing library data are as follows: The validated products were purified to generate a purified PCR product sample set. The concentration and purity of the purified PCR product sample set are calculated to generate product quality assessment data; Downstream sequencing compatibility screening was performed based on product quality assessment data, and a list of qualified samples was extracted. Sequencing libraries were constructed from the list of qualified samples to generate raw sequencing library data.
7. The high-throughput sequencing detection method for pathogenic microorganisms based on multiplex PCR detection according to claim 1, characterized in that, The specific steps of step S4 are as follows: Next-generation sequencing (NGS) is performed on the raw sequencing library data to generate NGS data. The second-generation sequencing data was filtered to generate the first-layer quality control dataset; The filtration process includes quality control (QC) screening, connector contamination sequence removal, and low-complexity sequence filtration. Align the first-level quality control dataset with the host reference genome and remove the host genome sequence to obtain the non-host sequence dataset; A second-level quality control report is generated by performing in-depth analysis on non-host sequence datasets; the in-depth analysis includes sequence length distribution analysis, GC content distribution analysis, and sequencing depth calculation. Quality assessment is performed based on the second-level quality control report, and the dataset to be compared is extracted.
8. The high-throughput sequencing detection method for pathogenic microorganisms based on multiplex PCR detection according to claim 1, characterized in that, The specific steps of step S5 are as follows: Obtain pathogen species-level reference database; The dataset to be compared is processed based on a pathogen species-level reference database to generate preliminary comparison results. Based on the primary alignment results, the number of sequence fragments, alignment coverage, sequence similarity, and abundance percentage are calculated to generate a species abundance statistical matrix. Set thresholds for pathogen sequence fragments; A detailed comparison of the species abundance statistical matrix was performed based on pathogen sequence fragment thresholds, and pathogen sequence fragments exceeding the thresholds were marked. Pathogen sequence fragments exceeding the threshold are screened based on their confidence threshold, and the final list of pathogenic microorganisms is output.
9. The high-throughput sequencing detection method for pathogenic microorganisms based on multiplex PCR detection according to claim 1, characterized in that, The specific steps for performing confidence threshold screening on pathogen sequence fragments exceeding the threshold and outputting the final list of pathogenic microorganisms are as follows: For pathogen sequence fragments exceeding the threshold, a clinical reference database association query is performed to generate clinically relevant annotation data; Quantitative calibration calculations were performed based on clinically relevant annotation data, and a reliability score was generated to produce a pathogen reliability score table. The reliability threshold is screened based on the pathogen reliability scoring table, and the final list of pathogenic microorganisms to be detected is output.
10. A high-throughput sequencing detection system for pathogenic microorganisms based on multiplex PCR detection, characterized in that, A method for performing high-throughput sequencing detection of pathogenic microorganisms based on multiplex PCR detection as described in claim 1, comprising: The internal control addition unit is used to quantitatively add internal control nucleic acid markers to the first microbial sample to be tested, and generate a standardized DNA template library. The PCR amplification unit is used to perform multiplex PCR amplification on a standardized DNA template library and to evaluate and determine the quality control results, generating a PCR quality control result table. The sequencing library unit is used to construct sequencing libraries based on PCR quality control results and generate raw sequencing library data. The host genome comparison unit is used to align the raw sequencing library data with the host reference genome and remove it, and extract the dataset to be compared. The pathogen species comparison unit is used to perform pathogen species comparison analysis based on the dataset to be compared, and output the final list of pathogenic microorganisms to be detected.
Citation Information
Patent Citations
Blood metagenome sequencing data analysis method and device and application thereof
CN110349630A
High-throughput sequencing detection method for pathogenic microorganisms with full-process quality control
CN111187813A
Method for detecting and identifying pathogens of children infectious diseases based on metagenomic sequencing
CN111394486A
High-throughput detection method, system and kit for determining microorganisms based on internal reference
CN116179664A
Quantitative reference for rapid detection of pathogenic microorganisms and detection method
CN116445588A