Cerebrospinal fluid responsibility pathogen data analysis method based on nanopore sequencing
By combining nanopore sequencing with FPT and Normalized-FPT values and biosafety level (BSL), the accuracy and sensitivity issues of identifying responsible pathogens in central nervous system infections have been resolved, achieving pathogen identification with high specificity and high sensitivity.
Patent Information
- Application Number
- CN202511360758.0
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-09-23
- Publication Date
- 2026-01-09
AI Technical Summary
Existing methods for detecting central nervous system infections suffer from low sensitivity, long detection cycles, reliance on prediction of specific pathogens, and difficulty in accurately identifying the responsible pathogen, especially in nanopore sequencing data where background microorganisms and environmental contaminants cause significant interference.
We employed a nanopore sequencing-based cerebrospinal fluid (BAN) data analysis method to normalize microbial abundance by introducing the FPT index and Normalized-FPT value, combined with biosafety level (BSL), and screened out responsible pathogens.
It significantly improves the accuracy and sensitivity of identifying responsible pathogens, especially under high background noise conditions, and is suitable for the precise diagnosis of central nervous system infections.
Smart Images

Figure CN121306243A_ABST
Abstract
Description
Technical Field
[0001] This invention relates to the fields of high-throughput sequencing, bioinformatics analysis, and molecular diagnostics, specifically to a method for analyzing cerebrospinal fluid pathogen data based on nanopore sequencing. Background Technology
[0002] Central nervous system infections are characterized by rapid onset and high mortality. Early and accurate identification of the culprit is crucial for guiding anti-infective treatment. Currently, commonly used detection techniques include traditional culture methods, PCR detection, and serological testing. However, these methods suffer from low sensitivity, long detection cycles, and reliance on the prediction of specific pathogens. Metagenomic high-throughput sequencing (mNGS), as a novel detection technology that does not rely on traditional microbial culture, can directly perform high-throughput sequencing of nucleic acids in cerebrospinal fluid samples and can simultaneously detect almost all pathogens with known genomes. Therefore, mNGS has enormous application potential in the detection of pathogens in central nervous system infections.
[0003] Currently, mNGS typically employs second-generation sequencing (SGS) platforms, such as those from Illumina and BGI. While these platforms offer high sequencing throughput, they often require simultaneous sequencing of large batches of samples, resulting in inflexible sample testing. The time from sample processing to result output is typically 24–48 hours or more. Secondly, SGS reads are typically 100–150 bases long, meaning the alignment results contain hundreds to thousands of microorganisms, posing a significant challenge to identifying the responsible pathogen from the second-generation sequencing data. Although the number of unique reads, coverage, and biosafety classification in the sequencing results can identify the responsible pathogen to some extent, its accuracy is only around 85% and needs improvement. Existing methods fail to adequately consider background microorganisms in cerebrospinal fluid sequencing and microorganisms from environmental contamination, leading to reduced efficiency in detecting responsible pathogens. Furthermore, the significant differences in microbial genome length and gene expression levels mean that relying solely on absolute sequence counts or sequence percentages for pathogen identification cannot effectively assess the microbial content in a sample, potentially overlooking the true responsible pathogen or misidentifying certain pathogens as responsible ones.
[0004] In recent years, nanopore sequencing has been widely used for pathogen detection due to its advantages of real-time single-molecule sequencing and long read lengths. Compared with SGS, nanopore sequencing can provide read lengths ranging from several kb to several Mb, which can improve the specificity of alignment and significantly reduce the number of microorganisms in the output results, theoretically further improving the accuracy of responsible pathogen analysis. However, even so, nanopore sequencing data of cerebrospinal fluid samples may still contain dozens of microorganisms. These microorganisms include responsible pathogens as well as contaminants from the environment and experimental background. How to accurately analyze the responsible pathogens from dozens of microorganisms remains a huge challenge in the analysis of cerebrospinal fluid nanopore sequencing data, which is related to whether nanopore sequencing technology can be successfully applied to the detection of clinical cerebrospinal fluid specimens. Based on the above problems, this invention proposes a method for the analysis of responsible pathogens in cerebrospinal fluid based on nanopore sequencing—BAN analysis method. Summary of the Invention
[0005] To address the shortcomings of existing technologies, this invention provides a method for analyzing cerebrospinal fluid pathogen data based on nanopore sequencing, thus solving the problems mentioned in the background section.
[0006] To achieve the above objectives, the present invention provides the following technical solution: a method for analyzing cerebrospinal fluid pathogen data based on nanopore sequencing, comprising the following specific steps:
[0007] S1. Sequencing data acquisition and data processing, as detailed below:
[0008] S101. Nucleic acid extraction and sequencing library construction (obtaining raw metagenomic / metagescriptomics data);
[0009] S102. Raw sequencing data processing: Separate sequencing data from different samples based on the barcode sequences carried in the primers or aptamers. After separation, use the NanoSeq analysis system to clean and compare the data of each sample to identify the microorganisms contained in the sequencing results and the corresponding number of reads.
[0010] The S2 and BAN analysis methods are implemented, comprehensively considering the differences in microbial genome length, amplification bias, gene expression levels, and potential pathogenicity risk. FPT (Fragments Per Thousand) is introduced as an abundance index to enable comparable analysis of microbial abundance within the same sample and between different samples. Furthermore, background microbial abundance is constructed using non-infected cerebrospinal fluid samples, and normalized FPT values are calculated to assess the bias of each microorganism relative to the background abundance, thereby improving the sensitivity and accuracy of pathogen identification. Finally, combined with Biosafety Level (BSL) classification, the most probable responsible pathogen is determined, achieving highly specific and sensitive pathogen identification.
[0011] The specific steps are as follows:
[0012] S201. Microbial Abundance Calculation (FPT): To ensure the comparability of microbial abundance within the same sample and across different samples, we construct a unified "metagenome" from all detected microbial reference genomes. Each microbial genome is considered a "gene" within this metagenome (similar to a single gene in transcriptome analysis). Within this framework, the relative abundance of microorganisms in a sample can be analogized to gene expression levels. Accordingly, referencing transcriptome standardization strategies, the number of fragments per thousand reads (FPT) is defined as the microbial abundance index, similar to TPM (Transcripts Per Million) in transcriptome sequencing analysis. FPT calculation simultaneously considers the number of aligned sequences, microbial genome length (or number of coding genes), and amplification bias, ensuring comparability of abundance across samples and species for different microorganisms.
[0013] S202. Introducing Biosafety Level (BSL) for pathogenic microorganism screening: In nanopore sequencing data of cerebrospinal fluid samples, multiple microorganisms can often be detected simultaneously. Besides potential responsible pathogens, this includes a large number of non-pathogenic microorganisms originating from experimental procedures or environmental background. Although these background microorganisms are not pathogenic, they may occupy a high abundance in the sequencing data. If not distinguished, they can easily interfere with the determination of the responsible pathogen. To improve the practicality and accuracy of sequencing results in clinical responsible pathogen determination, this invention further introduces Biosafety Level (BSL) as a pathogenicity assessment indicator into the BAN analysis framework. The BSL system classifies microorganisms according to their potential harm to human health and is an internationally recognized standard for stratifying microbial pathogenicity, divided into four levels (BSL-1 to BSL-4). The higher the level, the greater the pathogenicity and transmission risk. The specific implementation steps are as follows: The applicant obtains data from the National Microbiology Data Center... The Biosafety Level (BSL) information of known pathogenic microorganisms was extracted and organized from the National Microbiological Center (NMDC, https: / / nmdc.cn / ), resulting in a total of 1150 microorganisms and their corresponding biosafety levels. In subsequent analysis, all identified microorganisms were matched with the database for biosafety level (BSL). Microorganisms with a BSL of "non-pathogenic" or unknown were directly excluded. The BAN method incorporates pathogenicity risk considerations, effectively improving the specificity and reliability of pathogen screening in high-background samples, and is particularly suitable for the clinical interpretation of samples with multi-source complex infections or microecological disorders.
[0014] S203. Calculation of Normalized-FPT (FPT): After removing non-pathogenic microorganisms, several to dozens of pathogenic microorganisms are still detected in nanopore sequencing data of cerebrospinal fluid samples. These microorganisms may include potential responsible pathogens, as well as "background microorganisms" introduced from environmental pollution, experimental procedures, or resident pathogens in cerebrospinal fluid. These background microorganisms often have high reproducibility and appear stably with similar abundance in multiple samples, thus occupying a certain proportion in the overall microbial spectrum, which poses a great challenge to the identification of microorganisms with low abundance but which may be responsible pathogens.
[0015] Accurately distinguishing background microorganisms from infectious pathogens that actually originate from the sample has always been a challenge in pathogen metagenomic detection. To address this issue, this invention proposes a normalized index for microbial abundance, the Normalized-FPT value, based on nanopore sequencing data from non-infectious cerebrospinal fluid samples, in order to improve the accuracy of identifying responsible pathogens.
[0016] S204. Strategy for determining responsible pathogens, as detailed below:
[0017] (1) In 41 cerebrospinal fluid samples based on nanopore metagenomic sequencing, the Normalized-FPT value showed a clear distinguishing feature between the infected and non-infected groups. Among them, the maximum Normalized-FPT value of the responsible pathogen in the infected samples was significantly higher than that in the non-infected samples (see Figure 1 This result further validates the ability and applicability of Normalized-FPT to identify responsible pathogens.
[0018] (2) In order to improve the accuracy of identification of responsible pathogens in complex sample backgrounds, this invention introduces a multi-level judgment strategy based on the Normalized-FPT value, and achieves accurate judgment by combining microbial abundance and biosafety classification.
[0019] Optional, for metagenomic sequencing, the details are as follows:
[0020] (1) The cerebrospinal fluid sample was separated into two parts, precipitate and supernatant, after centrifugation, and pretreated separately;
[0021] (2) The precipitate was treated with DNA enzyme and then subjected to cell lysis. The supernatant was concentrated by ultrafiltration, and the two parts were then combined to extract total DNA.
[0022] (3) After the obtained DNA is purified by magnetic beads, it is amplified and library constructed using a kit, and finally the sequencing operation is completed.
[0023] Optional, for metagenomic sequencing, the details are as follows:
[0024] (1) Cerebrospinal fluid samples were extracted with whole RNA using a micro RNA extraction kit. The obtained RNA was reverse transcribed using random primers, and the cDNA obtained was amplified and purified using a single-strand library preparation kit.
[0025] (2) The purified nucleic acid was used to construct a library and perform sequencing by using a kit in the ligation sequencing method.
[0026] Optionally, the system described in step S102 provides two operating modes: metagenomic analysis and metagenomic analysis; each analysis mode includes the following steps:
[0027] (1) Data preprocessing: quality screening of raw sequencing reads to filter out short sequences with a length less than a set threshold; for metagenomic data, the filtering threshold is set to 500 bp; for metagenomic transcriptome data, it is set to 300 bp; this step is achieved by calling seqkit, which effectively removes low-quality sequences and improves the reliability of subsequent alignments.
[0028] (2) Remove human-derived sequences. The preprocessed sequences were compared with the human reference genome (hg38) using the software minimap2. Sequences derived from the host were removed, and only non-human-derived sequences were retained for subsequent pathogen identification.
[0029] (3) Pathogen alignment and identification: The remaining non-human sequences will be aligned with the constructed high-coverage microbial reference database. Our existing database integrates publicly released microbial genome information from NCBI, covering 7,609 viral genomes, 10,891 bacterial genomes, 216 fungal and protozoan genomes, and 100 worm genomes. The software minimap2 is used for species alignment, and the number of matching sequences and sequencing depth for each species are counted to achieve microbial classification and identification. To improve the accuracy of the analysis results, the software SAMtools is further called to remove non-optimal alignments marked as supplementary or secondary alignments in the BAM file, and only the main alignment results are retained for subsequent species annotation and abundance calculation. The calculated coverage and depth information are used as the core input variables for species abundance FPT and screening ranking in the BAN analysis method, providing a basis for subsequent determination of responsible pathogens.
[0030] Optionally, the FPT calculation in step S201 specifically includes:
[0031] (1) PCR amplification bias correction (eReads calculation): Considering that the PCR amplification stage is easily affected by factors such as template amount, GC content, and primer binding efficiency, resulting in uneven amplification of some microbial fragments; in order to reduce such systematic bias, the raw sequence number (eReads) is first normalized using the vertical sequencing depth to obtain the corrected effective number of reads (eReads).
[0032] The calculation formula is as follows:
[0033]
[0034] (2) FPT value standardization calculation: The eReads of the species are further corrected according to the genome length of the species, and normalized proportionally according to the number of microorganisms detected in each specimen to obtain the final FPT value:
[0035] For metagenomic data (DNA sequencing):
[0036]
[0037] For metatranscriptome data (RNA sequencing):
[0038]
[0039] Where n represents the total number of microorganisms identified in the sample, and i is the i-th type of microorganism;
[0040] (3) In order to improve the stability of the analysis and eliminate weak interference signals, this method defines microorganisms with FPT values less than 2 as noise items or removes them; the remaining microorganisms enter the screening process of "candidate pathogens" and perform subsequent normalized-FPT calculation and sorting to assist in the determination of the responsible pathogen.
[0041] Optionally, the Normalized-FPT value mentioned in step S203 further incorporates the pathogen abundance characteristics of non-infected cerebrospinal fluid samples based on the original FPT index. This method establishes a "reference abundance distribution" of background pathogens by statistically calculating the average abundance and standard deviation of each pathogen in multiple non-infected samples. Based on this, for a specific pathogen appearing in the test sample, its Normalized-FPT value is obtained by normalizing its actual measured FPT value in the sample with the reference background value. The higher the value, the more likely the microorganism is to have abnormal enrichment characteristics in the sample, suggesting that it may be the responsible pathogen. Conversely, if the Normalized-FPT value is within the background fluctuation range, it suggests that it may be an environmental or reagent background microorganism.
[0042] The formula for calculating the Normalized-FPT value of pathogens in the target sample is as follows:
[0043]
[0044] Among them, FPT i The mean(FPT) value represents the FPT value of microorganism i in the target sample. i u ) represents the average FPT value of microorganism i in a non-infected cerebrospinal fluid sample, SDof FPT i u The standard deviation of FPT for microorganism i in non-infected samples, 5 × 10⁻⁶ -3 It is the smallest non-zero value among the standard deviations of FPT for all microorganisms in non-infected cerebrospinal fluid samples, used to prevent the numerator from becoming infinite after dividing by zero, so as to enhance the comparability of values.
[0045] After normalization calculation, the abundance of the responsible pathogen in the sample was significantly higher than its level in the non-infectious background, thus ranking higher in the Normalized-FPT ranking relative to other pathogens, which facilitates the accurate identification of the responsible pathogen from the complex microbial spectrum.
[0046] Optionally, the specific determination process for precise screening in step 204 is as follows:
[0047] Priority rule: Normalized-FPT≥10 and ranked first;
[0048] If a pathogen has a Normalized-FPT value greater than or equal to 10 in the target sample and ranks first among all candidate pathogens, it can be directly identified as the responsible pathogen of the sample. This threshold setting reflects its significant abundance enrichment advantage in the sample and is a quantitative representation of a strong infection signal.
[0049] Auxiliary rule: When Normalized-FPT values are similar, BSL (Balanced Score) is introduced for determination;
[0050] If multiple pathogens in a sample have similar Normalized-FPT values and no significant difference in their ranking, making it difficult to identify them definitively based solely on abundance, a Biosafety Level (BSL) should be introduced as an auxiliary criterion. In this case, pathogens with higher BSL levels should be given priority as the responsible pathogens, reflecting their high-risk characteristics in terms of clinical pathogenicity and transmission risk.
[0051] This invention provides a method for analyzing cerebrospinal fluid pathogen data based on nanopore sequencing, which has the following beneficial effects:
[0052] This invention proposes a computational analysis process that integrates microbial abundance, microbial safety classification, and background normalization, which has the following significant technical advantages and beneficial effects:
[0053] (1) The FPT index is proposed to achieve accurate assessment and normalization of microbial abundance. Traditional pathogen abundance assessment mainly relies on the number of sequencing sequences or the relative proportion of sequences. However, this method is prone to significant bias when dealing with microorganisms with large differences in genome size, expression level and PCR amplification efficiency. To overcome this problem, this invention introduces the FPT index into cerebrospinal fluid metagenomic and metatranscriptomic data for the first time to standardize the abundance of different microorganisms in the sample. This index effectively improves the accuracy of species abundance estimation and the comparability between samples by correcting for variables such as genome length, amplification bias and so on. In addition, the Normalized-FPT strategy is further introduced to integrate background microbial data in non-infected cerebrospinal fluid samples and establish a background abundance threshold, which significantly improves the ability to distinguish responsible pathogens.
[0054] (2) Integrating Biosafety Level (BSL) to optimize pathogen ranking: This invention innovatively incorporates the Biosafety Level (BSL) defined by the National Microbiology Science Data Center into the pathogen screening and ranking process as an auxiliary basis for judgment, thereby reducing the interference of a large number of non-pathogenic microorganisms in the test results; In practical applications, when faced with multiple pathogens with relatively similar abundance, microorganisms with higher biosafety levels are given priority as the responsible pathogens. This strategy is particularly suitable for mixed infection scenarios and overcomes the limitations of traditional algorithms that are dominated by abundance and ignore virulence risk in pathogen ranking.
[0055] (3) The introduction of Normalzid-FPT significantly improved the accuracy of causative pathogen detection. Analysis of cerebrospinal fluid samples from 41 metagenomically sequenced cases showed that, relying solely on FPT for causative pathogen identification, only 10 out of 19 confirmed infection samples had the causative pathogen ranked first, with the remaining 4 ranking second. The overall sensitivity was 52.6%. Figure 4 A and 4C); however, after introducing a background database constructed from non-infected samples, reanalysis using the Normalized-FPT index revealed that the responsible pathogen was accurately identified in 18 infected samples, ranking first in abundance, and the sensitivity was significantly improved to 94.7% ( Figure 4 B and 4C); meanwhile, no pathogen misidentification occurred in any of the 22 non-infected samples, achieving 100% specificity; the effectiveness of this method was further validated in cerebrospinal fluid samples from another 18 metagenomic sequencing data; under the FPT analysis method, the responsible pathogen ranked first in 8 out of 10 infected samples, i.e., the sensitivity was 80% (B and 4C); Figure 4 (D and 4F); After introducing the Normalized-FPT index, the responsible pathogens in 9 infected specimens were accurately identified, all ranking first, with only 1 case (sample I16) failing to meet the judgment criteria, and the overall sensitivity increased to 90%. Figure 4 E and 4F); Meanwhile, no false positive results were found in any of the 8 non-infected samples, and the specificity still reached 100%;
[0056] In summary, this method effectively solves the problem of identifying responsible pathogens in cerebrospinal fluid specimens based on nanopore sequencing for the first time. It exhibits extremely high specificity and sensitivity in different types of nanopore sequencing data, and is particularly suitable for accurate identification of responsible pathogens under high background noise conditions such as central nervous system infections. It has good clinical application prospects and promotion value. Attached Figure Description
[0057] Figure 1The bar chart shows the maximum Normalized-FPT distribution in infected and non-infected cerebrospinal fluid samples of this invention. Each row in the chart represents a cerebrospinal fluid sample, and each point represents the maximum Normalized-FPT value detected in that sample.
[0058] Figure 2 Here is a flowchart of the metagenomic BAN analysis process for this invention;
[0059] Figure 3 This is a flowchart of the metatranscriptome BAN analysis process for this invention;
[0060] Figure 4 This diagram illustrates how the BAN analysis method significantly improves the sensitivity and specificity of identifying responsible pathogens. Detailed Implementation
[0061] The technical solutions of the present invention will be clearly and completely described below with reference to the accompanying drawings in the embodiments of the present invention. Obviously, the described embodiments are only some embodiments of the present invention, and not all embodiments.
[0062] In the description of this invention, unless otherwise stated, "a plurality of" means two or more; the terms "upper," "lower," "left," "right," "inner," "outer," "front end," "rear end," "head," "tail," etc., indicate the orientation or positional relationship based on the orientation or positional relationship shown in the accompanying drawings, and are only for the convenience of describing the invention and simplifying the description, and do not indicate or imply that the device or element referred to must have a specific orientation, or be constructed and operated in a specific orientation, and therefore should not be construed as a limitation of the invention. Furthermore, the terms "first," "second," "third," etc., are used for descriptive purposes only and should not be construed as indicating or implying relative importance.
[0063] In the description of this invention, it should be noted that, unless otherwise explicitly specified and limited, the terms "connected" and "linked" should be interpreted broadly. For example, they can refer to a fixed connection, a detachable connection, or an integral connection; they can refer to a mechanical connection or an electrical connection; they can refer to a direct connection or an indirect connection through an intermediate medium. Those skilled in the art can understand the specific meaning of the above terms in this invention based on the specific circumstances.
[0064] Please see Figures 1 to 4 A method for analyzing cerebrospinal fluid pathogen data based on nanopore sequencing includes the following specific steps:
[0065] S1. Sequencing data acquisition and data processing, as detailed below:
[0066] S101. Nucleic acid extraction and sequencing library construction (obtaining raw metagenomic / metagescriptomics data);
[0067] For metagenomic sequencing, the details are as follows:
[0068] (1) The cerebrospinal fluid sample was separated into two parts, precipitate and supernatant, after centrifugation, and pretreated separately;
[0069] (2) The precipitate was treated with DNA enzyme and then subjected to cell lysis. The supernatant was concentrated by ultrafiltration, and the two parts were then combined to extract total DNA.
[0070] (3) After the obtained DNA was purified by magnetic beads, it was amplified and library constructed using a kit, and finally the sequencing operation was completed.
[0071] For metagenomic sequencing, the details are as follows:
[0072] (1) Cerebrospinal fluid samples were extracted with whole RNA using a micro RNA extraction kit. The obtained RNA was reverse transcribed using random primers, and the cDNA obtained was amplified and purified using a single-strand library preparation kit.
[0073] (2) The purified nucleic acids were used to construct a library and perform sequencing using a kit in a ligation sequencing manner;
[0074] S102. Raw sequencing data processing: Separate sequencing data between different samples based on the barcode sequence carried in the primers or aptamers; After separation, use the NanoSeq analysis system to clean and compare the data of each sample to identify the microorganisms contained in the sequencing results and the corresponding number of reads and sequencing depth.
[0075] The system offers two operating modes: metagenomic analysis and metagenomic analysis; each analysis mode includes the following steps:
[0076] (1) Data preprocessing: quality screening of raw sequencing reads to filter out short sequences with a length less than a set threshold; for metagenomic data, the filtering threshold is set to 500 bp; for metagenomic transcriptome data, it is set to 300 bp; this step is achieved by calling seqkit, which effectively removes low-quality sequences and improves the reliability of subsequent alignments.
[0077] (2) Remove human-derived sequences. The preprocessed sequences were compared with the human reference genome (hg38) using the software minimap2. Sequences derived from the host were removed, and only non-human-derived sequences were retained for subsequent pathogen identification.
[0078] (3) Pathogen alignment and identification: The remaining non-human sequences will be aligned with the constructed high-coverage microbial reference database. Our existing database integrates publicly released microbial genome information from NCBI, covering 7,609 viral genomes, 10,891 bacterial genomes, 216 fungal and protozoan genomes, and 100 worm genomes. The software minimap2 is used for species alignment, and the number of matching sequences and sequencing depth for each species are counted to achieve microbial classification and identification. To improve the accuracy of the analysis results, the software SAMtools is further called to remove non-optimal alignments marked as supplementary or secondary alignments in the BAM file, and only the main alignment results are retained for subsequent species annotation and abundance calculation. The calculated coverage and depth information are used as the core input variables for species abundance FPT and screening ranking in the BAN analysis method, providing a basis for subsequent determination of responsible pathogens.
[0079] The S2 and BAN analysis methods are implemented, comprehensively considering differences in microbial genome length, amplification bias, gene expression levels, and potential pathogenicity risk. FPT is introduced as an abundance index to enable comparable analysis of microbial abundance within the same sample and between different samples. Furthermore, background microbial abundance is constructed using non-infected cerebrospinal fluid samples, and normalized FPT values are calculated to assess the bias of each microorganism relative to the background abundance, thereby improving the sensitivity and accuracy of pathogen identification. Finally, combined with biosafety level (BSL) classification, the most probable responsible pathogen is determined, achieving highly specific and sensitive pathogen identification.
[0080] The specific steps are as follows:
[0081] S201. Microbial Abundance Calculation (FPT): To ensure the comparability of microbial abundance within the same sample and across different samples, we construct a unified "metagenome" from all detected microbial reference genomes. Each microbial genome is considered a "gene" within this metagenome (similar to a single gene in transcriptome analysis). Within this framework, the relative abundance of microorganisms in a sample can be analogized to gene expression levels. Accordingly, referencing transcriptome standardization strategies, the number of fragments per thousand reads (FPT) is defined as the microbial abundance index, similar to TPM (Transcripts Per Million) in transcriptome sequencing analysis. FPT calculation simultaneously considers the number of aligned sequences, microbial genome length (or number of coding genes), and amplification bias, ensuring comparability of abundance across samples and species for different microorganisms.
[0082] FPT calculation specifically includes:
[0083] (1) PCR amplification bias correction (eReads calculation): Considering that the PCR amplification stage is easily affected by factors such as template amount, GC content, and primer binding efficiency, resulting in uneven amplification of some microbial fragments; in order to reduce such systematic bias, the raw sequence number (eReads) is first normalized using the vertical sequencing depth to obtain the corrected effective number of reads (eReads).
[0084] The calculation formula is as follows:
[0085]
[0086] (2) FPT value standardization calculation: The eReads of the species are further corrected according to the genome length of the species, and normalized proportionally according to the number of microorganisms detected in each specimen to obtain the final FPT value:
[0087] For metagenomic data (DNA sequencing):
[0088]
[0089] For metatranscriptome data (RNA sequencing):
[0090]
[0091] Where n represents the total number of microorganisms identified in the sample, and i is the i-th type of microorganism;
[0092] (3) To improve the stability of the analysis and eliminate weak interference signals, this method defines microorganisms with FPT values less than 2 as noise items or removes them; the remaining microorganisms enter the screening process of "candidate pathogens" and perform subsequent normalized-FPT calculation and sorting to assist in the determination of the responsible pathogen.
[0093] S202. Introducing Biosafety Level (BSL) for pathogenic microorganism screening: In nanopore sequencing data of cerebrospinal fluid samples, multiple microorganisms can often be detected simultaneously. Besides potential responsible pathogens, this includes a large number of non-pathogenic microorganisms originating from experimental procedures or environmental background. Although these background microorganisms are not pathogenic, they may occupy a high abundance in the sequencing data. If not distinguished, they can easily interfere with the determination of the responsible pathogen. To improve the practicality and accuracy of sequencing results in clinical responsible pathogen determination, this invention further introduces Biosafety Level (BSL) as a pathogenicity assessment indicator into the BAN analysis framework. The BSL system classifies microorganisms according to their potential harm to human health and is an internationally recognized standard for stratifying microbial pathogenicity, divided into four levels (BSL-1 to BSL-4). The higher the level, the greater the pathogenicity and transmission risk. The specific implementation steps are as follows: The applicant obtains data from the National Microbiology Data Center... The Biosafety Level (BSL) information of known pathogenic microorganisms was extracted and organized from the National Microbiological Center (NMDC, https: / / nmdc.cn / ), resulting in a total of 1150 microorganisms and their corresponding biosafety levels. In subsequent analysis, all identified microorganisms were matched with the database for biosafety level (BSL). Microorganisms with a BSL of "non-pathogenic" or unknown were directly excluded. The BAN method incorporates pathogenicity risk considerations, effectively improving the specificity and reliability of pathogen screening in high-background samples, and is particularly suitable for the clinical interpretation of samples with multi-source complex infections or microecological disorders.
[0094] S203. Calculation of Normalized-FPT (FPT): After removing non-pathogenic microorganisms, several to dozens of pathogenic microorganisms are still detected in nanopore sequencing data of cerebrospinal fluid samples. These microorganisms may include potential responsible pathogens, as well as "background microorganisms" introduced from environmental pollution, experimental procedures, or resident pathogens in cerebrospinal fluid. These background microorganisms often have high reproducibility and appear stably with similar abundance in multiple samples, thus occupying a certain proportion in the overall microbial spectrum, which poses a great challenge to the identification of microorganisms with low abundance but which may be responsible pathogens.
[0095] Accurately distinguishing background microorganisms from infectious pathogens that actually originate from the sample has always been a challenge in pathogen metagenomic detection. To address this issue, this invention proposes a normalized index for microbial abundance, the Normalized-FPT value, based on nanopore sequencing data from non-infectious cerebrospinal fluid samples, in order to improve the accuracy of identifying responsible pathogens.
[0096] The Normalized-FPT value, based on the original FPT index, further incorporates the pathogen abundance characteristics of non-infected cerebrospinal fluid samples. This method establishes a "reference abundance distribution" of background pathogens by statistically calculating the average abundance and standard deviation of each pathogen in multiple non-infected samples. Based on this, for a specific pathogen appearing in the test sample, its Normalized-FPT value is obtained by normalizing its actual measured FPT value in the sample with the reference background value. A higher value indicates that the microorganism is more likely to have abnormal enrichment characteristics in the sample, suggesting it may be the responsible pathogen; conversely, if the Normalized-FPT value is within the background fluctuation range, it suggests it may be an environmental or reagent background microorganism.
[0097] The formula for calculating the Normalized-FPT value of pathogens in the target sample is as follows:
[0098]
[0099] Among them, FPT i Men (FPT) represents the FPT value of microorganism i in the target sample. i u ) represents the average FPT value of microorganism i in a non-infected cerebrospinal fluid sample, SDof FPT i u The standard deviation of FPT for microorganism i in non-infected samples, 5 × 10⁻⁶ -3 It is the smallest non-zero value among the standard deviations of FPT for all microorganisms in non-infected cerebrospinal fluid samples, used to prevent the numerator from becoming infinite after dividing by zero, so as to enhance the comparability of values.
[0100] After normalization calculation, the abundance of the responsible pathogen in the sample was significantly higher than its level in the non-infectious background, thus ranking higher in the Normalized-FPT ranking relative to other pathogens, which facilitates the accurate identification of the responsible pathogen from the complex microbial spectrum.
[0101] S204. Strategy for determining responsible pathogens, as detailed below:
[0102] (1) In 41 cerebrospinal fluid samples based on nanopore metagenomic sequencing, the Normalized-FPT value showed a clear distinguishing feature between the infected and non-infected groups. Among them, the maximum Normalized-FPT value of the responsible pathogen in the infected samples was significantly higher than that in the non-infected samples (see Figure 1 This result further validates the ability and applicability of Normalized-FPT to identify responsible pathogens.
[0103] (2) To improve the accuracy of identifying responsible pathogens in complex sample backgrounds, this invention introduces a multi-level judgment strategy based on the Normalized-FPT value, and achieves accurate judgment by combining microbial abundance and biosafety classification. The specific judgment process is as follows:
[0104] Priority rule: Normalized-FPT≥10 and ranked first;
[0105] If a pathogen has a Normalized-FPT value greater than or equal to 10 in the target sample and ranks first among all candidate pathogens, it can be directly identified as the responsible pathogen of the sample. This threshold setting reflects its significant abundance enrichment advantage in the sample and is a quantitative representation of a strong infection signal.
[0106] Auxiliary rule: When Normalized-FPT values are similar, BSL (Balanced Score) is introduced for determination;
[0107] If multiple pathogens in a sample have similar Normalized-FPT values and no significant difference in their ranking, making it difficult to identify them definitively based solely on abundance, a Biosafety Level (BSL) should be introduced as an auxiliary criterion. In this case, pathogens with higher BSL levels should be given priority as the responsible pathogens, reflecting their high-risk characteristics in terms of clinical pathogenicity and transmission risk.
[0108] Example 1 (Virus): Detection of the responsible pathogen in a cerebrospinal fluid specimen clinically diagnosed with varicella-zoster virus infection.
[0109] This embodiment selects a cerebrospinal fluid sample numbered I1, which comes from a patient clinically diagnosed with encephalitis caused by varicella-zoster virus (VZV) infection. First, total nucleic acid was extracted from the sample, a nanopore sequencing library was constructed, and metagenomic sequencing was performed using a nanopore platform. A total of 923,626 raw sequences were obtained. Preliminary quality control was performed, including the removal of low-quality and short-read sequences, and the sample was split according to the index sequence. Subsequently, the data was aligned to the human reference genome to remove human-derived sequences. The alignment results showed that 94.55% of the reads were from the host genome. After removing human-derived sequences, the number of effective reads remaining for microbial analysis was 50,370 (see Table 1). The BAN analysis method proposed in this invention was used for systematic analysis of the microbiome. Non-human reads were aligned to our constructed local microbial genome database, and the number of aligned sequences and sequencing depth for each species were counted, and the FPT value was calculated. Based on the BAN method... The procedure identified 75 microorganisms, including 10 potential pathogens with an FPT value ≥2. Notably, VZV had a low absolute sequence number and initial relative abundance in this sample. Traditional methods using sequence number or relative proportion as indicators might not have been able to identify it as a responsible pathogen. However, by employing a Normalized-FPT analysis strategy and introducing a background microbial abundance spectrum constructed from non-infected cerebrospinal fluid samples to normalize the FPT value, the results showed that VZV's normalized-FPT value was significantly higher than that of other microorganisms, ranking first among all species and exceeding the pathogenicity threshold (Normalized-FPT>10), thus meeting the screening criteria for responsible pathogens (Table 2). In summary, this embodiment demonstrates that even under low abundance signal conditions, the BAN analysis method proposed in this invention can still accurately identify the true pathogenic microorganism VZV, fully verifying the high sensitivity, high specificity, and good practicality of this method in clinical sample pathogen screening. (Note: Raw sequencing data for samples are detailed in Table 1, and pathogens with FPT values greater than or equal to 2 in the samples are detailed in Table 2.) Table 1: Sequencing Information for I1 Samples
[0110]
[0111] Table 2: BAN Analysis Results
[0112]
[0113]
[0114] Example 2 (Fungi): Detection of the responsible pathogen in a cerebrospinal fluid sample clinically diagnosed with Cryptococcus neoformans infection.
[0115] This embodiment selects a cerebrospinal fluid sample numbered I6, which comes from a patient with encephalitis clinically diagnosed with Cryptococcus neoformans infection. First, nucleic acids are extracted from this cerebrospinal fluid sample, and a nanopore sequencing library is constructed. Metagenomic sequencing is performed using a nanopore sequencing platform. The initial sequencing data volume is 437,651 sequences. During preliminary quality control, short-read and low-quality sequences are removed. Then, human sequences are removed using an alignment algorithm, resulting in 89.69% of reads aligned to the human genome. After removing human sequences, the remaining effective reads for microbial alignment are 45,133 (Table 3). Subsequently, the BAN analysis method described in this invention is used to analyze the microbial genome. Microbial composition was analyzed; the remaining non-human sequencing data were compared to a pre-defined microbial database, and the sequence number and sequencing depth of each species were counted accordingly; according to the standard procedure of the BAN method, a total of 69 microorganisms were identified in this sample, of which 5 were potential pathogens with an FPT value greater than or equal to 2; Cryptococcus neoformans had a high abundance in this sample, ranking first in terms of effective sequence number, relative proportion, and FPT value. After correction for non-infected samples, its Normalized-FPT value exceeded the judgment threshold, ranking first and far higher than other pathogens (Table 4); therefore, this embodiment successfully applied the BAN analysis method of the present invention to accurately identify the responsible pathogen Cryptococcus neoformans in the sample.
[0116] (Note: Raw sequencing data of the samples are detailed in Table 3, and pathogens with FPT values greater than or equal to 2 in the samples are detailed in Table 4)
[0117] Table 3: Sequencing Information for I6 Specimens
[0118]
[0119] Table 4: BAN Analysis Results
[0120]
[0121] Example 3 (Bacteria): Detection of the responsible pathogen in a cerebrospinal fluid specimen clinically diagnosed with Streptococcus pneumoniae infection.
[0122] This embodiment selects a cerebrospinal fluid sample numbered I7, which comes from a patient clinically diagnosed with encephalitis caused by Streptococcus pneumoniae infection. First, nucleic acids are extracted from this cerebrospinal fluid sample, and a nanopore sequencing library is constructed. Metagenomic sequencing is performed using a nanopore sequencing platform. The initial sequencing data volume is 153,474 sequences. During preliminary quality control, short-read and low-quality sequences are removed. Then, human sequences are removed using an alignment algorithm, resulting in 95.98% of reads aligned to human genomes. After removing human sequences, the remaining effective reads for microbial alignment are 6,173 (Table 5). Subsequently, the microbial composition is analyzed using the BAN analysis method described in this invention. The remaining non-human sequencing data is aligned to a pre-defined microbial database, and the sequence count and sequencing depth for each species are calculated accordingly. According to the standard BAN method procedure, 35 microorganisms were identified in this sample. Nine potential pathogens with FPT ≥ 2 were identified. Notably, after calculating the Normalized-FPT value using background microbial data from non-infected samples, two pathogens had values greater than the threshold (Normalized-FPT > 10): Streptococcus cristatus ranked first and the responsible pathogen Streptococcus pneumoniae ranked second. The difference in their Normalized-FPT values was not significant, suggesting the possibility of co-infection. However, Streptococcus pneumoniae has a biosafety level of BSL-2, which is higher than the BSL-1 of Streptococcus cristatus (Table 6). Therefore, based on the comprehensive assessment strategy, Streptococcus pneumoniae was accurately identified as the responsible pathogen.
[0123] Table 5: Sequencing Information for I7 Specimens
[0124]
[0125] Table 6: BAN Analysis Results
[0126]
[0127]
[0128] This invention aims to provide a method for analyzing responsible pathogens, which can accurately identify the responsible pathogens of central nervous system infections from nanopore sequencing results. Although nanopore sequencing read lengths can reach several kb, sequencing results in cerebrospinal fluid still contain dozens of pathogens. Accurately identifying the responsible pathogens of infection is extremely important for subsequent clinical diagnosis and treatment, but it faces great challenges. Existing cerebrospinal fluid pathogen analysis methods do not fully consider background microorganisms and microorganisms from environmental pollution in the sequencing results. They often rely on the absolute number and relative abundance of microbial sequences for pathogen identification, and lack an effective exclusion mechanism for non-pathogenic microorganisms, which can easily lead to the neglect of responsible pathogens or the misidentification of certain pathogens as responsible pathogens.
[0129] Therefore, there is currently no effective analytical method to identify responsible pathogens from cerebrospinal fluid nanopore sequencing data.
[0130] This invention proposes a BAN analysis method, which is based on the core idea of "abundance estimation + background normalization + biosafety level filtering". It comprehensively considers factors such as microbial genome length, PCR amplification bias, and the proportion of reads for each microorganism, and introduces FPT as an abundance calculation index to make the abundance of the same microorganism comparable across different samples. Secondly, it obtains the background microbial species and abundance from non-infected cerebrospinal fluid samples using nanopore sequencing, and calculates the normalized-FPT value of the microorganisms in the sample to be identified through normalization. Finally, it ranks the microorganisms based on this value and combines it with the microbial biosafety level to determine the responsible pathogen. This analysis method can accurately identify pathogenic pathogens in complex microbial signals, avoid interference from non-pathogenic background microorganisms, and significantly improve the sensitivity and specificity of detection.
[0131] The above description is only a preferred embodiment of the present invention, but the scope of protection of the present invention is not limited thereto. Any equivalent substitutions or modifications made by those skilled in the art within the scope of the technology disclosed in the present invention, based on the technical solution and inventive concept of the present invention, should be covered within the scope of protection of the present invention.
Claims
1. A method for analyzing cerebrospinal fluid pathogen data based on nanopore sequencing, characterized in that, Accurately identifying the responsible pathogens for central nervous system infections from nanopore metagenomic or metatranscriptomic sequencing data includes the following specific steps: S1. Sequencing data acquisition and data processing, as detailed below: S101. Nucleic acid extraction and sequencing library construction to obtain raw metagenomic / metagetranscriptomics data; S102. Raw sequencing data processing: Separate sequencing data from different samples based on the barcode sequences carried in the primers or aptamers. After separation, use the NanoSeq analysis system to clean and compare the data of each sample to identify the microorganisms contained in the sequencing results and the corresponding number of reads. The S2 and BAN analysis methods are implemented, comprehensively considering differences in microbial genome length, amplification bias, gene expression levels, and potential pathogenicity risk. FPT is introduced as an abundance index to enable comparable analysis of microbial abundance within the same sample and between different samples. Furthermore, background microbial abundance is constructed using non-infected cerebrospinal fluid samples, and normalized FPT values are calculated to assess the bias of each microorganism relative to the background abundance, thereby improving the sensitivity and accuracy of pathogen identification. Finally, combined with biosafety level classification, the most likely responsible pathogen is determined, achieving highly specific and sensitive pathogen identification. The specific steps are as follows: S201. Microbial Abundance Calculation: To ensure the comparability of microbial abundance within the same sample and between different samples, all detected microbial reference genomes are constructed into a single "metagenary system," where the genome of each microorganism is considered a "gene" within this system. Within this framework, the relative abundance of microorganisms in a sample can be analogized to gene expression levels. Accordingly, referencing transcriptome standardization strategies, the number of fragments of a particular species per thousand reads is defined as the microbial abundance index, a concept similar to TPM (Transcripts Per Million) in transcriptome sequencing analysis. FPT calculation simultaneously considers the number of aligned sequences, microbial genome length, and amplification bias, ensuring the comparability of abundance across samples and species for different microorganisms. S202. Biosafety levels are introduced for the screening of pathogenic microorganisms. In the nanopore sequencing data of cerebrospinal fluid samples, multiple microorganisms were detected simultaneously. In addition to potential responsible pathogens, a large number of non-pathogenic microorganisms from experimental operations or environmental background were also detected. In order to improve the practicality and accuracy of sequencing results in the clinical determination of responsible pathogens, biosafety levels are further introduced as pathogenicity determination indicators in the BAN analysis framework. The BSL system classifies microorganisms based on their potential harm to human health; S203. The calculation of normalized-FPT (Normalized-FPT) for pathogens: Based on nanopore sequencing data of non-infected cerebrospinal fluid samples, a normalized index for microbial abundance, the Normalized-FPT value, is proposed to improve the accuracy of identification of responsible pathogens. S204. Strategy for determining responsible pathogens, as detailed below: (1) In 41 cerebrospinal fluid samples based on nanopore metagenomic sequencing, the Normalized-FPT value showed a clear distinguishing feature between the infected and non-infected groups. The maximum Normalized-FPT value of the responsible pathogen in the infected samples was significantly higher than that in the non-infected samples. This result further validated the ability and applicability of Normalized-FPT to identify the responsible pathogen. (2) In order to improve the accuracy of identification of responsible pathogens in complex sample backgrounds, this invention introduces a multi-level judgment strategy based on the Normalized-FPT value, and achieves accurate judgment by combining microbial abundance and biosafety classification.
2. The method according to claim 1, characterized in that, For metagenomic sequencing, the details are as follows: (1) The cerebrospinal fluid sample was separated into two parts, precipitate and supernatant, after centrifugation, and pretreated separately; (2) The precipitate was treated with DNA enzyme and then subjected to cell lysis. The supernatant was concentrated by ultrafiltration, and the two parts were then combined to extract total DNA. (3) After the obtained DNA is purified by magnetic beads, it is amplified and library constructed using a kit, and finally the sequencing operation is completed.
3. The method according to claim 1, characterized in that, Specifically regarding metagenomic sequencing, the details are as follows: (1) Cerebrospinal fluid samples were extracted with whole RNA using a micro RNA extraction kit. The obtained RNA was reverse transcribed using random primers, and the cDNA obtained was amplified and purified using a single-strand library preparation kit. (2) The purified nucleic acid was used to construct a library and perform sequencing by using a kit in the ligation sequencing method.
4. The method for analyzing cerebrospinal fluid pathogen data based on nanopore sequencing according to claim 1, characterized in that, The system described in step S102 provides two operating modes: metagenomic analysis and metagenomic analysis; each analysis mode includes the following steps: (1) Data preprocessing: quality screening of raw sequencing reads to filter out short sequences with a length less than a set threshold; for metagenomic data, the filtering threshold is set to 500 bp; for metagenomic transcriptome data, it is set to 300 bp; this step is achieved by calling seqkit, which effectively removes low-quality sequences and improves the reliability of subsequent alignments. (2) Remove human-derived sequences. The preprocessed sequences were compared with the human reference genome (hg38) using the software minimap2. Sequences derived from the host were removed, and only non-human-derived sequences were retained for subsequent pathogen identification. (3) Pathogen alignment and identification: The remaining non-human sequences will be aligned with the constructed high-coverage microbial reference database. Our existing database integrates publicly released microbial genome information from NCBI, covering 7,609 viral genomes, 10,891 bacterial genomes, 216 fungal and protozoan genomes, and 100 worm genomes. The software minimap2 is used for species alignment, and the number of matching sequences and sequencing depth for each species are counted to achieve microbial classification and identification. To improve the accuracy of the analysis results, the software SAMtools is further called to remove non-optimal alignments marked as supplementary or secondary alignments in the BAM file, and only the main alignment results are retained for subsequent species annotation and abundance calculation. The calculated coverage and depth information are used as the core input variables for species abundance FPT and screening ranking in the BAN analysis method, providing a basis for subsequent determination of responsible pathogens.
5. The method for analyzing cerebrospinal fluid pathogen data based on nanopore sequencing according to claim 1, characterized in that, The FPT calculation described in step S201 specifically includes: (1) PCR amplification deviation correction: Considering that the PCR amplification stage is easily affected by factors such as template amount, GC content, and primer binding efficiency, some microbial fragments are amplified unevenly; in order to reduce such systematic deviations, the original sequence number is first normalized using vertical sequencing depth to obtain the corrected effective sequence number. The calculation formula is as follows: (2) FPT value standardization calculation: The eReads of the species are further corrected according to the genome length of the species, and normalized proportionally according to the number of microorganisms detected in each specimen to obtain the final FPT value: For metagenomic data: For metatranscriptome data: Where n represents the total number of microorganisms identified in the sample, and i is the i-th type of microorganism; (3) To improve the stability of the analysis and eliminate weak interference signals, this method defines microorganisms with an FPT value less than 2 as noise items and removes them; the remaining microorganisms enter the screening process of "candidate pathogens" and perform subsequent normalized-FPT calculation and sorting to assist in the determination of the responsible pathogen.
6. The method for analyzing cerebrospinal fluid pathogen data based on nanopore sequencing according to claim 1, characterized in that, The Normalized-FPT value mentioned in step S203, based on the original FPT index, further incorporates the pathogen abundance characteristics of non-infected cerebrospinal fluid samples. By statistically calculating the average abundance and standard deviation of each pathogen in multiple non-infected samples, a "reference abundance distribution" of background pathogens is established. On this basis, for a specific pathogen appearing in the test sample, its Normalized-FPT value is obtained by normalizing its actual measured FPT value in the sample with the reference background value. The higher the value, the more likely the microorganism is to have abnormal enrichment characteristics in the sample, suggesting that it may be the responsible pathogen. Conversely, if the Normalized-FPT value is within the background fluctuation range, it suggests that it may be an environmental or reagent background microorganism. The formula for calculating the Normalized-FPT value of pathogens in the target sample is as follows: Among them, FPT i The mean(FPT) value represents the FPT value of microorganism i in the target sample. i u ) represents the average FPT value of microorganism i in a non-infected cerebrospinal fluid sample, SDof FPT i u The standard deviation of FPT for microorganism i in non-infected samples, 5 × 10⁻⁶ -3 It is the smallest non-zero value among the standard deviations of FPT for all microorganisms in non-infected cerebrospinal fluid samples, used to prevent the numerator from becoming infinite after dividing by zero, so as to enhance the comparability of values. After normalization calculation, the abundance of the responsible pathogen in the sample was significantly higher than its level in the non-infectious background, thus ranking higher in the Normalized-FPT ranking relative to other pathogens, which facilitates the accurate identification of the responsible pathogen from the complex microbial spectrum.
7. The method for analyzing cerebrospinal fluid pathogen data based on nanopore sequencing according to claim 1, characterized in that, The specific determination process for precise screening in step 204 is as follows: Priority rule: Normalized-FPT≥10 and ranked first; If a pathogen has a Normalized-FPT value greater than or equal to 10 in the target sample and ranks first among all candidate pathogens, it can be directly identified as the responsible pathogen of the sample. This threshold setting reflects its significant abundance enrichment advantage in the sample and is a quantitative representation of a strong infection signal. Auxiliary rule: When Normalized-FPT values are similar, BSL (Balanced Score) is introduced for determination; If multiple pathogens in a sample have similar Normalized-FPT values and no significant difference in their ranking, making it difficult to identify them definitively based solely on abundance, it is necessary to further introduce biosafety levels as an auxiliary criterion. In this case, pathogens with higher BSL levels should be given priority as the responsible pathogens, reflecting their high-risk characteristics in terms of clinical pathogenicity and transmission risk.