Standardized method and system for pathogen metagenome background bacterium level identification

By constructing a standardized method for identifying background bacteria in pathogen metagenomics and utilizing a dual determination mechanism of FC difference and standardized PG quantity, the problem of contamination control in mNGS detection was solved, enabling accurate identification and quantification of background bacteria and improving the accuracy and reliability of detection results.

CN121905301APending Publication Date: 2026-04-21PEOPLES HOSPITAL OF HENAN PROV +2
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202512044555.7
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-12-31
Publication Date
2026-04-21

AI Technical Summary

Technical Problem

Existing metagenomic high-throughput sequencing (mNGS) technology for pathogens faces challenges in contamination control during the detection process. It lacks standardized contamination thresholds, cannot quantitatively analyze contamination from different sources, and struggles to distinguish between low-level infection and contamination, leading to false positive or false negative results.

Method used

This paper provides a standardized method for identifying background bacteria at the pathogen metagenomic level. By constructing a microbial relative abundance matrix and a P-value matrix, hierarchical clustering is used to screen core background bacteria, a historical reference library is constructed, and a dual judgment mechanism is adopted to identify background bacteria by combining the FC fold difference index and standardized PG quantity. The empirical background bacteria are integrated to form a final marker set for signal classification and judgment.

Benefits of technology

It enables precise identification and quantification of background bacteria, ensuring the accuracy and reliability of test results. It can distinguish between background bacteria and pathogenic bacteria, reducing false positive or false negative results in low biomass samples.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN121905301A_ABST
    Figure CN121905301A_ABST
Patent Text Reader

Abstract

The invention discloses a standardization method and system for pathogen metagenome background bacterium level identification, and relates to the technical field of clinical pathogenic microorganism detection, the standardization method comprises the following steps: screening core background flora based on metagenome sequencing data; constructing a historical reference library; screening a background bacterium marker; evaluating background bacterium signal intensity; and determining the background bacterium level. According to the standardization method and the judgment system for pathogen metagenome background bacterium level identification, clinical specimen types are not distinguished, a background bacterium judgment method based on the standardized pg amount is introduced, and it is ensured that background bacterium detection amounts of different samples have comparability; and meanwhile, the accuracy of a detection result and the reliability of clinical interpretation can be improved by utilizing a background bacterium identification mechanism with dual-index collaboration, a historical reference library and a dynamic updating mechanism.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to the field of clinical pathogen detection technology, and in particular to a standardized method and system for identifying background bacteria at the metagenomic level of pathogens. Background Technology

[0002] In recent years, the incidence of infectious diseases has shown a significant upward trend, placing higher demands on etiological diagnostic techniques. Traditional pathogen detection methods, including isolation and culture, immunological detection, and real-time quantitative PCR, have shown significant limitations in clinical practice due to factors such as limited culture conditions, insufficient antibody specificity, or narrow target range. Statistics show that traditional methods can only provide a definitive etiological diagnosis for less than 30% of infectious cases. This situation has spurred the rapid development of next-generation sequencing technology in the field of clinical microbiology testing.

[0003] Metagenomic Next-Generation Sequencing (mNGS) technology breaks through the limitations of traditional culture methods by directly sequencing the whole genome of clinical specimens, achieving unbiased detection of pathogens such as bacteria, fungi, viruses, and parasites. This technology has the following significant advantages: (1) broad detection spectrum, covering tens of thousands of pathogens in a single test; (2) higher detection rate for difficult and rare pathogens; (3) particularly suitable for the identification of mixed infections and difficult-to-culture microorganisms; and (4) unique value in the diagnosis of acute and critical infections and infections in immunocompromised patients. Based on these advantages, mNGS technology has been recommended as an important pathogen diagnostic method by guidelines such as the "Expert Consensus on the Clinical Application of Chinese Metagenomic Next-Generation Sequencing Technology for Detecting Infectious Pathogens".

[0004] However, the clinical application of mNGS technology still faces significant challenges. The detection process of this technology is complex, involving several key steps: (1) wet experimental part: including sample pretreatment, host nucleic acid removal, nucleic acid extraction, library construction and sequencing, etc.; (2) dry experimental part: covering the construction of localized microbial databases, bioinformatics analysis and interpretation of clinical reports, etc. It is worth noting that during the wet experimental operation, several steps may introduce exogenous microbial contamination, mainly including: (1) reagent background bacteria: microbial nucleic acids present in sample processing reagents, nucleic acid extraction reagents and library construction reagents; (2) experimental operation contamination: pollutants introduced during personnel operation; (3) environmental microorganisms: environmental bacterial contamination caused by incomplete laboratory disinfection; (4) contamination carried by consumables: microorganisms carried during the production or storage of experimental consumables. The presence of these sources of contamination seriously interferes with the accuracy of the detection results and increases the difficulty of interpreting the reports.

[0005] Currently, clinical laboratories primarily monitor contamination by setting up negative controls (NCs). However, this method has significant limitations: (1) a lack of standardized contamination thresholds; (2) an inability to quantitatively analyze contamination from different sources; (3) difficulty in distinguishing between low-level infection and contamination; and (4) limited ability to differentiate between colonizing and contaminating bacteria. In particular, when testing low-biomass samples, contaminating nucleic acids may mask the true pathogen signal, leading to false positive or false negative results.

[0006] Therefore, addressing the urgent need for contamination control in mNGS testing and considering the shortcomings of current technology, this paper proposes developing a standardized metagenomic background bacteria identification method and establishing a scientific quantitative contamination assessment system. This will have significant clinical application value in improving the accuracy of mNGS testing and reducing the difficulty of report interpretation. Solving this technical challenge will significantly enhance the reliability and practicality of mNGS in the diagnosis of infectious diseases. Summary of the Invention

[0007] To address the above problems, this invention provides a standardized method and system for identifying pathogenic metagenomic background bacteria.

[0008] In a first aspect, the present invention provides a standardized method for identifying pathogenic metagenomic background bacteria, comprising the following steps: Screening of core background microbial communities based on metagenomic sequencing data: A microbial relative abundance matrix was constructed based on NC samples from metagenomic sequencing over a period of time. The microbial correlation matrix and P-value matrix were calculated using software. After filtering out microbial relationship pairs with P-values ​​less than 0.01, hierarchical clustering was used to select the microorganisms in the largest cluster as the main core background microorganisms. Constructing a historical reference library: The reference library is constructed according to four periods: the past 1 month, 2 months, 3 months, and 1 year, and the host-free / non-host-free treatment type. The data structure is formed by detection rate statistics, filtering of strongly positive cross-contamination microorganisms, calculation of standardized PG amount, filtering of outlier box plots, and calculation of statistical indicators. Background marker screening: Artificial NC control specimens were constructed. For clinical specimens, colonizing bacteria were iteratively filtered by standardized PG amount to obtain initial marker bacteria. The initial marker bacteria were then corrected by FC fold difference index and integrated with empirical background bacteria to form the final marker set. Background bacterial signal intensity was assessed: signal grading was performed based on the percentage of initial marker bacteria from the top 5 genera, the ratio of initial marker bacteria to Share bacteria, nucleic acid extraction concentration, and the total number of bacterial / fungal species detected. Determining the background microbial level: Through a dual determination mechanism of FC difference index and standardized PG index, combined with the priority determination rule for pathogenic microorganisms and the manual review and labeling rule, the final determination result of background / prominent microorganisms is output.

[0009] Furthermore, the step of constructing a microbial relative abundance matrix based on metagenomic sequencing NC samples over a period of time, calculating a microbial correlation matrix and a P-value matrix using software, and filtering microbial relationship pairs with P-values ​​less than 0.01, followed by selecting microorganisms from the largest cluster as the main core background bacteria using hierarchical clustering, includes the following process: A microbial relative abundance matrix was constructed based on NC specimens from metagenomic sequencing within one year. The microbial correlation matrix and p-value matrix were calculated using FastSpar software. 1000 random datasets were generated using the bootstrap method and the p-value matrix was calculated. After filtering out microbial relationship pairs with p-values ​​< 0.01, hierarchical clustering was performed using the hclust function in R language. Microorganisms in the largest cluster were selected as the main core background bacteria, satisfying that the detection amount of this bacterial community in more than 95% of the specimens is ≥ 30% of the total detection amount of the specimens.

[0010] Furthermore, the steps of constructing a reference library based on four periods (1 month / 2 months / 3 months / 1 year) and host-free / non-host-free treatment types, and forming a data structure through detection rate statistics, filtering out strongly positive cross-contaminating microorganisms, calculating standardized PG levels, filtering out outliers using box plots, and calculating statistical indicators include the following processes: NC specimens were categorized into four time periods (1 month, 2 months, 3 months, and 1 year) and host-free / non-host-free treatment types. After excluding NC specimens with less than 5 detected species, the following procedure was performed: a) Detection rate statistics: Calculate the detection frequency and number of each microorganism in the negative control; b) Strong positive cross-contamination filtering: If the max_RN of a certain microorganism in the same batch is ≥10000 and the number of sequences detected in the negative control is <max_RN×0.005%, then the microorganism is removed; c) Standardized PG quantity calculation: Set the total amount of core background bacteria to 100pg, calculate OPR=100pg / Common_tRN based on the total number of detected sequences Common_tRN, and derive PG_spe=Spe_RN×OPR; d) Outlier filtering: The up_bound value is calculated using the box plot method as 75th percentile + 1.5 × IQR, and a triple filtering rule is applied. e) Statistical indicator calculation: Calculate the mean, median, standard deviation, percentiles, and up-bound; f) Data structuring: Create a reference library containing statistical indicators such as wet experimental treatment type, period, microbial name, and 95th percentile and 99th percentile.

[0011] Furthermore, the steps of constructing artificial NC control specimens, for clinical specimens, to obtain initial marker bacteria through standardized PG iterative filtering of colonizing bacteria, and then integrating empirical background bacteria to form the final marker set after secondary correction by FC fold difference index, include the following processes: a) Construction of artificial negative controls: Calculate the Pearson correlation coefficient between recent negative controls and reference specimens (R≥0.7 and P≤0.05), screen homogeneous negative controls, calculate the normalized PG levels of each microorganism in homogeneous NC specimens, remove strongly positive contaminants, and calculate the average normalized PG level of each microorganism based on its detection frequency to construct artificial negative controls; b) Initial marker screening: Core background bacteria co-detected by clinical specimens and artificial negative control are sorted in descending order of clinical specimen sequence number. Iterative judgment is made: if the standardized PG amount is ≤ the historical reference library threshold, it is judged as background bacteria; otherwise, it is excluded until the judgment is completed. c) FC multiple correction: The ratio of the number of sequences in clinical specimens to those in artificial negative control is calculated. When the number of initial markers is >3, box plot analysis is performed to determine FC_bound. Markers with FC ≤ FC_bound and negative control detection rate ≥ 0.3 are screened and empirical background bacteria (Propionibacterium acnes, Komagataella phaffii) are merged to form the final marker set.

[0012] Furthermore, the step of signal grading based on the proportion of initial marker bacteria within the top 5 genera, the ratio of initial marker to Share bacteria, nucleic acid extraction concentration, and the total number of detected bacterial / fungal species includes the following processes: Based on the percentage of the top 5 genera, the ratio of initial marker to Share bacteria, and the concentration of extracted nucleic acid: Based on the percentage of initial marker bacteria within the top 5 genera, the ratio of initial marker to Share bacteria, nucleic acid extraction concentration, and the total number of bacterial / fungal species detected: a) Normal signal: The proportion of the initial marker in the TOP5 genera is ≥0.6 and the ratio of the initial marker to the Share bacteria is ≥0.8 and the extraction concentration is ≤0.5. The signal is judged as strong (<0.5) or very strong (<0.25) based on the proportion of Share bacteria. b) Weak signal: The proportion of the initial marker in the TOP5 genera is ≤0.2 or the ratio of the initial marker to Share bacteria is <0.5; c) Extremely weak signal: The proportion of initial markers in the TOP5 genera is ≤0.2, the extraction concentration is ≥0.5 and the number of initial markers for bacteria is ≤1, or the number of initial markers for fungi is 0.

[0013] Furthermore, the step of outputting the final judgment result of background / prominent status through the dual judgment mechanism of FC difference index and standardized PG index, combined with the pathogenic microorganism priority judgment rule and manual review marking rule, includes the following process: a) FC index determination: Bacterial / fungal FC ≥ 2 × FC_bound is considered prominent, and viral / parasitic FC ≥ 5 is considered prominent; b) PG index determination: Select the reference library cutoff value based on the detection frequency and frequency count. If the detection frequency in the past month is ≥0.07 and the frequency count is ≥6, the determination result of the past month shall be given priority. Otherwise, the indicators of the past 2 months / 3 months / 1 year shall be determined until a determination result is selected. Pathogenic microorganisms shall directly use the results of the past month. c) Combination of dual indicators: When the indicators are consistent, the judgment is made directly; when FC is prominent / PG is in the background, the pathogenic bacteria need to be manually reviewed; when FC is in the background / PG is prominent, the pathogenic bacteria and the prominence in the past month are judged as prominent; when there is no judgment for FC, it is consistent with PG.

[0014] Furthermore, the three rules for outlier filtering are as follows: microorganisms with a detection frequency of <60 times are removed if the PG amount is greater than up_bound; microorganisms with strong positive results in the same batch and a PG amount greater than up_bound are removed; and all outliers are removed if the proportion of abnormal samples is <15%. And / or, the selection rule for the cutoff value of the reference library is as follows: when the detection frequency is ≥0.07, the minimum value of the quartile array is taken for extremely weak signals, the third minimum value is taken for slightly weak signals, the second minimum value is taken for normal / slightly strong signals, and the maximum value is taken for extremely strong signals; when the detection frequency is <0.07, the minimum value is taken for extremely weak signals when the detection frequency is >6, the third minimum value is taken for slightly weak / normal / slightly strong signals, and the maximum value is taken for extremely strong signals; the minimum value is taken for all signals when the detection frequency is ≤6. And / or, the dynamic update mechanism includes automatically updating the reference library data according to a preset period, constructing it in combination with wet experimental treatment type, and using a dual threshold of detection frequency and frequency number to determine the cutoff value selection strategy.

[0015] Furthermore, the method is applicable to the identification of background microbial levels of bacteria, fungi, viruses, and parasites.

[0016] Secondly, based on the same inventive concept, the present invention provides a determination system for implementing the standardized method for identifying pathogen metagenomic background bacteria level as described in the first aspect, comprising: The core background microbial community module is used to construct a relative abundance matrix of microorganisms based on metagenomic sequencing NC samples over a period of time. The software calculates the microbial correlation matrix and P-value matrix. After filtering out microbial relationship pairs with P-values ​​less than 0.01, hierarchical clustering is used to select the microorganisms in the largest cluster as the main core background microorganisms. The historical reference library construction module is used to build reference libraries according to four periods: the past 1 month, 2 months, 3 months, and 1 year, and the host-free / non-host-free treatment type. The data structure is formed by detection rate statistics, filtering of strongly positive cross-contamination microorganisms, calculation of standardized PG amount, filtering of outlier box plots, and calculation of statistical indicators. The background bacteria marker module is used to construct artificial NC control specimens. For clinical specimens, the initial marker bacteria are obtained by iteratively filtering colonizing bacteria through standardized PG amount. The initial marker bacteria are then corrected by the FC fold difference index and integrated with empirical background bacteria to form the final marker set. The background bacteria signal intensity assessment module is used to classify signals based on the percentage of the initial marker bacteria from the top 5 genera, the ratio of the initial marker to the Share bacteria, the nucleic acid extraction concentration, and the total number of bacterial / fungal species detected. The background microbial level determination module is used to output the final determination result of background / prominent microorganisms through a dual determination mechanism of FC multiplier index and standardized PG index, combined with the pathogenic microorganism priority determination rule and manual review and labeling rule.

[0017] Thirdly, based on the same inventive concept, the present invention provides a computer-readable storage medium storing computer-executable instructions, which, when executed by a processor, are used to implement a standardized method for identifying pathogenic metagenomic background bacteria as described in any one of the first aspects.

[0018] The technical solutions provided in the embodiments of the present invention have at least the following advantages compared with the prior art: This invention provides a standardized method and system for identifying background bacteria levels in pathogen metagenomics. Specifically, this invention addresses the introduction of background microorganisms from reagents or the experimental environment during metagenomic high-throughput sequencing (mNGS) detection. It proposes a standardized method and system for identifying background bacteria levels in pathogen metagenomics without distinguishing between clinical specimen types. This not only introduces a background bacteria determination method based on standardized pg levels, ensuring comparability of background bacteria detection levels across different samples, but also improves the accuracy of detection results and the reliability of clinical interpretation by utilizing a dual-indicator synergistic background bacteria identification mechanism, a historical reference library, and a dynamic update mechanism. Specifically: 1) This invention introduces a background bacteria determination method based on standardized pg levels, ensuring the comparability of background bacteria detection levels among different samples. By comparing with historical negative control samples, the standardization process eliminates deviations caused by operational and environmental factors, ensuring accurate quantification of background bacteria levels. This method addresses the shortcomings of existing technologies in lacking quantitative determination and standardization, enabling precise assessment of the degree of background bacteria contamination.

[0019] 2) In this invention, the screening and identification of background bacteria not only relies on the PG standardization method but also incorporates the Fold Change (FC) index, forming a dual judgment mechanism. Specifically: PG index: Based on the standardized amount (pg) of background bacteria, background bacteria are screened and confirmed by comparing the amount of background bacteria in historical negative control samples. FC index: By comparing the number of bacterial sequences detected in clinical specimens and negative control samples, the Fold Change is calculated to further confirm the difference between background bacteria and true pathogens. The combination of these two indicators makes the screening of background bacteria markers more accurate, effectively distinguishing background bacteria from true pathogens, and ensuring the accuracy of background bacteria detection results in low biomass samples.

[0020] 3) This invention constructs a reference library using historical negative control samples and incorporates a dynamic update mechanism to ensure that the background bacteria determination method is based on the latest data. The historical reference library provides stable background bacteria data, and dynamic updates ensure that the method reflects the latest environmental changes and experimental data, thereby improving the stability and accuracy of background bacteria determination.

[0021] In summary, this invention addresses the challenge of identifying background bacteria introduced by reagents or the environment in metagenomic sequencing (mNGS). It innovatively constructs a standardized system for determining and evaluating background bacteria, enabling the quantification of microorganisms detected in clinical and negative control specimens, thus ensuring comparability of data across different specimens. Furthermore, this invention integrates various indicator data, particularly in the background bacteria marker screening process, combining the FC difference index and standardized PG index for background bacteria determination. This allows for the rapid and accurate differentiation of the background level (prominent / background) of detected bacteria in the specimen. For bacteria with low bacterial loads, it also provides suggestions for manual review, assisting in precise clinical interpretation. This provides an important tool for promoting the standardization and automation of clinical mNGS report interpretation within the industry. Attached Figure Description

[0022] The accompanying drawings, which are incorporated in and form part of this specification, illustrate embodiments consistent with the invention and, together with the description, serve to explain the principles of the invention.

[0023] To more clearly illustrate the technical solutions in the embodiments of the present invention or the prior art, the drawings used in the description of the embodiments or the prior art will be briefly introduced below. Obviously, for those skilled in the art, other drawings can be obtained based on these drawings without creative effort.

[0024] Figure 1 A complete flowchart of a standardized method for identifying pathogenic metagenomic background bacteria at the level provided in an embodiment of the present invention.

[0025] Figure 2 This is an enlarged flowchart of the current negative control monitoring and artificial NC construction module in the standardized method for identifying pathogenic metagenomic background bacteria level provided in the embodiments of the present invention.

[0026] Figure 3 This is an enlarged flowchart of the common background bacteria C1 and historical reference library construction module in the standardized method for identifying background bacteria at the pathogen metagenomic level provided in the embodiments of the present invention.

[0027] Figure 4 This is an enlarged flowchart of the initial marker screening module in the standardized method for identifying pathogen metagenomic background bacteria level provided in the embodiments of the present invention.

[0028] Figure 5 This is an enlarged flowchart of the final marker screening module in the standardized method for identifying pathogen metagenomic background bacteria level provided in the embodiments of the present invention.

[0029] Figure 6 This is an enlarged flowchart of the background signal level assessment module in the standardized method for identifying background bacteria level of pathogen metagenomics provided in an embodiment of the present invention.

[0030] Figure 7 This is an enlarged flowchart of the FC / PG dual-indicator collaborative determination of background bacteria level module in the standardized method for identifying background bacteria level of pathogen metagenomics provided in the embodiments of the present invention.

[0031] Figure 8 This is an enlarged flowchart of the background bacteria level comprehensive determination module in the standardized method for identifying background bacteria levels in pathogen metagenomics provided in this embodiment of the invention.

[0032] Figure 9 This is a partial result of the historical reference library in Embodiment 2 of the present invention.

[0033] Figure 10 This is a partial data result of the manual NC detection table in Embodiment 2 of the present invention.

[0034] Figure 11 After filtering out bacteria in all three indicators of Example 2 of this invention, which were determined to be background bacteria, the final data results of the form were interpreted. Detailed Implementation

[0035] To make the objectives, technical solutions, and advantages of the embodiments of the present invention clearer, the technical solutions of the embodiments of the present invention will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some embodiments of the present invention, not all embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those skilled in the art without creative effort are within the scope of protection of the present invention.

[0036] The present invention will be further illustrated below with reference to specific embodiments. It should be understood that these embodiments are for illustrative purposes only and are not intended to limit the scope of the invention. Experimental methods in the following embodiments, unless otherwise specified, are generally performed according to national standards. If no corresponding national standard exists, then generally accepted international standards, conventional conditions, or conditions recommended by the manufacturer are followed.

[0037] Example 1 This example provides a standardized method for identifying pathogen metagenomic background bacteria at the level, such as... Figure 1 , Figure 2 , Figure 3 , Figure 4 , Figure 5 , Figure 6 , Figure 7 and Figure 8 As shown, it includes the following steps: 1. Screening of key background bacteria in negative control Based on the assumption that the detection rate of background bacteria in each sample remains relatively constant over a certain period, a major core background bacterial group can be screened from the microbial detection profile, ensuring that the detection rate of these bacteria reaches more than 30% of the total detection rate in more than 95% of the samples. Specifically, NC samples that underwent metagenomic sequencing within one year were collected, and the relative abundance matrix of detected microorganisms in all samples was calculated. The correlation matrix between each detected microorganism was calculated using FastSpar (version: V1.0.0, parameters: --iterations 50 --exclude_iterations 20). Based on the bootstrap method, the fastspar_bootstrap (version: V1.0.0, parameters: --number 1000) software was used to first obtain 1000 random dataset matrices; then, FastSpar software was used to calculate the correlation of each matrix; finally, the fastspar_pvalues ​​(version: V1.0.0, parameters: --permutations 1000) software was used to calculate the p-value matrix between microorganisms.

[0038] After filtering out microbial relationship pairs with P values ​​less than 0.01 from the calculated correlation and P-value matrices, a new core microbial correlation matrix is ​​generated. Hierarchical clustering is then performed using the R hclust function, and the microorganisms in the largest cluster are selected as the main core background bacteria.

[0039] The main core background bacteria obtained from the screening are shown in Table 1.

[0040] Table 1 2. Methods for constructing historical reference libraries Species detection data from historical negative control specimens were collected and divided into four periods: the past 1 month, 2 months, 3 months, and 1 year, according to a pre-defined time frame. These were further categorized into host-de-host negative control and non-host-de-host negative control groups based on the wet experimental treatment method. After removing specimens with fewer than 5 detected species, historical reference libraries for each period and treatment type were constructed, including the following steps: ① Detection rate statistics: Calculate the detection frequency and number of detections of each microorganism in the negative control specimens; ② Filtering out strongly positive cross-contamination microorganisms: If a microorganism has a strongly positive detection with a sequence number exceeding 10,000 in the same batch of specimens (denoted as max_RN), and the number of detected sequences of the microorganism in the negative control is less than max_RN × 0.005%, it is determined to be a source of strongly positive cross-contamination and is removed from the species detection list of the corresponding negative control. ③ Calculation of raw standardized PG amount: The total amount of major core background bacteria in each negative control specimen is set at 100 pg. The standardized PG amount (OPR) per unit sequence is calculated based on the total number of detected sequences of major core background bacteria (Common_tRN). The calculation formula is: OPR = 100 pg / Common_tRN. Then, the standardized PG amount of each detected species is calculated: PG_spe = Spe_RN × OPR; ④ Outlier filtering: For the standardized PG dataset of each microorganism, the upper bound of outliers is calculated using the box plot method: up_bound = 75th percentile + 1.5 × (75th percentile - 25th percentile); a. For microorganisms detected less than 60 times, remove those with standardized PG levels higher than up_bound; b. Microorganisms in the same batch that have strong positive results (max_RN ≥ 10000) and whose standardized PG levels are higher than up_bound should be removed; c. If the proportion of samples with a standardized PG level higher than up_bound for a certain microorganism is less than 15% of the total detected samples, then all abnormal values ​​are removed. ⑤ Statistical database construction indicators: After filtering out outliers, calculate the mean, median, minimum, maximum, standard deviation, mean to median ratio, 95th percentile, 99th percentile, and up_bound for each microorganism; ⑥ Historical reference library composition: Construct a data structure that includes wet experimental treatment method, period, microbial name, original detection rate and frequency, mean, median, minimum, maximum, standard deviation, ratio of mean to median, 95th percentile, 99th percentile and upper bound of outliers to form a historical reference library.

[0041] 3. Background bacteria marker screening ① Construction of artificial negative control specimens Using the current batch of negative control specimens as reference specimens, the Pearson correlation coefficient between recent negative control specimens and reference specimens was calculated using the `cor_test` function in R language. Specimens with a correlation coefficient R ≥ 0.7 and a significance P ≤ 0.05 were selected as homogeneous negative controls. For both reference and homogeneous specimens, the mean number of standardized sequences and detection rate of each detected microorganism were calculated. The standardized PG levels of each microorganism were calculated according to steps 1-③. After removing microorganisms with strong positive cross-contamination, artificial negative control specimens were constructed.

[0042] ② Initial marker screening based on standardized PG levels: First, initial markers were screened from species co-detected in clinical specimens and artificial negative controls that belonged to the major core background bacteria. The specific process was as follows: a) Calculate the initial standardized PG level: The total PG level of all major core background bacteria co-detected in clinical specimens and artificial negative controls was set constant. Based on this, the initial standardized PG level was calculated according to the number of sequences detected in the clinical specimens for each bacterium. b) Iterative filtering of colonizing or infectious bacteria: The co-detected bacteria were sorted in descending order of the number of sequences detected in the clinical specimens, and an iterative method was used for judgment. In each round of judgment, only the species with the highest sequence number was judged: if its standardized PG level did not exceed the preset threshold of the corresponding species in the historical reference library, it was judged as a background bacterium; otherwise, it was judged as a colonizing or infectious bacterium and excluded. This iterative process was repeated until all species were judged, and the set of all species judged as background bacteria was used as the initial markers.

[0043] ③ Marker quadratic correction based on sequence number fold difference (FC) index For the initial marker set, a secondary correction is performed based on the sequence number difference between it and the artificial negative control: a) Difference calculation: For each initial marker, the ratio of its detected sequence number in clinical specimens to that in the artificial negative control is calculated to obtain the difference index. b) Determining the final marker set: Based on the number of initial markers (calculated separately for bacteria and fungi), different screening strategies are applied: If the number is greater than 3, a box plot analysis is performed on the difference index to calculate the difference analysis abnormality boundary value (FC_bound). Initial markers within the normal range and empirical background bacteria (Propionibacterium acnes and Komagataella phaffii) with a detection frequency of not less than 0.3 in the artificial negative control are selected to constitute the final markers; if the number is less than 3, the initial markers and empirical background bacteria (Propionibacterium acnes and Komagataella phaffii) are directly merged as the final markers.

[0044] 4. Background bacterial signal assessment Background bacterial signals were assessed separately for bacteria and fungi. The assessment involved several steps: First, the background bacterial signal was assumed to be normal. Then, initial marker bacteria were screened from the top 5 genera, and their proportion was calculated. If the proportion was ≥0.6, further analysis was needed to determine the ratio of initial marker bacteria to shared bacteria and the nucleic acid extraction concentration. If the ratio was ≥0.8 and the extraction concentration was ≤0.5, the signal was assessed as strong (ratio <0.5) or very strong (ratio <0.25) based on the ratio of shared bacteria to bacterial or fungal species. If the proportion was <0.6, further analysis was required. If the proportion of initial markers in the top 5 genera was ≤0.2, or the ratio of initial markers to shared bacteria was <0.5, the signal was assessed as weak. When the extraction concentration was ≥0.5, and the number of initial bacterial markers was less than or equal to 1, or acne was not detected / acne was detected but not in the background, the signal was considered very weak. Furthermore, if the number of initial fungal markers was 0, the background signal was also assessed as very weak.

[0045] 5. Determination of background bacterial levels ① Based on the FC multiple index: For bacteria co-detected in clinical specimens and artificial NC, the FC multiple value can be calculated. For bacteria or fungi, if FC ≥ 2 * FC_bound (calculated in 4-③-b), it is judged as prominent; otherwise, it is background. For viruses or parasites, if FC ≥ 5, it is prominent; otherwise, it is background.

[0046] ②Based on standardized PG metrics: a. Standardized PG index for each period is determined separately: For the historical reference libraries of each detected bacterium for the past 1 month, 2 months, 3 months, and 1 year, the standardized PG level of the bacterium detected in the clinical specimen is first compared with the cutoff of the reference library. If the PG level is ≥ the cutoff, the bacterium is considered prominent; otherwise, it is considered background.

[0047] Reference library cutoff selection rules: First, sort the 99th percentile, 95th percentile, up_bound, and 0 values ​​of the reference library from largest to smallest. Second, when the detection frequency is ≥ 0.07, the cutoff is the minimum of the four values ​​when the background bacterial signal is extremely weak, the third smallest when the background bacterial signal is slightly weak, the second smallest when the background bacterial signal is normal or slightly strong, and the maximum when the background bacterial signal is extremely strong. When the detection frequency is < 0.07, the microbial detection frequency is further determined. If the detection frequency is > 6, the cutoff is the minimum when the background bacterial signal is extremely weak, the third smallest when the background bacterial signal is slightly weak, normal, or slightly strong, and the maximum when the background bacterial signal is extremely strong; if the detection frequency is ≤ 6, all cutoffs are the minimum values.

[0048] b. Comprehensive Judgment of Standardized PG Indicators for Each Period: Based on the background bacterial level judgment results of the four periods in section a, a final judgment based on PG indicators is given. If the detection frequency in the past month is ≥0.07 and the number of detections is ≥6, the judgment of the past month is given priority as the final result; otherwise, the judgments for the past two months, past three months, and past year are made separately according to this rule until the final judgment result is determined. Additionally, for pathogenic bacteria, pathogenic fungi, or valued bacteria / fungi (valued bacteria label), the judgment result of the past month is directly selected as the final result; ③Combined judgment and manual review of FC and PG indicators: a. The FC indicator and PG indicator results are consistent (highlight / background): The overall judgment is highlight / background, and no manual review is required; b. FC indicator is considered prominent, PG indicator is considered background: If it is pathogenic bacteria / fungus, it is considered prominent overall and requires manual review. Non-pathogenic bacteria / fungus are considered background overall and do not require manual review. c. If the FC indicator is judged as background and the PG indicator as prominent: If it is pathogenic bacteria / fungus and has been judged as prominent in the past month, it will be judged as prominent overall and requires manual review; if it has been judged as background in the past month, it will be judged as background overall and does not require manual review. Non-pathogenic bacteria / fungus will be judged as background overall and does not require manual review. d.FC No judgment: The overall judgment result is consistent with the PG indicator judgment, and no manual review is required.

[0049] Example 2 In this example, the standardized method for identifying the background bacteria level of pathogen metagenome provided in the above Embodiment 1 is applied to collect clinical specimens to verify the background bacteria determination system, including the following steps: 1. Collection of clinical specimens and negative control specimens In this invention, 2137 clinical specimens were collected from cooperative hospitals from October 2024 to March 2025 in the recent 6 months, including 1384 cases of peripheral blood, 464 cases of bronchoalveolar lavage fluid, 138 cases of cerebrospinal fluid, 63 cases of sputum, 20 cases of pleural effusion, and 67 cases of other specimen types. All 2137 specimens were subjected to mNGS metagenomic sequencing, totaling 713 batches. After sequencing, the report interpreters interpreted the results. A total of 6459 pathogenic microorganisms were reported in 2137 specimens, including 2043 bacteria, 806 fungi, 3501 viruses, and 9 parasites. 1297 negative control specimens from the clinical specimen sequencing laboratory in the recent 1 year and 3 months were collected, including 378 host-depleted negative controls and 919 non-host-depleted negative controls.

[0050] 2. Construction of a common background bacteria library Based on the 1297 historical negative control specimens collected in step 1, a relative abundance matrix of detected bacteria was generated. The FastSpar (version: V1.0.0, parameters: --iterations 50 --exclude_iterations 20) software was used to calculate the correlation matrix between detected microorganisms; based on the bootstrap method, the fastspar_bootstrap (version: V1.0.0, parameters: --number 1000) software was used to preferentially obtain 1000 random dataset matrices; then the FastSpar software was used to calculate the correlation of each matrix respectively; finally, the fastspar_pvalues (version: V1.0.0, parameters: --permutations 1000) software was used to calculate the P-value matrix between microorganisms. After filtering out the microorganism relationship pairs with P-values less than 0.01, a core microorganism correlation matrix was regenerated, and the R hclust function was used for hierarchical clustering. The hierarchical clustering results were divided into 2 clusters, and the microorganisms in the larger cluster were selected as common background bacteria. Finally, a background bacteria library containing 186 common background bacteria was constructed (Table 1 in Embodiment 1).

[0051] 3. Construction of a historical reference library and artificial NC for each batch of clinical specimens ① Construction of the historical reference library: Based on 378 host-depleted negative controls and 919 non-host-depleted negative controls, after filtering out unqualified negative control specimens (number of detected species < 5) and contaminated microorganisms (strong positive cross-contaminated bacteria, RN < max_RN * 0.005%), the detection rates and detection frequencies of each bacterium in the recent 1 month / recent 2 months / recent 3 months / recent 1 year were respectively counted. According to the PG standardization method: PG = Spe_RN * 100pg / Common_tRN Where Spe_RN is the number of detected bacterial sequences, and Common_tRN is the total number of bacterial sequences co-detected in the NC specimen and the common background bacterial library.

[0052] Calculate the standardized PG amount for all detected bacteria. For each detected bacteria, perform boxplot analysis. Abnormal PG amounts that are abnormal in the boxplot analysis and have strong positives (max_RN≥1W) or a detection frequency <60 in the same batch, or abnormal in the boxplot analysis and have an outlier percentage ≤15%, are removed. Calculate the mean, median, mean / median ratio, 95th percentile, 99th percentile, and up-bound, and construct historical reference libraries for different host-de-hosted / non-host-de-hosted periods over the past 1 month, 2 months, 3 months, and 1 year.

[0053] ② Constructing Artificial NCs: Using the current batch of negative controls as ref_NCs, the Pearson correlation coefficient between recent (default 5 days) NCs and ref_NCs was calculated using the cor_test function in R. NCs with R values ​​≥ 0.7 and P values ​​≤ 0.05 were selected as homogeneous samples for ref_NCs. For both ref_NCs and homogeneous NCs, the mean normalized sequence number and detection rate of all detected bacteria were calculated. Simultaneously, the normalized PG content of all detected bacteria was calculated using the method in Example 1. After removing bacteria with strong positive cross-contamination, artificial NCs were constructed.

[0054] Background marker bacteria were screened for each clinical specimen, and background bacterial signals were evaluated. Background marker screening: For Shared bacteria that are co-detected in clinical specimens and manual NC and belong to the main core background bacteria, the standardized PG amount of each Shared bacteria can be calculated according to the method in Example 1: ① Initial screening of candidate markers: The standardized PG values ​​of Share bacteria are compared with the cutoff values ​​of the historical reference library over the past month. Each bacteria is then assessed to determine if it is a background marker. If the top-ranked bacteria is identified as prominent, it is removed, and the standardized PG values ​​of all Share bacteria are recalculated and background marker assessment is performed. If a bacteria is identified as background, the top-ranked bacteria are assessed again, and this process continues until all bacteria have been assessed. Finally, the background markers are determined as the initial markers.

[0055] ② Marker selection based on FC multiple correction: Calculate the FC multiple of the initial marker (the ratio of the number of sequences detected in clinical specimens to those detected in manual NC). The final marker can be confirmed based on the FC and the empirical background bacteria list. The specific rule is that if the number of initial markers is >3, bacteria within the normal range and with a detection frequency ≥0.3 in manual NC are screened through boxplot analysis of the FC multiple; if it is ≤3, both the initial marker and the empirical background bacteria are the final markers. Calculate the upper limit of the 95% CI of the final marker bacteria, FC_bound (i.e., μ + 2*ó).

[0056] ③ Calculate the final standardized PG amount: Based on the final marker, the standardized PG amount of each sequence in the clinical specimen can be calculated according to the method in Example 1, thereby obtaining the standardized PG amount of all detected bacteria.

[0057] Background bacterial signal assessment: Based on the percentage of initial markers from the top 5 genera (calculated once per genus), the ratio of initial marker bacteria to shared bacteria, nucleic acid extraction concentration, the ratio of shared bacteria to the number of detected bacteria or fungi, and the number of initial marker bacteria, the bacterial and fungal background bacterial signals in clinical specimens can be assessed into five levels: very weak, weak, normal, strong, and very strong. The specific rule is: the default assumption is that the background bacterial signal is normal. Then, initial marker bacteria are screened from the top 5 genera, and their percentage is calculated. If the percentage is ≥0.6, it indicates that the initial marker bacteria are near the top of the clinical specimen detection spectrum, and the signal may be strong or very strong. In this case, further analysis is needed to determine the ratio of initial marker bacteria to shared bacteria and the nucleic acid extraction concentration. If the ratio is ≥0.8 and the extraction concentration is ≤0.5, the signal can be assessed as strong (ratio <0.5) or very strong (ratio <0.25) based on the ratio of shared bacteria to the number of bacterial or fungal species. If the percentage is <0.6, further analysis is required. If the proportion of the initial marker in the top 5 genera is ≤0.2, or the ratio of the initial marker to Share bacteria is <0.5, the signal is assessed as weak. If the extraction concentration is ≥0.5, and the number of initial bacterial markers is less than or equal to 1, or acne is not detected, or acne is detected but not in the background, the signal is assessed as extremely weak. In addition, if the number of initial fungal markers is 0, the background signal is also assessed as extremely weak.

[0058] 5. Determine the background bacterial levels for each detected bacterium in clinical specimens. ① FC (Fluorescence) index determination: For co-detected bacteria in clinical specimens and artificial NC (Nuclear Non-Nuclear) assays, the FC (Fluorescence) ratio of the detected sequences can be calculated. If the FC ratio of bacteria and fungi is ≥ 2 * FC_bound (4-②), it is considered prominent; otherwise, it is considered background. If the FC ratio of viruses and parasites is ≥ 5, it is considered prominent; otherwise, it is considered background.

[0059] ② Standardized PG index determination: Determination is given for R1M / R2M / R3M / R1Y stages. The cutoff selection process is as follows: select from 0, N95, N99, and up_bound. When the detection frequency is ≥0.07, the minimum value is selected for no background / very weak, the third smallest value is selected for weak, the second smallest value is selected for normal / strong, and the maximum value is selected for very strong. When the detection frequency is <0.07 and the detection frequency is >6, the minimum value is selected for no background / very weak, the third smallest value is selected for weak / normal / strong, and the maximum value is selected for very strong. Clinical specimen species standardized PG count ≥ cutoff is considered prominent, otherwise it is considered background.

[0060] ③ Final comprehensive judgment and manual review labeling: When the FC and PG indicators are consistent, the comprehensive judgment is also consistent, and no manual review is required; if the FC is judged as prominent and the PG is background, the P label needs to be manually reviewed and judged as prominent, while non-P labels are judged as background and no review is required; if the FC is background and the PG is prominent, the P label needs to be judged according to the R1M judgment, while non-P labels are judged as background and no review is required; if the FC is not judged, the comprehensive judgment is consistent with the PG and no review is required.

[0061] 6. The personnel interpreting the comparison report reported the pathogens, and the performance of the background bacteria level determination is as follows: ① The consistency statistics with human interpretation are shown in Table 2.

[0062] Table 2 ②After determining the background bacterial level, the metagenomic detection form ratio can be simplified as shown in Table 3.

[0063] Table 3 Based on the above test results and Figure 9 ( Figure 9 (Results from a portion of the historical reference library) Figure 10 ( Figure 10 (Partial data results from the manual NC detection table) Figure 11 ( Figure 11(After filtering out bacteria identified as background bacteria by three indicators in step 5, the final interpretation of the data in the form shows that this invention innovatively constructs a standardized background bacteria identification and evaluation system to address the challenge of identifying background bacteria introduced by reagents or the environment in pathogen metagenomic sequencing (mNGS). This system can quantify the microorganisms detected in clinical and negative control specimens, making the data from different specimens comparable. Furthermore, this invention integrates various indicator data, particularly in the background bacteria marker screening process, combining the FC difference index and the standardized PG index for background bacteria identification. This allows for the rapid and accurate differentiation of the background level (prominent / background) of detected bacteria in the specimen. For bacteria with low loads, manual verification suggestions are also provided to assist in precise clinical interpretation, improving the accuracy of test results and the reliability of clinical interpretation.

[0064] Various embodiments of the present invention may exist in the form of a range; it should be understood that the description in the form of a range is merely for convenience and brevity and should not be construed as a hard limitation on the scope of the invention; therefore, it should be considered that the range description has specifically disclosed all possible subranges and single numerical values ​​within that range. For example, it should be considered that the range description from 1 to 6 has specifically disclosed subranges such as from 1 to 3, from 1 to 4, from 1 to 5, from 2 to 4, from 2 to 6, from 3 to 6, etc., and single numbers within the range, such as 1, 2, 3, 4, 5, and 6, regardless of the range. Furthermore, whenever a numerical range is referred to herein, it means including any referenced number (fraction or integer) within the range referred to.

[0065] The above description is merely a specific embodiment of the present invention, enabling those skilled in the art to understand or implement the invention. Various modifications to these embodiments will be readily apparent to those skilled in the art, and the general principles defined herein may be implemented in other embodiments without departing from the spirit or scope of the invention. Therefore, the present invention is not to be limited to the embodiments shown herein, but is to be accorded the widest scope consistent with the principles and novel features claimed herein.

Claims

1. A standardized method for identifying pathogenic metagenomic background bacteria, characterized in that, Includes the following steps: Screening of core background microbial communities based on metagenomic sequencing data: A microbial relative abundance matrix was constructed based on NC samples from metagenomic sequencing over a period of time. The microbial correlation matrix and P-value matrix were calculated using software. After filtering microbial relationship pairs with P-values ​​less than 0.01, hierarchical clustering was used to select the microorganisms in the largest cluster as the main core background microorganisms. Constructing a historical reference library: The reference library is constructed according to four periods: the most recent 1 month, 2 months, 3 months, and 1 year, and the host-free / non-host-free treatment type. The data structure is formed by detection rate statistics, filtering of strongly positive cross-contamination microorganisms, calculation of standardized PG amount, filtering of outlier box plots, and calculation of statistical indicators. Background marker screening: Artificial NC control specimens were constructed. For clinical specimens, colonizing bacteria were iteratively filtered by standardized PG amount to obtain initial marker bacteria. The initial marker bacteria were then corrected by FC fold difference index and integrated with empirical background bacteria to form the final marker set. Assess background bacterial signal intensity: Signal grading was performed based on the percentage of initial markers within the top 5 genera, the ratio of initial markers to the number of shared bacteria, nucleic acid extraction concentration, and the total number of detected bacterial / fungal species. Determining the background microbial level: Through a dual determination mechanism of FC difference index and standardized PG index, combined with the priority determination rule for pathogenic microorganisms and the manual review and labeling rule, the final determination result of background / prominent microorganisms is output.

2. The standardized method for identifying pathogen metagenomic background bacteria according to claim 1, characterized in that, The steps of constructing a microbial relative abundance matrix based on metagenomic sequencing NC samples over a period of time, calculating the microbial correlation matrix and P-value matrix using software, filtering microbial relationship pairs with P-values ​​less than 0.01, and selecting microorganisms from the largest cluster as the main core background bacteria using hierarchical clustering include the following processes: A microbial relative abundance matrix was constructed based on NC specimens from metagenomic sequencing within one year. The microbial correlation matrix and p-value matrix were calculated using FastSpar software. 1000 random datasets were generated using the bootstrap method and the p-value matrix was calculated. After filtering out microbial relationship pairs with p-values ​​< 0.01, hierarchical clustering was performed using the hclust function in R language. Microorganisms in the largest cluster were selected as the main core background bacteria, satisfying that the detection amount of this bacterial community in more than 95% of the specimens is ≥ 30% of the total detection amount of the specimens.

3. The standardized method for identifying pathogen metagenomic background bacteria according to claim 2, characterized in that, The steps involved in constructing a reference library based on four periods (1 month, 2 months, 3 months, and 1 year) and host-free / non-host-free treatment types, and forming a data structure through detection rate statistics, filtering out strongly positive cross-contaminating microorganisms, calculating standardized PG levels, filtering out outliers using box plots, and calculating statistical indicators, include the following processes: NC specimens were categorized into four time periods (1 month, 2 months, 3 months, and 1 year) and host-free / non-host-free treatment types. After excluding NC specimens with less than 5 detected species, the following procedure was performed: a) Detection rate statistics: Calculate the detection frequency and number of each microorganism in the negative control; b) Strong positive cross-contamination filtering: If the max_RN of a certain microorganism in the same batch is ≥10000 and the number of sequences detected in the negative control is <max_RN×0.005%, then the microorganism is removed; c) Standardized PG quantity calculation: Set the total amount of core background bacteria to 100pg, calculate OPR=100pg / Common_tRN based on the total number of detected sequences Common_tRN, and derive PG_spe=Spe_RN×OPR; d) Outlier filtering: The up_bound value is calculated using the box plot method as 75th percentile + 1.5 × IQR, and a triple filtering rule is applied. e) Statistical indicator calculation: Calculate the mean, median, standard deviation, percentiles, and up-bound; f) Data structuring: Create a reference library containing statistical indicators such as wet experimental treatment type, period, microbial name, and 95th percentile and 99th percentile.

4. The standardized method for identifying pathogen metagenomic background bacteria according to claim 3, characterized in that, The steps for constructing artificial NC control specimens, specifically for clinical specimens, include the following: initial marker bacteria are obtained by iteratively filtering colonizing bacteria using standardized PG levels; secondary correction is performed using FC fold differences; and empirical background bacteria are integrated to form the final marker set. a) Construction of artificial negative controls: Calculate the Pearson correlation coefficient (R≥0.7 and P≤0.05) between recent negative controls and reference specimens. After screening for homogeneous negative controls, calculate the standardized PG levels of each microorganism in the homogeneous NC specimens. After removing strongly positive contaminants, calculate the average standardized PG level of each microorganism based on its detection frequency to construct artificial negative controls; b) Initial marker screening: Sort the core background bacteria co-detected by clinical specimens and artificial NC in descending order of clinical specimen sequence number. Iteratively determine: if the standardized PG level is ≤ the historical reference library threshold, it is judged as a background bacteria; otherwise, it is excluded until the determination is completed. c) FC correction: Calculate the ratio of the number of sequences in clinical specimens to those in artificial negative control. When the initial number of markers is >3, perform box plot analysis to determine FC_bound. Screen markers with FC≤FC_bound and negative control detection rate≥0.3, and combine empirical background bacteria to form the final marker set.

5. The standardized method for identifying pathogen metagenomic background bacteria according to claim 4, characterized in that, The step of signal classification based on the proportion of initial marker bacteria within the top 5 genera, the ratio of initial marker to Share bacteria, nucleic acid extraction concentration, and the total number of detected bacterial / fungal species includes the following process: Based on the percentage of initial marker bacteria within the top 5 genera, the ratio of initial marker to Share bacteria, nucleic acid extraction concentration, and the total number of bacterial / fungal species detected: a) Normal signal: The proportion of the initial marker in the TOP5 genera is ≥0.6 and the ratio of the initial marker to the Share bacteria is ≥0.8 and the extraction concentration is ≤0.

5. The signal is judged as strong (<0.5) or very strong (<0.25) based on the proportion of Share bacteria. b) Weak signal: The proportion of the initial marker in the TOP5 genera is ≤0.2 or the ratio of the initial marker to Share bacteria is <0.5; c) Extremely weak signal: The proportion of initial markers in the TOP5 genera is ≤0.2, the extraction concentration is ≥0.5 and the number of initial markers for bacteria is ≤1, or the number of initial markers for fungi is 0.

6. The standardized method for identifying pathogen metagenomic background bacteria according to claim 5, characterized in that, The steps for outputting the final judgment result of background / prominent status through the dual judgment mechanism of FC multiplier index and standardized PG index, combined with the pathogenic microorganism priority judgment rule and manual review and marking rule, include the following process: a) FC index determination: Bacterial / fungal FC ≥ 2 × FC_bound is considered prominent, and viral / parasitic FC ≥ 5 is considered prominent; b) PG index determination: Select the reference library cutoff value based on the detection frequency and frequency count. If the detection frequency in the past month is ≥0.07 and the frequency count is ≥6, the determination result of the past month shall be given priority. Otherwise, the indicators of the past 2 months / 3 months / 1 year shall be determined until a determination result is selected. Pathogenic microorganisms shall directly use the results of the past month. c) Combination of dual indicators: Direct judgment when indicators are consistent; When FC is highlighted / PG background, pathogens require manual review; When FC background / PG prominence occurs, pathogenic bacteria that have been prominenced for the past month are considered to be prominence. When FC has no decision, it is the same as PG.

7. The standardized method for identifying pathogen metagenomic background bacteria according to any one of claims 1 to 6, characterized in that, The three rules for outlier filtering are as follows: microorganisms with a detection frequency of <60 times are removed if the PG amount is greater than up_bound; microorganisms with strong positive results in the same batch and a PG amount greater than up_bound are removed; and all outliers are removed if the proportion of abnormal samples is <15%. And / or, the selection rule for the cutoff value of the reference library is as follows: when the detection frequency is ≥0.07, the minimum value of the quartile array is taken for extremely weak signals, the third minimum value is taken for slightly weak signals, the second minimum value is taken for normal / slightly strong signals, and the maximum value is taken for extremely strong signals; when the detection frequency is <0.07, the minimum value is taken for extremely weak signals when the detection frequency is >6, the third minimum value is taken for slightly weak / normal / slightly strong signals, and the maximum value is taken for extremely strong signals; the minimum value is taken for all signals when the detection frequency is ≤6. And / or, the dynamic update mechanism includes automatically updating the reference library data according to a preset period, constructing it in combination with wet experimental treatment type, and using a dual threshold of detection frequency and frequency number to determine the cutoff value selection strategy.

8. The standardized method for identifying pathogen metagenomic background bacteria according to any one of claims 1 to 6, characterized in that, The method is applicable to the identification of background microbial levels of bacteria, fungi, viruses, and parasites.

9. A determination system for implementing the standardized method for identifying pathogen metagenomic background bacteria level according to any one of claims 1 to 8, characterized in that, include: The core background microbial community module is used to construct a relative abundance matrix of microorganisms based on metagenomic sequencing NC samples over a period of time. The software calculates the microbial correlation matrix and P-value matrix. After filtering out microbial relationship pairs with P-values ​​less than 0.01, hierarchical clustering is used to select the microorganisms in the largest cluster as the main core background microorganisms. The historical reference library construction module is used to build reference libraries according to four periods: the past 1 month, 2 months, 3 months, and 1 year, and the host-free / non-host-free treatment type. The data structure is formed by detection rate statistics, filtering of strongly positive cross-contamination microorganisms, calculation of standardized PG amount, filtering of outlier box plots, and calculation of statistical indicators. The background bacteria marker module is used to construct artificial NC control specimens. For clinical specimens, the initial marker bacteria are obtained by iteratively filtering colonizing bacteria through standardized PG amount. The initial marker bacteria are then corrected by the FC fold difference index and integrated with empirical background bacteria to form the final marker set. The background bacteria signal intensity assessment module is used to classify signals based on the percentage of the initial marker bacteria from the top 5 genera, the ratio of the initial marker to the Share bacteria, the nucleic acid extraction concentration, and the total number of bacterial / fungal species detected. The background microbial level determination module is used to output the final determination result of background / prominent microorganisms through a dual determination mechanism of FC multiplier index and standardized PG index, combined with the pathogenic microorganism priority determination rule and manual review and labeling rule.

10. A computer-readable storage medium, characterized in that, The computer-readable storage medium stores computer-executable instructions, which, when executed by a processor, are used to implement the standardized method for identifying pathogenic metagenomic background bacteria as described in any one of claims 1 to 8.