Method for identifying target microorganism in organism and application

Through the method of combining Kraken2 and internal standard technology, the problem of interference between background microorganisms and homologous species in second-generation sequencing of metagenomes is solved, which improves the sensitivity and specificity of respiratory pathogen identification and reduces the false positive rate.

CN120126575APending Publication Date: 2025-06-10DAAN GENE CO LTD +2
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202311676894.1
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2023-12-07
Publication Date
2025-06-10

AI Technical Summary

Technical Problem

The prior art is susceptible to interference from background microorganisms and homologous species when using metagenomic second-generation sequencing (mNGS) for respiratory pathogen identification, resulting in high false positive rates and insufficient sensitivity and specificity.

Method used

By using Kraken2 for library construction and annotation, combined with internal standard technology and homogenization treatment, the interference of biological sequences and background contaminated sequences was removed, the number of sequences of the target microorganisms was calculated and the positive judgment value was identified.

Benefits of technology

It effectively reduces the false positive rate, improves the sensitivity and specificity of respiratory pathogen identification, and ensures the accuracy of the detection results.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120126575A_ABST
    Figure CN120126575A_ABST
Patent Text Reader

Abstract

The invention discloses a method for identifying a target microorganism in a living body, which comprises the following steps: by taking a biological sample from the living body as a to-be-detected sample and taking sequencing data in a FastQ format in a sequencing result of the to-be-detected sample as original data, removing a sequence of the living body, performing species identification and homogenization treatment, and deducting the interference of other microorganism sequences and the interference of the NC background sequence from the homogenized data to obtain the sequence number of the target microorganism, judging by combining a positive judgment value, and identifying whether the organism contains the target microorganism or not. By utilizing the method, interference caused by background pollution and other species sequence information in a metagenome next-generation sequencing result can be removed, generation of false positive in a detection result is reduced, and detection sensitivity and specificity are improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the field of biomedical technologies, and in particular, to a method and application for identifying target microorganisms in an organism. Background Art

[0002] Respiratory tract infections are the deadliest infectious diseases in the world. Data from the Institute for Health Metrics and Evaluation's 2019 Global Burden of Disease Study showed that the incidence and case fatality rate of lower respiratory tract infections (LRTIs) ranked 4th among adults and children, and the number of deaths due to respiratory tract infections reached 2.6 million in 2019.

[0003] According to the China Antimicrobial Surveillance Network (CHINET) to evaluate the market capacity of different sample types, the proportion of respiratory system infections in clinical practice is as high as 37.8%. Pathogens detected in relatively large numbers in samples of respiratory system infections include Klebsiella pneumoniae, Acinetobacter baumannii, Pseudomonas aeruginosa, Staphylococcus aureus, Haemophilus influenzae, Streptococcus pneumoniae, Stenotrophomonas maltophilia, Escherichia coli, and Aspergillus fumigatus, etc.

[0004] There is a wide variety of pathogens causing respiratory system infections, and pathogen diagnosis is difficult. Only by making a rapid and accurate judgment of the pathogen and conducting targeted anti-infective treatment according to the type of pathogen can the best clinical prognosis be achieved. Traditional diagnosis and treatment methods mainly include the identification of biochemical indicators and the culture of pathogens, etc. The accuracy of the method operation is insufficient and the cycle is long, making it difficult to provide accurate auxiliary diagnosis for doctors. As a result, most drug treatments are empirical, which not only speeds up the cycle of pathogen resistance but also affects the timely diagnosis of patients.

[0005] Next-generation sequencing technology (NGS), also known as high-throughput or massively parallel sequencing technology, can simultaneously and independently sequence thousands or billions of DNA fragments. The application of NGS in clinical microbiological testing is called metagenomic next-generation sequencing (mNGS). This method does not rely on traditional microbial culture and can unbiasedly extract all nucleic acids in a sample for high-throughput sequencing. By combining bioinformatics analysis, after removing human sequences, it is compared with a pathogen database to obtain the species information of pathogenic microorganisms. The advantage of mNGS in the field of infectious disease diagnosis is that it can detect pathogens that cannot be detected by other conventional methods, including unculturable, rare, or even new pathogenic microorganisms. However, the interpretation of mNGS reports is often subjective at present, and there is a lack of unified clinical criteria for interpreting mNGS, including sequence thresholds, assessment of sensitivity and specificity.

[0006] When using metagenomic next-generation sequencing (mNGS) to identify respiratory pathogens, the resulting mNGS reports are difficult to fully distinguish human sources from respiratory pathogens (Klebsiella pneumoniae, Acinetobacter baumannii, Pseudomonas aeruginosa, Staphylococcus aureus, Haemophilus influenzae, Streptococcus pneumoniae, Stenotrophomonas maltophilia, Escherichia coli, and Aspergillus fumigatus). Moreover, due to the interference of background microorganisms and homologous species, false positives are likely to occur, leading to inaccurate identification of respiratory pathogens.

[0007] Therefore, there is an urgent need for a method suitable for identifying respiratory pathogens that can effectively remove the interference of background microorganisms and homologous species to improve the sensitivity and specificity of respiratory pathogen identification. Summary of the Invention

[0008] The purpose of the present invention is to overcome the above deficiencies of the prior art and provide a method and application for identifying target microorganisms in an organism.

[0009] The first object of the present invention is to provide a method for identifying target microorganisms in an organism.

[0010] The second object of the present invention is to provide a device for identifying microorganisms in an organism.

[0011] The third object of the present invention is to provide the application of the above device in identifying microorganisms in an organism.

[0012] The fourth object of the present invention is to provide the application of the above device in the preparation of products for guiding the medication of patients with respiratory infections.

[0013] To achieve the above objects, the present invention is realized through the following solutions:

[0014] A method for identifying target microorganisms in an organism, using a biological sample derived from the organism as a test sample, comprising the following steps:

[0015] S1. Based on the microbial sequences in the NT database and GenBank database, use Kraken2 to build a library to obtain a species classification database DB tax ;

[0016] S2. For each microbial reference sequence S target respectively use wgsim to simulate at a data volume of not less than 50× to obtain each microbial simulated sequence Seq s , and use the species classification database DB tax obtained in step S1 in combination with Kraken2 to annotate each microbial simulated sequence Seq s respectively, and obtain the number of all annotated sequences Read in each microbial simulated sequence Seq s ​x_x模拟 , and the number of sequences Read annotated as the target microorganism in each microbial simulation sequence Seq s y_x模拟 y_x模拟 , and calculate the weight b of the target microorganism sequences in each microbial simulation sequence Seq s k k , to obtain a homologous coefficient database; wherein, b k = Read y_x模拟 / Read x_x模拟 ;

[0017] S3. Using the cells of the organism without any microorganisms as a negative control sample, mixing the internal standard buffer and the sample to be detected at a volume ratio of 1:1 - 2 to obtain the sample to be detected after adding the internal standard; mixing the internal standard buffer and the negative control sample at a volume ratio of 1:1 - 2 to obtain the negative control sample after adding the internal standard;

[0018] Sequencing the sample to be detected after adding the internal standard and the negative control sample after adding the internal standard respectively to obtain the sequencing data of the sample to be detected after adding the internal standard in FastQ format and the sequencing data of the negative control sample after adding the internal standard in FastQ format, as the original data 1 and the original data 2;

[0019]

[0019]

[0020] wherein, each milliliter of the internal standard buffer contains 0.2 - 0.4 ng of the internal standard sequence, and the internal standard sequence is a sequence with a base length of not less than 200 bp and a similarity to the target microorganism reference sequence < 90%;

[0021]

[0021] spike-in_sample spike-in_NC spike-in_NC ;

[0022] And respectively removing the sequences aligned in the pre - processed data of the sample to be detected and the pre - processed data of the negative control sample to obtain the pre - processed data of the sample to be detected after removing the internal standard and the pre - processed data of the negative control sample after removing the internal standard;

[0023] S5. Using the species classification database DB obtained in step S1 taxAnnotate the preprocessed data of the sample to be detected obtained in step S4 after removing the internal standard by combining with Kraken2 to obtain the sequence numbers of each microorganism in the preprocessed data of the sample to be detected after removing the internal standard, and perform normalization processing based on the raw_data in the original data 1 shown in step S3 to obtain the sequence numbers of each microorganism in the sample to be detected after normalization. Among them, the sequence number of the target microorganism in the sample to be detected after normalization is denoted as y, and the sequence numbers of non-target microorganisms are denoted as x 1 ~x k , and the number of types of non-target microorganisms is k;

[0024] Use the species classification database DB obtained in step S1 tax Combine Kraken2 to annotate the preprocessed data of the negative control sample obtained in step S4 after removing the internal standard to obtain the sequence number of the target microorganism in the preprocessed data of the negative control sample after removing the internal standard, denoted as the background contamination sequence number Reads pollutants_NC ;

[0025] S6. Calculate the sequence number b of the target microorganism in the sample to be detected after deducting non-target microorganisms according to formula I 0 ;

[0026] Formula I: b 0 =y-(b 1 x 1 +b 2 x 2 +...+b k x k +e);

[0027] Among them, y is the sequence number y of the target microorganism in the sample to be detected after normalization obtained in step S5, x 1 ~x k are the sequence numbers of non-target microorganisms obtained in step S5, b 1 ~b k are the weights of the target microorganism sequences in the simulated sequences of each microorganism obtained in step S2, b k x k represents the sequence number of the target microorganism contained in the sequence numbers of non-target microorganisms obtained in step S5;

[0028] Select the sequence numbers of the microorganisms belonging to the same genus as the target microorganism among the sequence numbers of each microorganism in the sample to be detected after normalization obtained in step S5, and use 20% of the maximum value of the selected sequence numbers as the threshold to calculate the sum of the sequence numbers less than the threshold in the selected sequence numbers, denoted as e;

[0029] S7. Use the internal standard sequence number Reads in the preprocessed data of the sample to be detected obtained in step S4 spike-in_sample and the internal standard sequence number Reads in the negative control samplespike-in_NC The number of background contamination sequences Reads obtained in step S5 pollutants_NC Combined with formula II, calculate the number of background contamination sequences Reads of the sample to be detected pollutants_sample Combined with formula III, calculate the number of sequences Reads of the target microorganism after quality control last_sample ;

[0030] Formula II: Reads spike-in_sample / Reads spike-in_NC =Reads pollutants_sample / Reads pollutants_NC ;

[0031] Formula III: Reads last_sample =b 0 -Reads pollutants_sample ;

[0032] S8. Compare the number of sequences Reads of the target microorganism after quality control obtained in step S7 with the positive judgment value. If Reads last_sample ≥positive judgment value, the host sample contains the target microorganism; if Reads last_sample <positive judgment value, the host sample does not contain the target microorganism; last_sample <positive judgment value, the host sample does not contain the target microorganism;

[0033] The positive judgment value is calculated by combining the positive samples and negative samples, calculating the number of sequences Reads of the target microorganism after quality control in the positive samples and negative samples according to the methods shown in steps S1-S7, and combining the Youden index.

[0034] Preferably, the sample to be detected is a biological sample from a patient with respiratory tract infection.

[0035] More preferably, the biological sample is digestive juice, tissue fluid and / or tissue.

[0036] More preferably, the sample to be detected is bronchoalveolar lavage fluid from a patient with respiratory tract infection.

[0037] Preferably, when using Kraken2 to construct a library in step S1, the k-mer length is 35bp.

[0038] Preferably, the species classification database DB tax in step S1 contains microbial kmer sequences.

[0039] Preferably, the target microorganism in step S2 is Klebsiella pneumoniae, Acinetobacter baumannii, Pseudomonas aeruginosa, Staphylococcus aureus, Haemophilus influenzae, Streptococcus pneumoniae, Stenotrophomonas maltophilia, Escherichia coli and / or Aspergillus fumigatus.

[0040] More preferably, the target microorganism in step S2 is Klebsiella pneumoniae, and the positive judgment value in step S8 is 112.5;

[0041] The target microorganism in step S2 is Acinetobacter baumannii, and the positive judgment value in step S8 is 78;

[0042] The target microorganism in step S2 is Pseudomonas aeruginosa, and the positive judgment value in step S8 is 86.5;

[0043] The target microorganism in step S2 is Staphylococcus aureus, and the positive judgment value in step S8 is 31;

[0044] The target microorganism in step S2 is Haemophilus influenzae, and the positive judgment value in step S8 is 19.5;

[0045] The target microorganism in step S2 is Streptococcus pneumoniae, and the positive judgment value in step S8 is 18.5;

[0046] The microorganism in step S2 is Stenotrophomonas maltophilia, and the positive judgment value in step S8 is 120;

[0047] The microorganism in step S2 is Escherichia coli, and the positive judgment value in step S8 is 6.5;

[0048] The microorganism in step S2 is Aspergillus fumigatus, and the positive judgment value in step S8 is 29.5.

[0049] Preferably, the reference sequence S of each microorganism in step S2 target is from the NCBI database.

[0050] Preferably, in step S3, the internal standard buffer is mixed with the sample to be detected at a volume ratio of 1:1, and the internal standard buffer is mixed with the negative control sample at a volume ratio of 1:1.

[0051] More preferably, each milliliter of the internal standard buffer contains 0.3 ng of the internal standard sequence.

[0052] Preferably, the sequencing in step S3 is metagenomic next-generation sequencing.

[0053] Preferably, the sequences of the organisms in the raw data 1 and raw data 2 obtained in step S3 are removed using the snap-aligner software in step S4 specifically as follows: The raw data 1 and raw data 2 are respectively aligned to the reference genome of the organism using the snap-aligner software with version 1.0.3 based on preset parameters, and the organism sequences in the raw data 1 and raw data 2 are respectively removed.

[0054] More preferably, removing the host sequence in the original data is to remove the sequences with a Flag value of 4 in the bam file of the alignment result.

[0055] Preferably, the alignment in step S4 is: using the BWA software with the version number Version: 0.7.17-r1198-dirty for alignment based on preset parameters.

[0056] Among them, the number of sequences with a Flag not equal to 4 in the bwa.bam file in the BWA alignment result is used as the number of aligned sequences.

[0057] Preferably, the positive judgment value in step S8 is calculated by combining the positive samples and negative samples and calculating the number of sequences of the target microorganism after quality control in the positive samples and negative samples using the method shown in steps S1 to S7, and the positive judgment value corresponding to the maximum Youden index is used as the positive judgment value of the target microorganism.

[0058] The present invention also claims a device for identifying microorganisms in an organism, the device comprising a data acquisition component, an identification component, and a result output component;

[0059] The data acquisition component is used to acquire the sequencing data 1 in FastQ format of the sample to be detected and the sequencing data 2 in FastQ format of the negative quality control sample;

[0060] The sample to be detected is a biological sample derived from the organism, and the negative quality control sample is a cell of the organism without any microorganisms;

[0061] The identification component uses the sequencing data 1 and sequencing data 2 obtained by the data acquisition component as input data, and executes the method as described in any of the above, to obtain the identification result of microorganisms in the organism;

[0062] The result output component is used to output the identification result of microorganisms in the organism obtained by the identification component.

[0063] The present invention also claims the application of the above device in identifying microorganisms in an organism.

[0064] Preferably, the sample to be detected is a biological sample derived from a patient with a respiratory tract infection.

[0065] More preferably, the biological sample is digestive juice, tissue fluid, and / or tissue.

[0066] Further preferably, the sample to be detected is bronchoalveolar lavage fluid from a patient with a respiratory tract infection.

[0067] Preferably, the microorganism is Klebsiella pneumoniae, Acinetobacter baumannii, Pseudomonas aeruginosa, Staphylococcus aureus, Haemophilus influenzae, Streptococcus pneumoniae, Stenotrophomonas maltophilia, Escherichia coli and / or Aspergillus fumigatus.

[0068] The present invention also claims the application of the above device in the preparation of a product for guiding the medication of respiratory tract infection patients, and the guidance for the medication of respiratory tract infection patients is to guide respiratory tract infection patients to use medication according to respiratory pathogens;

[0069] The respiratory pathogens are Klebsiella pneumoniae, Acinetobacter baumannii, Pseudomonas aeruginosa, Staphylococcus aureus, Haemophilus influenzae, Streptococcus pneumoniae, Stenotrophomonas maltophilia, Escherichia coli and / or Aspergillus fumigatus.

[0070] Respiratory tract infection patients can use the above device to identify the types of respiratory pathogens in their bronchoalveolar lavage fluid and carry out targeted medication based on the types of respiratory pathogens.

[0071] Compared with the prior art, the present invention has the following beneficial effects:

[0072] The present invention provides a method for identifying a target microorganism in an organism. The method uses a biological sample derived from the organism as a sample to be detected, uses the sequencing data in FastQ format in the sequencing results of the sample to be detected as the original data, removes the organism sequence and then conducts species identification and normalization processing. After deducting the interference of the sequences of other microorganisms and the interference of the NC background sequence from the data after normalization processing, the number of sequences of the target microorganism is obtained and combined with a positive judgment value for judgment to identify whether the organism contains the target microorganism. Using this method can remove the interference of background contamination and the sequence information of other species in the metagenomic next-generation sequencing results, reduce the generation of false positives in the detection results, and improve the sensitivity and specificity of the detection. BRIEF DESCRIPTION OF THE DRAWINGS

[0073] Figure 1 It is a flowchart corresponding to the method for identifying 9 respiratory pathogens in respiratory tract infection patients shown in Example 1. DETAILED DESCRIPTION OF THE EMBODIMENTS

[0074] The following further elaborates the present invention in detail with reference to the accompanying drawings of the specification and specific embodiments. The embodiments are only used to explain the present invention and are not used to limit the scope of the present invention. The test methods used in the following embodiments are all conventional methods unless otherwise specified; the materials, reagents, etc. used are all reagents and materials that can be obtained from commercial channels unless otherwise specified.

[0075] Example 1 A method for identifying 9 respiratory pathogens in respiratory tract infection patients

[0076] Taking the alveolar lavage fluid of patients with respiratory system infections as the sample to be tested, a method for identifying 9 respiratory pathogens was established. The flow chart is as Figure 1 shown; the 9 respiratory pathogens include Klebsiella pneumoniae, Acinetobacter baumannii, Pseudomonas aeruginosa, Staphylococcus aureus, Haemophilus influenzae, Streptococcus pneumoniae, Stenotrophomonas maltophilia, Escherichia coli, and Aspergillus fumigatus.

[0077] Taking the detection of Klebsiella pneumoniae as an example, the alveolar lavage fluid of M patients with respiratory system infections with known positive and negative results of Klebsiella pneumoniae was selected as the test sample (the actual positive and negative results of Klebsiella pneumoniae in the M test samples were: Y positive sample numbers and Z negative sample numbers, M = Y + Z, positive samples were samples containing Klebsiella pneumoniae, and negative samples were samples not containing Klebsiella pneumoniae), as follows:

[0078] 1. Establish a species classification database

[0079] Download the NT database (i.e., the NCBI nucleic acid sequence database) sequences from NCBI. Using the Karken2 software, with a k-mer length of 35 bp, a database was built based on the microbial sequences in the NT database and GenBank database to obtain the species classification database DB tax ; and confirm the information of Klebsiella pneumoniae in the species classification database DB tax in.

[0080] 2. Establish a homology coefficient database

[0081] Based on the microbial detection information (mNGS) of 1000 samples (alveolar lavage fluid of respiratory infection patients), the top 10 microorganisms with the highest positive intensity in the microbial detection information of each sample and the microorganisms that appeared in more than 100 samples were summarized as the common microbial set X in the respiratory tract.

[0082] Use wgsim (version: 1.15 - 9 - g4be6986) to simulate the reference sequence S target of each microorganism x in the common microbial set X in the respiratory tract with a data volume of 50× to obtain the simulated sample sequence Seq s of each microorganism x; among them, the reference sequence S target of each microorganism x in the microbial set X in the respiratory tract

[0083] Utilize the species classification database DB tax constructed in step 1 and combine it with the Karken2 software to annotate Seq s to obtain the number of all annotated sequences Read x_x模拟 in Seq sThe number of sequences Read annotated as Klebsiella pneumoniae in y_x模拟 , and calculate the proportion b of the number of Klebsiella pneumoniae sequences in each microorganism x according to formula I k , obtain the weight of the number of Klebsiella pneumoniae sequences in each microorganism x, and obtain a homologous coefficient database; Formula I: b k = Read y_x模拟 / Read x_x Simulation.

[0084] 3. Sample addition of internal standard and sequencing

[0085] Add internal standard buffer to the test sample and negative control sample (NC, hela cells) respectively according to a volume ratio of 1:1 to obtain the test sample after adding internal standard and the negative control sample after adding internal standard; Each milliliter of the internal standard buffer contains 0.3 ng of the internal standard sequence, and the nucleotide sequence of the internal standard sequence is shown in SEQ ID NO: 1.

[0086] Perform metagenomic next-generation sequencing on the test sample after adding internal standard and the negative control sample after adding internal standard respectively, and obtain the sequencing data of the test sample after adding internal standard (raw data 1) in FastQ format in the metagenomic next-generation sequencing results and the sequencing data of the negative control sample after adding internal standard (raw data 2) in FastQ format in the metagenomic next-generation sequencing results.

[0087] 4. Data preprocessing

[0088] Use snap-aligner (version: 1.0.3) based on preset parameters to align the raw data 1 and raw data 2 obtained in step S3 to the human hg38 genome respectively, remove the human-derived sequences aligned in the raw data, and obtain the preprocessed data of the test sample and the preprocessed data of the negative control sample respectively;

[0089] Use BWA software (version number: Version: 0.7.17-r1198-dirty) based on preset parameters to perform BWA alignment on the preprocessed data of the test sample and the preprocessed data of the negative control sample with the internal standard sequence whose nucleotide sequence is shown in SEQ ID NO: 1 in step 3 respectively, and obtain the number of sequences aligned in the preprocessed data of the test sample and the number of sequences aligned in the preprocessed data of the negative control sample respectively, as the number of internal standard sequence Reads in the preprocessed data of the test sample spike-in_sample and the number of internal standard sequence Reads in the preprocessed data of the negative control sample spike-in_NC ;

[0090] And remove the aligned sequences in the preprocessed data of the test sample and the preprocessed data of the negative control sample respectively to obtain the preprocessed data of the test sample after removing the internal standard and the preprocessed data of the negative control sample after removing the internal standard;

[0091] (The number of sequences with a Flag not equal to 4 in the bwa.bam file of the BWA alignment results of the detection sample preprocessing data and the negative control sample preprocessing data is recorded as the number of aligned sequences)

[0092] 5. Species annotation and normalization

[0093] Use the species classification database DB constructed in step 1 tax Combine Kraken2 to annotate the preprocessing data of the detection sample after removing internal standards in step 4, obtain the number of sequences of each microorganism in the preprocessing data of the detection sample after removing internal standards, and divide the number of sequences of each microorganism by raw_data in the original data 1 and then multiply by 10 7 , to obtain the normalized detection sample and the number of sequences of each microorganism after 10M normalization; where the number of sequences of Klebsiella pneumoniae in the normalized detection sample is denoted as y, and the number of sequences of each microorganism other than Klebsiella pneumoniae is x 1 ~x k , and the number of non-target microorganism species is k

[0094] Use the species classification database DB constructed in step 1 tax Combine Kraken2 to annotate the preprocessing data of the negative control sample after removing internal standards in step 4, obtain the number of sequences of Klebsiella pneumoniae in the preprocessing data of the negative control sample after removing internal standards, and denote it as the NC background contamination sequence number Reads pollutants_NC .

[0095] 6. Species carryover deduction

[0096] Calculate the number of sequences of Klebsiella pneumoniae in the detection sample after species deduction according to formula II 0 ;

[0097] Formula II: b 0 =y-(b 1 x 1 +b 2 x 2 +...+b k x k +e);

[0098] Among them, y is the number of sequences of Klebsiella pneumoniae in the normalized detection sample obtained in step 5; x 1 ~x k are the numbers of sequences of each microorganism other than Klebsiella pneumoniae obtained in step 5; b 1 ~b k are the weights of the number of sequences of Klebsiella pneumoniae in each microorganism x obtained in step S2, b k xk Denotes the number of Klebsiella pneumoniae sequences contained in the number of sequences of each microorganism other than Klebsiella pneumoniae obtained in step 5;

[0099] Select the number of sequences of microorganisms belonging to the same species as Klebsiella pneumoniae among the number of sequences after 10M normalization of each microorganism obtained in step S5, and use 20% of the maximum value of the selected number of sequences as the threshold, and calculate the sum of the number of sequences less than the threshold in the selected number of sequences, denoted as e.

[0100] 7. NC background interference deduction

[0101] Use the number of internal standard sequences Reads in the preprocessed data of the test sample obtained in step 4 spike-in_sample and the number of internal standard sequences Reads in the preprocessed data of the negative control sample spike-in_NC , the number of NC background contamination sequences Reads obtained in step S5 pollutants_NC , and calculate the number of background contamination sequences Reads of the test sample by combining formula III pollutants_sample ;

[0102] Formula III: Reads spike-in_sample / Reads spike-in_NC =Reads pollutants_sample / Reads pollutants_NC ;

[0103] According to formula IV, subtract the number of background contamination sequences Reads of the test sample from the number of Klebsiella pneumoniae sequences b 0 in the test sample after species deduction obtained in step 6 pollutants_sample , to obtain the number of sequences Reads of Klebsiella pneumoniae after quality control in the test sample last_sample ;

[0104] Formula IV: Reads last_sample =b 0 -Reads pollutants_sample .

[0105] 8. Setting of positive threshold

[0106] Set different positive thresholds (0 ≤ positive threshold ≤ Reads last_sample ), compare the Reads last_sample obtained in step 7 with the positive threshold to obtain the positive and negative test results of Klebsiella pneumoniae in the test sample; among them, Reads last_sample < positive threshold, denoted as a negative sample; Reads last_sample ≥ positive threshold, denoted as a positive sample.

[0107] Compare the detected positive and negative results of Klebsiella pneumoniae in the test sample with the actual positive and negative results of Klebsiella pneumoniae in the test sample to obtain true positive (TP), false positive (FP), true negative (TN), and false negative (FN) results, and calculate sensitivity, specificity, and Youden index.

[0108] Among them, sensitivity = TP / (TP + FN); specificity = TN / (TN + FP); Youden index = sensitivity + specificity - 1.

[0109] Calculate the sensitivity, specificity, and Youden index corresponding to each positive threshold, and select the positive threshold corresponding to the maximum Youden index as the positive judgment value of Klebsiella pneumoniae; the positive judgment value of Klebsiella pneumoniae is 112.5.

[0110] 9. Judgment method

[0111] Compare the Reads obtained in step 7 last_sample with the positive judgment value of Klebsiella pneumoniae obtained in step 8. If Reads last_sample ≥ 112.5, it is judged as positive, indicating that the test sample contains Klebsiella pneumoniae; if Reads last_sample < 112.5, it is judged as negative, indicating that the test sample does not contain Klebsiella pneumoniae.

[0112] Replace Klebsiella pneumoniae with the remaining 8 respiratory pathogens (Acinetobacter baumannii, Pseudomonas aeruginosa, Staphylococcus aureus, Haemophilus influenzae, Streptococcus pneumoniae, Stenotrophomonas maltophilia, Escherichia coli, and Aspergillus fumigatus) respectively, and according to the above method (steps 1 to 8) and combined with the actual positive and negative results of respiratory pathogens in the test sample, obtain the positive judgment values of the 8 respiratory pathogens respectively.

[0113] Among them, the positive judgment value of Acinetobacter baumannii is 78; the positive judgment value of Pseudomonas aeruginosa is 86.5; the positive judgment value of Staphylococcus aureus is 31; the positive judgment value of Haemophilus influenzae is 19.5; the positive judgment value of Streptococcus pneumoniae is 18.5; the positive judgment value of Stenotrophomonas maltophilia is 120; the positive judgment value of Escherichia coli is 6.5; the positive judgment value of Aspergillus fumigatus is 29.5.

[0114] According to the method shown in step 9 and combined with the positive judgment values of the 8 respiratory pathogens, judge whether the test sample contains 8 respiratory pathogens (Acinetobacter baumannii, Pseudomonas aeruginosa, Staphylococcus aureus, Haemophilus influenzae, Streptococcus pneumoniae, Stenotrophomonas maltophilia, Escherichia coli, and Aspergillus fumigatus).

[0115] Example 2 Verification of a method for identifying 9 respiratory pathogens in patients with respiratory tract infections

[0116] 1. Experimental methods

[0117] The bronchoalveolar lavage fluid of 389 patients with respiratory system infections was selected as the verification sample 1 for identifying Acinetobacter baumannii, including 32 positive samples and 357 negative samples;

[0118] The bronchoalveolar lavage fluid of 552 patients with respiratory system infections was selected as the verification sample 2 for identifying Aspergillus fumigatus, including 27 positive samples and 525 negative samples;

[0119] The bronchoalveolar lavage fluid of 281 patients with respiratory system infections was selected as the verification sample 3 for identifying Haemophilus influenzae, including 29 positive samples and 252 negative samples;

[0120] The bronchoalveolar lavage fluid of 523 patients with respiratory infections was selected as the verification sample 4 for identifying Klebsiella pneumoniae, including 35 positive samples and 488 negative samples;

[0121] The bronchoalveolar lavage fluid of 383 patients with respiratory infections was selected as the verification sample 5 for identifying Pseudomonas aeruginosa, including 33 positive samples and 350 negative samples;

[0122] The bronchoalveolar lavage fluid of 401 patients with respiratory infections was selected as the verification sample 6 for identifying Staphylococcus aureus, including 21 positive samples and 380 negative samples;

[0123] The bronchoalveolar lavage fluid of 390 patients with respiratory infections was selected as the verification sample 7 for identifying Stenotrophomonas maltophilia, including 33 positive samples and 357 negative samples;

[0124] The bronchoalveolar lavage fluid of 469 patients with respiratory infections was selected as the verification sample 8 for identifying Streptococcus pneumoniae, including 22 positive samples and 447 negative samples;

[0125] The bronchoalveolar lavage fluid of 278 patients with respiratory infections was selected as the verification sample 9 for identifying Escherichia coli, including 31 positive samples and 247 negative samples.

[0126] Specifically, the experimental group was as follows: The verification samples 1 to 9 were respectively identified according to the method shown in Example 1, and combined with the positive judgment value obtained in step 8 of Example 1, the positive and negative results of Acinetobacter baumannii in verification sample 1, the positive and negative results of Aspergillus fumigatus in verification sample 2, the positive and negative results of Haemophilus influenzae in verification sample 3, the positive and negative results of Klebsiella pneumoniae in verification sample 4, the positive and negative results of Pseudomonas aeruginosa in verification sample 5, the positive and negative results of Staphylococcus aureus in verification sample 6, the positive and negative results of Stenotrophomonas maltophilia in verification sample 7, the positive and negative results of Streptococcus pneumoniae in verification sample 8, and the positive and negative results of Escherichia coli in verification sample 9 were obtained.

[0127] Compare the positive and negative results of verification samples 1 to 9 with their actual positive and negative results, obtain the true positive, false positive, true negative, and false negative results of each verification sample after identification according to the method shown in Example 1, and calculate the sensitivity and specificity.

[0128] The difference between the control group and the experimental group is that: the identification method of the control group does not include "Steps 6 and 7" in the method shown in Example 1, and the remaining processing steps are the same. Identify verification samples 1 to 9 to obtain the positive threshold, positive and negative results, sensitivity, and specificity results of the control group.

[0129] 2. Experimental results

[0130] The identification results of the experimental group and the control group for verification samples 1 to 9 are shown in Table 2, and the sensitivity and specificity results of the experimental group and the control group are shown in Table 3.

[0131] Table 2 Identification results of the experimental group and the control group for verification samples 1 to 9

[0132]

[0133] Table 3 Sensitivity and specificity results of the experimental group and the control group

[0134]

[0135]

[0136] The results show that: when identifying 9 respiratory pathogens (Klebsiella pneumoniae, Acinetobacter baumannii, Pseudomonas aeruginosa, Staphylococcus aureus, Haemophilus influenzae, Streptococcus pneumoniae, Stenotrophomonas maltophilia, Escherichia coli, and Aspergillus fumigatus) according to the method shown in Example 1 (experimental group), compared with identifying 9 respiratory pathogens according to the identification method that does not perform "Steps 6 and 7" in the method shown in Example 1 (control group), the sensitivity and specificity for the identification of 9 respiratory pathogens are both significantly improved, and the number of samples with false positive (FP) identification results is also significantly reduced.

[0137] The results indicate that: the method shown in Example 1 can effectively remove the interference of background microorganisms and homologous species during the identification of 9 respiratory pathogens in patients with respiratory tract infections, effectively reduce the detection of false positive samples, and has excellent sensitivity and specificity.

[0138] Example 3 An identification system for identifying 9 respiratory pathogens in patients with respiratory tract infections

[0139] An identification system for identifying 9 respiratory pathogens in patients with respiratory tract infections includes a data acquisition component, an identification component, and a result output component;

[0140] The data acquisition component is used to obtain the FastQ file in the off-machine data after sequencing the bronchoalveolar lavage fluid of patients with respiratory tract infections and the FastQ file 1 in the off-machine data of negative control sample sequencing; the identification component uses the FastQ file and the FastQ file 1 obtained by the data acquisition component as input data, and implements the identification method shown in Example 1 to obtain the identification results of 9 respiratory pathogens; the result output component is used to output the identification results of 9 respiratory pathogens obtained by the identification component.

[0141] Finally, it should be noted that the above embodiments are only used to illustrate the technical solutions of the present invention rather than limit the protection scope of the present invention. For those of ordinary skill in the art, other different forms of changes or modifications can be made based on the above description and ideas. It is not necessary and impossible to enumerate all implementation manners here. Any modifications, equivalent replacements, and improvements made within the spirit and principle of the present invention shall be included in the protection scope of the claims of the present invention.

Claims

1. A method for identifying a target microorganism in an organism, characterized in that, using a biological sample derived from the organism as a sample to be detected, comprising the following steps: S1. Based on the microbial sequences in the NT database and the GenBank database, use Kraken2 to build a database to obtain the species classification database DB tax ; S2. Reference sequences S of each microorganism target Use wgsim to simulate at least 50× data volume for each microorganism to obtain the simulated sequences Seq of each microorganism s , and use the species classification database DB obtained in step S1 tax Combine Kraken2 to annotate each simulated microorganism sequence Seq s respectively, to obtain the number of all annotated sequences Read in each simulated microorganism sequence Seq s , and the number of sequences Read annotated as the target microorganism in each simulated microorganism sequence Seq x_x模拟 , and calculate the weight b of the target microorganism sequences in each simulated microorganism sequence Seq s , to obtain the homologous coefficient database; where, b y_x模拟 =Read s / Read k ; k ; y_x模拟 ; x_x模拟 ; S3. Using the cells of the organism without any microorganisms as a negative control sample, mixing the internal standard buffer with the sample to be detected at a volume ratio of 1:1 to 2 to obtain the sample to be detected after adding the internal standard; mixing the internal standard buffer with the negative control sample at a volume ratio of 1:1 to 2 to obtain the negative control sample after adding the internal standard; Sequencing the sample to be detected after adding the internal standard and the negative control sample after adding the internal standard respectively to obtain the sequencing data of the sample to be detected after adding the internal standard in FastQ format and the sequencing data of the negative control sample after adding the internal standard in FastQ format, as the original data 1 and the original data 2; Wherein, each milliliter of the internal standard buffer contains 0.2 - 0.4 ng of the internal standard sequence, and the internal standard sequence is a sequence with a base length of not less than 200 bp and a similarity to the target microorganism reference sequence < 90%; S4. Using the snap-aligner software to remove the sequences of the organism in the original data 1 and the original data 2 obtained in step S3 to obtain the preprocessed data of the sample to be detected and the preprocessed data of the negative control sample respectively; Compare the preprocessed data of the sample to be detected and the preprocessed data of the negative control sample with the internal standard sequence in step S3 respectively, and obtain the number of sequences aligned in the preprocessed data of the sample to be detected and the number of sequences aligned in the preprocessed data of the negative control sample respectively, as the number of internal standard sequence Reads in the preprocessed data of the sample to be detected spike-in_sample and the number of internal standard sequence Reads in the preprocessed data of the negative control sample spike-in_NC ; And removing the aligned sequences in the preprocessed data of the sample to be detected and the preprocessed data of the negative control sample respectively to obtain the preprocessed data of the sample to be detected after removing the internal standard and the preprocessed data of the negative control sample after removing the internal standard; S5. Use the species classification database DB obtained in step S1 tax Combine with Kraken2 to annotate the preprocessed data of the sample to be detected after removing internal standards obtained in step S4, obtain the sequence numbers of each microorganism in the preprocessed data of the sample to be detected after removing internal standards, and perform normalization processing based on the raw_data in the original data 1 shown in step S3 to obtain the sequence numbers of each microorganism in the sample to be detected after normalization. Among them, the sequence number of the target microorganism in the sample to be detected after normalization is denoted as y, and the sequence numbers of non-target microorganisms are denoted as x 1 ~x k , and the number of species of non-target microorganisms is k; The species classification database DB obtained in step S1 tax Combine Kraken2 to annotate the preprocessed data of the negative control sample after removing internal standards obtained in step S4, and obtain the number of sequences of the target microorganism in the preprocessed data of the negative control sample after removing internal standards, denoted as the background contamination sequence number Reads pollutants_NC ; S6. Calculate the number of sequences b of the target microorganism in the sample to be detected after deducting the non-target microorganism according to Formula I 0 ; Formula I: b 0 = y - (b 1 x 1 + b 2 x 2 +... + b k x k + e); Among them, y is the number of sequences of the target microorganism in the homogenized sample to be detected obtained in step S5, and x 1 ~x k is the number of sequences of non-target microorganisms obtained in step S5, and b 1 ~b k is the weight of the target microorganism sequence in each microorganism simulation sequence obtained in step S2, and b k x k represents the number of sequences of the target microorganism contained in the number of sequences of non-target microorganisms obtained in step S5; Select the number of sequences of the microorganisms belonging to the same genus as the target microorganism among the number of sequences of each microorganism in the normalized sample to be detected obtained in step S5, and use 20% of the maximum value of the selected number of sequences as the threshold, and calculate the sum of the number of sequences less than the threshold in the selected number of sequences, denoted as e; S7. Use the number of internal standard sequence Reads in the preprocessed data of the sample to be detected obtained in step S4 spike-in_sample and the number of internal standard sequence Reads in the negative control sample spike-in_NC , the number of background contamination sequence Reads obtained in step S5 pollutants_NC , and calculate the number of background contamination sequence Reads of the sample to be detected by combining formula II pollutants_sample , and calculate the number of sequences Reads of the target microorganism after quality control by combining formula III last_sample ; Formula II: Reads spike-in_sample / Reads spike-in_NC =Reads pollutants_sample / Reads pollutants_NC ; Formula III: Reads last_sample = b 0 - Reads pollutants_sample ; S8. Compare the number of reads of the target microorganism after quality control obtained in step S7, i.e., Reads last_sample with the positive judgment value. If Reads last_sample ≥ the positive judgment value, it indicates that the target microorganism is present in the host sample; if Reads last_sample < the positive judgment value, it indicates that the target microorganism is not present in the host sample; The positive judgment value is obtained by calculating the number of sequences of the target microorganism after quality control in the positive sample and the negative sample by combining the positive sample and the negative sample according to the method shown in steps S1 - S7, and combining the Youden index.

2. The method according to claim 1, characterized in that, The sample to be detected is a biological sample derived from a patient with a respiratory tract infection.

3. The method according to claim 1, characterized in that, The target microorganism in step S2 is Klebsiella pneumoniae, Acinetobacter baumannii, Pseudomonas aeruginosa, Staphylococcus aureus, Haemophilus influenzae, Streptococcus pneumoniae, Stenotrophomonas maltophilia, Escherichia coli and / or Aspergillus fumigatus.

4. The method according to claim 3, characterized in that, The target microorganism in step S2 is Klebsiella pneumoniae, and the positive judgment value in step S8 is 112.5; The target microorganism in step S2 is Acinetobacter baumannii, and the positive judgment value in step S8 is 78; The target microorganism in step S2 is Pseudomonas aeruginosa, and the positive judgment value in step S8 is 86.5; The target microorganism in step S2 is Staphylococcus aureus, and the positive judgment value in step S8 is 31; The target microorganism in step S2 is Haemophilus influenzae, and the positive judgment value in step S8 is 19.5; The target microorganism described in step S2 is Streptococcus pneumoniae, and the positive judgment value described in step S8 is 18.5; The microorganism described in step S2 is Stenotrophomonas maltophilia, and the positive judgment value described in step S8 is 120; The microorganism described in step S2 is Escherichia coli, and the positive judgment value described in step S8 is 6.5; The microorganism described in step S2 is Aspergillus fumigatus, and the positive judgment value described in step S8 is 29.

5.

5. The method according to claim 1, wherein, the sequencing described in step S3 is metagenomic next-generation sequencing.

6. The method according to claim 1, wherein, the alignment described in step S4 is: using the BWA software with version number Version:0.7.17-r1198-dirty to perform alignment based on preset parameters.

7. A device for identifying microorganisms in an organism, wherein, the device includes a data acquisition component, an identification component, and a result output component; the data acquisition component is used to acquire the sequencing data 1 in FastQ format of the sample to be detected and the sequencing data 2 in FastQ format of the negative control sample; the sample to be detected is a biological sample derived from the organism, and the negative control sample is a cell of the organism without any microorganisms; the identification component uses the sequencing data 1 and sequencing data 2 acquired by the data acquisition component as input data, and executes the method described in any one of claims 1 to 6 to obtain the identification result of microorganisms in the organism; the result output component is used to output the identification result of microorganisms in the organism obtained by the identification component.

8. Use of the device according to claim 7 in identifying microorganisms in an organism.

9. The use according to claim 8, wherein, the microorganisms are Klebsiella pneumoniae, Acinetobacter baumannii, Pseudomonas aeruginosa, Staphylococcus aureus, Haemophilus influenzae, Streptococcus pneumoniae, Stenotrophomonas maltophilia, Escherichia coli, and / or Aspergillus fumigatus.

10. Use of the device according to claim 7 in the preparation of a product for guiding the medication of patients with respiratory tract infections, wherein guiding the medication of patients with respiratory tract infections is to guide patients with respiratory tract infections to use medications according to respiratory tract pathogens; the respiratory tract pathogens are Klebsiella pneumoniae, Acinetobacter baumannii, Pseudomonas aeruginosa, Staphylococcus aureus, Haemophilus influenzae, Streptococcus pneumoniae, Stenotrophomonas maltophilia, Escherichia coli, and / or Aspergillus fumigatus.