A system and method for predicting the sensitivity of Klebsiella pneumoniae to ertapenem
Through the Exp(-k) power value calculation method and second-generation high-throughput sequencing technology, the sensitivity of Klebsiella pneumoniae to ertapenem is quickly and accurately predicted, and the problem of time-consuming and cost-effectiveness of traditional detection methods is solved, and simple and stable identification of drug sensitivity is achieved, which assists in the rational use of drugs in clinical practice.
Patent Information
- Application Number
- CN202310131972.3
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2023-02-06
- Publication Date
- 2025-08-22
- Estimated Expiration
- 2043-02-06
AI Technical Summary
The prior art is difficult to quickly and accurately predict the sensitivity of Klebsiella pneumoniae to ertapenem. The traditional detection methods are time-consuming and costly, and cannot meet the rapid diagnosis and treatment needs of clinical severe infections.
The Exp(-k) power value calculation method was used to calculate the copy number of KPC-1, marA, nalC, sul1, CTX-M-15, and APH(3')-Ia genes in Klebsiella pneumoniae strains, and the gene copy number was obtained using the second-generation high-throughput sequencing technology, and the sensitivity of Klebsiella pneumoniae to ertapenem was predicted through computer programs.
It achieves rapid and accurate prediction of Klebsiella pneumoniae drug sensitivity, reduces detection costs, improves operation ease and detection stability, assists in clinical rational use of drugs, and reduces medical costs.
Smart Images

Figure CN116110511B_ABST
Abstract
Description
Technical Field
[0001] The present invention belongs to the field of bioinformatics, and in particular relates to a system and method for predicting the sensitivity of Klebsiella pneumoniae to ertapenem. Background Art
[0002] Antibiotics were once humanity's "secret weapon" against numerous diseases. The discovery of a series of antibiotics in the late 19th and early 20th centuries significantly increased human life expectancy. In recent years, the continued use of antibiotics has led to increasing overuse, leading to increasing clinical antibiotic resistance and adverse reactions, and placing a heavy burden on the global economy. Effectively controlling the misuse of antibiotics in medical settings is a crucial component of addressing the global problem of antibiotic resistance.
[0003] Pathogens are microorganisms that can invade the human body and cause infection or even infectious diseases. These include bacteria, viruses, fungi, parasites, mycoplasmas, chlamydiae, rickettsiae, and spirochetes. Microbial samples vary widely. Intestinal samples include feces and mucous membranes; fluid samples include urine, blood, cerebrospinal fluid, saliva, sputum, bronchoalveolar lavage fluid, and amniotic fluid; swab samples include oral, genital, and skin samples; and others include tissue, liver, eye, and placenta samples.
[0004] The genus Klebsiella, scientifically known as Klebsiella Trevisan and classified at the genus level, is a straight bacterium with a diameter of 0.3 to 1.0 μm and a length of 0.6 to 6.0 μm. It can be found singly, in pairs, or in short chains. Currently reported species (strains) of this genus include Klebsiella pneumoniae, Klebsiella aerogenes, Klebsiella oxytoca, Klebsiella quasipneumoniae, Klebsiella variicola, and Klebsiella michiganensis.
[0005] Klebsiella pneumoniae, the model species (strain) of the genus Klebsiella, is widely present in the environment and easily colonizes the respiratory and intestinal tracts of patients, causing infections in multiple sites such as the digestive tract, respiratory tract, and bloodstream. It is a common opportunistic pathogen that causes pneumonia in humans and is also a common nosocomial drug-resistant bacterium. According to a study by the Second Military Medical University, the resistance rate to ertapenem among carbapenem-resistant K. pneumoniae isolated from 2014 to 2017 was 62.5% (252 / 403).
[0006] Ertapenem is a carbapenem antibiotic primarily used to treat aerobic Gram-negative bacterial infections. Its characteristics include: definite efficacy and broad-spectrum antimicrobial activity against common Gram-positive and Gram-negative bacteria, including extended-spectrum β-lactamase-producing Enterobacteriaceae and anaerobic bacteria; rapid bactericidal activity, enabling rapid and low resistance screening; and a low risk of developing resistant bacteria after treatment.
[0007] Bacterial drug susceptibility testing is currently the most commonly used method for detecting bacterial resistance in clinical and laboratory settings, both domestically and internationally. Methods include the disc method, agar dilution, broth dilution, and concentration gradient methods. With the exception of the disc method, all other methods can yield relatively accurate minimum inhibitory concentrations (MICs) of drugs. Bacterial drug susceptibility testing first requires obtaining pure cultures, making it inapplicable to difficult-to-cultivate or non-cultivable bacteria. It is also time-consuming and sometimes fails to meet the current clinical needs for rapid diagnosis and symptomatic treatment of severe and acute infections. Traditional methods for detecting and identifying pathogens fail to meet the comprehensive requirements of broad coverage, rapidity, and accuracy. Diagnosis and treatment of infectious diseases are primarily based on empirical, targeted methods. Clinicians and patients urgently need innovative testing methods that can more comprehensively, accurately, and rapidly identify infectious pathogens, assist in diagnosis and rationalize drug treatment, shorten treatment courses, reduce mortality, and lower medical costs.
[0008] With the promotion of emerging technologies such as PCR, whole genome sequencing, microfluidics, and the VITEK-2compact fully automated bacterial identification and drug susceptibility system, the exploration of new technologies for bacterial resistance detection has deepened, and various new methods for bacterial resistance detection are becoming increasingly mature. While the VITEK-2compact fully automated bacterial identification and drug susceptibility system is simple and rapid, its accuracy in strain identification and drug susceptibility assessment is affected by the sample condition and strain culture conditions, and its cost is relatively high.
[0009] Therefore, there is an urgent need in the art to develop a method and system that can quickly, accurately and cost-effectively predict the sensitivity of Klebsiella pneumoniae strains to ertapenem. Summary of the Invention
[0010] In view of the above-mentioned deficiencies and needs of the existing technology in this field, the present invention aims to provide a system and method for predicting the sensitivity of Klebsiella pneumoniae to ertapenem.
[0011] The technical solutions of the present invention are as follows:
[0012] A system for predicting the sensitivity of Klebsiella pneumoniae to ertapenem is provided with: a computing unit; the computing unit includes: a computer-readable storage medium on which a computer program is stored; the computer program is characterized in that when executed by a processor, an Exp(-k) power value calculation method is implemented; the Exp(-k) power value calculation method includes the following calculation steps:
[0013] S1: Calculate the k value according to the following formula I:
[0014] Formula I:
[0015] S2: Find the power of Exp(-k) with the natural constant e as base and -k as exponent;
[0016] In Formula 1:
[0017] C1 is the copy number of the KPC-1 gene in the Klebsiella pneumoniae strain to be predicted,
[0018] C2 is the copy number of the marA gene in the Klebsiella pneumoniae strain to be predicted,
[0019] C3 is the copy number of the nalC gene in the Klebsiella pneumoniae strain to be predicted,
[0020] C4 is the copy number of the sul1 gene in the Klebsiella pneumoniae strain to be predicted,
[0021] C5 is the copy number of the CTX-M-15 gene in the Klebsiella pneumoniae strain to be predicted,
[0022] C6 is the copy number of the APH(3')-Ia gene in the Klebsiella pneumoniae strain to be predicted.
[0023] The system for predicting the sensitivity of Klebsiella pneumoniae to ertapenem is characterized by further being provided with: a result output unit; the calculation unit transmits the calculated Exp(-k) power value to the result output unit, and the result output unit identifies the Exp(-k) power value and outputs the result;
[0024] Preferably, the natural constant e=2.718281828459045.
[0025] When the result output unit identifies that the Exp(-k) power value is less than 1, it outputs the drug resistance result R;
[0026] The result output unit outputs the sensitive result S when it recognizes that the power value of Exp(-k) is ≥1;
[0027] The result output unit and the calculation unit are connected by a data path, and the Exp(-k) power value calculated by the calculation unit is transmitted to the result output unit via the data path;
[0028] Preferably, the sensitive result S indicates that the Klebsiella pneumoniae to be predicted is sensitive to ertapenem; the drug-resistant result R indicates that the Klebsiella pneumoniae to be predicted is resistant to ertapenem.
[0029] The system for predicting the sensitivity of Klebsiella pneumoniae to ertapenem is characterized by further comprising: an experimental unit and a data input unit;
[0030] The experimental unit and the data input unit are connected by a data path; the experimental unit outputs the experimental results, which are transmitted to the data input unit via the data path and converted into independent variable data;
[0031] The data input unit is connected to the calculation unit via a data path; the independent variable data is transmitted to the calculation unit via the data path.
[0032] The independent variable data include: the values of C1, C2, C3, C4, C5, and C6;
[0033] Preferably, the experimental results include: the copy number of the KPC-1 gene in the Klebsiella pneumonia strain to be predicted, the copy number of the marA gene in the Klebsiella pneumonia strain to be predicted, the copy number of the nalC gene in the Klebsiella pneumonia strain to be predicted, the copy number of the sul1 gene in the Klebsiella pneumonia strain to be predicted, the copy number of the CTX-M-15 gene in the Klebsiella pneumonia strain to be predicted, and the copy number of the APH(3')-Ia gene in the Klebsiella pneumonia strain to be predicted.
[0034] A method for predicting the sensitivity of Klebsiella pneumoniae to ertapenem, comprising:
[0035] S1: Calculate the k value according to the following formula I:
[0036] Formula I:
[0037] S2: Find the power of Exp(-k) with the natural constant e as base and -k as exponent;
[0038] In Formula 1:
[0039] C1 is the copy number of the KPC-1 gene in the Klebsiella pneumoniae strain to be predicted,
[0040] C2 is the copy number of the marA gene in the Klebsiella pneumoniae strain to be predicted,
[0041] C3 is the copy number of the nalC gene in the Klebsiella pneumoniae strain to be predicted,
[0042] C4 is the copy number of the sul1 gene in the Klebsiella pneumoniae strain to be predicted,
[0043] C5 is the copy number of the CTX-M-15 gene in the Klebsiella pneumoniae strain to be predicted,
[0044] C6 is the copy number of the APH(3′)-Ia gene in the K. pneumoniae strain to be predicted;
[0045] The predicted result corresponding to the Exp(-k) power value <1 is that Klebsiella pneumoniae is resistant to ertapenem, and the predicted result corresponding to the Exp(-k) power value ≥1 is that Klebsiella pneumoniae is sensitive to ertapenem.
[0046] The natural constant e=2.718281828459045;
[0047] The copy numbers of KPC-1, marA, nalC, sul1, CTX-M-15, and APH(3')-Ia genes in the predicted Klebsiella pneumoniae strains were obtained by second-generation high-throughput sequencing.
[0048]
[0049] Preferably, the genomic contigs are the longest contigs fragments assembled from sequencing results using SPAdes v3.13.0 assembly software;
[0050] The genome contigs depth is the depth of genome contigs calculated by SPAdes v3.13.0 assembly software;
[0051] The depth of the contigs where the gene is located refers to the sum of the depths of the gene on each contig with a copy of the gene;
[0052] Preferably, each contig with a copy of the gene is annotated by comparing the cds and protein sequences of the gene with the CARD database using blat (v.36) software and diamond (v2.0.4.142) software;
[0053] Preferably, the depth of a gene on each contig having a copy of the gene is calculated using SPAdes v3.13.0 assembly software.
[0054] The beneficial effects of the present invention are as follows:
[0055] Based on the prediction system and method of the present invention, after conventional microbial sample processing, DNA extraction and other necessary sequencing steps can be performed. Bioinformatics analysis is then used to determine the status of Klebsiella pneumoniae prediction system-related features in the sample. This feature status information is then imported into the system to predict the sample's drug sensitivity. Compared to traditional methods, this method offers advantages such as ease of operation, shortened detection time, and accurate species identification.
[0056] To effectively evaluate the performance of a prediction system, a dataset not used in the prediction system's development is required. This independent dataset is called a test set. System prediction performance evaluation methods include F1-score, precision, recall, and confusion matrix.
[0057] The method of the present invention also has the following advantages:
[0058] The present invention evaluated the accuracy of the system using a test set, achieving an average accuracy of 0.933, an F1-score of 0.889, and a recall score of 0.813. This method is less susceptible to subjective factors, such as operator input, and exhibits excellent detection stability. Furthermore, it enables rapid and accurate identification of infectious pathogens and prediction of drug susceptibility in test samples, aiding diagnosis and enabling rational, standardized medication and treatment. Furthermore, it boasts high throughput and reduces medical costs. BRIEF DESCRIPTION OF THE DRAWINGS
[0059] Figure 1 Schematic diagram of the structure of the drug resistance prediction system provided in some embodiments of the present invention (in the dotted box) and its workflow diagram.
[0060] Figure 2 Schematic diagram of the structure of the drug resistance prediction system (in the dotted box) and its workflow diagram provided for other embodiments of the present invention. DETAILED DESCRIPTION
[0061] In order to facilitate understanding of the present invention, the present invention will be described more comprehensively in the following examples.
[0062] Unless otherwise defined, all technical and scientific terms used herein have the same meaning as those commonly understood by those skilled in the art of the present invention. The terms used in this specification of the present invention are only for the purpose of describing specific embodiments and are not intended to limit the present invention.
[0063] Unless otherwise specified, the reagents used in the following examples are commercially available.
[0064] Sources of biological materials
[0065] The 194 samples used in the experimental examples of the present invention were pure cultures of Klebsiella pneumoniae isolated from clinical blood cultures and were obtained from Peking Union Medical College Hospital, Chinese Academy of Medical Sciences.
[0066] All tested bacterial species (strains) were identified as Klebsiella (scientific name: Genus Klebsiella, Latin name: Klebsiella Trevisan, systematic taxonomic level: Genus) by mass spectrometry MALDI-TOF MS.
[0067] On the Illumina Novaseq NGS sequencing platform, these strains included 144 cases of Klebsiella pneumoniae, 23 cases of Klebsiella aerogenes, 9 cases of Klebsiella oxytoca, 8 cases of Klebsiella quasipneumoniae, 7 cases of Klebsiella variicola, and 3 cases of Klebsiella michiganensis, all of which were reported species or strains of the genus Klebsiella.
[0068] The above-mentioned bacterial species or strains can be obtained from common Klebsiella pneumoniae pneumonia cases or from the applicant's laboratory. The applicant commits to distributing the strains to the public for verification of the technical effects of the present invention within 20 years from the filing date of this invention.
[0069] The first group of embodiments, the drug resistance prediction system of the present invention
[0070] This group of embodiments provides a system for predicting the sensitivity of Klebsiella pneumoniae to ertapenem. All embodiments in this group have the following common features: Figure 1 and Figure 2 As shown, the system for predicting the sensitivity of Klebsiella pneumoniae to ertapenem is provided with: a computing unit; the computing unit includes: a computer-readable storage medium having a computer program stored thereon; characterized in that when the computer program is executed by a processor, an Exp(-k) power value calculation method is implemented; the Exp(-k) power value calculation method is calculated according to the following calculation steps:
[0071] S1: Calculate the k value according to the following formula I:
[0072] Formula I:
[0073] S2: Find the power of Exp(-k) with the natural constant e as base and -k as exponent;
[0074] In Formula 1:
[0075] C1 is the copy number of the KPC-1 gene in the Klebsiella pneumoniae strain to be predicted,
[0076] C2 is the copy number of the marA gene in the Klebsiella pneumoniae strain to be predicted,
[0077] C3 is the copy number of the nalC gene in the Klebsiella pneumoniae strain to be predicted,
[0078] C4 is the copy number of the sul1 gene in the Klebsiella pneumoniae strain to be predicted,
[0079] C5 is the copy number of the CTX-M-15 gene in the Klebsiella pneumoniae strain to be predicted,
[0080] C6 is the copy number of the APH(3')-Ia gene in the Klebsiella pneumoniae strain to be predicted.
[0081] In some embodiments of the present invention, the value of the natural constant e is 2.718281828459045.
[0082] In a more specific embodiment, the above genes are all genes reported in the art, as follows:
[0083] The KPC-1 gene is the KPC-1 gene described in the article “Novel Carbapenem-Hydrolyzing b-Lactamase, KPC-1, from a Carbapenem-Resistant Strain of Klebsiella pneumoniae”.
[0084] The marA gene is the marA gene described in the article “Correlation of the expression of acrB and the regulatory genes marA, soxS and ramA with antimicrobial resistance in clinical isolates of Klebsiella pneumoniae endemic to New York City”.
[0085] The nalC gene is the nalC gene recorded in the article “Mutations in NalC induce MexAB-OprM overexpression resulting in high level of aztreonam resistance in environmental isolates of Pseudomonas aeruginosa”.
[0086] The sul1 gene is the sul1 gene described in the article “Co-occurrence of Klebsiella variicola and Klebsiellapneumoniae Both Carrying blaKPC from a Respiratory Intensive Care Unit Patient”.
[0087] The CTX-M-15 gene is the CTX-M-15 gene described in the article “High prevalence of CTX-M-15-producing Klebsiellapneumoniae isolates in Asian countries: diverse clones and clonaldissemination”.
[0088] The APH(3')-Ia gene is the APH(3')-Ia gene described in the article "Carbapenem-Resistant Klebsiella pneumoniae Strains Exhibit Diversity in Aminoglycoside-Modifying Enzymes, Which Exert Differing Effects on Plazomicin and Other Agents".
[0089] In a further embodiment, Figure 1 and Figure 2 As shown, the system for predicting the sensitivity of Klebsiella pneumoniae to ertapenem is further provided with: a result output unit; the result output unit outputs a sensitive result or a drug-resistant result; the sensitive result means that the Klebsiella pneumoniae to be predicted is sensitive to ertapenem; the drug-resistant result means that the Klebsiella pneumoniae to be predicted is resistant to ertapenem;
[0090] When the Exp(-k) power value is less than 1, the result output unit outputs the drug resistance result R;
[0091] When the Exp(-k) power value is ≥1, the result output unit outputs the sensitive result S;
[0092] Preferably, the result output unit and the calculation unit are connected by a data path;
[0093] Preferably, the Exp(-k) power value calculated by the calculation unit is transmitted to the result output unit via a data path.
[0094] In a further embodiment, Figure 1 As shown, the system for predicting the sensitivity of Klebsiella pneumoniae to ertapenem also includes: an experimental unit and a data input unit;
[0095] The experimental unit and the data input unit are connected by a data path; the experimental unit outputs the experimental results which are transmitted to the data input unit via the data path and converted into independent variable data;
[0096] The data input unit is connected to the computing unit via a data path; the independent variable data is transmitted to the computing unit via the data path;
[0097] In a more specific embodiment, the data path is a data transmission medium known to those skilled in the computer and electronics fields. The data path can be a wired or wireless form, for example, a wired path, a line, or a wireless path, a Wi-Fi connection, a wireless channel, etc.
[0098] Preferably, the independent variable data include: numerical values of C1, C2, C3, C4, C5, and C6;
[0099] Preferably, the experimental results include: the copy numbers of KPC-1, marA, nalC, sul1, CTX-M-15, and APH(3')-Ia genes in the Klebsiella pneumoniae strain to be predicted.
[0100] The copy number of a known gene in a known strain is conventionally known to those skilled in the art of molecular biology and bioinformatics through conventional techniques (e.g., sequencing and bioinformatics analysis). The KPC-1, marA, nalC, sul1, CTX-M-15, and APH(3')-Ia genes involved in the experimental results output by the experimental unit of the prediction system of the present invention are all genes reported in the art, and their gene information and primary structure sequences can be queried through the NCBI website or other known bioinformatics databases. The copy number of each of the above genes in the Klebsiella pneumoniae strain to be predicted can be determined by performing whole genome sequencing.
[0101] In other specific embodiments, the copy numbers of rmtB, AAC(6')-Ib-cr6, KPC-1, APH(3")-Ib, FosA5, and SHV-25 genes in the Klebsiella pneumoniae strain to be predicted are obtained by a second-generation high-throughput sequencing method.
[0102] In a more specific embodiment,
[0103]
[0104] Preferably, the genomic contigs are the longest contigs fragments assembled from sequencing results using SPAdes v3.13.0 assembly software;
[0105] The genome contigs depth is the depth of genome contigs calculated by SPAdes v3.13.0 assembly software;
[0106] The depth of the contigs where the gene is located refers to the sum of the depths of the gene on each contig with a copy of the gene;
[0107] Preferably, each contig with a copy of the gene is annotated by comparing the cds and protein sequences of the gene with the CARD database using blat (v.36) software and diamond (v2.0.4.142) software;
[0108] Preferably, the depth of a gene on each contig having a copy of the gene is calculated using SPAdes v3.13.0 assembly software.
[0109] The second-generation high-throughput sequencing method has a conventional technical meaning well known to those skilled in the art, and obtaining the gene copy number using the second-generation high-throughput sequencing method is a conventional technical means well known to those skilled in the art.
[0110] In some specific embodiments, the specific method for calculating the gene copy number is as follows:
[0111] The strain was sequenced using the second-generation high-throughput sequencing method. The average sequencing depth was about 150x, and the approximate sequencing volume for Klebsiella pneumoniae was about 1G. The contigs depth calculated during the assembly process using SPAdes (v3.13.0) assembly software was used as the standard, and the longest contigs fragment was defined as the genomic fragment. The contigs were predicted using prokka software (1.14.6) to obtain all gene cds and protein sequences on the contigs. The cds and protein sequences were compared with the CARD database using blat (v.36) and diamond (v2.0.4.142) software, respectively. Sequences with a similarity greater than 90% were positive sequences, and the annotation results of all drug-resistant genes were obtained. The copy number of all genes on the contigs was calculated according to formula II as follows:
[0112] Formula II:
[0113] If a gene has two or more genomic copies in different contigs or on the same contig, the final gene copy number is equal to the sum of all calculated copy numbers of the gene. The calculation method is as follows:
[0114] Assuming that there is only one copy of the KPC-1 gene in all contigs, the copy number of the KPC-1 gene is:
[0115]
[0116] Assuming that the KPC-1 gene has two copies in one contig and no copies in other contigs, the copy number of the KPC-1 gene is:
[0117]
[0118] Assuming that the KPC-1 gene has one copy in contig 1 and contig 2 and no copy in other contigs, the copy number of the KPC-1 gene is:
[0119]
[0120] In a more specific embodiment, the result output unit, the experiment unit, and the data input unit are all provided with a computer-readable storage medium on which a computer program is stored.
[0121] In some embodiments, the computer program on the computer-readable storage medium of the result output unit, when executed by the processor, implements a method for comparing the size of the Exp(-k) power value with 1 and outputs the result;
[0122] The method for comparing the Exp(-k) power value with 1 and outputting the result refers to:
[0123] When the Exp(-k) power value is less than 1, the result output unit outputs the drug resistance result R;
[0124] When the Exp(-k) power value is ≥1, the result output unit outputs the sensitive result S.
[0125] In other embodiments, the computer program on the computer-readable storage medium of the experimental unit implements a gene copy number calculation method when executed by a processor;
[0126] The gene copy number calculation method is a conventional technical means well known to those skilled in the art, and is specifically calculated according to the following steps:
[0127] S1: The depth of genome contigs calculated by SPAdes v3.13.0 assembly software was maximized to obtain genome contigs;
[0128] S2: The cds and protein sequences of a gene were compared with the CARD database using blat (v.36) and diamond (v2.0.4.142) software, and then annotated to obtain the contigs containing the gene copies.
[0129] S3: SPAdes v3.13.0 assembly software was used to calculate the depth of the gene on each contig containing the gene copy;
[0130] S4: Calculate the sum of the depths of the gene on each contig with a copy of the gene to obtain the depth of the contig where the gene is located;
[0131] S5: Calculate the copy number of the gene according to the following formula:
[0132]
[0133] In some embodiments, the computer program on the computer-readable storage medium of the data input unit implements dimensionless processing of the copy number of the gene when executed by the processor.
[0134] The dimensionless processing refers to removing the data dimension or data unit of the gene copy number to obtain a dimensionless value, that is, the independent variable data. Generally speaking, the data dimension or data unit of the gene copy number is: copy, individual or copies.
[0135] In other embodiments, Figure 2As shown, the system for predicting the sensitivity of Klebsiella to ertapenem may not be provided with a data input unit. The experimental unit is directly connected to the calculation unit via a data path, so that the gene copy number or independent variable data calculated by the experimental unit can be directly input into the calculation unit to calculate the Exp(-k) power value.
[0136] Group 2 Example: Method for predicting the resistance of Pneumococcus to ertapenem of the present invention
[0137] This group of embodiments provides a method for predicting the sensitivity of Klebsiella pneumoniae to ertapenem. This group of embodiments has the following common features: the method comprises the following steps:
[0138] S1: Calculate the k value according to the following formula I:
[0139] Formula I:
[0140] S2: Find the power of Exp(-k) with the natural constant e as base and -k as exponent;
[0141] In Formula 1:
[0142] C1 is the copy number of the KPC-1 gene in the Klebsiella pneumoniae strain to be predicted,
[0143] C2 is the copy number of the marA gene in the Klebsiella pneumoniae strain to be predicted,
[0144] C3 is the copy number of the nalC gene in the Klebsiella pneumoniae strain to be predicted,
[0145] C4 is the copy number of the sul1 gene in the Klebsiella pneumoniae strain to be predicted,
[0146] C5 is the copy number of the CTX-M-15 gene in the Klebsiella pneumoniae strain to be predicted,
[0147] C6 is the copy number of the APH(3')-Ia gene in the Klebsiella pneumoniae strain to be predicted.
[0148] The predicted result corresponding to the Exp(-k) power value <1 is that Klebsiella pneumoniae is resistant to ertapenem, and the predicted result corresponding to the Exp(-k) power value ≥1 is that Klebsiella pneumoniae is sensitive to ertapenem.
[0149] In Formula I above, e, a mathematical constant, is the base of the natural logarithm function, also known as the natural constant, natural base, or Euler's number. It is an infinite, non-repeating decimal with the conventional technical meaning commonly understood by those skilled in the art of mathematics. Its value is approximately: e = 2.71828182845904523536...
[0150] In some embodiments of the present invention, the value of the natural constant e is 2.718281828459045.
[0151] In some specific embodiments, the copy numbers of KPC-1, marA, nalC, sul1, CTX-M-15, and APH(3')-Ia genes in the Klebsiella pneumoniae strain to be predicted are obtained by a second-generation high-throughput sequencing method.
[0152] In a more specific embodiment,
[0153]
[0154] Preferably, the genomic contigs are the longest contigs fragments assembled from sequencing results using SPAdes v3.13.0 assembly software;
[0155] The genome contigs depth is the depth of genome contigs calculated by SPAdes v3.13.0 assembly software;
[0156] The depth of the contigs where the gene is located refers to the sum of the depths of the gene on each contig with a copy of the gene;
[0157] Preferably, each contig with a copy of the gene is annotated by comparing the cds and protein sequences of the gene with the CARD database using blat (v.36) software and diamond (v2.0.4.142) software;
[0158] Preferably, the depth of a gene on each contig having a copy of the gene is calculated using SPAdes v3.13.0 assembly software.
[0159] The second-generation high-throughput sequencing method has a conventional technical meaning well known to those skilled in the art, and obtaining the gene copy number using the second-generation high-throughput sequencing method is a conventional technical means well known to those skilled in the art.
[0160] In some specific embodiments, the specific method for calculating the gene copy number is as follows:
[0161] The strain was sequenced using the second-generation high-throughput sequencing method. The average sequencing depth was about 150x, and the approximate sequencing volume for Klebsiella pneumoniae was about 1G. The contigs depth calculated during the assembly process using SPAdes (v3.13.0) assembly software was used as the standard, and the longest contigs fragment was defined as the genomic fragment. The contigs were predicted using prokka software (1.14.6) to obtain all gene cds and protein sequences on the contigs. The cds and protein sequences were compared with the CARD database using blat (v.36) and diamond (v2.0.4.142) software, respectively. Sequences with a similarity greater than 90% were positive sequences, and the annotation results of all drug-resistant genes were obtained. The copy number of all genes on the contigs was calculated according to formula II as follows:
[0162] Formula II:
[0163] If a gene has two or more genomic copies in different contigs or on the same contig, the final gene copy number is equal to the sum of all calculated copy numbers of the gene. The calculation method is as follows:
[0164] Assuming that there is only one copy of the KPC-1 gene in all contigs, the copy number of the KPC-1 gene is:
[0165]
[0166] Assuming that the KPC-1 gene has two copies in one contig and no copies in other contigs, the copy number of the KPC-1 gene is:
[0167]
[0168] Assuming that the KPC-1 gene has one copy in contig 1 and contig 2 and no copy in other contigs, the copy number of the KPC-1 gene is:
[0169]
[0170] Experimental Examples, Performance Evaluation of the Prediction System and Prediction Method of the Present Invention
[0171] The prediction system of the present invention was evaluated using 194 clinical samples. The comparison of the broth microdilution classification results of the 194 clinical samples and the system prediction results is shown in Table 1. In the table below, S represents sensitive and R represents resistant.
[0172] Table 1
[0173]
[0174]
[0175]
[0176]
[0177]
[0178] The test result data generates a confusion matrix as shown in Table 2:
[0179] Table 2
[0180]
[0181] Assume that TP (True Positive) represents the number of true positives, FP (False Positive) represents the number of false positives, FN (False Negative) represents the number of false negatives, and TN (Ture Negative) represents the number of true negatives. Precision refers to the proportion of positive samples determined by the classifier to be positive. Recall refers to the proportion of positive samples predicted to be positive. Accuracy refers to the proportion of samples that the classifier correctly judges to be correct. F1-score is the harmonic mean of precision and recall, with a maximum of 1 and a minimum of 0. The calculation results of each indicator are as follows:
[0182]
[0183]
[0184]
[0185]
[0186] The above-described embodiments merely illustrate the implementation methods of the present invention. While the descriptions are relatively specific and detailed, they should not be construed as limiting the scope of the patent. It should be noted that a person skilled in the art would be able to make various modifications and improvements without departing from the spirit of the present invention, all of which fall within the scope of protection of the present invention. Therefore, the scope of protection of the patent for this invention shall be determined by the appended claims.
Claims
1. A system for predicting the sensitivity of Klebsiella pneumoniae to ertapenem, comprising: a calculation unit and a result output unit; the calculation unit comprises: A computer-readable storage medium having a computer program stored thereon; wherein the computer program, when executed by a processor, implements a method for calculating an Exp(-k) power value; the method for calculating an Exp(-k) power value comprises the following calculation steps: S1: Calculate the k value according to the following formula I: Formula I: ; S2: Find the power of Exp(-k) with the natural constant e as base and -k as exponent; In Formula 1: C1 is the copy number of the KPC-1 gene in the Klebsiella pneumoniae strain to be predicted, C2 is the copy number of the marA gene in the Klebsiella pneumoniae strain to be predicted, C3 is the copy number of the nalC gene in the Klebsiella pneumoniae strain to be predicted, C4 is the copy number of the sul1 gene in the Klebsiella pneumoniae strain to be predicted, C5 is the copy number of the CTX-M-15 gene in the Klebsiella pneumoniae strain to be predicted, C6 is the copy number of the APH(3')-Ia gene in the Klebsiella pneumoniae strain to be predicted; the calculation unit transmits the calculated Exp(-k) power value to the result output unit, and the result output unit identifies the Exp(-k) power value and outputs the result.
2. The system for predicting the sensitivity of Klebsiella pneumoniae to ertapenem according to claim 1 is characterized in that: The natural constant e=2.718281828459045.
3. The system for predicting the sensitivity of Klebsiella pneumoniae to ertapenem according to claim 1 is characterized in that: When the result output unit identifies that the Exp(-k) power value is less than 1, it outputs the drug resistance result R; The result output unit outputs a sensitive result S when it recognizes that the power value of Exp(-k) is ≥1.
4. The system for predicting the sensitivity of Klebsiella pneumoniae to ertapenem according to claim 1, characterized in that: The result output unit and the calculation unit are connected via a data path, and the Exp(-k) power value calculated by the calculation unit is transmitted to the result output unit via the data path.
5. The system for predicting the sensitivity of Klebsiella pneumoniae to ertapenem according to claim 3, characterized in that: The sensitive result S indicates that the Klebsiella pneumoniae to be predicted is sensitive to ertapenem; the drug-resistant result R indicates that the Klebsiella pneumoniae to be predicted is resistant to ertapenem.
6. A system for predicting the sensitivity of Klebsiella pneumoniae to ertapenem according to any one of claims 1 to 5, characterized in that: Also provided are: experimental unit and data input unit; The experimental unit and the data input unit are connected by a data path; the experimental unit outputs the experimental results, which are transmitted to the data input unit via the data path and converted into independent variable data; The data input unit is connected to the calculation unit via a data path; the independent variable data is transmitted to the calculation unit via the data path.
7. The system for predicting the sensitivity of Klebsiella pneumoniae to ertapenem according to claim 6, characterized in that: The independent variable data include: numerical values of C1, C2, C3, C4, C5, and C6.
8. The system for predicting the sensitivity of Klebsiella pneumoniae to ertapenem according to claim 6, characterized in that: The experimental results include: the copy number of the KPC-1 gene in the Klebsiella pneumonia strain to be predicted, the copy number of the marA gene in the Klebsiella pneumonia strain to be predicted, the copy number of the nalC gene in the Klebsiella pneumonia strain to be predicted, the copy number of the sul1 gene in the Klebsiella pneumonia strain to be predicted, the copy number of the CTX-M-15 gene in the Klebsiella pneumonia strain to be predicted, and the copy number of the APH(3')-Ia gene in the Klebsiella pneumonia strain to be predicted.
9. A method for predicting the sensitivity of Klebsiella pneumoniae to ertapenem, characterized in that: include: S1: Calculate the k value according to the following formula I: Formula I: ; S2: Find the power of Exp(-k) with the natural constant e as base and -k as exponent; In Formula 1: C1 is the copy number of the KPC-1 gene in the Klebsiella pneumoniae strain to be predicted, C2 is the copy number of the marA gene in the Klebsiella pneumoniae strain to be predicted, C3 is the copy number of the nalC gene in the Klebsiella pneumoniae strain to be predicted, C4 is the copy number of the sul1 gene in the Klebsiella pneumoniae strain to be predicted, C5 is the copy number of the CTX-M-15 gene in the Klebsiella pneumoniae strain to be predicted, C6 is the copy number of the APH(3′)-Ia gene in the K. pneumoniae strain to be predicted; The predicted result corresponding to the Exp(-k) power value <1 is that Klebsiella pneumoniae is resistant to ertapenem, and the predicted result corresponding to the Exp(-k) power value ≥1 is that Klebsiella pneumoniae is sensitive to ertapenem.
10. The method for predicting the sensitivity of Klebsiella pneumoniae to ertapenem according to claim 9, characterized in that: The natural constant e=2.718281828459045.
11. The method for predicting the sensitivity of Klebsiella pneumoniae to ertapenem according to claim 9, characterized in that: The copy numbers of KPC-1, marA, nalC, sul1, CTX-M-15, and APH(3')-Ia genes in the predicted Klebsiella pneumoniae strains were obtained by second-generation high-throughput sequencing.
12. The method for predicting the sensitivity of Klebsiella pneumoniae to ertapenem according to claim 9, characterized in that: The copy number of the gene in the Klebsiella pneumoniae strain to be predicted = .
13. The method for predicting the sensitivity of Klebsiella pneumoniae to ertapenem according to claim 12, characterized in that: The genome contigs are the longest contigs fragments obtained by assembling the sequencing results using SPAdes v3.13.0 assembly software; The genome contigs depth is the depth of genome contigs calculated by SPAdes v3.13.0 assembly software; The depth of the contigs where the gene is located refers to the sum of the depths of the gene on each contig with a copy of the gene.
14. The method for predicting the sensitivity of Klebsiella pneumoniae to ertapenem according to claim 13, characterized in that: Each contig containing a copy of the gene was annotated after comparing the cds and protein sequences of the gene with the CARD database using blat (v.36) and diamond (v2.0.4.142) software.
15. The method for predicting the sensitivity of Klebsiella pneumoniae to ertapenem according to claim 13, characterized in that: The depth of a gene in each contig with a copy of the gene was calculated using the SPAdes v3.13.0 assembly software.
Citation Information
Patent Citations
Method and system for predicting sensitivity of klebsiella pneumoniae to ceftriaxone
CN114898800A
Method and system for predicting sensitivity of klebsiella pneumoniae to cefepime
CN114898808A