Computational method for rapid and non-invasive detection of sepsis based on long read single molecule sequencing
Long-read, single-molecule sequencing of cell-free DNA allows for rapid and non-invasive detection of bacterial pathogens and resistance, addressing the inefficiencies of current methods and enabling timely antibiotic therapy.
Patent Information
- Authority / Receiving Office
- WO · WO
- Patent Type
- Applications
- Current Assignee / Owner
- GENETON
- Filing Date
- 2024-11-18
- Publication Date
- 2026-05-21
AI Technical Summary
Current methods for detecting bacterial resistance in sepsis are time-consuming and invasive, delaying effective antibiotic therapy, which is critical in sepsis treatment.
A method involving long-read, single-molecule sequencing of cell-free DNA from blood samples, followed by quality control, filtering, mapping, and analysis for bacterial species and resistance markers, enabling rapid and non-invasive identification of pathogens and their resistance profiles.
Enables rapid detection of bacterial pathogens and their resistance within hours, facilitating timely and targeted antibiotic therapy, reducing the risk of antibiotic resistance and improving patient outcomes.
Abstract
Description
[0001] Computational method for rapid and non-invasive detection of sepsis based on long read single molecule sequencing
[0002] Field of the invention
[0003] The present invention relates to a method and system to predict bacterial resistance based on sequencing data. This invention also relates to analysis of genetic markers to estimate bacterial resistance to different types of antibiotics with high accuracy. The present invention belongs to the field of molecular biology and biotechnology, more specifically to the field DNA sequencing, bioinformatics and genomics.
[0004] Background of the invention
[0005] It is known that the bacterial genome encompasses all genetic material within a bacterium, primarily organized into a single, circular chromosome, although some species may possess multiple chromosomes or, in rare instances, linear chromosomes. The relatively compact structure of bacterial genomes, compared to eukaryotes, allows for efficient regulation and replication, with genome sizes ranging from 500 kilobases (kb) to over 10 megabases (Mb) in larger species. This compact genome plays a critical role in identifying pathogenic bacteria, especially in clinical contexts such as sepsis, where rapid detection of the causative agent is paramount.
[0006] Sepsis, a life-threatening response to infection, is often triggered by bacterial pathogens. Accurate and timely detection of the bacterial source is crucial for effective treatment. In bacterial infections leading to sepsis, the genome encodes essential genes for processes such as metabolism, replication, and virulence. Virulence factors, often harbored on plasmids, are key to the pathogenicity of bacteria and detecting these genetic elements can help identify the source of sepsis. Plasmids, which are small, independently replicating DNA molecules, often carry genes for antibiotic resistance and virulence, facilitating bacterial adaptation and survival in hostile environments, such as during an immune response.
[0007] In sepsis cases, understanding the genomic landscape of the infecting bacteria is essential, particularly in determining whether resistance genes are present. These genes can either reside on plasmids or be integrated into the bacterial chromosome, often in regions like integrons or resistance islands. Such regions contribute to bacterial survival and virulence, directly impacting the progression of sepsis. Antibiotic resistance complicates the treatment of sepsis, as resistant bacteria are harder to eliminate. Bacteria acquire resistance through mutations or the transfer of resistance genes, which are often associated with sepsis-causing pathogens. Identifying the specific bacterial species and its resistance profile is crucial to guiding appropriate antimicrobial therapy and controlling the spread of resistance. To this end, laboratory tests are employed to pinpoint both the pathogen and its resistance mechanisms.
[0008] The process of treating sepsis lege artis begins with the isolation of bacteria from clinical samples, typically through blood cultures, which remain the gold standard for diagnosing sepsis. Blood cultures allow for the growth of the bacterial pathogen under controlled laboratory conditions, providing crucial information about the causative agent. Once the bacteria are isolated, antibiotic susceptibility testing (AST) is performed. AST assesses how the bacteria respond to a range of antibiotics, offering precise insights into which drugs are most effective for treating the infection. This cultivation-based approach is critical in tailoring the treatment to the specific pathogen, ensuring that antibiotics are used effectively and appropriately.
[0009] While waiting for culture results, empirical antibiotic therapy is often initiated. In sepsis, time is of the essence, and delays in treatment can lead to worse outcomes. Therefore, broad-spectrum antibiotics are administered based on clinical guidelines and the suspected source of infection. These empiric antibiotics aim to cover a wide range of possible pathogens until the specific causative agent is identified through cultures. Once the AST results are available, the empirical therapy is refined, or "de-escalated," to target the pathogen with the most effective and narrow-spectrum antibiotic. This step minimizes the overuse of broad-spectrum antibiotics, reducing the risk of promoting antibiotic resistance while ensuring that the patient receives optimal treatment.
[0010] Whole genome sequencing is a powerful tool to fully understand the bacteria’s DNA and identify resistance genes. The process includes preparing the DNA for sequencing, running it through a sequencing machine, and then assembling the sequences to map out the bacteria’s entire genome. This allows scientists to spot both mutations and resistance genes carried on mobile genetic elements. After sequencing, bioinformatics tools are used to analyze the genome, predominantly by comparing the sequenced DNA to known genes coding antibiotic resistance. Some common bioinformatics tools include ResFinder, which compares the bacterial DNA to a database of known antibiotic resistance genes, helping to quickly identify any present resistance genes; CARD (Comprehensive Antibiotic Resistance Database), which contains detailed information on resistance genes and their mutations, aiding in understanding how these genes work and what antibiotics they resist; and ARG-ANNOT, a tool specialized in finding antibiotic resistance genes specifically in bacterial genomes by comparing the DNA to a set of resistance gene sequences. These tools allow scientists to match the bacteria’s DNA to known resistance gene. In some cases, bioinformatic tools can also identify mutations in the DNA that are responsible for resistance. This approach allows clinicians to administer precise antibiotics, omitting the lengthy step of bacteria cultivation and at the same time reducing the risk of ineffective antibiotic administration.
[0011] Traditional laboratory techniques include phenotypic susceptibility testing, such as disk diffusion, broth microdilution, and agar dilution, which typically take 18 to 48 hours to assess the minimum inhibitory concentration (MIC) of antibiotics. Automated systems like Vitek® and Phoenix™ provide results on average within 9 to 12 hours by automating phenotypic testing with standardized protocols.
[0012] Molecular techniques, including PCR, are widely used for detecting specific resistance genes like mecA for MRSA. This method takes 2 to 6 hours. Real-time PCR reduces the time to 1 to 2 hours by allowing real-time monitoring of DNA amplification. DNA sequencing, such as for identifying mutations in the rpoB gene in Mycobacterium tuberculosis, can take several hours to a day, depending on the platform. DNA microarrays enable the simultaneous detection of multiple resistance genes, providing a comprehensive view in 4 to 8 hours.
[0013] Bioinformatic techniques involve the use of genomic databases like CARD (Comprehensive Antibiotic Resistance Database) and ResFinder, which help in identifying resistance genes from genomic data. Whole genome sequencing (WGS) provides a complete profile of bacterial resistance, though the process can take anywhere from a few hours to a couple of days. Comparative genomics uses bioinformatic tools to compare bacterial genomes and identify resistance mechanisms in a few hours to days, while metagenomics, used to identify resistance genes in environmental or clinical samples, can take several days due to the complexity of the data.
[0014] Next-generation sequencing (NGS), including Illumina and Oxford Nanopore, offers rapid, large-scale sequencing of bacterial genomes. NGS-based resistance detection can take from 24 hours to a few days, depending on the depth of sequencing and bioinformatic analysis.
[0015] Newer CRISPR-based diagnostics, such as SHERLOCK and DETECTR, have emerged to detect resistance genes in 1 to 2 hours, making them valuable for point-of-care diagnostics in clinical settings. The patent US10988792B2 describes methods for determining antimicrobial resistance, specifically focusing on detecting and quantifying resistance in microorganisms directly from clinical samples like blood or bodily fluids. The primary technique involves using resistance-determining affinity ligands that bind to the microorganisms under specific conditions. These ligands facilitate the differentiation between resistant and non-resistant strains by comparing the extent of ligand binding. The process includes contacting the microorganism with the ligand, separating the bound complex from unbound components, and quantitatively measuring the ligand bound to the microorganisms. This method can rapidly identify resistance, typically within a timeframe of about 240 minutes or less,
[0016] Summary of the invention
[0017] For the purposes of this description, the invention will be further described using the terms set forth in the following text. Other technical and scientific terms used herein have the same meaning as commonly understood by the persons skilled in the art of medicine, molecular genetics, molecular biology, bioinformatics and machine learning.
[0018] Definitions and General Techniques
[0019] As used herein, the term "genome" refers to the complete set of genetic instructions found in an organism, r
[0020] As used herein, the term "reference genome" refers to a representative example of a species' genome that serves as a baseline for mapping sequencing reads.
[0021] As used herein, the term "DNA sequencing" refers to a process of determining the precise order of nucleotides within a DNA molecule.
[0022] As used herein, the term "sequencing reads" refers to a read refers to the sequences of nucleotides obtained from fragmenting and decoding segments of DNA, which are then used to reconstruct the original sequence or align to a reference genome. The read must be long enough, typically at least 30-35 base pairs, to act as a sequence tag that can be clearly mapped to a specific location on a reference genome.
[0023] As used herein, the term "QC" or "quality control" refers to the process of assessing the accuracy and integrity of DNA sequencing reads. It involves checking the data for errors, ensuring that the sequences are of high quality, and filtering out problematic or low-quality reads. As used herein, the term "trimming" refers to the process of removing low-quality or unnecessary segments from the ends of DNA reads.
[0024] As used herein, the term "mapping" refers to the alignment of the sequence information from NGS (i.e. DNA fragment the genomic position of which is unknown) with a matching sequence in reference to the human genome. This can be done several ways. Reads that do not map uniquely (map to several positions) are usually excluded from the analysis. The alignment is usually done by computer algorithms well known to the persons skilled in the art of molecular biology and bioinformatics.
[0025] The term "SAM / BAM file" refers to a file stores aligned sequencing reads either in a text format (SAM) or a compressed binary format (BAM). For each read, it details the position on the reference genome, the mapping and sequencing quality, the location of the paired read in paired-end sequencing, among other data. It’s a standard format for storing aligned reads and each file's reference genome information is included in its header.
[0026] The term "FASTQ file" refers to a file contains all the sequencing reads, along with their quality scores, and is the standard format for storing such data. It is typically compressed to conserve disk space. Most modern mapping software programs can accept this format as input.
[0027] A first aspect of the present invention provides a method for the rapid and non-invasive detection of sepsis and identification of bacterial antibiotic resistance comprising the following steps:
[0028] a. collecting a biological sample from a patient suspected of having sepsis;
[0029] b. extracting cell-free DNA from the biological sample;
[0030] c. preparing a sequencing-ready library from the extracted DNA;
[0031] d. performing long-read, single-molecule sequencing on the prepared library;
[0032] e. processing the sequencing data to obtain raw sequencing reads;
[0033] f. performing quality control on the sequencing reads, wherein low-quality reads are removed;
[0034] g. filtering the remaining sequencing reads to isolate bacterial genomic sequences;
[0035] h. mapping the filtered bacterial sequences against a reference set of bacterial genomes associated with sepsis;
[0036] i. identifying bacterial species present in the sample based on the mapped sequences; j. analyzing the bacterial genome sequences for antibiotic resistance markers using statistical methods. The procedure initiates with the collection and systematic processing of blood samples from individuals. Each specimen is prepared for sequencing via a series of standard biochemical procedures.
[0037] Initially, DNA is isolated from the biological sample employing distinct biochemical and physical techniques. This DNA is then formatted into a sequencing-ready library, followed by the sequencing process.
[0038] Post-sequencing, the individual’s genomic data is converted into sequencing reads, typically stored in a FASTQ format. These reads undergo quality assessments, with substandard reads being discarded by means of trimming. The sequencing reads are meticulously filtered to exclusively retain the bacterial genome. These reads are subsequently mapped against the arbitrary number of most prevalent polymers, usually 15. The outcome of this mapping is stored in a .bam file. Concurrently, the bacterial genomic FASTQ file is analyzed using the Taxonomic classification tools and Antimicrobiotic resistance identification tools. Results of these three analyses are stored in a text file serving as a report for the clinician, upon which an effective treatment can be administered.
[0039] According to another aspect of the invention, a computer system configured to perform the afore mentioned method, is proposed. These method steps can be implemented as modules and submodules within a computer system, which may include computing devices, servers, and communication means (e.g., LAN, internet) for data exchange with other systems and databases. The computing devices and servers preferably have a CPU, GPU, RAM, non-volatile storage (e.g., hard disk), network interfaces, and peripheral devices such as a keyboard and display. Software programs and data are loaded into RAM for processing by the CPU or GPU, generating results for display, output, transfer, or storage.
[0040] The modules and submodules may be preferably implemented as computer programs or procedures in common programming languages and executed by the CPU or GPU as object or bytecode. They may also be implemented in hardware, such as integrated circuits or ROM components, enabling each device and server to function as a specialized computer. These programs may be stored on various memory media like HDD, SSD, flash drives, RAM, ROM, etc.
[0041] The computer system designed to process pathogenic cause of sepsis samples may include modules for sequencing read processing, variant calling, age-related marker analysis, model training and testing, and new sample classification. Additionally, a computer program product with computer- readable instructions can be loaded and executed to perform these operations. Such a computer program product represents another aspect of the present invention.
[0042] The computer system may either be a single system handling all computations or a server distributing tasks across several computing nodes. Each node performs specific computations and sends the results back to the server.
[0043] Examples
[0044] Example 1. Preparation of a Nanopore sequencing library.
[0045] Peripheral blood is collected from septic patient using a sterile syringe, with 10 mL drawn and transferred to EDTA tubes to prevent clotting. To separate the plasma or serum from the blood cells, the sample is centrifuged at 1500 xg for 10 minutes at 4°C. Following centrifugation, the plasma or serum is carefully pipetted into a sterile tube, ensuring minimal disturbance to the buffy coat.
[0046] Next, cell-free DNA (cfDNA) is extracted from the plasma or serum using a commercial kit QIAamp Circulating Nucleic Acid Kit. After extraction, the cfDNA is eluted in a 50 pL of nuclease-free water for subsequent analysis.
[0047] Quality control of the cfDNA is performed to ensure the integrity and quantity of the extracted DNA. Quantification is carried out using a fluorometric assay using the Qubit dsDNA HS Assay. Additionally, the integrity and fragment size of the cfDNA are assessed using an automated electrophoresis system, like the Agilent Bioanalyzer or TapeStation.
[0048] Example 2, Obtaining information about the presence of the Escherichia coli pathogen and its resistance to beta-lactam antibiotics.
[0049] After sequencing the blood sample from the septic patient, a total of 9 804 reads were obtained using Nanopore sequencing. Of these, 126 reads were identified as belonging to the pathogen Escherichia coli (E. coli). The sequencing run was carried out on a nanopore flow cell, with real-time monitoring of the data stream. The raw sequencing reads were basecalled using Guppy, resulting in a FASTQ file, and subsequent data quality control was carried out by tools FastQC and Porechop ABI. After filtering of eukaryotic reads using Minimap2 and Samtools, remaining reads were mapped using the same tools to 15 most common pathogens causing sepsis resulting in sorted BAM file. By analyzing the file using the Qualimap tool, the presence of Escherichia coli was revealed. This pathogen was detected in 126 reads out of the total 9 804, indicating a significant presence of the bacterium in the bloodstream of the septic patient. This was further confirmed by Taxonomic classification tool Kraken2. Antibiotic resistance of sequenced pathogen was analyzed using AMRFinderPlus and ResFinder and resistance to beta-lactam antibiotics was discovered.
[0050] Example 3, Reporting information on present pathogens and their antibiotic resistance.
[0051] A Computer system in this example is a standalone computing server without additional computing nodes. The server comprises of Intel Core i5 - 7300HQ 2.50 GHz, HyperX 8GB DDR4 2666 MHz CL16 FURY series, Western Digital 2TB Ultrastar DC HA210 SATA HDD and a LAN connection to external databases.
[0052] Sequencing data enters the computer system together with the human reference genome and 15 genomes of most common pathogens causing sepsis. The quality control of the sequencing reads is performed by the Sequencing Quality Control module incorporating FastQC and Porechop ABI tool, followed by mapping step of the Mapping module performed by the Minimap2 and Samtools tools. Various analyses, some of which were described in the invention summary, were performed in the Analysis module. This example results in a comprehensive report for the clinician with information on present pathogens and their antibiotic resistance.
[0053] Industrial applicability
[0054] The method and system according to the present invention can be used in various fields, including clinical research, medicine, public health, agriculture and sports science.
Claims
Claims1. A computer-implemented method for the rapid and non-invasive detection of sepsis and identification of bacterial antibiotic resistance, the method comprising:a. collecting a biological sample from a patient suspected of having sepsis; b. extracting cell-free DNA from the biological sample;c. preparing a sequencing-ready library from the extracted DNA;d. performing long-read, single-molecule sequencing on the prepared library; e. processing the sequencing data to obtain raw sequencing reads;f. performing quality control on the sequencing reads, wherein low-quality reads are removed;g. filtering the remaining sequencing reads to isolate bacterial genomic sequences; h. mapping the filtered bacterial sequences against a reference set of bacterial genomes associated with sepsis;i. identifying bacterial species present in the sample based on the mapped sequences; j. analyzing the bacterial genome sequences for antibiotic resistance markers using statistical methods.
2. The method of claim 1, wherein the quality control of sequencing reads comprises trimming the low-quality ends of the sequencing reads and removing substandard reads.
3. The method of according to any one of the previous claims, wherein the bacterial genome sequences are mapped against a reference set comprising the genomes of the most common bacterial species known to cause sepsis.
4. The method according to any one of the previous claims, further comprising analyzing the bacterial genome for the presence of antibiotic resistance genes by comparing the sequences against a database of known resistance genes.
5. The method of claim 4, wherein the antibiotic resistance genes are identified through bioinformatic analysis, wherein the analysis identifies genetic mutations or the presence of resistance genes responsible for antibiotic resistance.
6. The method according to any one of the previous claims, further comprising generating a report containing information about the identified bacterial species and their antibioticresistance profiles, wherein the report is provided to a clinician for therapeutic decisionmaking.
7. A computer system comprising computing means configured to perform the method according to any one of the previous claims.
8. A computer program comprising instructions which, if executed by a computer, ensure the implementation of the method according to any one of the claims 1 to 6.
9. A computer data medium comprising program instructions which, when executed by a computer, ensure the implementation of the method according to any one of the claims 1 to 6.