Microbial DNA analysis for disease classification
By using masked microbial reference genome and terminal motif analysis, the problems of difficult mycobacterial culture and low nucleic acid detection sensitivity in TB diagnosis have been solved, enabling rapid and accurate detection of TB.
Patent Information
- Application Number
- CN202480064070.5
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Priority Date
- 2023-10-24
- Filing Date
- 2024-10-24
- Publication Date
- 2026-05-15
AI Technical Summary
Existing technologies for TB diagnosis suffer from difficulties in mycobacterial culture and low sensitivity of microbial nucleic acid detection, resulting in unmet needs for rapid diagnosis and treatment.
By using masked microbial reference genomes and terminal motif analysis, the severity of specific microbial diseases in biological samples can be determined. This includes analyzing sequence reads and correlations of cell-free DNA molecules, and using masked microbial reference genomes for alignment with human reference genomes to improve detection accuracy.
It improves the sensitivity and accuracy of TB diagnosis, enabling rapid differentiation between TB-positive and TB-negative individuals, thus meeting the needs of rapid diagnosis and treatment.
Smart Images

Figure CN122055461A_ABST
Abstract
Description
[0001] Cross-reference to related applications
[0002] This application claims priority to U.S. Provisional Application No. 63 / 545,540, filed October 24, 2023, entitled "Analysis of Microbial DNA For Disease Classification," and is a non-provisional application of that application, the entire contents of which are incorporated herein by reference for all purposes. Background Technology
[0003] Pleural fluid refers to the fluid accumulation between the two layers of the pleura. In healthy individuals, the pleural cavity contains a small amount of fluid (approximately 10 mL to 20 mL), which includes low levels of white blood cells, proteins, and nucleic acids. Pleural effusion refers to the excessive accumulation of fluid in the pleural cavity, which can be caused by infection, malignancy, or inflammatory conditions.
[0004] Tuberculosis (TB) is one of the leading infectious diseases that cause millions of deaths each year. TB is caused by infection with a group of mycobacterial species (Mycobacterium tuberculosis complex, MTBC). They are characterized by more than 99% nucleotide similarity and identical 16S rRNA sequences at the nucleotide level with the representative pathogenic species Mycobacterium tuberculosis (MTB) (Brosch et al. Int J Med Microbiol. 2000; 290(2):143-52). Not included in MTBC and Mycobacterium leprae ( Mycobacterium leprae Mycobacteria in this group are called nontuberculous mycobacteria (NTM).
[0005] Tuberculosis (TB) diagnosis remains a challenge in disease management due to the difficulty of culturing mycobacteria. Microbial culture using saliva or other bodily fluids has long been considered the gold standard for TB diagnosis; however, positive results can take weeks, failing to meet the need for rapid TB diagnosis and treatment (Moore et al. Diagn Microbiol Infect Dis. 2005; 52(3):247-54; Chang et al. Sci Rep. 2022; 12(1):16972). To address this issue, molecular diagnostic methods for detecting mycobacterial DNA have been proposed (e.g., Xpert MTB / RIF), providing rapid results (Ismail et al. PLoS One. 2015; 10(11):e0141851). However, this method has suboptimal sensitivity (approximately 50% to 70%) due to the low levels of mycobacterial nucleic acid in the sample (Yu et al. PLoS One. 2021; 16(6):e0253658; Pan et al. PLoS One. 2021; 16(6):e0253879). Summary of the Invention
[0006] Methods, systems, and apparatus are provided for determining the level of disease caused by a specific microorganism in a biological sample of an object. The biological sample may include cell-free DNA of bacteria and cell-free DNA of the object.
[0007] In one example technique, a masked microbial reference genome can be used to determine the amount of free DNA molecules corresponding to a specific microbial species associated with a specific microbial disease. Masking can remove regions shared with one or more other species. Such masking can filter out DNA molecules that are incorrectly identified as originating from a specific microbial species, increasing the accuracy of disease severity determination.
[0008] In another example technique, terminal motifs from both the object and a cell-free DNA fragment from a specific microbial species are used. The correlation between the amount of a set of terminal sequence motifs from the object and that from the specific microbial species can be determined. For TB, the two sets of amounts are substantially more correlated for positive objects compared to negative objects.
[0009] One general aspect includes a method for analyzing a biological sample to determine the grade of a specific microbial disease in the biological sample of an object. The method may include analyzing cell-free DNA molecules from the biological sample to obtain sequence reads. A masked microbial reference genome associated with a specific microbial species may be stored. The masked microbial reference genome may be generated from a microbial reference genome of the specific microbial species. The microbial reference genome may include (1) specific regions identified as species-specific and (2) non-specific genomic regions shared with one or more other species. The masked microbial reference genome may be generated by removing the non-specific genomic regions from the microbial reference genome. The method may also include aligning the sequence reads to the masked microbial reference genome to identify a group of cell-free DNA molecules as originating from the specific microbial species. The method may also include determining the amount of this group of cell-free DNA molecules. The method may also include classifying the grade of the specific microbial disease of the object based on a comparison of the amount with a reference value.
[0010] Another general aspect includes methods for analyzing biological samples to determine the grade of a specific microbial disease in the biological sample of an object. The methods may include analyzing cell-free DNA molecules from the biological sample to obtain sequence reads. Analysis of cell-free DNA molecules may include determining terminal sequence motifs at at least one end of the cell-free DNA molecules. The methods may also include identifying a first set of cell-free DNA molecules as originating from the object by comparing the sequence reads to a human reference genome. The methods may also include identifying a second set of cell-free DNA molecules as originating from a specific microbial species associated with a specific microbial disease by comparing the sequence reads to a microbial reference genome. The methods may also include using the sequence reads of the first set of cell-free DNA molecules to determine a first quantity for each of the terminal sequence motifs in the first set of cell-free DNA molecules, thereby obtaining multiple first quantities. The methods may also include using the sequence reads of the second set of cell-free DNA molecules to determine a second quantity for each of the terminal sequence motifs in the second set of cell-free DNA molecules, thereby obtaining multiple second quantities. The methods may also include measuring correlation values between the multiple first quantities and the multiple second quantities. The methods may also include determining a classification of the specific microbial disease grade of the object based on a comparison of the correlation values to a reference value.
[0011] These and other embodiments of this disclosure are described in detail below. For example, other embodiments relate to systems, devices, and computer-readable media associated with the methods described herein.
[0012] A better understanding of the characteristics and advantages of the embodiments of this disclosure can be obtained by referring to the following detailed description and accompanying drawings.
[0013] Brief description of the attached figures
[0014] Figure 1This is an example illustration of the lungs and pleural cavity of a patient with pleural effusion.
[0015] Figure 2 Illustrations of the Mycobacterium tuberculosis complex (MTBC) and mycobacterial species are shown.
[0016] Figure 3 An overview of the clinical information of the experimental samples and subjects is displayed.
[0017] Figure 4 Example MTB whole-genome capture probe is shown.
[0018] Figure 5A The number of MTBC-derived DNA fragments in the pleural fluid sample is shown. Figure 5B Normalized MTBC abundance of MTBC-derived DNA fragments in pleural fluid samples is shown.
[0019] Figure 6 The number of DNA fragments that matched TB and non-TC species in pleural fluid samples was shown when Bowtie was used.
[0020] Figure 7 The number of DNA fragments classified as originating from MTBC, unclassified Mycobacterium genus, and nontuberculous Mycobacterium genus (NTM) is shown. Figure 7 It also shows the number of DNA fragments classified as originating from MTBC on a log-scale scale.
[0021] Figure 8A The ROC curves for MTBC abundance are shown in countries where mycobacteria are not endemic. Figure 8B The ROC curves for MTBC abundance are shown in countries where mycobacteria are endemic.
[0022] Figure 9 The first method for MTB reference genome masking is shown.
[0023] Figure 10 A second method for generating a masked MTB reference genome is shown.
[0024] Figure 11A The results are shown using the Kraken2 software for a larger set of samples: 23 TB and 55 non-TB. Figure 11B The results using the masked MTB genome are shown.
[0025] Figure 12 This is a flowchart illustrating a method for analyzing biological samples to determine the level of a specific microbial disease in a biological sample of an object, according to an embodiment of the present disclosure.
[0026] Figures 13A to 13BThe fragment size distribution of human nuclear DNA in plasma and pleural fluid according to embodiments of the present disclosure is shown. Figure 13A The distribution of fragment sizes of human nuclear DNA in plasma (blue) and pleural fluid (red) is shown. Figure 13B The cumulative frequency of human cell nuclear DNA size in plasma (blue) and pleural fluid (red) is shown.
[0027] Figure 14 The motif ranking of tetramer terminal motifs from patient plasma and pleural fluid is shown.
[0028] Figure 15A The ranking of tetramer terminal motifs of plasma nuclear DNA from the two patients was shown. Figure 15B The ranking of tetramer terminal motifs of nuclear DNA from pleural fluid cells in the two patients was shown.
[0029] Figure 16A The motif O / E ratio of plasma nuclear DNA was shown for the two patients. Figure 16B The motif O / E ratio of nuclear DNA in the pleural fluid cells of the two patients was shown.
[0030] Figure 17A The correlation matrix, representing the correlation coefficients of plasma samples, is shown. Figure 17B The correlation matrix, representing the correlation coefficients of pleural fluid samples, is shown.
[0031] Figure 18 The comparison of correlation coefficients between plasma and pleural fluid samples according to embodiments of the present disclosure is shown.
[0032] Figure 19A The correlation between human nuclear DNA and MTBC DNA in terms of the O / E ratio of tetramer terminal motifs was shown for TB-positive patients. Figure 19B The correlation between human nuclear DNA and MTBC DNA in terms of the O / E ratio of tetramer terminal motifs (excluding CGNN motifs) was shown for TB-positive patients.
[0033] Figure 20A The correlation between human nuclear DNA and MTBC DNA in terms of the O / E ratio of dimer terminal motifs (excluding CG motifs) was shown for TB-positive patients. Figure 20B The correlation between human nuclear DNA and MTBC DNA in terms of the O / E ratio of dimer terminal motifs (excluding CG motifs) was shown for TB-negative patients.
[0034] Figure 21AThis is a 2D plot showing the correlation coefficient and MTBC abundance between human nuclear DNA and MTBC DNA in pleural fluid samples. Figure 21B It is a ROC analysis that shows the performance of abundance and correlation coefficient in distinguishing between TB samples and non-TB samples.
[0035] Figure 22A The correlation between human nuclear DNA and MTBC DNA from TB-positive patients was shown regarding the frequency of dimer terminal motifs (excluding CG). Figure 22B This shows the correlation between human nuclear DNA and MTBC DNA in TB-negative patients regarding the frequency of dimer terminal motifs (excluding CG).
[0036] Figure 23A The comparison of frequency correlation coefficients between TB and non-TB group samples is shown. Figure 23B The correlation coefficients of the O / E ratios between the TB and non-TB groups are shown.
[0037] Figure 24A Principal component analysis (PCA) shows the O / E ratio of dimer terminal motifs (excluding CG motifs) of MTBC DNA. Figure 24B Principal component analysis (PCA) shows the frequency of dimer terminal motifs (excluding CG motifs) in MTBC DNA.
[0038] Figure 25 ROC curves are provided, showing the performance of machine learning models trained on MTBC 2-mer motif frequencies or motif O / E ratios in distinguishing between TB and non-TB samples.
[0039] Figures 26A to 26C The fragment size distribution of human cell nuclear DNA (blue) and MTBC (red) in the TB sample is shown.
[0040] Figure 27 This is a flowchart illustrating a method for analyzing biological samples to determine the level of a specific microbial disease in a biological sample of an object, according to an embodiment of the present disclosure.
[0041] Figure 28A A graph showing the number of MTBC fragments in TB and non-TB objects was obtained using nanopore sequencing without masking the MTB reference genome. Figure 28B A graph showing the number of MTBC fragments in TB and non-TB objects was determined using nanopore sequencing and by masking the MTB reference genome.
[0042] Figures 29A to 29F The comparison between nanopores and Illumina in terms of size and terminal motif analysis is shown.
[0043] Figure 30 A measurement system according to an embodiment of the present disclosure is shown.
[0044] Figure 31 A block diagram of an example computer system that can be used with systems and methods according to embodiments of this disclosure is shown.
[0045] the term
[0046] “ organize "This corresponds to a group of cells that work together as a functional unit. More than one type of cell can be found in a single tissue. Different types of tissues can be composed of different types of cells (e.g., liver cells, alveolar cells, or blood cells), but can also correspond to tissues from different organisms (mother and fetus) or to tissues corresponding to healthy cells and tumor cells."
[0047] “ biological samples"A biological sample" refers to any sample obtained from an object (such as a human or other animal, such as a pregnant woman, an individual with cancer or other disease, or an individual suspected of having cancer or other disease, an organ transplant recipient, or an object suspected of having a disease process involving an organ (such as the heart in a myocardial infarction, the brain in a stroke, or the hematopoietic system in anemia) and containing one or more nucleic acid molecules of interest. Biological samples can be bodily fluids, such as blood, plasma, serum, urine, vaginal fluid, fluid from hydrocele (e.g., testicular swelling), vaginal douches, pleural effusion, ascites, etc. Cerebrospinal fluid, saliva, sweat, tears, sputum, bronchoalveolar lavage fluid, breast milk, aspirated fluid from different parts of the body (e.g., thyroid, breast), intraocular fluid (e.g., aqueous humor), etc. Fecal samples may also be used. In various embodiments, biological samples enriched with cell-free DNA (e.g., plasma samples obtained via centrifugation) may contain a majority of cell-free DNA, for example, more than 50%, 60%, 70%, 80%, 90%, 95%, or 99% cell-free DNA. Centrifugation protocols may include, for example, 3,000... The liquid fraction is obtained by centrifugation at, for example, 30,000 g for 10 minutes, followed by centrifugation at 30,000 g for 10 minutes to remove residual cells. As part of the analysis of the biological sample, a statistically significant number of cell-free DNA molecules in the biological sample can be analyzed (e.g., to provide accurate measurement results). In some embodiments, at least 1,000 cell-free DNA molecules are analyzed. In other embodiments, at least 10,000, 50,000, 100,000, 500,000, 1,000,000, or 5,000,000 or more cell-free DNA molecules can be analyzed. At least the same number of sequence reads can be analyzed. Any quantity described herein can be any of the numbers listed above. Example sample sizes can include 30, 50, 100, 200, 300, 500, 1,000, 5,000, or 10,000 or more nanograms, or 1, 2, 3, 4, 5, 6, 7, 8, 9, or 10 ml.
[0048] the term" Comparison "", control sample "", Background Sample "", refer to "", Reference Sample "", normal "and" normal samples"This term can be used interchangeably to generally describe samples that do not have a specific disease condition or are otherwise healthy. In one example, a template-free control (NTC) sample with contaminating DNA can be considered a reference sample. In another example, a reference sample is a sample extracted from an uninfected object. Reference samples can be obtained from the object or from a database. A reference generally refers to a reference genome, which is used to map sequence readings obtained from sequencing samples from an object. A reference genome generally refers to a haploid or diploid genome that can be aligned and compared with sequence readings from a biological sample. For a haploid genome, there is only one nucleotide at each locus. For a diploid genome, heterozygous loci can be identified, such loci having two alleles, where either allele can allow for a match on alignment with that locus. A reference genome (e.g., by including one or more microbial genomes) can be a reference microbial genome corresponding to a specific microbial species."
[0049] “ Reference genome The “reference sequence” or “reference sequence” can be the entire genome sequence of a reference organism, one or more portions of a reference genome that may be continuous or non-contiguous, a consistent sequence of multiple reference organisms, a compiled sequence based on different components of different organisms, or any other suitable reference sequence. As an example, the reference genome / sequence length can be at least 1,000, 10,000, 50,000, 100,000, 500,000, 1,000,000, 5,000,000, 10,000,000, 50,000,000, 100,000,000, 500,000,000, one billion, or three billion nucleotides, such as the complete human genome or a duplicated human genome. The reference may also include information about reference variations known to be found in a population of organisms.
[0050] As used in this article, the phrase " healthy "This usually refers to an object in good health. Such an object shows the absence of any malignant or non-malignant disease." healthy individuals "May have other diseases or conditions unrelated to the measured condition (which is not usually considered 'healthy')."
[0051] As used herein, a “fragment” (e.g., a DNA or RNA fragment) may refer to a portion of a polynucleotide or polypeptide sequence comprising at least three consecutive nucleotides. Nucleic acid fragments may retain the biological activity and / or some characteristics of the parent polypeptide. Nucleic acid fragments may be double-stranded or single-stranded, methylated or unmethylated, intact or nicked, and complexed or uncomplexed with other macromolecules (e.g., lipid particles, proteins). Nucleic acid fragments may be linear or circular. Tumor-derived nucleic acids may refer to any nucleic acid released from tumor cells, including pathogen nucleic acids from pathogens within tumor cells. As part of a biological sample analysis, a statistically significant number of fragments may be analyzed; for example, at least 1,000 fragments may be analyzed. As other examples, at least 5,000, 10,000, 50,000, 100,000, 500,000, 1,000,000, or 5,000,000 or more fragments may be analyzed, and such fragments may be selected randomly or according to one or more criteria.
[0052] “ Sequence readings"Sequence reads" refers to a sequence of nucleotides obtained from any part or all of a nucleic acid molecule. For example, a sequence read can be a short nucleotide sequence (e.g., 20 to 150 nucleotides) sequenced from a nucleic acid fragment, a short nucleotide sequence at one or both ends of a nucleic acid fragment, or the sequencing of an entire nucleic acid fragment present in a biological sample. Sequence reads can be obtained in various ways, such as using sequencing technologies or using probes (e.g., in hybridization arrays or capture probes that can be used in microarrays) or amplification technologies (e.g., polymerase chain reaction (PCR) or linear or isothermal amplification using a single primer). Example sequencing technologies include high-throughput parallel sequencing, targeted sequencing, Sanger sequencing, sequencing by ligation, ion semiconductor sequencing, and single-molecule sequencing (e.g., sequencing using nanopores), or single-molecule real-time sequencing (e.g., from Pacific...). (Biosciences). This type of sequencing can be random sequencing or targeted sequencing (e.g., enriching such regions by using capture probes that hybridize with a specified region or by amplifying certain regions). Example probe-based techniques include real-time PCR and digital PCR (e.g., droplet digital PCR). As part of biological sample analysis, a statistically significant number of sequence reads can be analyzed, for example, at least 1,000 sequence reads can be analyzed. As other examples, at least 50... 00, 10,000, 50,000, 100,000, 500,000, 1,000,000, or 5,000,000 or more sequence readings. Furthermore, the number of sequence readings determined for embodiments of this disclosure may be at least 1,000, 5,000, 10,000, 50,000, 100,000, 100,000, 500,000, 1,000,000, or 5,000,000.
[0053] the term" Mapping "or" Alignment "Associating a sequence with a location or coordinates (e.g., genomic coordinates) in a reference (e.g., a reference genome) that has a known reference sequence, wherein the sequence is similar to the known reference sequence at that location in the reference. This can be done according to..." Mapping quality This is used to measure or report similarity. In one example of mapping quality used in this paper, the mapping quality X of a sequence relative to the reported position or coordinates in the reference indicates that the probability of the sequence mapping to a different position is no greater than 10^(-X / 10). For example, a mapping quality of 30 means that the probability of the sequence mapping to another position is less than 0.1%.
[0054] the term" The DNA of the microorganisms that cause the infection "" refers to DNA molecules derived from one or more microorganisms known to cause infection in organisms (such as humans).
[0055] Sequence readings may include "..." associated with the end of the segment. terminal sequence The terminal sequence can correspond to the outermost N bases of the fragment, such as 1 to 30 bases at the end of the fragment. If the sequence read corresponds to the entire fragment, the sequence read can include two terminal sequences. When paired-end sequencing provides two sequence reads corresponding to the ends of a fragment, each sequence read can include one terminal sequence.
[0056] “ sequence motif "Terminal motif" can refer to a short, recurring pattern of bases in a DNA fragment (e.g., a free DNA fragment). A sequence motif can appear at the end of a fragment and is therefore part of or includes the terminal sequence. "Terminal motif" (also called "terminal sequence motif") can refer to a sequence motif of a terminal sequence that may preferentially appear at the ends of DNA fragments of a particular type of tissue. A terminal motif can also appear just before or after the end of a fragment and thus still correspond to the terminal sequence. Nucleases may have a specific cleavage preference for a particular terminal motif and a second preferred cleavage preference for a second terminal motif. The number of nucleotides (nt) at the end of the fragment used for analysis can be, for example (but not limited to), 1 nt, 2 nt, 3 nt, 4 nt, 5 nt, 6 nt, 7 nt, 8 nt, 9 nt, and 10 nt. nt or higher. In some embodiments, the fragment terminal motif may be defined by one or more nucleotides spanning a position near the fragment end. The fragment terminal motif may be defined by one or more nucleotides at a genomic locus aligned around the fragment end in a reference genome. Various numbers of motifs may be used, for example, at least 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, 20, 30, 40, 50, 60, 70, 80, 90, 100, 150, 200, 250, or 256 terminal motifs.
[0057] “ Sequence motif pairs "or" Terminal motif pairs"Terminal motifs" can refer to a pair of terminal motifs on a specific DNA fragment. For example, a DNA fragment with an A at the 5' end of one strand and an A at the 5' end of the other strand can be defined as a sequence motif pair with A<>A. As another example, a DNA fragment with an A at the 5' end of one strand and a T at the 3' end of the same strand can be defined as a sequence motif pair with A<>T, which would correspond to an A<>A fragment defined using the 5' ends of both strands. Sequence motifs of other lengths can be used. Different pairings of terminal motifs can be referred to as different types of fragments. Terminal motif pairs can include terminal motifs of the same length, for example, both being monomers or both being dimers, but can also include terminal motifs of different lengths, for example, one end being a dimer and the other end consisting of monomers. Terminal motif pairs can also include, for example, one or more bases at the end of a DNA fragment, as determined by alignment with a reference genome. Such cases can be named t|A, where T appears exactly before the cleavage site at the 5' end and A appears after the cleavage site.
[0058] “ Terminal motif spectrum "Terminal motifs" can refer to the relationship of the terminal sequences (e.g., 1 to 30 bases) of cell-free DNA fragments (also called DNA segments) in a sample. Various relationships can be provided, such as the amount of cell-free DNA fragments with a specific terminal sequence (terminal motif), or the relative frequency of cell-free DNA fragments with a specific terminal sequence compared to one or more other terminal sequences. In some cases, other types of parameters (e.g., size) are used to determine terminal motif profiles. For example, terminal motif profiles can be provided in various ways that describe the amount of cell-free DNA fragments with one or more specific terminal sequences for a given size (single length or size range).
[0059] “ relative frequency "(also referred to simply as "frequency")" can refer to a proportion (e.g., percentage, fraction, or concentration). In particular, the relative frequency of a particular terminal motif (e.g., A, CG, TAG, etc.) or terminal motif pair (e.g., A<>A) can provide the proportion of free DNA fragments having that terminal motif or that particular terminal motif pair.
[0060] terminal motifs Expected frequency "The expected frequency can be determined based on a reference sequence within a reference genome region, for example, how many times a particular terminal motif appears in the reference genome region. The exact expected frequency will depend on the sequence of the region and can be normalized; for example, the size of the region can be defined as the total number of k-mer terminal motifs in that region. The expected frequency can provide information about whether the measured frequency is higher than the expected frequency, because some regions may have more CpG sites than others."
[0061] “ O / E ratio The O / E ratio refers to the ratio of the observed frequency to the expected frequency of a particular terminal motif, which can be used for downstream analysis. In the O / E ratio, O is the observed frequency (i.e., a normalized quantity) of a specific set of one or more k-mer terminal motifs. The frequency can be determined using any of the normalization techniques described herein. For example, the observed frequency can be determined as the percentage of fragments having one of that specific set of k-mer terminal motifs among all k-mer terminal motifs (e.g., tricramer terminal motifs).
[0062] the term" Sequencing depth "x" refers to the number of times a locus is covered by sequence reads aligned with that locus. Loci can be as small as nucleotides, as large as chromosome arms, or as large as the entire genome. Sequencing depth can be expressed as 50x, 100x, etc., where "x" refers to the number of times a locus is covered by sequence reads. Sequencing depth can also be applied to multiple loci or the entire genome, in which case x can refer to the average number of times the locus, haploid genome, or whole genome is sequenced, respectively. Ultra-deep sequencing can refer to a sequencing depth of at least 100x.
[0063] “ Separation value "Separation value corresponds to the difference or ratio involving two values, such as the difference or ratio of two proportions or two methylation levels. A separation value can be a simple difference or ratio. As an example, the direct ratio of x / y is a separation value, and x / (x+y) is also a separation value. Separation values can include other factors, such as a multiplication constant. As another example, a difference or ratio can be a function of the following values, such as the difference or ratio of the natural logarithms (ln) of two values. Separation values can include both differences and ratios. Separation values can be compared to a threshold to determine whether the separation between two values is statistically significant."
[0064] As the term "as used in this article" parameter "This refers to a numerical value that characterizes the numerical relationship between quantitative datasets and / or between quantitative datasets. For example, the ratio (or a function of the ratio) between a first quantity of a first nucleic acid sequence and a second quantity of a second nucleic acid sequence is a parameter. Parameters can be used to determine any classification described herein, such as any classification concerning fetal, cancer, or transplant analysis." Separation value " is a parameter (also called a measure) that provides a measure of the sample's variation between different categories (states), and therefore can be used to determine different categories.
[0065] “ Correlation valueA correlation value is an example of a separating value between two sets of values, such as the separating value between pairs of corresponding values. A set of values can form a vector, which can represent multidimensional data points. As an example, a correlation value can be a set of differences between each pair of values. Such values can be standardized, for example, by the number of pairs. A correlation coefficient is a type of correlation value, such as the Pearson coefficient. Correlation values include, but are not limited to, the Pearson correlation coefficient, Spearman's rank correlation, Phi correlation, Kendall's rank correlation, Jaccard similarity, cosine similarity, etc. Correlation values can be implicitly determined using machine learning models (e.g., clustering, PCA, SVM, or neural networks). Such models can take two sets of values and provide a score based on the correlation between the two sets of values.
[0066] As the term "as used in this article" Classification A plus sign ("+") refers to any number or other character associated with a specific characteristic of a sample. For example, the symbol "+" (or the word "positive") can indicate that a sample is classified as having a missing or amplified component. Classifications can be binary (e.g., positive or negative) or have more levels of classification (e.g., scales from 1 to 10 or 0 to 1), including probabilities. Different techniques used to determine a classification can be combined, for example by majority voting or by requiring all initial / intermediate classifications to be the same (e.g., positive), to obtain a final classification from the initial or intermediate classifications used for each of the different techniques.
[0067] the term" Cutoff value "and" threshold"Cutoff value" refers to a predetermined numerical value used in the operation. For example, a cutoff size can refer to a size that fragments exceeding which will be excluded. As another example, a threshold can be a value above or below which a specific classification applies. Any of these terms can be used in any of these situations. A cutoff value or threshold can be a "reference value" or derived from a value that represents a specific classification or distinguishes two or more classifications. A cutoff value can be predetermined with or without reference to the characteristics of the sample or object. For example, a cutoff value can be selected based on the age or sex of the test subject. A cutoff value can be selected after and based on the output of the test data. For example, a specific cutoff value can be used when the sequencing of a sample reaches a certain depth. As another example, a reference object with known classifications and measurements of one or more conditions (e.g., methylation level, statistical size value, or count) can be used to determine a reference level. Distinguishing between different conditions and / or categories of conditions (e.g., whether an object has a condition). A reference value can be chosen as a representative of a category (e.g., the mean) or a value between two clusters of the measure (e.g., chosen to obtain desired sensitivity and specificity). As another example, a reference value can be determined based on statistical simulation of a sample. Any of these terms can be used in either of these contexts. Such reference values can be determined in various ways, as those skilled in the art will understand. For example, a measure can be determined for two different groups of objects with different known categories, and a reference value can be chosen as a representative of a category (e.g., the mean) or a value between two clusters of the measure (e.g., chosen to obtain desired sensitivity and specificity). As another example, a reference value can be determined based on statistical simulation of a sample. Specific cutoff values, thresholds, reference values, etc., can be determined based on the desired accuracy (e.g., sensitivity and specificity).
[0068] “ Microbial disease levels"Disease grade can refer to the presence, quantity, extent, or severity of a disease associated with the subject, as well as the response to treatment. Examples of microbial diseases include tuberculosis (caused by the Mycobacterium tuberculosis complex, MTBC) and staphylococcal infection (caused by Staphylococcus bacteria). The grade can be zero. The subject's health status can be considered as a classification of being disease-free. Disease grades can be numbers or other markers, such as symbols, letters, and colors. Disease grades can be used in a variety of ways. For example, screening can check for the presence of a disease in someone who was previously unaware of having one. Assessment can investigate someone who has been diagnosed with a disease to monitor its progression over time, investigate the availability of treatments, etc." Effectiveness or prognosis. In one implementation, prognosis may be expressed as the probability of patient death, or the probability of disease progression after a specified period or time. Testing may refer to 'screening' or to checking whether someone has indicative characteristics of the disease (such as symptoms or other positive tests) has the disease. Diseases can be caused by various types of microorganisms, including bacteria and other microorganisms. Grading can also indicate the type of infection, such as tuberculosis, anthrax, tetanus, leptospirosis, pneumonia, cholera, botulism, and Pseudomonas infections. In some cases, disease grading refers to a condition related to an organism's response to microorganisms, including sepsis, bacteremia, and septicemia.
[0069] “ Machine learning models"(ML model)" can refer to a software module configured to run on one or more processors to provide categorical or numerical characteristics of one or more samples. An ML model can include various parameters (e.g., functions for coefficients, weights, thresholds, functional attributes, such as activation functions). As an example, an ML model can include at least 10, 100, 1,000, 5,000, 10,000, 50,000, 100,000, or one million parameters. Sample data (e.g., training samples) can be used to generate an ML model to make predictions on test data. Various numbers of parameters can be used. The training samples, for example, at least 10, 100, 1,000, 5,000, 10,000, 50,000, 100,000, or at least 200,000 training samples. One example is an unsupervised learning model. Another example type of model is supervised learning that can be used with embodiments of this disclosure. Example supervised learning models may include different methods and algorithms, including analytical learning, statistical models, artificial neural networks (e.g., including convolutional layers and / or variable layers), boosting (meta-algorithms), Bayesian statistics, etc. Statistics, case-based reasoning, decision tree learning, inductive logic programming, Gaussian process regression, genetic programming, data grouping, kernel estimator, learning automata, learning classification systems, minimum information length (decision trees, decision graphs, etc.), multilinear subspace learning, naive Bayes classifier, maximum entropy classifier, conditional random fields, nearest neighbor algorithm, probabilistic approximate correct learning (PAC), ripple down rule, knowledge acquisition methods, symbolic machine learning algorithms, sub-symbolic machine learning algorithms, minimum complexity machine (MCM), random forest, ordered classification, data preprocessing, handling imbalanced datasets, statistical relation learning, or Proaftn (a multi-criteria classification algorithm), or a combination of these types. Models may include linear regression, logistic regression, deep recurrent neural networks (e.g., Long Short-Term Memory, LSTM), hidden Markov models, etc. Supervised learning models can be trained using various methods, including Hidden Models (HMM), Linear Discriminant Analysis (LDA), k-means clustering, density-based spatial clustering with noise (DBSCAN), random forest algorithms, support vector machines (SVM), or any model described herein. Various cost / loss functions and optimization techniques can be used to define the error from known labels (e.g., least squares and absolute difference from known classifications), such as backpropagation, steepest descent, conjugate gradients, and Newton and quasi-Newton techniques.
[0070] The terms "about" or "approximately" may mean within an acceptable range of error for a particular value as determined by those skilled in the art, the range of error depending in part on how the value is measured or determined, i.e., limitations of the measurement system. For example, according to practice in the art, "about" may mean within one or more standard deviations. Alternatively, "about" may mean a range of up to 20%, up to 10%, up to 5%, or up to 1% of a given value. Alternatively, particularly with respect to biological systems or methods, the terms "about" or "approximately" may mean within a certain order of magnitude of the value, within 5 times, or more preferably within 2 times. If a particular value is described in this application and claims, unless otherwise stated, it should be assumed that the term "about" means within an acceptable range of error for that particular value. The term "about" may have the meaning as commonly understood by those skilled in the art. The term "about" may mean ±10%. The term "about" may mean ±5%.
[0071] As used in this disclosure, the term "ROC" or "ROC curve" can refer to a receiver operating characteristic curve. An ROC curve can be a graphical representation of the performance of a binary classifier system. For any given method, an ROC curve can be generated by plotting sensitivity and specificity at various threshold settings. The sensitivity and specificity of a method for detecting the presence of tumors in a subject can be determined at various concentrations of tumor-derived DNA in a plasma sample of the subject. Furthermore, by providing at least one of three parameters (e.g., sensitivity, specificity, and threshold settings), an ROC curve can determine the value or expected value of any unknown parameter. Unknown parameters can be determined using a curve fitted to the ROC curve. For example, by providing the concentration of tumor-derived DNA in the sample, the expected sensitivity and / or specificity of the detection can be determined. The term "AUC" or "ROC-AUC" can refer to the area under the receiver operating characteristic curve. This metric, taking into account the sensitivity and specificity of the method, can provide a measure of the diagnostic utility of the method. ROC-AUC can range from 0.5 to 1.0, where values closer to 0.5 may indicate limited diagnostic utility (e.g., lower sensitivity and / or specificity), while values closer to 1.0 indicate greater diagnostic utility (e.g., higher sensitivity and / or specificity). See, for example, Pepe et al, “Limitations of the Odds Ratio in Gauging the Performance of a Diagnostic, Prognostic, or Screening Marker,” Am. J. Epidemiol 2004, 159 (9): 882-890, which is incorporated herein by reference in its entirety. Additional methods for characterizing diagnostic utility include the use of likelihood functions, odds ratios, information theory, predicted values, calibration (including goodness of fit), and reclassification measures. For example, examples of these methods are outlined in Cook’s “Use and Misuse of the Receiver Operating Characteristics Curve in Risk Prediction”, Circulation 2007, 115:928-935, which is incorporated herein by reference in its entirety.
[0072] Where a range of values is provided, it should be understood that, unless the context explicitly specifies otherwise, each intermediate value between the upper and lower limits of that range, accurate to the tenths of the lower limit unit, is also specifically disclosed. Embodiments of this disclosure cover each smaller range between any stated values or intermediate values within the stated range and any other stated values or intermediate values within the stated range. The upper and lower limits of these smaller ranges may independently include or exclude those ranges (e.g., the range may be greater than or less than a specified number), and each range including any boundary, no boundary, or two boundaries within a smaller range is also covered by this disclosure, subject to any specifically excluded boundaries within the stated range. Where a stated range includes one or both boundaries, ranges excluding any one or both of those included boundaries are also included in this disclosure.
[0073] Standard abbreviations can be used, such as bp, base pair; kb, kilobase; pl, picoliter; s or sec, second; min, minute; h or hr, hour; aa, amino acid; nt, nucleotide, etc.
[0074] Unless otherwise defined, all technical and scientific terms used herein have the same meaning as commonly understood by one of ordinary skill in the art to which this disclosure pertains. While any methods and materials similar to or equivalent to those described herein may be used in practice and testing of embodiments of this disclosure, some potential and exemplary methods and materials may now be described.
[0075] Detailed description
[0076] Metagenomic next-generation sequencing (mNGS) offers an unbiased approach to detecting a broad range of pathogens in clinical samples and has been proposed for the diagnosis of infectious diseases (Oreskovic et al. J. Clin. Microbiol. 2021; 59(8):e0007421; Oreskovic et al. Int. J. Infect. Dis. 2021;112, 330–337). However, metagenomic sequencing analyses for tuberculosis diagnosis are again limited by low concentrations of MTB and a background of contaminated nontuberculous mycobacterial DNA, which may have sequences very similar to pathogenic MTB (Chang et al. SciRep. 2022; 12(1):16972). Such problems can also affect other diseases and microorganisms other than MTB.
[0077] To filter out measurements associated with nontuberculous mycobacterial DNA, some implementations can mask genomic regions shared with other microorganisms, allowing measurements of DNA fragments to be more accurately attributed to a specific target microorganism. Shared regions can be identified by comparing reference genomes of different species, for example, by segmenting a species' reference genome using a K-mer and then comparing the K-mer with one or more reference genomes of other species (alignment).
[0078] Additionally or alternatively, fragment terminal motifs of mycobacterial and host-derived nucleic acids are used to detect specific microbial diseases. Metrics determined based on these fragment terminal motifs can be used to differentiate pathogen-derived signals from environmental contamination or microbial classification errors in similar microbial genome sequences, further reducing false positive rates while maintaining high assay sensitivity. In summary, this innovative approach combining MTB capture probes and nucleic acid terminal motif analysis offers promising improvements in tuberculosis diagnosis.
[0079] In some implementations, a whole-MTB genome capture probe set can be used. This set of whole-MTB genome capture probes can substantially enrich mycobacterial nucleic acid fragments in the sample. A novel analytical method has enabled highly sensitive detection of MTBC-derived sequences in pleural fluid samples from patients with tuberculous pleurisy.
[0080] I. Lung Sample
[0081] Figure 1 This is an example illustration of the lungs and pleural cavity of a patient with pleural effusion according to an embodiment of this disclosure. Pleural fluid 110 refers to the fluid accumulation located between the two layers of pleura. In healthy individuals, the pleural cavity contains a small amount of fluid (approximately 10 to 20 mL), which contains low levels of leukocytes, proteins, and nucleic acids. Pleural effusion refers to the excessive accumulation of fluid in the pleural cavity, which can be caused by infection, malignancy, or inflammatory conditions.
[0082] Pleural fluid is one example of a biological sample that can be used to determine the classification of a specific microbial disease level. Other examples are provided, for instance, in the terminology section. Saliva, cerebrospinal fluid, urine, and peritoneal dialysis fluid may be used, for example.
[0083] II. Mycobacterium tuberculosis complex (MTBC)
[0084] Mycobacteria can cause tuberculosis (TB), but not all mycobacteria cause TB. Therefore, populations with a background of non-tuberculous mycobacterial DNA may make accurate diagnosis of TB difficult. Other microorganisms and their associated diseases may also present similar challenges.
[0085] A. Examples of Mycobacterial Species
[0086] Figure 2 Illustrations of the Mycobacterium tuberculosis complex (MTBC) and mycobacterial species according to embodiments of this disclosure are shown. The Mycobacterium tuberculosis complex (MTBC) refers to a group of genetically related mycobacterial species that can cause tuberculosis in humans or other animals. MTBCs belong to the mycobacterial family and have genomic sequences similar to other species that do not cause tuberculosis. Therefore, mycobacterial species can be divided into two categories: MTBCs and non-tuberculous mycobacteria (NTMs). Specifically, MTBCs are characterized by more than 99% nucleotide similarity to the representative pathogenic species Mycobacterium tuberculosis (MTB) and have the same 16S rRNA sequence (Brosch et al. Int J Med Microbiol. 2000; 290(2):143-52). When an individual has tuberculosis, the bacterial burden of DNA from MTB bacteria may be low.
[0087] B. Example Technology
[0088] Typically, when analyzing pleural fluid and mycobacterial tuberculosis (MTBC) for diagnosis, biopsies of the patient's lung membranes are collected for microbial culture. Microbial culture uses saliva or other bodily fluids for TB diagnosis, but positive detection can take several weeks, which is insufficient for rapid TB diagnosis and treatment. Recently, studies have been conducted on obtaining pleural fluid and measuring Mycobacterium tuberculosis DNA, but the sensitivity or levels of mycobacterial DNA have been shown to be low. Additionally, molecular diagnostic methods for detecting mycobacterial DNA (e.g., Xpert MTB / RIF) have been proposed to provide rapid test results. However, this method has suboptimal sensitivity due to the low levels of mycobacterial nucleic acids in the sample, and it can potentially be improved.
[0089] 1. Targeted Approach
[0090] To detect and analyze MTBC in pleural fluid, pleural fluid samples were collected from 20 patients with TB infection (TB group) or patients without TB infection (non-TB group). Pleural fluid samples were collected using a needle inserted into the patient's pleural fluid, with volumes ranging from 10 mL to 20 mL aspirated. MTB DNA molecules were enriched from a pleural fluid DNA library for targeted sequencing of the pleural fluid samples. In some embodiments, enrichment was performed using a hybridization capture probe system. MTB capture probes were designed to cover the entire bacterial genome. In some embodiments, as a reference, capture probes targeting human autosomal regions were also included in the capture reaction.
[0091] Figure 3An overview of the clinical information of the experimental samples and subjects according to the embodiments of this disclosure is shown. Of the 20 patients, 6 had tuberculous pleurisy (TB group), confirmed by positive TB culture of pleural tissue and / or TB PCR of pleural fluid. Fourteen patients without tuberculous pleurisy (non-TB group) were recruited as negative controls.
[0092] To address issues related to low levels of mycobacterial nucleic acid in samples or low abundance of MTBC DNA in pleural fluid, MTB whole-genome capture probes can be used.
[0093] Figure 4 An example MTB whole-genome capture probe according to an embodiment of this disclosure is shown. The required number of probes can be determined by dividing the target genome size by the probe size (e.g., 5 Mb bp divided by 80 bp). As an example, the probe may comprise 80 to 120 base pairs. The probe / primer is used to amplify the target MTB genome.
[0094] In some implementations, since sequencing depth can influence the number of TB reads detected, to normalize MTBC abundance against background human DNA, human whole-exome targeted capture probes can be mixed with MTBC probes, for example, at a specified concentration ratio prior to the experiment. Various concentration ratios can be used, such as at least 1000:1, 500:1, 200:1, 100:1, 50:1, or 10:1, where the amount of MTBC probes is higher than the amount of human genome probes.
[0095] like Figure 4 As shown, DNA fragment 410 with aptamers hybridizes with probe 420 for targeted capture. PCR amplification and / or sequencing can be performed as needed. Targeted capture of the MTB genome can focus on various sizes of the reference genome, for example, regions with genome sizes of approximately 0.5 Mb, 1 Mb, 2 Mb, 3 Mb, 4 Mb, or 5 Mb. Therefore, some implementation schemes can be used for targeted capture sequencing of pleural fluid cfDNA or other free samples.
[0096] In some implementations, CRISPR-based enrichment strategies for targeted sequencing (e.g., CRISPR / dCas9-based systems) can be used.
[0097] 2. Other technologies
[0098] In other implementations, whole-genome or random sequencing can be used. Therefore, targeting techniques are not required. Furthermore, for targeting techniques, PCR or other amplification techniques can be used with or without sequencing. For example, digital PCR or real-time PCR can be used in at least some implementations. Various types of amplification and / or sequencing techniques can be used.
[0099] III. Addressing the difficulties in detecting mycobacterial DNA abundance in pleural fluid
[0100] The abundance of MTBC fragments may be inaccurate when non-TB mycobacterial DNA is present.
[0101] A. Abundance of MTBC
[0102] The abundance of MTBCs can be determined by aligning sequence reads to different reference genomes, and the number of fragments aligned with MTBC species can be counted. Various techniques and standards can be used to determine how to assign sequence reads to specific species and / or genera. These techniques include taxonomic techniques such as Kraken, Kraken2, Megablast, Centrifuge, KrakenUniq, and MetaPhlAn. Additionally, alignment tools (e.g., bowtie2, bowtie, bwa, soap, and minimap2) can be used alone in conjunction with a cutoff value for mapping quality.
[0103] 1. Taxonomic labels (e.g., using Kraken2)
[0104] Kraken2 is a system for assigning taxonomic tags to short DNA sequences, typically obtained through metagenomic studies. A bioinformatics analysis workflow analyzes short DNA sequences from bacteria to determine the genus and species levels for each sequence read. Taxonomic tags can then be used to determine the abundance of MTBCs or other specific microbial species.
[0105] Taxonomic classification software can use databases of different microorganisms (e.g., different reference genomes of different microorganisms). Reference genomes in microbial databases come from different microbial species. Sequences are aligned at the species level. However, sequences may be aligned with the same mapping quality to multiple reference genomes of different species within the same genus. In this case, the sequences may only be classified and labeled at the genus level.
[0106] In another example, a top-down approach can be used to first attempt to align sequence readings to a genus-level reference genome (called the genus reference genome). The software can then attempt to align the readings to one or more other reference genomes further down the taxonomic tree to see if the readings can be mapped to a reference genome for a specific species (called the species reference genome). If the mapping quality improves, the species can be identified. For example, if the sequence has sufficient specificity, the DNA fragment can be classified as a specific species. However, if the sequence aligns to multiple species with the same quality, it indicates that there is no specific, unique alignment. In this case, the sequence can be classified, for example, at the genus level or a higher level.
[0107] Therefore, after targeted capture and sequencing, for example, as regarding Figure 4 The DNA fragments discussed can be aligned to a reference genome (e.g., the MTBC genome) and the human genome. DNA fragments aligned to the MTBC genome are defined as MTBC fragments. DNA fragments aligned to mycobacterial genus but not assigned to any further level (e.g., species level) are defined as unclassified mycobacterial fragments. DNA fragments aligned to nontuberculous mycobacterial species or other nontuberculous genera within the mycobacterial family are referred to as nontuberculous mycobacterial fragments. The number of MTBC DNA fragments can be counted and further normalized by the number of human DNA fragments detected in the same sample, for example, MTBC abundance (readings per million reads, RPM).
[0108] Figure 5A Multiple MTBC-derived DNA fragments were shown in the pleural fluid sample. Figure 5B Normalized MTBC abundance of MTBC-derived DNA fragments in pleural fluid samples is shown. For Figure 5B Standardization is achieved by the number of human DNA fragments detected in the same sample.
[0109] Figures 5A to 5B The data were based on samples from 20 patients. Of these 20 patients, 6 had TB infection and 14 did not. Since both TB culture and TB PCR results were available, the corresponding results were used together as the gold standard. Commercially available PCR was used to obtain TB PCR results, using a single-region marker. Figures 5A to 5B As shown, although MTBC abundance in the TB group is generally higher than that in the non-TB group, there is overlap between the two groups that may be caused by high background non-TB data. Figure 5A As shown, nearly 100 MTBC fragments can be detected in four non-TB group samples.
[0110] 2. Bowtie2
[0111] As another example of genus / species classification, the Bowtie alignment tool was used. Sequence reads were aligned to the human reference genome and the MTBC reference genome. Mapping quality was determined using a mapping quality cutoff of 30. Mapping quality is the probability that a read will align to an incorrect position (potentially adjusted based on base calls, e.g., via software like Phred). A mapping quality of 30 can be equated to a 0.1% alignment error probability. The conversion from mapping quality (MQ) to alignment error rate (E) can be based on the formula: E = 10 -MQ / 10 Various cutoff values can be used, such as 5, 10, 15, 20, 30, or higher. A larger sample group was analyzed: 23 cases of TB and 55 cases of non-TB.
[0112] Figure 6 The number of DNA fragments in pleural fluid samples aligned with MTBC and non-MTBC species using Bowtie2 is shown. It can be seen that there is considerable overlap between TB cases and non-TB cases. This significant overlap can be attributed to the genetic similarity of the MTB genome to other microbes. Even though sequence reads can be mapped to microbacterial tuberculosis genomes with very high mapping quality, a large number of misclassified fragments still exist.
[0113] B. Reading counts for each object (MTBC and nontuberculous mycobacteria)
[0114] For different microorganisms, the Kraken2 results were further analyzed on a per-subject basis. The results showed that there was high background nontuberculous mycobacterial DNA contamination in the samples, which might have been misclassified as MTBC sequences using the Kraken2 method.
[0115] Figure 7 Figure 710 shows the number of DNA fragments classified as originating from MTBC, unclassified Mycobacterium, and nontuberculous Mycobacterium (NTM). As previously discussed, sequencing reads from 20 patients could be aligned to the human genome. Reads that could not be aligned to the human genome were then aligned to a microbial database or a valid microbial reference genome. The microbial database includes different TB species. Figure 710 shows the number of MTBC fragments detected in each sample without standardization. Figure 710 shows that, generally, TB samples have a higher TB ratio (i.e., a higher number of fragments aligned to the MTBC reference genome).
[0116] Figure 710 shows the counts of DNA fragments derived from unclassified *Mycobacterium* and other nontuberculous *Mycobacterium* genera. The four non-TB samples with high TB abundance (marked with black arrow 730) tended to have more nontuberculous *Mycobacterium* DNA fragments. The high detection load of MTBC DNA may be due to high background nontuberculous *Mycobacterium* DNA contamination in the four samples, which was misclassified as MTBC sequences due to high genomic similarity, as illustrated in one study (Chang et al. Sci Rep. 2022; 12(1):16972). Readings that could be aligned to *Mycobacterium* genera but not further to species were indicated as “unclassified *Mycobacterium*”.
[0117] Figure 720 shows the log-scale number of DNA fragments classified as originating from MTBCs. Figure 720 shows MTBC abundance normalized by the number of detectable human reads. Overlap can be observed. Similar to Figure 710, TB samples tend to have higher MTBC abundance in terms of the number of MTBC fragments when compared with non-TB samples.
[0118] A. Non-local controls and local controls
[0119] In countries where mycobacteria are endemic, the accuracy of using MTBC abundance to detect TB is severely reduced, as shown in Chang et al. Sci Rep. 2022; 12(1):16972.
[0120] Figure 8A The ROC curves for MTBC abundance are shown in countries where mycobacteria are not endemic. Figure 8B The ROC curves for MTBC abundance are shown in countries where mycobacteria are endemic.
[0121] like Figure 8B As shown, due to the low abundance of MTBC and the high sequence similarity between mycobacterial genomes, the nontuberculous mycobacterial background cannot be distinguished from true disease signals. In this study, plasma, urine, and oral swab samples were collected from patients in different regions, namely endemic and non-endemic TB areas. Chang performed whole-genome sequencing.
[0122] Chang's inferential diagnostic performance is limited by the low burden of Mycobacterium tuberculosis and the background of non-tuberculous mycobacterial DNA. Local controls have been shown to confound with genuine tuberculosis samples from local areas. Performance is affected by the background of non-tuberculous mycobacterial DNA from local controls.
[0123] IV. Masking of shared areas in MTBC
[0124] The above description illustrates the difficulty of distinguishing contaminated readings (e.g., non-TB mycobacterial DNA fragments) when making comparisons.
[0125] A. Identify shared areas
[0126] The genome of Mycobacterium tuberculosis (MTB) shares high sequence similarity with the genomes of other non-tuberculous mycobacterial species or other bacterial genera. This can introduce taxonomic misclassification, especially when sequencing reads are short. To address this issue, we introduce masking techniques, which can be implemented in various ways. Masking techniques can identify and mask non-MTB-specific genomic regions, preventing non-MTB reads from being misclassified as originating from MTB. Masking can be performed on any targeted microbial genome, where background DNA fragments from similar microorganisms are filtered out.
[0127] To identify shared regions, the target microbial genome (e.g., the MTB reference genome) is compared to a background reference genome (the non-target microbial genome). Regions from one genome can be compared to regions from another genome to identify shared regions. Two example methods are described below.
[0128] Figure 9 A first method for masking the MTB reference genome is shown. A set of overlapping K-mers 905 of length K are generated from the target microbial genome 902 to be analyzed (i.e., the MTB). For example, the first method, such as using a sliding window of length K, can cut the MTB genome reference into overlapping K-mers. Various K values can be used, such as 5, 10, 15, 20, 22, 24, 26, 28, 30, or 32 base pairs. Other values can be used, such as any value in the range of 20 to 35 or lower or higher.
[0129] In step 910, these K-mers are aligned with reference genomes of multiple non-targeted microbial species (e.g., bacterial genomes other than those of the Mycobacterium tuberculosis complex). The K-mers in the first group 912 are those that can be aligned with these bacterial genomes, while the K-mers in the second group 914 are those that cannot be aligned with those genomes.
[0130] In step 920, the unmapped K-mers (MTB-specific K-mers) are re-aligned with the MTB reference genome to identify MTB-specific regions, or more generally, target-specific regions potentially for applications other than TB.
[0131] In step 930, regions in the MTB reference genome not covered by any K-mers can be masked with the "N" character. Alternatively, aligned (mapped) K-mers can be re-aligned to directly identify shared non-specific regions.
[0132] Figure 10 A second method for generating a masked MTB reference genome is shown.
[0133] In step 1010, a set of overlapping K-mers 1005 of length K are generated from the non-targeted microbial genome 1001.
[0134] In step 1020, these K-mers are aligned with the genome of the target microorganism to be analyzed (e.g., MTB). Regions in the MTB reference genome covered by K-mers are masked with an "N" character (step 1030). Alternatively, unaligned (unmapped) K-mers can be re-aligned to directly identify shared non-specific regions.
[0135] The masked MTB reference genome 1040 can be used to definitively identify MTB-derived sequences. Confounding sequences are masked by a high penalty for fuzzy alignments on the masked regions. Alignment tools used here include, but are not limited to: bowtie2, bowtie, bwa, soap, minimap2, etc. For alignment tools, if we re-align the readings against the masked MTB reference genome 1040, the readings cannot be aligned with regions containing the N character because only MTB-specific regions or, more generally, target-specific regions for other microbial / disease targets are used.
[0136] For either method, aligned or unaligned K-mers can be used to identify target-specific regions. Depending on whether the K-mers originate from the target genome or the background non-target genome, aligned or unaligned reads can be used to identify target-specific or non-target-specific regions. If a target-specific region (K-mer) is identified, the remaining regions (K-mers) can be masked as non-target-specific regions.
[0137] B. Comparison
[0138] Figure 11A The results are shown using Kraken2 software on a large sample: 23 cases of TB and 55 cases of non-TB. Significant overlap was observed in the number of identified MTBC fragments, which is consistent with... Figure 5A resemblance.
[0139] Figure 11BResults using the masked MTB genome are shown. The abundance of MTBC fragments is indicated by the number of MTBC fragments, determined by aligning sequence reads to MTBC-specific regions using Bowtie2. A mapping quality of 30 or higher is used. Sequences with a mapping quality of 30 or higher are maintained. Those skilled in the art will understand that other mapping quality values can be used, for example, depending on the desired sensitivity and specificity. And other alignment software with corresponding thresholds for the mapping quality determined for the specific alignment software used can be used.
[0140] It is evident that using only one sample to identify a non-zero (only one) number of MTBC fragments results in better separation. Therefore, 100% accuracy can be achieved. The Bowtie2 alignment tool was used, but any alignment tool, as understood by those skilled in the art, can also be used. Furthermore, the abundance of microbial DNA associated with specific microbial diseases can be normalized.
[0141] C. Using a masked reference genome
[0142] Figure 12 This is a flowchart illustrating a method 1200 for analyzing biological samples to determine the level of a specific microbial disease in the target biological sample. Specific microbial diseases can be associated with specific microbial species (e.g., bacterial species). For example, Mycobacterium tuberculosis is associated with TB, and Staphylococcus aureus is associated with staphylococcal infection. Method 120 can filter out sequence reads that may originate from similar species but not from the target microbial species, thereby removing noise and achieving improved accuracy.
[0143] Method 1200 and any of the methods described herein can be performed wholly or partially using a computer system comprising one or more processors configured to perform the steps. Therefore, some embodiments relate to computer systems configured to perform the steps of any of the methods described herein, potentially having different components that perform the respective steps or groups of steps.
[0144] At box 1210, multiple cell-free DNA molecules from a biological sample are analyzed to obtain sequence reads. Cell-free DNA molecules can be analyzed by receiving the corresponding sequence reads and analyzing them via a computer. In any of the methods described in this disclosure, various techniques can be used for such analyses, and may include performing assays. For example, sequencing can be used for analysis, such as high-throughput parallel sequencing, targeted sequencing, and single-molecule sequencing (e.g., using nanopore sequencing or using real-time single-molecule sequencing (e.g., from Pacific Biosciences)). In some cases, capture probes that bind to a portion or the entire genome of a microorganism are used to enrich DNA molecules from the microorganism in the biological sample. Example PCR techniques include real-time PCR and digital PCR (e.g., droplet digital PCR). Analysis may include physical steps of performing such assays and receiving measurement data obtained from such assays, or may only include receiving measurement data.
[0145] In some implementations, targeted sequencing may use capture probes for a microbial reference genome and capture probes for a human reference genome, wherein the concentration of capture probes for the microbial reference genome is higher than that for the human reference genome, for example, at the ratio described herein.
[0146] At box 1220, a masked microbial reference genome of a specific microbial species associated with a specific microbial disease is stored. The masked microbial reference genome can be generated from a microbial reference genome of the specific microbial species. The microbial reference genome may include (1) specific regions identified as species-specific and (2) non-specific genomic regions shared with one or more other species. Specific and non-specific regions can be identified as described herein, for example, for Figure 9 and Figure 10 And the corresponding description. A masked microbial reference genome can be generated by removing non-specific genomic regions from a microbial reference genome. One or more other species may include the object and / or other microorganisms that may have a similar reference genome, for example, 10%, 20%, 30%, 40%, 50%, 60%, 70%, 80%, or 90% similarity.
[0147] At box 1230, sequence reads are aligned to a masked microbial reference genome to identify a group of multiple cell-free DNA molecules as originating from a specific microbial species. As those skilled in the art will understand, any alignment tool (e.g., Bowtie, bwa, etc.) can be used. When the specific microbial disease is TB, the specific microbial species could be MTBC. Other microbial diseases may be associated with another target microbial species.
[0148] Aligning cell-free DNA molecules can include determining genomic locations within a reference genome. For example, one or more sequence reads of a DNA molecule (e.g., paired reads at the ends or reads of the entire molecule) can be aligned or attempted to be aligned using any of a variety of alignment techniques understood by those skilled in the art. Sequence reads can be aligned with multiple microbial reference genomes in a taxonomic tree to identify the group of cell-free DNA molecules as originating from a specific microbial species, where multiple microbial reference genomes include masked microbial reference genomes.
[0149] Alignment can be with some or all of a masked microbial reference genome. Alignment software with output mapping quality can be used to align sequence reads with a masked microbial reference genome. When the mapping quality is greater than a threshold, such as 30 or other values described herein, the sequence reads are identified as originating from a specific microbial species.
[0150] As another example, probe-based techniques can identify DNA molecules as originating from specific locations, for example, by emitting a specific color against a specific probe corresponding to a specific genomic location. Location determination can be applied to some or all of a reference genome, for example, if only a portion of the genome is analyzed. As an example, the amount of genome analyzed can be greater than 0.01%, 0.1%, 1%, 5%, 10%, or 50%. Such analyses can be performed using other methods described herein.
[0151] At box 1240, the quantity of a set of multiple free DNA molecules is determined. The quantity can be an absolute quantity or a normalized quantity. For example, the quantity can be normalized by the total number of reads obtained, such as the number of reads per million. Another example is the use of normalization, such as the number of sequence reads identified as originating from the object by comparison with a human reference genome. Such reads from the object can be nuclear or mitochondrial, which can be determined using the corresponding reference genome.
[0152] At box 1250, the classification of the subject's specific microbial disease level is determined based on a comparison of the quantity with a reference value. Reference values can be selected using measurements obtained from one or more reference samples of known classifications (e.g., disease-positive or disease-negative). For example, reference values can be determined using a first batch of training samples from subjects known to have the specific microbial disease and a second batch of training samples from subjects known not to have the specific microbial disease. Figure 11B This displays a measurement chart for such a reference sample. Example reference values can be 2, 3, 4, or 5, for... Figure 11B The training set shown will provide 100% accuracy.
[0153] As described above, one or more non-specific genomic regions can be identified in various ways. For example, method 1200 can identify non-specific genomic regions shared with one or more other species by comparing a microbial reference genome of a specific microbial species with one or more other reference genomes of one or more other species. In some embodiments, identifying non-specific genomic regions may include dividing one or more other reference genomes into a set of K-mers and aligning this set of K-mers with the microbial reference genome to identify the non-specific genomic region. As an example, K can be 20 to 35. The non-specific genomic region may correspond to a subset of this set of K-mers aligned with the microbial reference genome. That is, a subset (part) of the K-mers aligned with the microbial reference genome.
[0154] In other embodiments, identifying one or more non-specific genomic regions includes dividing the microbial reference genome into a set of K-mers and aligning these K-mers with one or more other reference genomes to identify the non-specific genomic regions. As an example, K can be 20 to 35. In some embodiments, the non-specific genomic regions may correspond to a subset of the K-mers aligned with one or more other reference genomes. In other embodiments, non-specific genomic regions are identified as regions that do not correspond to a subset of the K-mers not aligned with one or more other reference genomes. That is, specific regions can be identified, and non-specific regions can be identified as the remainder of the microbial reference genome.
[0155] V. Terminal motif correlation between host and microbial DNA
[0156] In addition to using the abundance of DNA fragments from the target microbial genome, or as an alternative, terminal sequence motifs (also known as end motifs) can be used. A terminal motif corresponds to a sequence at either end of a DNA fragment, such as the 5' end of a dimer fragment. Other numbers of bases can be used, and one or both strands can be used. The terminology section provides a further detailed description of terminal sequence motifs.
[0157] The following description shows the amount (e.g., relative frequency, such as rank) of terminal motifs of host DNA (e.g., human DNA, cell nucleus, and / or mitochondria) associated with the target microbe DNA in a sample (e.g., a pleural sample), and that this can be used to indicate that the host has a disease associated with the target microbe.
[0158] A. Fragmentation analysis of human DNA in pleural fluid and plasma
[0159] The fragmented terminal motif characterization of pleural fluid cfDNA has not yet been investigated. Whole-genome sequencing (Illumina platform) was performed on 14 pairs of pleural fluid and plasma samples from 7 patients with pleural effusion. For each sample, at least 30 million DNA fragments were sequenced.
[0160] 1. Fragment size of human cell nuclear cfDNA in pleural fluid.
[0161] Figures 13A to 13B The fragment size distribution of human cell nuclear DNA in plasma and pleural fluid according to embodiments of the present disclosure is shown. Figure 13A The size distribution of human nuclear DNA fragments in plasma 1320 (blue) and pleural fluid 1310 (red) is shown. Figure 13B The cumulative frequency of human cell nuclear DNA size is shown in plasma 1340 (blue) and pleural fluid 1330 (red).
[0162] like Figure 13A As shown, pleural fluid samples had a higher frequency or proportion of short cfDNA than plasma samples. Furthermore, pleural fluid samples exhibited a lower level of 166 bp peaks, which were predominant in plasma samples. Specifically, plasma samples showed higher frequency peaks targeting the 100 bp to 200 bp range. Pleural fluid samples exhibited a more pronounced 10 bp periodic distribution pattern, with frequency peaks preceding the highest 166 bp. Figure 13B This displays the cumulative frequency of human cell nuclear DNA in plasma and pleural fluid samples. (Compared to...) Figure 13A Similarly, pleural fluid samples had a higher accumulation frequency of short cfDNA fragments than plasma samples.
[0163] 2. Terminal motif analysis of human cell nuclear cfDNA in pleural fluid
[0164] Figure 14 The motif ranking of tetramer terminal motifs from patient plasma and pleural fluid is displayed. The ranking uses the amount (e.g., relative frequency) of DNA fragments having each of a set of terminal motifs, and then sorts the terminal motifs by amount (ranking). Other exemplary amounts of terminal motifs may be relative frequency (e.g., observed frequency) and the O / E ratio of observed frequency to expected frequency. Figure 14 Use the ranking determined by the O / E ratio.
[0165] A set of parameter values (e.g., magnitude, frequency, O / E ratio, rank, etc.) for the terminal motifs of a given sample can comprise the terminal motif spectrum of that sample. Such a terminal motif spectrum can be a vector representing multidimensional data points and can be compared, for example, as... Figure 14The comparisons can be shown or otherwise made. Comparisons of terminal motif spectra can provide correlation values, such as the distance between the corresponding vectors representing the terminal motif spectra.
[0166] The motif ranking of tetramers can be determined by analyzing the terminal motifs of human cell nuclear cfDNA in paired plasma and pleural fluid samples. For example... Figure 14 As shown, the ranking of tetrameric terminal motifs (256 in total) based on the O / E ratio is compared in plasma (x-axis) and pleural fluid (y-axis) samples from the same patient. The O / E ratio is the ratio of the observed frequency to the expected frequency of a given terminal motif. In the O / E ratio, O is the observed frequency of a particular set of terminal motifs measured in the sequenced DNA fragment (e.g., normalized by the total number of readings), and E is the expected frequency of the terminal motif determined based on a reference genome sequence. The frequency can be determined using any of the normalization techniques described herein. For example, the observed frequency can be determined as the percentage of fragments having one of a particular set of k-mer terminal motifs (e.g., tricramical terminal motifs). The expected frequency of a terminal motif can be determined based on a reference sequence within a region of the reference genome (e.g., how many times a particular terminal motif appears in a region of the reference genome).
[0167] Different colors represent the first base of the tetrameric motif. Point 1410 has the first base C. Point 1420 has the first base T. Point 1430 has the first base G. Point 1440 has the first base A. N represents an A, T, C, or G base. For the same patient, pleural fluid and plasma are generally correlated in terms of terminal motif profiles, but variations still exist. As an example, T-terminal human fragments (e.g., fragments with TNNN-terminal motifs) are preferentially increased in pleural fluid compared to paired plasma samples. In other words, TNNN-terminal motifs are excessively present in the pleural fluid cfDNA of some patients. Previous analyses and studies have also shown that DNASE1 can be preferentially cleaved at the T-terminus, which could indicate a higher concentration of DNASE1 in pleural fluid relative to plasma.
[0168] Terminal motifs in multiple plasma and pleural fluid samples were also analyzed. By comparing two plasma samples from different patients, the terminal motif profiles were highly correlated in terms of motif ranking or O / E ratio (Pearson r = 0.998). Figure 15A and Figure 16A However, for paired pleural fluid samples from the same two patients, the correlation decreased (Pearson r = 0.832). Figure 15B and Figure 16B ).
[0169] Figure 15AThe image shows the rank of tetramer terminal motifs in the plasma nuclear DNA of two patients according to an embodiment of this disclosure. Both patients had a medical condition causing pleural effusion. Specifically, the x-axis represents the motif rank of the tetramer terminal motif in the plasma sample of the first patient, while the y-axis represents the motif rank of the tetramer terminal motif in the plasma sample of the second patient. Figure 15A As shown, the tetramer terminal motif ranking of plasma nuclear DNA from the first and second patients was highly correlated, with some variation between the two patients. The tetramer terminal motif ranking of plasma nuclear DNA from the first and second patients had a Pearson r of 0.998.
[0170] Figure 15B The image shows the rank of tetramer terminal motifs in the nuclear DNA of pleural fluid cells from two patients according to embodiments of this disclosure. Both patients had medical conditions causing pleural effusion. Specifically, the x-axis represents the motif rank of the tetramer terminal motif in the pleural fluid sample of the first patient, while the y-axis represents the motif rank of the tetramer terminal motif in the plasma sample of the second patient. Figure 15B As shown, especially with Figure 15A The tetramer terminal motif rankings of the plasma nuclear DNA from patients 1 and 2 showed a relatively small correlation with those of the pleural fluid nuclear DNA from patients 1 and 2. The tetramer terminal motif rankings of the pleural fluid nuclear DNA from patients 1 and 2 had a Pearson r of 0.832.
[0171] Differences in correlation may be of concern. The amount of terminal motifs can vary because pleural fluid from different patients may have different nuclease profiles. For example, due to inherent characteristics or mechanisms in pleural fluid samples from different patients, the ranking of pleural fluid samples from different patients tends to show a less relevant relationship compared to plasma samples from different patients.
[0172] Figure 16A The motif O / E ratio of nuclear DNA in plasma cells from two patients is shown. Both patients had a medical condition causing pleural effusion. Specifically, the x-axis represents the motif O / E ratio of tetramer terminal motifs in the plasma sample of the first patient, while the y-axis represents the motif O / E ratio of tetramer terminal motifs in the plasma sample of the second patient. Figure 16A As shown, the O / E ratio of the tetramer terminal motifs of plasma nuclear DNA from the first and second patients was highly correlated, with a Pearson r of 0.998.
[0173] Figure 16BThe motif O / E ratio of nuclear DNA in pleural fluid cells from two patients is shown. In some embodiments, both patients have a medical condition causing pleural effusion. Specifically, the x-axis represents the motif O / E ratio of tetramer terminal motifs in the pleural fluid sample from the first patient, while the y-axis represents the motif O / E ratio of tetramer terminal motifs in the pleural fluid sample from the second patient. Figure 16B As shown, especially with Figure 16A The correlation between the O / E ratio of plasma nuclear DNA motifs in the first and second patients and the O / E ratio of pleural fluid nuclear DNA motifs in the first and second patients was small, with a Pearson r of 0.832.
[0174] As the correlation matrix shows, a similar correlation pattern was also found across all seven patients with medical conditions that caused pleural effusion.
[0175] Figure 17A The correlation matrix shows the correlation coefficients for plasma samples from seven patients. In this example, the correlation coefficients are Pearson's r, although other types of correlation values can also be used. Figure 17A As shown, the correlation coefficients of plasma samples among the seven patients were generally in the range of 0.9 to 1.
[0176] Figure 17B A correlation matrix is shown, representing the correlation coefficients of pleural fluid samples according to embodiments of this disclosure. For example... Figure 17B As shown, the correlation coefficients of pleural fluid samples among the seven patients were generally in the range of 0.7 to 1. Figure 17B Showing with Figure 17A Compared to the overall decrease in correlation.
[0177] Figure 18 A comparison of correlation coefficients between plasma and pleural fluid samples according to embodiments of this disclosure is shown. Figure 18 As shown, the correlation coefficient between different pleural fluid samples was significantly lower than that between different plasma samples, which indicates that the terminal motif profile of cfDNA in pleural fluid is more varied than that in plasma samples.
[0178] B. Terminal motif analysis of MTBC and host DNA in pleural fluid
[0179] We also analyzed the terminal motifs of MTBC DNA in the pleural fluid.
[0180] 1. Terminal motif correlation analysis using O / E ratio
[0181] Figure 19AThis section shows the correlation between human nuclear DNA and MTBC DNA regarding the O / E ratio of tetramer terminal motifs in TB-positive patients. MTBC DNA in this section refers to those aligned to the unmasked MTB genome. Human nuclear DNA does not include mitochondrial DNA, but one or both of these DNAs can be used. Therefore, the nuclear DNA mentioned below also applies to mitochondrial DNA or a combination of both. This analysis is based on the terminal motif profiles of human nuclear DNA and MTBC DNA from a pleural fluid sample of a confirmed TB-positive patient.
[0182] like Figure 19A As shown, the terminal motif profiles of MTBC DNA are highly correlated with those of nuclear DNA, with a Pearson r of 0.84. Some data points (e.g., the CGNN motif) are exceptions to this high correlation. Based on previous research, it is known that CG methylation is approximately 60% to 70% in the human genome, while in bacteria, the level of CG methylation is considerably lower. Furthermore, previous studies have shown that DNASE1L3 has a high preference for CG when it is methylated. Therefore, the CG motif is expected to be preferentially present in human DNA rather than in MTBC. Therefore, excluding the CG motif could further increase the correlation.
[0183] Figure 19B This study shows the correlation between human nuclear DNA and MTBC DNA regarding the O / E ratio of tetramer terminal motifs (excluding the CGNN motif) in TB-positive patients. As discussed above, although methylation at CG sites can affect motifs, this type of methylation at CG sites only affects motifs in human nuclear DNA and not in MTBC DNA. Therefore, excluding the CGNN motif increases the correlation when determining the correlation between human nuclear DNA and MTBC DNA.
[0184] like Figure 19B As shown, with Figure 19A In contrast, excluding CGNN motifs to increase the correlation improved the Pearson r-value from 0.84 to 0.91. Since CpG methylation is rare in bacteria (Phelan et al. Sci Rep. 2018; 8(1):160), no preference for CGNN motifs was observed in MTBC DNA. Therefore, given the difference in total methylation between the two species, there is a weak correlation in CGNN motif cleavage preference between human nuclear DNA and MTBC DNA.
[0185] Furthermore, focusing the analysis on dimer motifs expands the detection coverage. For example, when the analysis focuses on tetramer terminal motifs, there can be a total of 256 different types of tetramer terminal motifs. Similarly, for dimer terminal motifs, there may only be 16 terminal motifs. Therefore, if 100 MTBC fragments can be detected and tetramer terminal motifs are analyzed, some tetramer terminal motifs may not be covered, and the values for these tetramer terminal motifs will be zero. On the other hand, when analyzing dimer terminal motifs, most of them can have values. For some patients, including non-TB patients, the number of TB readings can be quite small. If only a very limited number of TB readings are available, the analysis will be affected and have high noise. Therefore, to eliminate high noise, the analysis can be focused on dimer terminal motifs.
[0186] Figure 20A This shows the correlation between human nuclear DNA and MTBC DNA regarding the O / E ratio of dimer terminal motifs (excluding CG motifs) in TB-positive patients. Figure 20A As shown, in TB-positive samples, the O / E ratio of the dimer terminal motifs of MTBC DNA was highly correlated with that of nuclear DNA (Pearson r = 0.91).
[0187] Figure 20B The correlation between human nuclear DNA and MTBC DNA regarding the O / E ratio of dimer terminal motifs (excluding the CG motif) in TB-negative patients is shown. Unlike the correlation between human nuclear DNA and MTBC DNA in TB-positive patients regarding the O / E ratio of dimer terminal motifs, the correlation between MTBC DNA and human nuclear DNA in TB-negative samples is relatively weak (Pearson r = 0.23). Therefore, the correlation coefficient alone can be used to distinguish between TB and non-TB samples.
[0188] To confirm that the correlation coefficient can be used to distinguish between TB and non-TB samples, we analyzed 20 patients. Of these 20 patients, 6 had TB infection and 14 did not. Since dimer-terminal motif analysis is less susceptible to noise from low sequencing depths, it can be applied to samples with a low number of MTBC fragments. The performance of the dimer-terminal motif correlation analysis is shown below.
[0189] Figure 21AThis is a 2D plot illustrating the correlation coefficients (between human nuclear DNA and MTBC DNA) and MTBC abundance in pleural fluid samples according to embodiments of this disclosure. The x-axis represents MTBC abundance, and the y-axis represents the correlation coefficient. Higher values indicate higher correlation. Red dot 2120 represents TB cases, while blue dot 2130 represents non-TB cases. By using the different correlation coefficients shown by TB and non-TB cases, the two groups (e.g., TB vs. non-TB) can be distinguished. For example, a cutoff value of any value from approximately 0.25 to 0.4 provides perfect separation between TB and non-TB cases.
[0190] Although based on the O / E ratio (e.g., as Figure 20A and 20B (as shown) to determine Figure 21A The correlation can be determined using different parameters. For example, frequency (e.g., the relative percentage of each terminal motif) or ranking analysis (e.g., frequency or O / E ratio) can be used instead of the O / E ratio for each terminal motif. As part of comparing nuclear DNA and MTB DNA, cluster analysis can be performed between the terminal motif profiles of nuclear DNA and MTBC DNA for TB detection to determine correlation values.
[0191] Measured parameters of nuclear DNA and MTBC DNA (e.g., frequency, O / E ratio, or rank of such values) can form a vector pair that can be compared pairwise. For example, the distance between two vectors can be determined, which can be treated as multidimensional data points. Correlation values can be implicitly determined using machine learning models (e.g., clustering, PCA, SVM, or neural networks). Such a model can receive two sets of values and provide a score (e.g., the probability of a microbial disease) based on the correlation between the terminal motif profiles of human and microbial DNA. Therefore, samples can be scored based on the amount of terminal motifs of human and microbial DNA input into the ML model, where the score reflects the difference between TB cases and non-TB cases.
[0192] Figure 21B This is an ROC analysis demonstrating the performance of abundance and correlation coefficient in distinguishing between TB and non-TB samples according to embodiments of this disclosure. The analysis shows that the correlation coefficient has an AUC of 1.0, and the MTBC abundance using the unmasked genome has an AUC of 0.94. Therefore, the correlation coefficient improves accuracy.
[0193] 2. Terminal fundamental sequence frequency
[0194] The data in the previous section used the O / E ratio. The data in this section use terminal motif frequencies, which are determined as the amount of terminal sequences with a specific terminal motif divided by the total number of terminal motifs determined from the free DNA fragment. The data indicate that terminal motif frequencies can also be used.
[0195] Figure 22A The correlation between human nuclear DNA and MTBC DNA from TB-positive patients is shown regarding the frequency of dimer terminal motifs (excluding CG). Figure 22A As shown, for TB-positive samples, the frequency of the dimer terminal motif of MTBC DNA was highly correlated with the frequency of nuclear DNA (Pearson r = 0.75).
[0196] Figure 22B The correlation between human nuclear DNA and MTBC DNA from TB-negative patients was shown regarding the frequency of dimer terminal motifs (CG exclusion). Figure 22B As shown, the correlation between MTBC DNA and human nuclear DNA was relatively weak in TB-negative samples compared to the correlation between human nuclear DNA and MTBC DNA in terms of dimer terminal motif frequencies in TB-positive patients (Pearson r = 0.30).
[0197] Figure 23A A comparison of frequency correlation coefficients between 20 TB and non-TB group samples is shown. The performance of correlation analysis using dimer motif frequencies (e.g., effective separation of TB and non-TB samples) confirms that frequency correlation coefficients can be used to distinguish between TB and non-TB samples. As shown in the figure, any cutoff value in the range of approximately 0.3 to 0.55 provides perfect distinction between TB and non-TB samples.
[0198] Figure 23B The comparison of correlation coefficients of O / E ratios among 20 TB and non-TB group samples according to embodiments of the present disclosure is shown. The performance of correlation analysis using dimer motif frequencies (e.g., effective separation of TB and non-TB samples) confirms that the correlation coefficients of O / E ratios can be used to distinguish between TB and non-TB samples.
[0199] 3. PCA Analysis and Machine Learning Models
[0200] Figure 24A Principal component analysis (PCA) was performed to show the O / E ratio of the dimer terminal motifs of MTBC DNA (excluding the CG motif). Figure 24A The provided clustering shows a moderate number of groups, TB 2420 and non-TB 2410. Figure 24BPrincipal component analysis (PCA) shows the frequencies of dimer terminal motifs (excluding CG motifs) in MTBC DNA. (Compared to...) Figure 24A resemblance, Figure 24B The results show moderate clustering of the TB2440 and non-TB 2430 groups. More extensive clustering is likely to occur using more components.
[0201] Figure 25 ROC curves are provided, showing the performance of a machine learning model trained according to embodiments of this disclosure based on MTBC 2-mer motif frequencies or motif O / E ratios in distinguishing between TB and non-TB samples. Figure 25 As shown, the machine learning model described in this paper exhibits good performance, with an AUC of 0.89 for the motif O / E ratio and 0.87 for the motif frequency.
[0202] In addition to the MTBC 2-mer motif frequency or motif O / E ratio, the machine learning model can also be trained on different parameters or any k-mer motif (e.g., 3-mer motif, 4-mer motif, etc.) as needed. In some implementations, the machine learning model may include a support vector machine (SVM) model, for example, using leave-one-out cross-validation.
[0203] As an example, the input to a machine learning (ML) model can be abundance values, O / E ratios, frequencies, etc., of the target microbial species and optionally host DNA (e.g., human cell nuclear DNA). The input can also include correlation values between or within the input parameters, which can be determined individually according to the ML model. The output of the machine learning model can be such correlation values from the input values. In some implementations, the correlation may not be a single value.
[0204] Therefore, various implementations may use only the terminal motifs of microbial DNA to determine the grade of a specific microbial disease, but may also use the terminal motifs of host DNA. Methods may analyze biological samples to determine the grade of a specific microbial disease in a biological sample of an object, wherein the biological sample includes cell-free DNA of a microorganism. Methods may include analyzing cell-free DNA molecules from a biological sample to obtain sequence reads. Analysis of cell-free DNA molecules may include determining a terminal sequence motif at at least one end of a cell-free DNA molecule; identifying a first set of cell-free DNA molecules as originating from a specific microbial species associated with a specific microbial disease by comparing the sequence reads to a microbial reference genome; using the sequence reads of the first set of cell-free DNA molecules, determining a first quantity for each of a set of terminal sequence motifs in the first set of cell-free DNA molecules, thereby obtaining multiple first quantities; and using the multiple first quantities to determine a classification of the grade of the specific microbial disease of the object. Figures 24A to 25As shown in B, this can be achieved by feeding multiple primary values into a machine learning model that provides the probability of a sample having or not having the disease. A higher probability (potentially above a threshold) can be used to determine the classification.
[0205] C. Fragment size of MTBC DNA fragments
[0206] Figures 26A to 26C The image shows the fragment size distribution of human cell nuclear DNA 2620 (blue) and MTBC 2610 (red) in the TB sample. Figures 26A to 26C As shown, MTBC DNA tends to be shorter than human cell nuclear DNA. For example, as Figure 26A As shown, MTBC DNA has a median fragment size of 113 bp, while human cell nuclear DNA has a median fragment size of 149 bp.
[0207] D. Using the method of terminal motif correlation
[0208] Figure 27 This is a flowchart illustrating a method for analyzing biological samples to determine the level of a specific microbial disease in a biological sample of an object, according to an embodiment of this disclosure.
[0209] At box 2710, cell-free DNA molecules from a biological sample are analyzed to obtain sequence reads. Box 2710 can be performed in a manner similar to box 1210 of method 1200. Analysis of cell-free DNA molecules may include determining a terminal sequence motif at at least one end of the cell-free DNA molecule.
[0210] At box 2720, a first set of cell-free DNA molecules is identified as originating from the object by comparing sequence reads with a human reference genome. Such alignment (mapping) can be performed using various software tools as described herein and understood by those skilled in the art. The first set of cell-free DNA molecules originating from the object may include mitochondrial DNA and / or nuclear DNA. Therefore, at least a portion of the first set of cell-free DNA molecules identified as originating from the object may include nuclear DNA.
[0211] At box 2730, a second group of cell-free DNA molecules was identified as originating from a specific microbial species associated with a specific microbial disease. Identification can be performed by comparing sequence reads to a microbial reference genome, which may or may not be a masked microbial reference genome.
[0212] In some implementations, a second group of cell-free DNA molecules can be identified by comparing sequence reads with multiple microbial reference genomes in a taxonomic tree, where the multiple microbial reference genomes include a microbial reference genome. Alternatively or additionally, the second group of cell-free DNA molecules can be identified by comparing sequence reads with microbial reference genomes using alignment software that outputs mapping quality. When the mapping quality is greater than a threshold, the sequence reads can be identified as originating from a specific microbial species.
[0213] At box 2740, a first quantity is determined for each of a set of terminal sequence motifs of the first group of cell-free DNA molecules, thus obtaining multiple first quantities. The first quantity can be determined using sequence readings of the first group of cell-free DNA molecules. The first quantity can be an absolute value or a normalized value, such as a relative frequency, like the percentage of the first group having a particular terminal sequence motif. Normalization can interpret the sequence background of the reference genome used, such as the O / E ratio.
[0214] As an example, the length of this set of terminal sequence motifs can be two bases (2-mer), three bases (3-mer), or four bases (4-mer). This set of terminal sequence motifs may not include CG terminal motifs, as described herein. As an example, this set of terminal sequence motifs may include at least 10, 11, 12, 13, 14, 15, 16, 64, or 256 different terminal sequence motifs.
[0215] At box 2750, a second quantity is determined for each of the set of terminal sequence motifs of the second group of cell-free DNA molecules, thereby obtaining multiple second quantities. The second quantity can be determined using sequence readings of the second group of cell-free DNA molecules. The second quantity can also be an absolute value or a standardized value, for example, as described for the first quantity.
[0216] The first and second quantities can be the ratio of the observed quantity to the expected quantity, referred to as O / E in this paper. For example, the expected quantity of the terminal sequence motif of this group can be determined based on a reference sequence of a human reference genome. Then, determining the classification may include standardizing each of the multiple first quantities with the expected quantity to obtain a standardized first quantity, which is used to measure the correlation value.
[0217] As a further example, any of these values can be used to determine the ranking, which can serve as a first quantity. Thus, the first quantity could be a ranking of each of the terminal sequence motifs in a first group of cell-free DNA molecules having the corresponding terminal sequence motif of that group. Similarly, the second quantity could be a ranking of each of the terminal sequence motifs in a second group of cell-free DNA molecules having the corresponding terminal sequence motif of that group.
[0218] At box 2760, a correlation value is measured to determine the correlation between a plurality of first quantities and a plurality of second quantities. Measuring the correlation value may include determining the difference between the corresponding first quantity and the corresponding second quantity for each of the terminal sequence motifs in the set. These differences may be summed and possibly standardized by the number of terminal motifs in the set.
[0219] In some implementations, the correlation value can be the Pearson correlation coefficient (r), which measures linear correlation. The Pearson correlation coefficient is a number from -1 to 1 that measures the strength and direction of the relationship between two variables. When one variable changes, the other variable changes in the same direction. Such correlation can be measured as the ratio between the product of the covariance of the two variables and their standard deviations. The Pearson correlation coefficient can be calculated using the following formula, where r is the correlation coefficient, x... i It is the set of values for the first variable. It is the average of the values of the first variable, y i It is the set of values for the second variable. It is the average value of the second variable.
[0220]
[0221] Correlation values include, but are not limited to, Pearson correlation coefficient, Spearman rank correlation, Phi correlation, Kendall rank correlation, Jaccard similarity, cosine similarity, etc.
[0222] At box 2770, the classification of an object's specific microbial disease level is determined based on a comparison of the correlation value with a reference value. The reference value can be selected using measurements obtained from one or more reference samples of known classifications (e.g., disease-positive or disease-negative). For example, the reference value can be determined using a first batch of training samples from objects known to have the specific microbial disease and a second batch of training samples from objects known not to have the specific microbial disease. Figure 21A A measurement chart for this type of reference sample is shown. Example reference values can be from 0.25 to 0.4.
[0223] In some implementations, machine learning models can be used to measure correlation values and determine the classification of a specific microbial disease level of an object. For example, multiple primary and secondary values can be input into the machine learning model, which can determine the correlation values as an intermediate step before outputting a classification, which can be a probability.
[0224] Similar to Method 1200, the specific microbial species can be a bacterial species, such as the Mycobacterium tuberculosis complex (MTBC). The specific microbial disease can be tuberculosis.
[0225] VI. Examples of Nanopores
[0226] We also used the aforementioned technology for nanopore sequencing. To validate the feasibility of the nanopore platform for detecting mycobacterial DNA in pleural fluid samples, we selected 10 MTB-targeted capture libraries and sequenced them on an Illumina platform. Two samples were confirmed to be TB-positive by culture or qPCR (TB group), and the other eight samples were negative by culture or qPCR (non-TB). We sequenced the libraries in a Nanopore PromethION run, which produced an average of 6.9 million raw sequencing reads per sample.
[0227] A. Masked MTBC genome
[0228] For the masking technique, we first aligned the raw sequencing reads to a human reference genome. Unaligned non-human reads were then used to detect MTBC reads using two different strategies. (1) Without a proposed masking strategy, bowtie2 was used to align non-human reads to both the human and MTB reference genomes. (2) bowtie2 was used to align non-human reads to both the human and masked MTB reference genomes.
[0229] Figures 28A to 28B This demonstrates the use of two different alignment methods, without masking ( Figure 28A ) and cover ( Figure 28B The method was used to determine the number of MTBC DNA fragments detected in TB-positive (TB) and TB-negative (non-TB) samples. Figure 28A A graph showing the number of MTBC fragments determined for both TB and non-TB subjects using nanopore sequencing without masking the MTB reference genome. Figure 28B A graph showing the number of MTBC fragments determined for both TB and non-TB subjects when using nanopore sequencing and masking the MTB reference genome.
[0230] As shown in the figure, MTBC sequences were successfully detected in two TB-positive samples using two strategies. However, the second strategy, which only involved masking, was not successfully detected. Figure 28B No MTBC readings were detected in any TB-negative (non-TB) samples. Furthermore, the separation between TB and non-TB samples was complete. Figure 28B The first strategy has no overlap and 100% sensitivity and specificity, while the second strategy has overlap between TB and non-TB objects, resulting in less than 100% sensitivity and / or specificity. Using the masking strategy, any reference value (cutoff value) from 0 to approximately 5 provides 100% accuracy.
[0231] B. Terminal motif
[0232] The following analysis was performed using TB-positive samples with the highest number of MTBC readings.
[0233] Figures 29A to 29F The comparison between nanopores and Illumina in terms of size and terminal motif analysis is shown. Figure 29A The size distribution of nuclear DNA in pleural fluid was shown. Figure 29B The size distribution of MTBC DNA in pleural fluid is shown. Generally, the size distribution is similar, except that the nanopore data provide longer DNA fragments for both the nucleus and MTBC cfDNA than those from Illumina.
[0234] Figure 29C The ranking of terminal motifs by the O / E ratio of nuclear DNA is shown. Figure 29D The ranking of terminal motifs by O / E ratio across MTBC DNA is shown. For both nuclear DNA and MTBC DNA, the terminal motif profiles determined from Illumina and nanopore data are highly correlated.
[0235] Figure 29E The correlation between human nuclear DNA and MTBC DNA in terms of the O / E ratio of dimer terminal motifs was shown using Illumina data. Figure 29F The correlation between human nuclear DNA and MTBC DNA in terms of the O / E ratio of dimer terminal motifs was demonstrated using nanopore data. Both Illumina and nanopore data showed a high correlation between dimer terminal motifs of nuclear DNA and MTBC DNA.
[0236] Therefore, this type of technology will be useful in various types of measurements.
[0237] VII. Example Implementation
[0238] The following are example implementations that can be used with various embodiments of this disclosure.
[0239] A. Classification of MTBC-derived sequencing reads
[0240] Paired end sequencing reads are aligned to a reference human genome (e.g., hg38). Reads not aligned to the human genome are then realigned to a microbial database comprising complete reference genomes from mycobacterial families and other microorganisms. The microbial origin (i.e., taxonomy) of these sequencing reads can be determined. The alignment procedure is performed using Kraken2 (Wood et al. Genome Biol. 2019; 20(1):257). Alignment can also be performed using other bioinformatics algorithms, including BLAST, FASTA, Bowtie, BWA, BFAST, SHRiMP, SSAHA2, NovoAlign, SOAP, etc., which can be used with a mapping quality threshold assigned to any given species or genus and can be used in taxonomy.
[0241] B. DNA fragment end motif analysis
[0242] Terminal motif determination can be based on the length of the k-mer used; for example, tetramer motifs correspond to the 4-nucleotide sequence at each 5' end of a DNA molecule (Watson and Crick). Terminal motif frequency is defined as the proportion of terminal motifs to the total number of terminal motifs.
[0243] Because DNA terminal motif frequencies are influenced by the sequence background of the reference genome being analyzed, some implementations can be normalized using the sequence background of the reference genome. To compare terminal preferences under different sequence backgrounds (e.g., between human DNA and mycobacterial DNA), a normalization method was applied. Terminal motif counting was performed within a region corresponding to the reference genome (e.g., masked or unmasked microbes or humans, which may, for example, be masked for repetitive regions). A sliding window of the same size as the terminal motif could be slid across this region to identify and count the occurrence of each terminal motif. For each terminal motif, the expected frequency E could be determined as the proportion of K-mers from the reference genome having that particular terminal motif.
[0244] The terminal motif frequency measured in a sample is calculated as the ratio of terminal motifs to the total number of terminal motifs, and is called the expected terminal motif frequency. The 5' terminal motif frequency of the sequenced DNA fragment is called the observed motif frequency (O). Alternatively, the 3' terminal motif frequency can be used. The observed terminal motif frequency is normalized to the expected terminal motif frequency (E), and the resulting value is defined as the O / E ratio. A higher O / E ratio value indicates a higher preference for terminal motifs. The terminal motif frequency and O / E ratio of human cell nuclear cfDNA can be determined using the same method.
[0245] C. DNA fragment size analysis
[0246] DNA fragment size can be determined by the number of nucleotides between the outermost genomic coordinates of paired-end sequencing reads. The fragment size of human cell nuclear DNA can be directly inferred from the alignment results. For MTBCs, classified MTBC reads are re-aligned against microbial databases using tools including Bowtie2, BWA, or SOAP. The size of the MTBC fragment is then determined.
[0247] VIII. Example Treatment
[0248] The implementation plan may further include treating the subject after determining the infection level classification. For example, treatment may be provided based on the predicted amount of microorganisms in the subject's biosample. In some cases, treatment is provided based on the type of tissue in which the infection occurs. Tissue type can be used to guide antibiotic treatment, antibiotics specific to drug-resistant strains, surgery, or any other form of treatment. Furthermore, the infection level can be used to determine the aggressiveness of any type of treatment, which may also be based on the disease level. For example, sepsis can be treated with antibiotics and blood pressure support medication. In some implementation plans, the more a parameter's value (e.g., quantity or size) exceeds a reference value, the more aggressive the treatment can be.
[0249] Examples of treatments for microbial infections include, but are not limited to, the following: antibiotics or antimicrobial agents (which may be specific to drug-resistant strains); antiviral agents; antiparasitic agents; and antifungal agents. In some cases, different types of drugs and treatments are provided based on the type of microbial species identified from the subject. For example, if Mycobacterium tuberculosis is found in the subject, drugs such as isoniazid (INH), rifampin (RIF), rifabutin, rifapentine (RPT), pyrazinamide (PZA), or any fluoroquinolone may be provided. In another example, if Clostridium botulinum is identified in the subject, antitoxin may be provided.
[0250] IX. Example System
[0251] Figure 30A measurement system 3000 according to one embodiment of the present disclosure is illustrated. The system includes a sample 3005, such as a free nucleic acid molecule (e.g., DNA and / or RNA of a host and / or microorganism), within an analytical device 3010, wherein an analysis 3008 can be performed on the sample 3005. For example, the sample 3005 can be contacted with reagents of the analysis 3008 to provide a signal of physical characteristics (e.g., sequence information of the free nucleic acid molecule) 3015. An example of the analytical device may be a flow channel comprising probes and / or primers or droplets through which the analysis moves (where the droplets comprise the analysis). The physical characteristics 3015 of the sample (e.g., fluorescence intensity, voltage, or current) are detected by a detector 3020. The detector 3020 can perform measurements at time intervals (e.g., periodic time intervals) to obtain data points constituting a data signal. In one embodiment, an analog-to-digital converter converts the analog signal from the detector into digital form at multiple times.
[0252] The analysis device 3010 and detector 3020 can form an analysis system, such as a sequencing system that performs sequencing according to the embodiments described herein. Data signal 3025 is transmitted from detector 3020 to logic system 3030. As an example, data signal 3025 can be used to determine the sequence and / or location of a nucleic acid molecule (e.g., DNA and / or RNA) in a reference genome. Data signal 3025 can include various measurements taken simultaneously, such as different colors of fluorescent dyes or different electrical signals for different molecules of sample 3005, and therefore data signal 3025 can correspond to multiple signals. Data signal 3025 can be stored in local memory 3035, external memory 3040, or storage device 3045. The analysis system can consist of multiple analysis devices and detectors.
[0253] The logic system 3030 may be a computer system, ASIC, microprocessor, graphics processing unit (GPU), etc., or may include a computer system, ASIC, microprocessor, graphics processing unit (GPU), etc. It may also include a display (e.g., a monitor, LED display, etc.) and user input devices (e.g., a mouse, keyboard, buttons, etc.), or be connected to a display and user input devices. The logic system 3030 and other components may be part of a stand-alone or network-connected computer system, or may be directly connected to or incorporated into a device including detector 3020 and / or analysis device 3010 (e.g., a sequencing device). The logic system 3030 may also include software executing in processor 3050. The logic system 3030 may include a computer-readable medium storing instructions for controlling the measurement system 3000 to perform any of the methods described herein. For example, the logic system 3030 may provide commands to a system including analysis device 3010 to perform sequencing or other physical operations. Such physical operations may be performed in a specific order, such as adding and removing reagents in a specific order. Such physical operations can be performed using robotic systems (such as robotic arms) that can be used to obtain samples and perform analysis.
[0254] The measurement system 3000 may also include a treatment device 3060 capable of providing treatment to an object. The treatment device 3060 can determine treatment and / or perform treatment. Examples of such treatments may include surgery, radiation therapy, chemotherapy, immunotherapy, targeted therapy, hormone therapy, and stem cell transplantation. A logic system 3030 may be connected to the treatment device 3060, for example, to provide results of the methods described herein. The treatment device may receive input from, for example, imaging devices and other devices (e.g., for controlling treatment, such as controlling a robotic system).
[0255] The measurement system 3000 may also include a reporting device 3055, which can display the results of any of the methods described herein (e.g., as determined using the measurement system). The reporting device 3055 can communicate with a reporting module within a logic system 3030, which can aggregate, format, and send reports to the reporting device 3055. The reporting module can display information determined using any of the methods described herein. The information can be displayed by the reporting device 3055 in any format that can be recognized and understood by a user of the measurement system 3000. For example, the information can be displayed by the reporting device 3055 in a format that is displayed, printed, or transmitted, or any combination thereof.
[0256] Any computer system mentioned in this article can utilize any suitable number of subsystems. Examples of such subsystems are shown in [the document / examples]. Figure 31In computer system 10. In some embodiments, the computer system includes a single computer device, wherein a subsystem may be a component of the computer device. In other embodiments, the computer system may include multiple computer devices having internal components, each of which is a subsystem. The computer system may include desktop and laptop computers, tablet computers, mobile phones, and other mobile devices.
[0257] Shown Figure 31 The subsystems in the display are interconnected via system bus 75. Additional subsystems, such as printer 74, keyboard 78, one or more storage devices 79, monitor 76 (e.g., display screen, such as LED) connected to display adapter 82, and other devices, are displayed. Peripheral devices and I / O devices connected to input / output (I / O) controller 71 can be connected to the computer system via a variety of devices known in the art, such as input / output (I / O) port 77 (e.g., USB, FireWire®). For example, I / O port 77 or external interface 81 (e.g., Ethernet, Wi-Fi, etc.) can be used to connect computer system 10 to a wide area network (e.g., the Internet), a mouse input device, or a scanner. The interconnection via system bus 75 allows central processing unit 73 to exchange information with each subsystem and control the execution of multiple instructions from system memory 72 or storage device 79 (e.g., fixed disk, such as hard disk drive, or optical disk), as well as information exchange between subsystems. System memory 72 and / or storage device 79 may include computer-readable media. Another subsystem is the data collection device 85, such as a camera, microphone, accelerometer, etc. Any data mentioned herein can be output from one component to another and can also be output to the user.
[0258] A computer system may include multiple identical components or subsystems, connected together, for example, via an external interface 81, an internal interface, or via a removable storage device that can be connected to and moved from one component to another. In some embodiments, the computer system, subsystem, or device may communicate via a network. In such cases, one computer may be considered a client and another a server, where each computer may be part of the same computer system. Clients and servers may each include multiple systems, subsystems, or components. In various embodiments, the method may involve a variety of numbers of clients and / or servers, including at least 10, 20, 50, 100, 200, 500, 1,000, or 10,000 devices. The method may include a variety of numbers of communication messages between devices, including at least 100, 200, 500, 1,000, 10,000, 50,000, 100,000, 500,000, or one million communication messages. Such communications can involve at least 1 MB, 10 MB, 100 MB, 1 GB, 10 GB, or 100 GB of data.
[0259] Aspects of the implementation may be implemented using hardware circuitry (e.g., application-specific integrated circuits or field-programmable gate arrays) and / or using computer software stored in a modular or integrated manner in memory with a general-purpose programmable processor in the form of logical control, and thus the processor may include memory storing software instructions for configuring the hardware circuitry and an FPGA or ASIC having configuration instructions. As used herein, the processor may include a single-core processor, a multi-core processor on the same integrated chip, or single board or network hardware, and multiple processing units on dedicated hardware. Based on this disclosure and the teachings provided herein, those skilled in the art will appreciate other ways and / or methods of implementing embodiments of this disclosure using hardware and combinations of hardware and software.
[0260] Any software component or function described in this application may be implemented in the form of software code using, for example, conventional or object-oriented techniques. This software code is executed by a processor using any suitable computer language (e.g., Java, C, C++, C#, Objective-C, Swift) or scripting language (e.g., Perl or Python). The software code may be stored as a series of instructions or commands on a computer-readable medium for storage and / or transmission. Suitable non-transitory computer-readable media may include random access memory (RAM), read-only memory (ROM), magnetic media (e.g., hard disk or floppy disk) or optical media (e.g., closed-disk CD, digital optical disc (DVD), or Blu-ray disc), flash memory, and similar devices. The computer-readable medium may be any combination of such devices. Additionally, the order of operations may be reconfigured. A process may terminate upon completion of its operations, but may have additional steps not included in the figure. A process may correspond to a method, function, program, subroutine, subroutines, etc. When a process corresponds to a function, its termination may correspond to the function returning to the calling function or the main function.
[0261] Such programs can also be encoded and transmitted using carrier signals suitable for transmission over wired, optical, and / or wireless networks (including the Internet) conforming to various protocols. Therefore, computer-readable media can be constructed using data signals encoded by such programs. Computer-readable media encoded with program code (e.g., as software) can be packaged with compatible devices or provided separately from other devices (e.g., for download via the Internet). Any such computer-readable media can reside on or within a single computer product (e.g., a hard disk drive, CD, or an entire computer system) and can reside on or within different computer products within a system or network. A computer system may include a monitor, printer, or other suitable display for providing a user with any of the results mentioned herein.
[0262] Any method described herein can be performed wholly or partially using a computer system including one or more processors configured to perform the steps described herein. Any operation performed via the processor (e.g., comparison, determination, contrast, computer calculation, computation) can be performed in real time. The term "..." real time"This can refer to a computational operation or process completed within a certain time limit. For example, the time limit could be 1 minute, 1 hour, 1 day, or 7 days. Therefore, an implementation can refer to a computer system configured to perform the steps of any method described herein, potentially having different components that perform the respective steps or groups of steps. Although presented in the form of numbered steps, the steps of the methods herein can be performed simultaneously or at different times or in different orders. Furthermore, a portion of these steps can be used in conjunction with a portion of other steps of other methods. Additionally, all or some of the steps can be optional. Moreover, any step of any method can be performed using modules, units, circuits, or other means of a system for performing these steps."
[0263] The specified details of a particular embodiment may be combined in any suitable manner without departing from the spirit and scope of the embodiments of this disclosure. However, other embodiments of this disclosure may refer to specific embodiments relating to each individual aspect or a particular combination of these individual aspects.
[0264] The foregoing description of exemplary embodiments of this disclosure has been presented for purposes of illustration and description. It is not intended to be exhaustive or to limit this disclosure to the precise form described, and many modifications and variations are possible in light of the foregoing teachings.
[0265] Unless explicitly stated otherwise, the use of “a,” “an,” or “the” is intended to mean “one or more.” Unless explicitly stated otherwise, the use of “or” is intended to mean “inclusive of or” rather than “exclusive of or.” Referring to the “first” component does not necessarily require the provision of the second component. Furthermore, unless explicitly stated otherwise, referring to the “first” or “second” component does not limit the referenced component to a particular location. The term “based on” is intended to mean “at least partially based on.”
[0266] The claims are drafted to exclude any optional elements. Therefore, this statement is intended to serve as a basis for the use of exclusive terms such as “solely,” “only,” and similar terms, or for the use of negative limitations, in conjunction with the description of the claimed elements.
[0267] All patents, patent applications, publications, and specifications mentioned herein are incorporated herein by reference in their entirety for all purposes. No prior art is acknowledged. In the event of any conflict between this application and the references provided herein, this application shall prevail.
Claims
1. A method for analyzing biological samples to determine the degree of disease caused by a specific microorganism in a biological sample of an object, said biological sample comprising cell-free DNA of the microorganism and cell-free DNA of the object, said method comprising: Analyze free DNA molecules from the biological sample to obtain sequence reads; The storage of a masked microbial reference genome of a specific microbial species associated with the specific microbial disease, wherein the masked microbial reference genome is generated from a microbial reference genome of the specific microbial species, the microbial reference genome comprising (1) specific regions identified as unique to the specific microbial species and (2) non-specific genomic regions shared with one or more other species, and wherein the masked microbial reference genome is generated by removing the non-specific genomic regions from the microbial reference genome; The sequence reads are compared with the masked microbial reference genome to identify a group of cell-free DNA molecules from the specific microbial species. Determine the amount of the aforementioned group of free DNA molecules; and The classification of the specific microbial disease of the object is determined based on a comparison of the quantity with a reference value.
2. The method of claim 1, further comprising: Non-specific genomic regions shared with one or more other species are identified by comparing the microbial reference genome of the specific microbial species with one or more other reference genomes of one or more other species.
3. The method of claim 2, wherein identifying the non-specific genomic region comprises: The one or more other reference genomes are divided into a group of K-mers, where K is 20 to 35; and The set of K-mers is compared with the microbial reference genome to identify the non-specific genomic regions.
4. The method of claim 3, wherein the non-specific genomic region corresponds to a subset of the set of K-mers aligned with the microbial reference genome.
5. The method of claim 2, wherein identifying the non-specific genomic region comprises: The microbial reference genome was divided into a group of K-mers, where K was 20 to 35; and The set of K-mers is compared with one or more other reference genomes to identify the non-specific genomic regions.
6. The method of claim 5, wherein the non-specific genomic region corresponds to a subset of the set of K-mers aligned with the one or more other reference genomes.
7. The method of claim 5, wherein the nonspecific genomic region is identified as a region that does not correspond to a subset of the set of K-mers that are not aligned with the one or more other reference genomes.
8. The method of claim 1, wherein the sequence readings are compared with multiple microbial reference genomes in a taxonomic tree to identify the set of cell-free DNA molecules as originating from the specific microbial species, and wherein the multiple microbial reference genomes include the masked microbial reference genome.
9. The method of claim 1, wherein the sequence reading is compared with the masked microbial reference genome using alignment software that outputs mapping quality, and wherein the sequence reading is identified as originating from the specific microbial species when the mapping quality is greater than a threshold.
10. The method of claim 1, wherein the quantity is standardized.
11. The method of claim 9, wherein the quantity is standardized by dividing by the total number of the sequence reads or the number of sequence reads from the genome of the object.
12. The method of claim 1, wherein one or more of the other types include the object.
13. A method for analyzing a biological sample to determine the degree of a disease caused by a specific microorganism in the biological sample of an object, the biological sample comprising cell-free DNA of the microorganism and cell-free DNA of the object, the method comprising: Analyzing cell-free DNA molecules from the biological sample to obtain sequence reads, wherein the analysis of cell-free DNA molecules includes: Determine the terminal sequence motif at at least one end of the free DNA molecule; By comparing the sequence readings with a human reference genome, the first group of cell-free DNA molecules was identified as originating from the object; By comparing the sequence reads with a microbial reference genome, the second group of cell-free DNA molecules were identified as originating from a specific microbial species associated with the specific microbial disease. Using the sequence readings of the first group of free DNA molecules, a first quantity of each of a set of terminal sequence motifs of the first group of free DNA molecules is determined, thereby obtaining multiple first quantities; Using sequence readings of the second group of free DNA molecules, a second quantity of each of the terminal sequence motifs of the second group of free DNA molecules is determined, thereby obtaining multiple second quantities; The correlation value measures the correlation between the plurality of first quantities and the plurality of second quantities; and The classification of the specific microbial disease of the object is determined based on the comparison between the correlation value and the reference value.
14. The method of claim 13, wherein the second group of cell-free DNA molecules is identified by comparing the sequence readings with multiple microbial reference genomes in a taxonomic tree, and wherein the multiple microbial reference genomes include the microbial reference genome.
15. The method of claim 13, wherein the microbial reference genome is a masked microbial reference genome.
16. The method of claim 13, wherein the second group of cell-free DNA is identified by comparing the sequence readings with the microbial reference genome using alignment software that outputs mapping quality, and wherein the sequence readings are identified as originating from the specific microbial species when the mapping quality is greater than a threshold.
17. The method of claim 13, wherein at least a portion of the first group of free DNA molecules identified as originating from the object comprises mitochondrial DNA.
18. The method of claim 13, wherein at least a portion of the first group of free DNA molecules identified as originating from the object comprises nuclear DNA.
19. The method of claim 13, wherein the first quantity is the relative frequency of the terminal sequence motif.
20. The method of claim 13, further comprising: Based on the reference sequence of the human reference genome, the expected quantity of the set of terminal sequence motifs is determined, wherein determining the classification includes standardizing each of the plurality of first quantities with the expected quantity to obtain a standardized first quantity, the standardized first quantity being used to measure the correlation value.
21. The method of claim 13, wherein the first quantity is a ranking of each of the set of terminal sequence motifs based on the abundance of the first set of free DNA molecules having the corresponding terminal motifs of the set, and wherein the second quantity is a ranking of each of the set of terminal sequence motifs based on the abundance of the second set of free DNA molecules having the corresponding terminal motifs of the set.
22. The method of claim 13, wherein the length of the set of terminal sequence motifs is two bases, three bases, or four bases.
23. The method of claim 22, wherein the set of terminal sequence motifs does not include CG terminal motifs.
24. The method of claim 22, wherein the set of terminal sequence motifs comprises at least 10 terminal sequence motifs.
25. The method of claim 13, wherein measuring the correlation value comprises determining the difference between a corresponding first quantity and a corresponding second quantity for each of the set of terminal sequence motifs.
26. The method of claim 13, wherein the machine learning model is used to measure the correlation value and determine the classification of the level of the specific microbial disease of the object.
27. The method of claim 26, wherein the plurality of first quantities and the plurality of second quantities are input into the machine learning model.
28. The method as claimed in any of the preceding claims, wherein the reference value is determined using a first batch of training samples from subjects known to have the specific microbial disease and a second batch of training samples from subjects known not to have the specific microbial disease.
29. The method as described in any of the preceding claims, wherein the specific microbial species is a bacterial species.
30. The method of claim 29, wherein the bacterial species is Mycobacterium tuberculosis complex (MTBC).
31. The method of claim 29, wherein the specific microbial disease is tuberculosis.
32. The method of any of the preceding claims, wherein analyzing the cell-free DNA molecules comprises receiving sequence reads obtained from targeted sequencing of the cell-free DNA molecules from the biological sample.
33. The method of claim 32, further comprising performing the targeted sequencing.
34. The method of claim 32, wherein the targeted sequencing uses capture probes of a microbial reference genome and capture probes of a human reference genome, wherein the concentration of the capture probes of the microbial reference genome is higher than the concentration of the capture probes of the human reference genome.
35. A computer product comprising a non-transitory computer-readable medium storing a plurality of instructions, which, when executed, cause a computer system to perform the method as described in any of the preceding claims.
36. A system comprising: The computer product as described in claim 35; and One or more processors configured to execute instructions stored on the computer-readable medium.
37. A system comprising means for performing any of the methods described above.
38. A system comprising one or more processors configured to perform any of the methods described above.
39. A system comprising modules that respectively perform the steps of any of the methods described above.