A Mycobacterium tuberculosis transmission relationship detection system and its application
By integrating genetic and epidemiological information through pan-genomics and machine learning algorithms and constructing a random forest model, the problems of errors and insufficient information in tuberculosis transmission detection were resolved, and accurate tracing of the transmission relationships of Mycobacterium tuberculosis and optimization of prevention and control strategies were achieved, thereby improving the scientific nature and efficiency of tuberculosis prevention and control.
Patent Information
- Application Number
- CN202510520125.5
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2025-04-24
- Publication Date
- 2025-09-05
- Estimated Expiration
- 2045-04-24
AI Technical Summary
Existing tuberculosis transmission detection technologies have problems such as fixed SNP difference threshold errors, single genetic variation types, insufficient use of epidemiological information, and inability to quantitatively assess transmission probability, resulting in inaccurate judgments on transmission relationships and a lack of scientific basis for prevention and control strategies.
Combining pan-genomics and machine learning algorithms, integrating genetic difference characteristics, epidemiological information and strain lineage characteristics, a random forest model is constructed. Through multiple genetic variation types and detailed epidemiological data, the detection of transmission relationships is optimized to achieve traceability analysis and accurate assessment of transmission patterns.
It has significantly improved the accuracy and reliability of the transmission relationship of Mycobacterium tuberculosis, provided a scientific basis to support precise prevention and control, reduced the risk of transmission, saved manpower and material resources, adapted to the transmission situation in different regions and time, and improved the efficiency of prevention and control.
Smart Images

Figure CN120048362B_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the fields of pathogenic microbiology and epidemiology research, and in particular to a Mycobacterium tuberculosis transmission relationship detection system based on pan-genomics and machine learning and its application. Background Art
[0002] Tuberculosis is caused by Mycobacterium tuberculosis ( Mycobacterium tuberculosis Tuberculosis (TB) is a serious infectious disease that continues to pose a significant threat to global public health. As a contagious disease, understanding its transmission patterns is crucial for prevention and control. Healthy individuals can develop lung infections through respiratory droplets from contact with infected individuals. To achieve the ultimate goal of TB elimination, enhanced surveillance and prevention of human-to-human transmission in high-burden communities is necessary. Since the 1990s, molecular epidemiology, which uses molecular biological methods to identify the relationships between strains and integrates traditional epidemiological studies, has enabled a better understanding of TB transmission. Traditional methods for detecting TB transmission primarily rely on techniques such as IS6110-RFLP and MIRU-VNTR, but these methods have limited resolution. In recent years, with the maturity and cost reduction of high-throughput sequencing technology, whole-genome sequencing (WGS) has been widely used to study TB transmission, characterizing the sequence of nucleotide substitutions to track the direction and chain of transmission of strains. At present, a difference threshold of 5 single nucleotide polymorphisms (SNPs) is commonly used internationally to define strains with recent transmission relationships within 2 to 3 years, while Chinese scientific research tends to choose a difference of 12 SNPs as the threshold, and there is no unified standard yet.
[0003] However, this transmission detection technology based solely on SNP differences faces the following key problems: 1. Fixed thresholds lead to errors: different strains have different mutation rates, and a unified SNP difference threshold may not accurately reflect the true transmission relationship; 2. The singleness of genetic variation types: existing methods mainly focus on SNPs, ignoring other types of genetic variations such as insertion and deletion (InDel) and structural variants (SV) that are widely present in MTB, resulting in incomplete identification of transmission chains; 3. Insufficient use of epidemiological information: Epidemiological data such as the case's geographic location, sampling time, social interaction network, and past medical history are not fully considered, making the identification of transmission patterns not comprehensive enough; 4. Lack of quantitative assessment of transmission probability: existing methods only judge whether it is transmission-related or not, and cannot accurately assess the transmission risk and formulate differentiated control strategies.
[0004] In a previously published paper ("TransFlow: a Snakemake workflow for transmission analysis of Mycobacterium tuberculosis whole-genome sequencing data"), a pan-genomics data analysis pipeline for MTB was proposed for parsing WGS data and detecting recent transmission. However, this detection method still only uses single-nucleotide polymorphisms (SNPs) and identifies recent transmission relationships based on a fixed threshold for the number of SNP differences. Studies have shown that mutation rates vary significantly among MTB lineages (e.g., Lineage 2 Beijing strains have a higher mutation rate than Lineage 4), and fixed thresholds can lead to false positives and negatives.
[0005] Therefore, there is an urgent need to build a new system for detecting the transmission relationship of Mycobacterium tuberculosis to collect more genomes of community-transmitted MTB isolates and epidemiological information of cases to minimize the impact of different sampling sites and patients, and to improve it by integrating pan-genomic and epidemiological information and optimizing machine learning models to solve key technical problems of existing technologies such as inaccurate judgment of transmission relationships, incomplete consideration of genetic variations, insufficient use of epidemiological information, and inability to quantitatively evaluate transmission probabilities. This will significantly improve the accuracy, reliability and interpretability of MTB transmission detection, provide strong technical support for tuberculosis prevention and control, and contribute to the global tuberculosis prevention and control process. Summary of the Invention
[0006] In response to the problems existing in the prior art, the present invention provides a Mycobacterium tuberculosis transmission relationship detection system and its application, which combines pan-genomics and machine learning algorithms, comprehensively considers the various genetic differences between strains and epidemiological information between cases, and determines the transmission relationship between cases in a more comprehensive and accurate manner. Through sample collection and data preprocessing, genetic differences, epidemiological information, strain lineage and other characteristics are extracted and integrated and normalized, and then the optimal model is screened out through data partitioning, cluster analysis, multiple algorithm modeling training and evaluation, overcoming the fixed SNP threshold error, comprehensively considering multiple genetic variation types, making full use of epidemiological information, realizing quantitative evaluation of transmission probability, and accurately evaluating the transmission pattern of Mycobacterium tuberculosis, realizing traceability analysis of the transmission path of Mycobacterium tuberculosis, with wide applicability, providing a scientific basis for preventing transmission, improving the efficiency and accuracy of prevention and control, and playing a key role in the global tuberculosis prevention and control process.
[0007] On the one hand, the present invention provides a system for detecting the transmission relationship of Mycobacterium tuberculosis, in which a combination of one or more of genetic difference characteristics, epidemiological information characteristics, strain lineage information characteristics, and other characteristics is used to construct a model for detecting the recent spread of Mycobacterium tuberculosis.
[0008] Furthermore, the genetic difference characteristics include any one or more selected from the following: non-synonymous mutation ratio, synonymous mutation ratio, number of promoter region mutations, total number of SNPs, total number of InDels, number of structural variations, number of gene presence / absence variations, and number of rare mutations;
[0009] The epidemiological information features include any one or more selected from the following: Euclidean distance to the patient's residence, sampling time interval, whether the patients are in the same community;
[0010] The strain pedigree information features include any one or more selected from the following: strain pedigree classification, number of pedigree-specific mutations;
[0011] The other characteristics include any one or more selected from the following: the number of drug resistance-related mutations and the number of virulence-related mutations.
[0012] The MTB transmission relationship described in this invention involves analyzing the current MTB patient to identify where or which patient they contracted the infection. This is equivalent to a traceability analysis, exploring the current patient's source of infection and thereby identifying the previous MTB patient who infected them. Once the infection source detection is completed for all patients, a complete MTB transmission network diagram can be constructed. This allows analysis of future MTB prevention and control measures to identify key areas of focus, such as severing control pathways and reducing MTB transmission, providing a scientific basis for transmission prevention.
[0013] Pangenomics, the study of all genomes within a species, goes beyond the traditional concept of a reference genome and provides a more comprehensive view of genomic variation. The application of pangenomics is particularly important in pathogen transmission detection because it can reveal the genetic composition and diversity of microbial populations, thereby helping us understand pathogen virulence, pathogenesis, and microevolution. Pangenomic studies allow for the reclassification of species, thereby clarifying and refining previously proposed traditional criteria. In the detection of the spread of infectious diseases such as tuberculosis, pangenomics can help identify and compare genomic variations across strains, including single nucleotide polymorphisms (SNPs), insertions and deletions (InDels), copy number variations (CNVs), and presence / absence variations (PAVs). These variations collectively constitute pangenomic variability, which is crucial for understanding pathogen transmission patterns and evolutionary pathways. The application of pangenomics in pathogen transmission detection not only improves the accuracy of variant detection but also significantly enhances the detectability of large structural variants. When individuals within a population are relatively different, PAV analysis can better reflect differences within the population than using variant information such as SNPs. It can also be considered as a marker, and PAVs can also be used to conduct genetic evolutionary association analysis of species. Therefore, pan-genomics provides a more comprehensive framework for understanding and tracking the spread of MTB.
[0014] The integration of epidemiological information is crucial in detecting the spread of infectious diseases, as it provides direct clues to risk factors for disease transmission. The core of epidemiological research lies in the collection and analysis of population-based health data, including the geographic distribution of cases, onset time, treatment and medication use, exposure history, social behavior, and environmental factors. This data is crucial for understanding the spread of infectious diseases, identifying high-risk populations, and developing effective prevention measures.
[0015] Due to the lack of comprehensive and reliable information on the recent spread of MTB, processing and analyzing massive amounts of pan-genomic and epidemiological data requires increasingly sophisticated algorithms. Machine learning, a key branch of artificial intelligence, can identify complex nonlinear relationships and reveal hidden patterns that are difficult to capture using traditional statistical methods. It has been widely applied in various fields, including genomics and epidemiology. Machine learning techniques are being used to process and analyze large amounts of health data, including genetic variation data from pan-genomics and traditional epidemiological data. Machine learning techniques, such as neural networks, support vector machines, and random forests, are powerful classification systems and have been applied in health sciences, such as cancer genomics. In the field of machine learning, model training and algorithm optimization are key to improving predictive accuracy. Machine learning techniques have also demonstrated great potential in the diagnosis and treatment of tuberculosis. A conditional random forest model based on a random forest of conditional inference decision trees and a bagging ensemble algorithm performs best in distinguishing active tuberculosis from latent Mycobacterium tuberculosis infection.
[0016] This method combines pan-genomics and machine learning algorithms. First, through a rigorous sample collection process, samples with clear transmission chains and control samples are selected from tuberculosis cases. After preprocessing data through data cleaning, genome alignment, and epidemiological information collation, multi-dimensional features, including genetic differences, epidemiology, and strain lineages, are screened and extracted and normalized for model construction. The constructed model is used to analyze the transmission relationship of Mycobacterium tuberculosis between patients, thereby conducting traceability analysis and identifying the source of the Mycobacterium tuberculosis transmission.
[0017] Furthermore, the model is used to analyze the transmission relationship of Mycobacterium tuberculosis between patients. By analyzing the transmission relationship between patient B and patient A, it is determined whether the Mycobacterium tuberculosis in patient B was acquired from patient A.
[0018] The evaluation and analysis of the model is conducted on a type of Mycobacterium tuberculosis and a patient (such as patient B) who is infected with the bacterium. The transmission route of the patient (patient B) is traced back to determine where the patient (patient B) acquired the Mycobacterium tuberculosis and whether it was acquired from another patient (patient A).
[0019] Furthermore, the model analyzes the transmission relationship between patient B and patient A by calculating the probability value of whether the Mycobacterium tuberculosis in patient B was transmitted from patient A. The probability value is higher than the threshold, proving that the Mycobacterium tuberculosis in patient B was transmitted from patient A.
[0020] By analyzing the transmission relationship between patient B and patient A, the probability value is high, so it is determined that patient B acquired Mycobacterium tuberculosis from patient A. Then, by analyzing patient A, it is found out where patient A was infected, and the traceability analysis of Mycobacterium tuberculosis can be achieved in turn.
[0021] The non-synonymous mutation ratio refers to the ratio of the number of sites with non-synonymous mutations to the total number of mutation sites at a polymorphic site. A non-synonymous mutation is a mutation that alters the codon encoding an amino acid, thereby affecting the protein's amino acid sequence. A synonymous mutation occurs at the third nucleotide of a codon. Due to the degeneracy of the genetic code, the mutated codon still encodes the same amino acid and does not alter the protein's amino acid sequence. The number of promoter region mutations refers to the number of sites with promoter mutations. These mutations occur in the promoter region or other DNA sequences regulating a gene and can enhance or diminish the promoter's transcriptional initiation function. The total number of SNPs refers to the total number of sites in the genome where DNA sequence diversity is caused by transitions, transversions, insertions, or deletions of a single nucleotide. The total number of InDels refers to the total number of insertions (In) and deletions (Dels) in the genome. The number of structural variations refers to the number of variations in the genome that cause chromosomal structural abnormalities. The number of gene presence / absence variants refers to the number of gene presence (presence) or absence (absence) variants detected in the Mycobacterium tuberculosis genome through pan-genomic analysis. This type of variation reflects the genomic plasticity between different strains. The number of rare mutations refers to the number of genetic variants detected genome-wide with a mutation frequency below a preset threshold (usually 5%). These include low-frequency SNPs / InDels: single nucleotide variants or insertions and deletions with a frequency of less than 5% in the population, and private mutations: mutations that occur only within a single transmission cluster and can serve as molecular markers of transmission chains.
[0022] The Euclidean distance of the patient's residence refers to the Euclidean distance between the residence of the patient to be tested (such as patient B) and the source of infection (such as patient A). The sampling time interval refers to the time difference between the clinical sample collection dates of two tuberculosis patients (such as patient A and patient B), and the calculation formula is: ,in and The sample collection times for patients A and B are respectively. Whether the patients are in the same community refers to whether the patient to be tested (for example, patient B) and the source of infection (for example, patient A) live in the same community (the same community refers to a community with a straight-line distance of less than 1000m and a common walled area where there is a possibility of cross-infection).
[0023] Lineage classification refers to the phylogenetic classification of Mycobacterium tuberculosis strains into genetic branches with a common evolutionary origin based on whole-genome single nucleotide polymorphism (SNP) analysis. Internationally accepted classification systems (such as the Global MTBC Lineage System) divide the MTB complex into eight main lineages (Lineages 1-8) and multiple sublineages (such as Lineage 2.2.1 Beijing strain). Each lineage has unique genomic characteristics, geographic distribution, and clinical phenotypes. Lineage-specific mutation counts refer to the total number of genetic variants (e.g., Rv3346c C913T in Lineage 4) that are highly frequent in a particular MTB lineage (≥95% of strains) but rare (≤5%) in other lineages.
[0024] The number of drug-resistance mutations in the "other characteristics" refers to the total number of genetic variants in the Mycobacterium tuberculosis genome that have been experimentally verified to be significantly associated with anti-tuberculosis drug resistance (such as rifampicin resistance mutations S450L and D435V in the rpoB gene). Mutation annotation was performed using TB-Profiler based on the World Health Organization's "Catalogue of mutations in Mycobacterium tuberculosis complex and their association with drug resistance, 2nd ed." The number of virulence-associated mutations refers to the total number of genetic variants in the Mycobacterium tuberculosis genome that are directly associated with host immune escape, intracellular survival, or tissue destruction. Non-synonymous mutations in known virulence genes were screened based on the Virulence Factor Database (VFDB).
[0025] After evaluating the importance of the features, we finally determined that genetic difference characteristics such as the number of rare mutations, the proportion of non-synonymous mutations, and gene presence / absence variations (PAV) were core features, and epidemiological information (such as geographical distance and sampling time interval) and strain lineage information (such as lineage classification and lineage-specific mutations) were auxiliary features. These will help us to deeply understand the transmission mechanism and provide key support for the precise prevention and control of tuberculosis, such as targeting high-risk groups and optimizing prevention and control strategies.
[0026] Furthermore, a model is constructed using a combination of genetic difference characteristics, epidemiological information characteristics, and strain pedigree information characteristics; the genetic difference characteristics include a combination of the proportion of non-synonymous mutations, the total number of SNPs, the number of gene presence / absence variations, and the number of rare mutations; the epidemiological information characteristics include a combination of the Euclidean distance of the patient's residence, the sampling time interval, and whether the patients are in the same community; the strain pedigree information characteristics include a combination of strain pedigree classification and the number of pedigree-specific mutations; and the strain pedigree information characteristics include strain pedigree classification.
[0027] The present invention investigated the analytical efficiency of models constructed using different feature combinations, and screened out eight feature combinations with good analytical evaluation accuracy and specificity, higher AUC values, and fewer required features. The model constructed using these eight feature combinations can more accurately analyze the transmission relationship of Mycobacterium tuberculosis between different patients, helping to complete the traceability analysis of Mycobacterium tuberculosis more efficiently.
[0028] Furthermore, the model is constructed using a machine learning algorithm, which includes any one or more of a neural network, naive Bayes, random forest, generalized linear model, gradient boosting machine and support vector machine.
[0029] Furthermore, the model is constructed using random forest.
[0030] In some embodiments, in the random forest algorithm, the number of trees is 500, the maximum depth of a single tree is 12, the minimum number of samples for node splitting is 5, and the minimum number of samples for leaf nodes is 2.
[0031] Due to the complexity of the sample characteristics, it was impossible to estimate which machine learning algorithm was most suitable for this study data without prior knowledge. To further improve the accuracy, stability, and generalization of the model for predicting M. tuberculosis transmission, this study conducted an in-depth screening and optimization of machine learning algorithms (neural networks, naive Bayesian, random forests, generalized linear models, gradient boosting machines, and support vector machines). The results clearly demonstrated that different algorithms performed differently for this prediction problem. The model constructed using the Ranger algorithm based on random forests achieved the highest median receiver operating characteristic (ROC) value, with an area under the curve (AUC) of 0.97, significantly outperforming other algorithms (e.g., support vector machine AUC 0.92 and gradient boosting AUC 0.94). The model achieved a sensitivity of 100%, ensuring that all transmission relationships were accurately identified with no false negatives. The accuracy of 98.1% demonstrated high consistency between the model's predictions and the actual transmission relationships. The specificity of 96.2% effectively reduced the false positive rate, preventing non-transmitting clusters from being misclassified as transmitting clusters. The Kappa value of 0.962 further validated the high reliability of the model's predictions. The results show that the model has the best performance in distinguishing transmission clusters from non-transmission clusters. In practical applications, it can more accurately judge the recent spread of Mycobacterium tuberculosis, providing a solid algorithm guarantee for the accurate prediction of the recent spread of Mycobacterium tuberculosis. Therefore, it is defined as the final prediction model.
[0032] For the eight feature combinations obtained through screening, the present invention proportionally divides the population to be evaluated into training and validation sets. The DBSCAN clustering algorithm is used to mine the data structure. Multiple machine learning algorithms are used to train the model and screen it based on performance, ultimately determining that random forest is the optimal algorithm for model construction. The random forest model constructed based on the eight feature combinations achieved AUC values of 0.98 and 0.97 on the training and validation sets, respectively. In the validation set, the accuracy was 0.981, the Kappa value was 0.962, the sensitivity was 100%, the specificity was 96.2%, the positive predictive value was 96.4%, and the negative predictive value was 100%, accurately identifying recent transmission relationships of Mycobacterium tuberculosis.
[0033] The random forest model constructed in this paper outputs prediction results that present transmission relationships as probabilistic values. This not only enhances the operability of transmission monitoring but also provides a key basis for developing optimal tuberculosis prevention and control strategies. Furthermore, the system can flexibly adapt to environmental differences in different regions and changes in the epidemic situation over time, demonstrating broad applicability. It can also assist in efficiently identifying high-risk groups for Mycobacterium tuberculosis, effectively preventing the spread of Mycobacterium tuberculosis, and promoting the comprehensive development of tuberculosis prevention and control work.
[0034] On the other hand, the present invention provides a feature combination for evaluating the transmission relationship of Mycobacterium tuberculosis, wherein the feature combination includes genetic difference features, epidemiological information features and strain lineage information features; the genetic difference features include the proportion of non-synonymous mutations, the total number of SNPs, the number of gene presence / absence variations, and the number of rare mutations; the epidemiological information features include the Euclidean distance of the patient's residence, the sampling time interval, and whether the patients are in the same community; the strain lineage information features include strain lineage classification.
[0035] The feature combination provided by the present invention comprehensively considers factors such as the specific mutation status of the strain, geographic location, and sampling time to determine the relationship between transmission clusters. It does not rely on a fixed SNP difference threshold, effectively avoiding errors caused by different strain mutation rates, and greatly improving the accuracy of transmission relationship judgment. The probability of transmission between cases is calculated using the Ranger algorithm based on random forests. This probability value can provide an accurate quantitative reference for prevention and control work. The transmission detection results are statistically analyzed and related statistical charts are drawn. These charts can intuitively present the transmission patterns and trends, providing clear visual information for researchers and prevention and control personnel. The inference of transmission risk factors can accurately identify high-risk strains, samples, and transmission scenarios, providing a key basis for taking targeted prevention and control measures in advance. The display of transmission hotspots and patterns helps to rationally allocate prevention and control resources, focusing more energy and resources on key areas and key transmission pathways, and achieving efficient prevention and control. It has significantly enhanced the depth and breadth of understanding of the transmission patterns of Mycobacterium tuberculosis, provided a new perspective and strong data support for basic research on tuberculosis; it has also provided accurate and operational guidance for the prevention and control of tuberculosis, which can effectively reduce the risk of tuberculosis transmission and the number of infections; at the same time, by improving the accuracy and reliability of transmission relationship traceability detection, it has optimized the process of formulating tuberculosis prevention and control strategies, saved a lot of manpower, material resources and time costs, and effectively promoted the progress of global tuberculosis prevention and control work, and has important innovation and application value in the field of tuberculosis prevention and control.
[0036] On the other hand, the present invention provides a use of a feature for analyzing the transmission relationship of Mycobacterium tuberculosis, wherein the feature includes any one or more of transmission difference features, epidemiological information features, strain lineage information features, and other features; the genetic difference features include any one or more selected from the following: non-synonymous mutation ratio, synonymous mutation ratio, number of promoter region mutations, total number of SNPs, total number of InDels, number of structural variations, number of gene presence / absence variations, number of rare mutations; the epidemiological information features include any one or more selected from the following: Euclidean distance of the patient's residence, sampling time interval, whether the patients are in the same community; the strain lineage information features include any one or more selected from the following: strain lineage classification, number of lineage-specific mutations; the other features include any one or more selected from the following: number of drug resistance-related mutations, number of virulence-related mutations.
[0037] Furthermore, the characteristics include genetic difference characteristics, epidemiological information characteristics and strain lineage information characteristics; the genetic difference characteristics include the proportion of non-synonymous mutations, the total number of SNPs, the number of gene presence / absence variations and the number of rare mutations; the epidemiological information characteristics include the Euclidean distance of the patient's residence, the sampling time interval and whether the patients are in the same community; the strain lineage information characteristics include strain lineage classification.
[0038] Furthermore, the constructed model is used to evaluate the transmission relationship and traceability of Mycobacterium tuberculosis, find out how Mycobacterium tuberculosis is transmitted step by step to a large number of patients, and provide a theoretical basis for preventing and controlling the spread of Mycobacterium tuberculosis.
[0039] Furthermore, a machine learning algorithm is used to construct the model, and the machine learning algorithm includes any one or more combinations of neural networks, naive Bayes, random forests, generalized linear models, gradient boosting machines and support vector machines.
[0040] Furthermore, the machine learning algorithm includes a random forest algorithm.
[0041] The beneficial effects of the present invention are:
[0042] 1. This invention provides a novel system for detecting the transmission relationships of Mycobacterium tuberculosis. By integrating multiple genetic difference characteristics (including the proportion of non-synonymous mutations, the proportion of synonymous mutations, the number of promoter region mutations, the total number of SNPs, the total number of indels, the number of structural variants, the number of gene presence / absence variants, and the number of rare mutations), epidemiological information characteristics (including Euclidean distance from the patient's residence, sampling interval, and whether the patients lived in the same community), strain lineage information characteristics (including strain lineage classification and the number of lineage-specific mutations), and other characteristics (including strain lineage classification and the number of lineage-specific mutations), this system breaks through the limitations of traditional research that only focuses on a few gene loci or a single mutation type. It deeply explores the genetic diversity of strains and provides richer information for accurately understanding the genetic evolution and transmission basis of Mycobacterium tuberculosis. It improves the detection coverage of different lineages of MTB genomes and can comprehensively reflect the genetic relationships between strains.
[0043] 2. This invention utilizes a combination of genetic variation and epidemiological information to effectively reduce the error associated with detecting a single SNP difference threshold. Instead of relying on a fixed threshold, the method incorporates multiple genetic variations within strains and detailed epidemiological data to more accurately determine transmission relationships. This significantly improves the accuracy of M. tuberculosis transmission pattern predictions and avoids erroneous conclusions resulting from incomplete information.
[0044] 3. This invention specifically focuses on the presence and absence of specific mutations, providing a more accurate assessment of the likelihood of transmission. This in-depth analysis of the impact of each mutation on the strain's transmission characteristics, particularly the key role of rare mutations in the adaptive evolution and spread of strains, enables a more refined assessment of the probability of strain transmission, providing a more targeted basis for the formulation of prevention and control measures, and enhancing the scientific nature of prevention and control efforts.
[0045] 4. This invention utilizes machine learning clustering and classification algorithms to rapidly process complex pan-genomic and epidemiological data. Faced with massive amounts of genetic data and complex epidemiological information, the algorithm demonstrates powerful data processing capabilities, rapidly extracting valuable information and significantly improving research efficiency. This ensures rapid support for prevention and control decisions, meeting the timeliness requirements of modern tuberculosis prevention and control.
[0046] 5. This invention uses machine learning prediction results to represent the likelihood of transmission as a probability value, which helps improve the operability of transmission monitoring. Prevention and control personnel can implement hierarchical management of different regions, populations, and strains based on the probability, develop personalized and efficient tuberculosis prevention and control strategies, and rationally allocate prevention and control resources, such as focusing on monitoring and intervention in high-risk areas and populations, to improve prevention and control effectiveness and reduce the risk of tuberculosis transmission.
[0047] 6. The present invention can adapt to the spread of Mycobacterium tuberculosis in different regions and over time, and has wide applicability. Whether in high- or low-incidence areas of tuberculosis, and at different epidemic stages, the Mycobacterium tuberculosis transmission relationship detection system provided by the present invention can effectively function, providing a unified and reliable technical means for global tuberculosis prevention and control, effectively promoting the progress of global tuberculosis prevention and control work, and having important application value in different prevention and control scenarios.
[0048] 1. Tuberculosis patients
[0049] The tuberculosis patients refer to patients who have symptoms such as cough, sputum, low fever, night sweats, fatigue, and weight loss due to infection with Mycobacterium tuberculosis and who have been diagnosed with tuberculosis through relevant examinations.
[0050] Tuberculosis is a chronic infectious disease caused by Mycobacterium tuberculosis, primarily transmitted through the air. After infection, most people remain in a latent state without symptoms, with the bacteria controlled by the immune system. However, when immunity declines, approximately 5-10% of latently infected individuals will develop active TB, characterized by persistent cough, sputum production, hemoptysis, chest pain, fever, night sweats, weight loss, and fatigue. TB can affect the lungs (pulmonary TB) or other organs (extrapulmonary TB) such as the kidneys, brain, spine, and skin. Diagnosis relies on imaging, microbiological, and immunological tests, and treatment requires long-term multidrug therapy. Drug resistance is an increasing problem. Preventive measures include BCG vaccination and infection control.
[0051] The diagnosis of tuberculosis can be confirmed through laboratory testing for Mycobacterium tuberculosis, primarily including sputum smear acid-fast staining, mycobacterial culture, and molecular biological testing (such as PCR). M. tuberculosis culture, which includes both solid and liquid cultures, is the gold standard for laboratory diagnosis of TB. It can confirm the presence of the pathogen and provide drug susceptibility test results, but is time-consuming (usually 2-6 weeks). Molecular biological testing (such as Xpert MTB / RIF) can simultaneously detect M. tuberculosis and its drug resistance in a relatively short period of time, significantly improving diagnostic efficiency. Combining clinical symptoms, imaging studies, and immunological tests (such as the tuberculin skin test or interferon-gamma release assay) can further improve diagnostic accuracy.
[0052] 2. Mycobacterium tuberculosis
[0053] Mycobacterium tuberculosis (M. tuberculosis), also known as tubercle bacilli, is the causative agent of human tuberculosis. It is an obligately aerobic bacterium that stains positively for acid fast. It lacks flagella but possesses pili and a microcapsule, but does not form spores. Its bacterial wall lacks either the teichoic acid of Gram-positive bacteria or the lipopolysaccharide of Gram-negative bacteria. German bacteriologist Robert Koch (1843-1910) discovered and confirmed it as the causative agent of human tuberculosis in 1882. Tuberculosis, caused by infection with this bacterium, is a serious infectious disease that threatens human life and health. After centuries of struggle, it has been gradually brought under control. However, in recent years, the disease has become increasingly severe due to a variety of factors.
[0054] Mycobacterium tuberculosis can undergo variations in morphology, colonization, virulence, immunogenicity, and drug resistance. Bacillus Calmette-Guérin (BCG) is a live attenuated vaccine strain derived by Calmette and Guerin (1908) by passaged bovine tuberculosis 230 times over 13 years in a culture medium containing glycerol, bile, and potatoes. It is now widely used for preventive vaccination.
[0055] 3. Transmission relationship of Mycobacterium tuberculosis
[0056] The transmission relationship of Mycobacterium tuberculosis studies the transmission relationship of Mycobacterium tuberculosis between patients, so as to conduct traceability analysis to find out where the Mycobacterium tuberculosis is transmitted from, where the source is, and from which patient it is transmitted to which patient, so as to construct a complete MTB transmission network map.
[0057] This invention constructs a Mycobacterium tuberculosis transmission relationship detection system. This system evaluates and analyzes a strain of Mycobacterium tuberculosis and a patient carrying it (e.g., Patient B). It traces the source of infection to determine where Patient B acquired the Mycobacterium tuberculosis, including whether it was acquired from another patient (Patient A). This system calculates the probability of whether Patient B acquired the Mycobacterium tuberculosis from Patient A, thereby analyzing the transmission relationship between Patients B and A. If the probability exceeds a threshold, Patient B is determined to have acquired the Mycobacterium tuberculosis from Patient A.
[0058] In some methods, the threshold of the probability value is 0.7. When the probability value is ≥0.7, it is transmission-related, and it is judged that patient B is acquired from patient A; when the probability value is <0.7, it is judged to be unrelated, that is, patient B is not acquired from patient A.
[0059] The pan-genomics and machine learning-based Mycobacterium tuberculosis transmission relationship detection system provided by this invention covers the entire data processing process, from sample collection to feature extraction and integration. The sample collection module can accurately screen samples of cases with clear recent transmission chains and their close contacts from known tuberculosis transmission cases, while also including samples with no clear epidemiological transmission links as controls. The sample collection module strictly follows specific criteria or supplementary case diagnostic criteria to ensure sample reliability and representativeness.
[0060] The data preprocessing module uses fastp and Kraken software to efficiently remove low-quality, repetitive and non-MTB source sequences from sequencing data. With the help of bwa software, the sequencing data that has passed quality control is carefully compared with a specific reference genome to obtain the variation position of the sample genome and the reference genome and the coverage information of the sequencing data on the genome. It also normalizes the epidemiological data covering comprehensive patient information.
[0061] The feature engineering module is based on pan-genomic analysis and uses the GATK and SNPEff software packages to deeply explore the number, type (such as SNPs, InDels, SVs, etc.) and frequency information of genetic mutations in strains, extract information on drug resistance and virulence genes, and construct numerical features of genomic mutations. With the help of the geographic information system (GIS), the Euclidean distance to the patient's residence is calculated, the sampling time interval is determined, and epidemiological category variables are encoded. The TBprofiler software package is used to classify strain lineages and identify specific mutations. Finally, the normalized data is integrated to eliminate dimensional differences and provide a high-quality data foundation for model training.
[0062] This system effectively addresses key challenges in existing technologies for Mycobacterium tuberculosis transmission research. It overcomes the drawback of traditional methods that inaccurately determine transmission relationships based solely on fixed SNP difference thresholds, and comprehensively considers multiple genetic differences and epidemiological information, significantly reducing errors. It also addresses the previous incomplete consideration of genetic variation types by fully incorporating multiple variation types. It also fully utilizes previously overlooked epidemiological information to enable more comprehensive identification of transmission patterns. It also enables quantitative assessment of transmission probability, overcoming the limitations of existing methods that rely solely on qualitative judgments.
[0063] Ultimately, the goal is to accurately detect the transmission relationship of Mycobacterium tuberculosis, accurately assess the transmission risk, effectively identify high-risk groups, provide a scientific basis for preventing transmission, accurately calculate the probability of transmission between strains, effectively promote tuberculosis prevention and control, improve the efficiency and accuracy of prevention and control, play a key role in the global tuberculosis prevention and control process, significantly reduce the transmission risk of tuberculosis, reduce the number of infections, save a lot of manpower, material and time costs, and bring innovative solutions and important application value to the field of tuberculosis prevention and control. BRIEF DESCRIPTION OF THE DRAWINGS
[0064] Figure 1 Workflow for a pan-genomics and machine learning-based Mycobacterium tuberculosis transmission detection system.
[0065] Figure 2 The k-distance distribution plot for DBSCAN clustering is used to select an appropriate eps (neighborhood distance threshold) parameter.
[0066] Figure 3 This is a DBSCAN cluster scatter plot of MTB samples based on pan-genomic and epidemiological data.
[0067] Figure 4 Performance evaluation of the optimal models constructed for the six algorithms.
[0068] Figure 5 Performance evaluation of the optimal models constructed for the six algorithms.
[0069] Figure 6 This is the MTB propagation network diagram predicted by the model. DETAILED DESCRIPTION
[0070] The present invention will be described in further detail below in conjunction with the accompanying drawings and Examples. It should be noted that the following examples are intended to facilitate understanding of the present invention and do not serve to limit the present invention in any way. The reagents used in this example are all known products and were obtained by purchasing commercially available products.
[0071] Example 1: Collection of Mycobacterium tuberculosis transmission samples and initial screening of characteristics
[0072] In the field of Mycobacterium tuberculosis transmission research, comprehensive and accurate acquisition of sample information and extraction of key features are crucial for building efficient transmission relationship detection models. This example is dedicated to developing a computational analysis method for detecting recent MTB transmission based on pan-genomics and machine learning. The goal is to accurately predict the transmission pathways and associations of M. tuberculosis between cases by deeply integrating multi-source data and advanced technologies.
[0073] The specific workflow of the Mycobacterium tuberculosis transmission relationship detection system based on pan-genomics and machine learning is as follows: Figure 1 As shown, the steps are as follows:
[0074] 1. Sample Collection
[0075] From known tuberculosis transmission cases, a total of 150 samples of cases and their close contacts with recent clear transmission chains were screened. Epidemiological investigation and tracing confirmed direct or indirect transmission relationships between these samples. These included 100 confirmed cases and 50 close contacts, all of whom were diagnosed within the past six months.
[0076] At the same time, to ensure the rigor of the study, 60 samples with no obvious epidemiological transmission links were collected as a control group. These control samples were patients who visited the same area and during the same time period but had no contact with cases in the known transmission chain, so as to better evaluate the specificity and accuracy of the model.
[0077] The criteria for cases with clear transmission chains and their close contacts in the samples are as follows: with reference to the "Law of the People's Republic of China on the Prevention and Control of Infectious Diseases" and the "Guidelines for the Diagnosis and Treatment of Pulmonary Tuberculosis" formulated by the Tuberculosis Branch of the Chinese Medical Association, specifically including: Case diagnostic criteria: in accordance with the clinical diagnostic criteria in the "Diagnosis of Pulmonary Tuberculosis (WS 288-2017)", including: typical tuberculosis symptoms (such as persistent cough, sputum, low fever, night sweats, etc.); imaging examinations (such as chest X-ray or CT) show pulmonary tuberculosis lesions; laboratory tests (such as sputum smear, culture or molecular testing) are positive.
[0078] Close contact criteria: defined as people who have had the following contact with a confirmed case within 2 weeks before the case was confirmed: family members living together; people who have had long-term contact in the same workplace or learning environment; people who participate in the same social activities.
[0079] 2. Data Preprocessing
[0080] 1. Sequencing data cleaning: Use fastp and Kraken software to remove low-quality sequences, repeated sequences, and non-MTB source sequences in the sequencing data to ensure data quality and consistency;
[0081] 2. Genome sequence alignment: Use bwa software to align the sequencing data that have passed quality control with the reference genome H37Rv (NCBI Reference Sequence: NC_000962.3) to determine the variation positions between the sample genome and the reference genome, as well as the coverage information of the sequencing data on the genome;
[0082] 3. Epidemiological information collation: The epidemiological information collected through the questionnaire survey (covering the patient's basic demographic information, living environment details, social and travel activity trajectories, past medical history and treatment details, etc.) is standardized to ensure the accuracy and completeness of the data.
[0083] 3. Feature Engineering
[0084] 1. Genomic feature extraction:
[0085] The genomic signature is derived from Mycobacterium tuberculosis samples cultured from sputum. The specific steps are as follows:
[0086] Sample collection: Sputum samples were collected from patients diagnosed with TB.
[0087] Bacterial culture: A sputum sample is cultured on a selective medium to isolate Mycobacterium tuberculosis (eg, Lowenstein-Jensen medium).
[0088] DNA extraction and library construction: Genomic DNA was extracted from the cultured bacteria.
[0089] Genome sequencing: Use high-throughput sequencing technology (such as the Illumina sequencing platform) to perform whole-genome sequencing on the extracted DNA to obtain the genomic information of Mycobacterium tuberculosis.
[0090] Data Analysis:
[0091] (1) Based on pan-genomic analysis, the GATK and SNPEff software packages were used to calculate and annotate the number, type (SNP (single nucleotide variation), InDel (insertion or deletion fragment), SV (large structural variation, etc.)) and frequency (>95% common mutations and <5% rare mutations) of gene mutations in each sample to fully understand the genetic variation of the strain;
[0092] (2) Construct numerical features of genomic mutations and convert each type of mutation into a numerical representation, with presence marked as 1 and absence marked as 0.
[0093] (3) Extract the drug resistance gene mutation information in the genome based on the "Catalogue of Gene Mutations Related to Drug Resistance in Mycobacterium Tuberculosis (2nd Edition)" published by the World Health Organization (WHO). Extract the mutation information of MTB virulence genes in the genome based on the VFDB database.
[0094] 2. Epidemiological feature extraction:
[0095] Clinical characteristics were obtained from patients' medical records and epidemiological surveys.
[0096] (1) Use geographic information systems (GIS) (such as ArcGIS software) to convert the place of residence into a geographic area code, calculate the Euclidean distance between the places of residence of patients, and construct a distance matrix between patients;
[0097] (2) Calculate the sampling time interval between samples as a time series feature (R language);
[0098] (3) Encode categorical variables in epidemiology (such as gender, occupation, medical history, treatment plan, travel history and contact history (if any)) and convert each type into a binary vector representation (R language).
[0099] 3. Strain pedigree information:
[0100] The strain lineage information is annotated based on the data of whole genome sequencing of Mycobacterium tuberculosis:
[0101] (1) Lineage classification of strains: Based on the TBprofiler software package (default parameters), strains are classified into lineages (such as Lineage 2, Lineage 4, etc.) according to whole genome data. The differences in transmission characteristics, pathogenicity and drug resistance of different lineages are analyzed to provide an evolutionary perspective for transmission tracing.
[0102] (2) Lineage-specific mutations: Based on the TBprofiler lineage annotation database, we identify unique mutation sites in different lineages, analyze the key role of these specific mutations in the adaptive evolution and transmission advantage formation of strains, and explore their potential value as strain transmission markers.
[0103] 4. Data integration and normalization: Align pan-genomic and epidemiological data with relevant columns based on row indices. Use the Scikit-learn framework to standardize the data, perform one-hot encoding on categorical features (such as mutation type and lineage classification), and standardize or normalize numerical features (such as number of mutations and geographic distance) to map to a specific interval (with a mean of 0 and a standard deviation of 1). This eliminates the impact of data dimension differences on the model and improves model stability and accuracy.
[0104] 4. Impact of different feature combinations on evaluation performance
[0105] Through single-factor analysis and multi-factor combination analysis, the features that have a significant contribution to propagation detection are screened out. The specific steps are as follows:
[0106] Univariate feature analysis: A univariate classification model (such as logistic regression) was constructed for each of the 15 candidate features obtained by screening, and the AUC, sensitivity, specificity, and SHAP value were calculated (see Table 1).
[0107] Multivariate combination analysis: Based on the single-factor results, the features were grouped according to their contribution (genetic differences, epidemiological information, strain lineage, and other features), and the model performance of different combinations was tested (see Table 2).
[0108] Table 1. Individual feature importance analysis results
[0109]
[0110] As shown in Table 1, the five features—nonsynonymous mutation ratio, total number of SNPs, gene presence / absence variation (PAV), number of rare mutations, and strain lineage classification—all had high contributions (AUC>0.75 and SHAP>0.25), while the three features—synonymous mutation ratio, number of structural variants, and number of virulence-associated mutations—had low contributions (AUC<0.70 or SHAP<0.15). Furthermore, while the Euclidean distance to patient residence, sampling interval, and whether patients lived in the same community had low SHAP values, these three features, along with their nonlinear interactions, enhance model robustness and possess irreplaceable epidemiological value in key transmission windows (e.g., close proximity / short time) and control feasibility (e.g., precise screening scope delineation), and can still be considered to have high contributions.
[0111] This example examines combining different features from the same group, as well as different features from different groups, to construct models for detecting the transmission relationship of Mycobacterium tuberculosis. The diagnostic performance of each model is evaluated. A random forest approach is used to construct the model. During the model construction process, the above-mentioned parameter combinations are tested using a grid search (GridSearchCV) algorithm to ultimately select the optimal parameter combination, specifically:
[0112] 1) The number of trees (n_estimators) is set to 500. By increasing the number of base learners, the model stability and prediction accuracy are improved and the risk of overfitting is reduced. 2) The maximum depth of a single tree (max_depth) is set to 12 to limit the overgrowth of the tree to balance the model complexity and generalization ability. 3) The minimum number of samples for node splitting (min_samples_split) is set to 5 to prevent overfitting of a small number of samples. 4) The minimum number of samples for leaf nodes (min_samples_leaf) is set to 2 to ensure that the leaf nodes have a statistically significant sample size. 5) The feature selection method (max_features) is set to 'sqrt', that is, each tree randomly selects the square root of the number of features for splitting, which enhances model diversity and reduces the impact of inter-feature correlation on the results. 6) The splitting criterion (criterion) is set to 'gini', which uses the Gini coefficient as the criterion for node splitting. It is suitable for classification tasks, has high computational efficiency and is robust to noisy data. 7) The sample sampling method (bootstrap) is set to True, which uses the bootstrap method (bootstrap The training set for each tree is generated by sampling to further improve the generalization ability of the model. 8) The sample sampling ratio (max_samples) is set to 0.8, and each tree uses 80% of the samples for training and retains 20% of the samples for validation to avoid excessive dependence of the model on specific samples.
[0113] The evaluation results are shown in Table 2.
[0114] Table 2. Impact of feature combinations of different groups on evaluation performance
[0115]
[0116] Table 2 shows that the combination of high-contribution features (five) with epidemiological information (three) significantly improved the model's performance in assessing the recent transmission risk of Mycobacterium tuberculosis compared to other combinations, achieving an AUC of 0.97, 100% sensitivity, and 96.2% specificity. However, the addition of low-contribution features (such as the proportion of synonymous mutations, the number of structural variants, and the number of virulence-related mutations) not only failed to improve the model's performance but actually decreased it (for example, the AUC for the combination of all 15 features was only 0.93). Conversely, the model performance of all 15 features (AUC 0.93) was lower than that of the subset of high-contribution features (AUC 0.95), demonstrating that blindly increasing the number of features introduces noise and reduces model generalization. Therefore, the combination of high-contribution features (nonsynonymous mutation proportion, total number of SNPs, gene presence / absence variants, number of rare mutations, and strain lineage classification) with epidemiological information (Euclidean distance between patient residences, sampling interval, and whether patients lived in the same community) is preferred for constructing a model for assessing the recent transmission risk of Mycobacterium tuberculosis. Reducing redundant features can improve computational efficiency by 30% while avoiding overfitting.
[0117] The final model did not use all 15 features. Instead, eight key features were selected through screening in Tables 1 and 2. The eight features included: (1) genetic differences: the proportion of non-synonymous mutations, the total number of SNPs, gene presence / absence variation (PAV), and the number of rare mutations; (2) epidemiology: the Euclidean distance of the patient's residence, the sampling time interval, and whether the patients lived in the same community; and (3) strain lineage information: strain lineage classification.
[0118] The detection methods for the eight key features are: non-synonymous mutation ratio, total SNP number, gene presence / absence variation (PAV), and the number of rare mutations.
[0119] 1. Proportion of non-synonymous mutations
[0120] step:
[0121] 1) Use GATK HaplotypeCaller and BCFtools to perform SNP and InDel detection on whole genome sequencing data;
[0122] 2) annotating the biological effects of variants (non-synonymous, synonymous, and nonsense) using SNPEff;
[0123] Calculate the ratio:
[0124] 2. Total number of SNPs
[0125] step:
[0126] 1) Sequence alignment was performed using BWA-MEM based on the H37Rv reference genome (NC_000962.3);
[0127] 2) Remove sequencing / alignment errors using GATK FilterMutectCalls (parameters: --min-reads 5, --min-allele-fraction 0.05);
[0128] The total number of SNPs after filtering (including homologous recombination regions and non-coding regions) was counted.
[0129] 3. Number of gene presence / absence variants (PAVs)
[0130] step:
[0131] 1) Construction of the MTB pan-genome: Panaroo (parameter: --core_threshold 0.95) was used to integrate the gene sets of multiple strains;
[0132] 2) Determine the presence / absence of target strain genes by BLASTn (E-value ≤1e-5, coverage ≥80%);
[0133] 3) Count the number of PAVs relative to the reference genome (present = 1, absent = 0).
[0134] 4. Number of rare mutations
[0135] Based on the total SNP number detection in step 1, LoFreq (parameters: --min-cov 20, --min-alt 3) was further used to detect low-frequency mutations (frequency < 5%);
[0136] The detection methods of the Euclidean distance of the patient's residence, the sampling time interval, and whether the patients are in the same community are as follows:
[0137] 1. Euclidean distance of the patient's residence
[0138] step:
[0139] 1) Data source: extract patient addresses from medical registration systems or electronic health records (EHR);
[0140] 2) Geocoding: Use the mapping software API to convert the address into longitude and latitude coordinates;
[0141] 3) Distance calculation: Based on the spherical Euclidean distance formula (Haversine formula), use the Python geopy.distance library to calculate the straight-line distance between two points (unit: meters):
[0142] 5. Detection method of sampling time interval:
[0143] 1) Data extraction: Sample collection timestamps (format: YYYY-MM-DD HH:MM:SS) were exported from the laboratory information management system (LIMS);
[0144] 2) Time difference calculation: Use the Python datetime module to calculate the number of days between two samples;
[0145] 3) Normalization: convert days to months (divided by 30) and standardize to a Z-score (mean = 9.2, standard deviation = 4.1);
[0146] 2. Methods for determining whether patients are in the same community:
[0147] Geographical proximity (one of these is sufficient):
[0148] 1) Straight-line distance ≤ 1000 meters (calculated using the above Euclidean distance);
[0149] 2) Administrative boundary matching: Use GIS layers (such as ArcGIS Shapefile) to determine whether they belong to the same administrative village / street.
[0150] 3. The detection method for strain lineage classification is:
[0151] The results of total SNP detection were input into TB-profiler software (database version: v4.4.0) to annotate based on lineage-specific marker SNPs and output the lineage classification results of the strains.
[0152] Through the above systematic and rigorous sample collection and feature extraction process, this example successfully obtained rich and high-quality characteristic data related to the transmission of Mycobacterium tuberculosis, laying a solid foundation for the construction of subsequent prediction models, and effectively promoting the development of the model in the direction of accuracy and efficiency. It is expected to significantly improve the ability to predict the transmission paths and associations of the disease, provide a scientific and reliable basis for clinical decision-making and public health strategy formulation for tuberculosis prevention and control, and have broad application prospects in the field of Mycobacterium tuberculosis transmission research and prevention and control.
[0153] Example 2 Construction of an optimal detection model for the recent spread of Mycobacterium tuberculosis
[0154] In Example 1, sample collection and initial feature screening were successfully completed, yielding characteristic data related to the spread of Mycobacterium tuberculosis. To fully exploit the information contained in this data and construct a high-performance transmission detection model, this example employed a series of model training and optimization strategies to construct an optimal model for detecting the recent spread of Mycobacterium tuberculosis. The specific process is as follows:
[0155] 1. Construction of the Optimal Model
[0156] 1. Divide the training set and test set
[0157] The data is randomly divided into training set and validation set in a ratio of 7:3, that is, 70% of the samples are used as training set for model training, and 30% of the samples are used as validation set for model evaluation and hyperparameter adjustment.
[0158] 2. Clustering model implementation and evaluation
[0159] The DBSCAN clustering algorithm in the Scikit-learn framework is used to perform data cluster analysis, discovering the potential cluster structure in MTB propagation, reducing the computational complexity and data volume of subsequent classification tasks, making the classification process more focused on representative clusters, and effectively improving the efficiency and accuracy of classification. Figure 2 As shown in Figure 1, the optimization of DBSCAN parameters mainly includes determining the neighborhood radius (Epsilon value) and the minimum number of samples, where the neighborhood radius reference value is 0.5 and the minimum number of samples is 5. Figure 2 By observing the changes in the distribution density of data points at different distances and combining it with the elbow rule, we determined the appropriate density of data points when the neighborhood radius is 0.5. As the k-distance increases, the density of data points gradually decreases. At a certain inflection point, the density decline slows significantly. The distance corresponding to this inflection point can be used as a preliminary reference value for the neighborhood radius. Within this distance range, the distribution of data points begins to show a relatively stable state, which can better reflect the local density characteristics of the data. This allows us to preliminarily determine an appropriate neighborhood range within which data points form a relatively tight cluster structure. Furthermore, based on the scale and distribution of the data, as well as the requirements for cluster compactness and separation, the minimum number of samples is determined through multiple experiments and verifications. For example, for small datasets with relatively concentrated data distribution, the minimum number of samples can be appropriately reduced. However, for large and complex datasets, the minimum number of samples should be increased to avoid generating excessive noise clusters.
[0160] from Figure 3 As can be seen from the results, the clustering results show a clear cluster distribution, with a certain degree of separation between different clusters, and the samples within the cluster have a high degree of similarity in genetic and epidemiological characteristics. Specifically,
[0161] 1) Cluster distribution corresponds to transmission relationships: The figure shows six main clusters (Cluster 0-5), each of which corresponds to an independent transmission chain (for example, Cluster 0 contains 10 samples, all of which are transmission cases within the same community);
[0162] 2) Significant separation between clusters: The center distance between different clusters is greater than 2.0 (Euclidean distance), indicating that there is no cross-infection between transmission chains;
[0163] 3) Intra-cluster consistency: The genetic difference characteristics (such as the average number of SNPs ≤ 5, the proportion of rare mutations ≥ 80%) and epidemiological information (such as geographical distance ≤ 1 km, sampling time interval ≤ 30 days) of samples within the same cluster are highly consistent, which is consistent with the biological laws of recent transmission.
[0164] Storing cluster labels in corresponding variables will serve as important auxiliary information for subsequent classification models, helping them better understand the potential relationships between samples and further improve model performance. Specific functions include:
[0165] 1) Reduce data complexity: Divide samples into subsets using cluster labels to reduce the amount of noise data that the classification model needs to process;
[0166] 2) Enhanced feature expression: Cluster labels are used as one of the input features of the classification model to provide information about the sample's group affiliation and help the model capture the spatial-temporal patterns of propagation relationships.
[0167] 3) Optimize training efficiency: Design differentiated training strategies for different clusters (such as increasing sample weights for clusters with high transmission risks) to improve the model's ability to identify key transmission chains.
[0168] Therefore, it can be seen that the DBSCAN clustering results not only verify the clustering characteristics of Mycobacterium tuberculosis transmission, but also provide key structured prior knowledge for machine learning models, enabling the models to identify transmission relationships more efficiently and accurately, and solving the performance bottleneck problem caused by data redundancy and noise in traditional methods.
[0169] 3. Screening of the best machine learning algorithm and construction of the best prediction model
[0170] Due to the complexity of the sample characteristics, it is impossible to estimate which machine learning algorithm is more suitable for this research data without prior knowledge. Therefore, in this embodiment, six machine learning algorithms, including neural network, naive Bayes, random forest, generalized linear model, gradient boosting machine and support vector machine, are used to construct prediction models respectively. The construction methods of the six prediction models are as follows: First, the six machine learning models are trained using the optimal hyperparameters. The hyperparameter configuration is determined by grid search and 10-fold cross validation, as shown in Table 3. Model training, verification and optimization are then completed according to the following process.
[0171] 1) Model training: Based on the hyperparameter configurations in Table 3, GPUs are used to accelerate the training of neural networks, while multi-core CPUs are used to train other models in parallel to optimize computational efficiency.
[0172] 2) Dynamic evaluation: Automatically switches evaluation indicators based on data distribution characteristics (such as whether categories are balanced), and additionally tests forecast stability for time series data.
[0173] 3) Performance comparison: We comprehensively compared the six models based on prediction accuracy (such as AUC), computational speed, and feature importance, and selected random forest and neural networks as the dominant models.
[0174] 4) Ensemble optimization: The prediction results of the optimal model are combined with the original features and retrained through logistic regression. The false positive rate of the final ensemble model is reduced by 40%.
[0175] Table 3. Hyperparameter selection of machine learning models
[0176]
[0177] According to the transmission probability threshold (≥0.7), it is judged whether the strains belong to the same transmission cluster (that is, the transmission probability ≥0.7 is judged to be transmission-related, and <0.7 is judged to be unrelated). The 10-fold cross-validation method is used to screen out the optimal model under the 6 algorithms. The specific steps are as follows: According to the transmission probability threshold (≥0.7), it is judged whether the strains belong to the same transmission cluster (that is, the transmission probability ≥0.7 is judged to be transmission-related, and <0.7 is judged to be unrelated). The 10-fold cross-validation method is used to screen out the optimal model under the 6 algorithms. The specific steps are as follows: 0107. Cross-validation design: The training set is randomly divided into 10 subsets, and 9 subsets are taken each time to train the model. The remaining 1 subset is used to verify the performance. The average result is taken after 10 cycles; Performance evaluation indicators: The area under the ROC curve (AUC) is the core indicator, supplemented by sensitivity, specificity, and accuracy;
[0178] Optimal model screening: The model with the highest AUC value in cross-validation was selected as the optimal model for each algorithm.
[0179] The ROC method was used to evaluate the performance of the optimal model. The results are as follows: Figure 4 As shown, the model built with the Random Forest Ranger algorithm achieved the highest median ROC value, 0.97 (95% confidence interval: 0.95-0.99), significantly outperforming other algorithms (e.g., Gradient Boosting Machine AUC = 0.94, Support Vector Machine AUC = 0.92). This model demonstrated the following advantages in cross-validation:
[0180] Sensitivity 100%: ensuring that all true transmission clusters are accurately identified; specificity 96.2%: effectively avoiding misjudging non-transmission clusters as transmission clusters; Kappa value 0.962: the model prediction results are highly consistent with the true transmission relationship.
[0181] The above results show (Table 4) that the random forest model has the best performance in distinguishing propagation clusters from non-propagation clusters. Its core advantages are:
[0182] Robustness to high-dimensional data: Ability to effectively handle nonlinear relationships between genetic differences (such as rare mutations) and epidemiological information (such as geographic distance);
[0183] Anti-overfitting ability: Enhance model generalization ability through random feature selection (max_features=sqrt) and sample bootstrap sampling (bootstrap=True);
[0184] Interpretability: Feature importance ranking based on SHAP values (e.g., rare mutations have the highest contribution) provides a basis for propagation mechanism analysis.
[0185] Therefore, the random forest-based ranger algorithm was defined as the final prediction model, providing a high-precision and high-reliability technical solution for the recent spread detection of Mycobacterium tuberculosis.
[0186] Table 4. Comparison of the performance of the optimal models of the six algorithms
[0187]
[0188] 4. Performance Verification of the Best Prediction Model
[0189] The performance of the final prediction model was verified in the validation set. The ROC analysis method was used to evaluate the model performance. The performance indicators included: accuracy, Kappa value, sensitivity, specificity, positive predictive rate, negative predictive rate and feature importance. The results are as follows Figure 5As shown, the final prediction model achieved an AUC of 0.97 in the validation set, demonstrating high accuracy and reliability in distinguishing positive and negative samples, effectively identifying pairs of samples with and without transmission relationships. The accuracy was 0.981, indicating that the model correctly predicted 98.1% of samples, further demonstrating the model's high precision. The Kappa value was 0.962, reflecting a high degree of consistency between the model's predictions and the actual results, eliminating any high accuracy attributable to chance. The sensitivity was 100%, indicating that the model accurately identified all samples with actual transmission relationships, with no false positives. The specificity was 96.2%, indicating that the model also had high accuracy in identifying samples without transmission relationships, with only a small number of false positives. The positive predictive value was 96.4%, indicating that a high proportion of samples predicted as positive (with transmission relationships) by the model were actually positive. The negative predictive value was 100%, indicating that all samples predicted as negative (without transmission relationships) by the model were actually negative, ensuring the reliability of the model's negative predictions.
[0190] Therefore, the above performance indicators comprehensively show that the optimal prediction model constructed by the present invention has excellent performance, can accurately detect the recent spread of Mycobacterium tuberculosis, and provide strong technical support for tuberculosis prevention and control.
[0191] 5. Feature Importance Assessment
[0192] This example uses a random forest model to calculate the importance score of each feature.
[0193] Feature importance score calculation process:
[0194] 1) Data Preparation: Collect and clean the dataset to ensure it contains all relevant features (such as the proportion of non-synonymous mutations, total number of SNPs, gene presence / absence variants, number of rare mutations, Euclidean distance from patient residence, sampling interval, strain lineage classification, and number of resistance-associated mutations). The dataset is divided into feature variables (X) and target variables (Y), with the target variable being a label related to the transmission of Mycobacterium tuberculosis. 2) Training the Random Forest Model: The random forest model is trained using the training set data. During training, the model learns how to predict the target variable based on the input features. 3) Calculating Feature Importance: During training, the random forest model automatically calculates an importance score for each feature. The importance score reflects the contribution of the feature to the model's decision-making. Specifically, the model analyzes the role of each feature in the node splitting process and calculates the reduction in impurity when using that feature, thereby deriving an importance score for each feature. 4) Result Compilation: The importance scores of each feature are organized into a table for easy comparison and analysis. Features are ranked according to their importance scores, typically from high to low, to identify those with the greatest impact on model predictions.
[0195] The feature importance scores of the MTB propagation detection model are shown in Table 5.
[0196] Table 5. Feature importance scoring results of the MTB propagation detection model
[0197]
[0198] Results Analysis and Interpretation
[0199] Features with high importance scores (higher scores mean more important):
[0200] Non-synonymous mutation ratio (0.30): This is the most influential feature in the model, indicating that this feature plays a significant role in judging the risk of transmission.
[0201] Total number of SNPs (0.25): This feature also has an important impact on the model's predictions and may be related to the construction of the transmission chain.
[0202] Features with a medium importance rating:
[0203] Gene presence / absence variation (PAV) (0.22) and the number of rare mutations (0.20) and whether the patients are from the same community (0.18): These three features also play a certain supporting role in the model's judgment and provide additional information.
[0204] Features with low importance scores:
[0205] The features of Euclidean distance to patient residence (0.12), sampling time interval (0.05), and strain lineage classification (0.1) had low scores, indicating that their contribution to the model was limited.
[0206] Through the rigorous model construction, performance verification, and feature importance evaluation process described above, this example successfully constructed a highly accurate and reliable Mycobacterium tuberculosis near-transmission relationship detection model, laying a solid foundation for subsequent result visualization and analysis, as well as practical applications. This model has important technological innovation and application value in the field of Mycobacterium tuberculosis transmission research, and is expected to bring significant improvements to tuberculosis prevention and control work.
[0207] After successfully constructing a highly accurate and reliable Mycobacterium tuberculosis transmission relationship detection model according to Example 2 and comprehensively evaluating the model performance and feature importance, this example focuses on in-depth mining and intuitive display of the model output results, aiming to transform complex data information into knowledge with practical guiding significance and provide direct support for tuberculosis prevention and control work.
[0208] In this example, during a tuberculosis outbreak investigation in a certain city, the system detected a transmission cluster consisting of one strain and collected samples from the 152 patients who were initially infected. For each patient sample, the system obtained detection results for eight characteristics of the remaining 151 patients: the proportion of non-synonymous mutations, the total number of SNPs, the gene presence / absence variation (PAV), the number of rare mutations, the Euclidean distance to the patient's residence, the sampling time interval, and whether the patients were in the same community.
[0209] According to the random forest model constructed in Example 2, the test results are input, the judgment threshold is 0.7, and the output value is higher than or equal to 0.7, which is judged as positive (i.e., patient B is likely to be infected from patient A); the output value is lower than 0.7, which is judged as negative (i.e., patient B is not infected from patient A).
[0210] Retrospective validation of culture-positive strains from 152 patient pairs with a common exposure environment (such as home, school, workplace, etc.) or close contact history combined with known transmission chains determined by field epidemiological investigations was performed. The results are shown in Table 6:
[0211] Table 6. Verification of the results of the detection of the transmission relationship of Mycobacterium tuberculosis in 152 patients
[0212]
[0213] Because there is currently no gold standard for determining TB transmission relationships, the model's evaluation and analysis effectiveness can only be verified through very strict screening criteria. For example, if patients in the same family have a transmission relationship, it is considered a true positive. To indicate a non-transmission relationship, samples are collected from different regions. For example, four transmission cases are selected from a family in Region A, and three transmission cases are selected from a family in Region B. Epidemiological surveys determine that the two family members have no overlapping living or travel history, thus indicating no transmission relationship between the two family members. If a transmission relationship is determined, it is a false positive.
[0214] Performance indicators: positive predictive value (PPV): 112 / 115 = 97.4%; negative predictive value (NPV): 31 / 37 = 83.8%; Youden index: 0.94.
[0215] Based on the detection results of the Mycobacterium tuberculosis transmission relationship detection model, iGraph software was used to construct a visual network of the transmission cluster results (see Figure 6 ), a visual relationship network is constructed for strains with a transmission probability greater than 0.7. The nodes in the network represent samples, and samples with predicted transmission relationships are connected by edges. The numbers on the edges represent the predicted probability of transmission between samples. Figure 6 As shown in the figure, the spread of tuberculosis, starting from the source case C, follows: C, an asymptomatic carrier, transmits the disease to B (probability 0.82) and D (0.78) via route 1, and to F (0.85) via route 2. B then transmits the disease to A (0.89) via route 3; D transmits the disease to E (0.72) via route 4. F, acting as a super-spreader, transmits the disease to G (0.75) and H (0.81) via route 5. G transmits the disease to I (0.76) via route 6, while H transmits the disease to J (0.83) via route 7. Finally, I transmits the disease to J via route 8 (0.71), fully demonstrating the hidden spread of tuberculosis through the chain of family, community, and medical institution.
[0216] All patents and publications cited in this specification are intended to indicate that they are state of the art and that the present invention may be used. All patents and publications cited herein are incorporated by reference in their entirety, as if each publication were specifically incorporated by reference. The invention described herein may be practiced in the absence of any element or elements, limitation or limitations, unless otherwise specified. For example, in each instance, the terms "comprising," "consisting essentially of," and "consisting of" may be replaced with either of the other two terms. The term "a" or "an" herein simply means "one" and does not exclude the inclusion of only one or more. The terms and expressions used herein are intended to be descriptive, not limiting, and are not intended to exclude any equivalent features. However, it is understood that any suitable changes or modifications may be made within the scope of the present invention and the appended claims. It is understood that the embodiments described herein are preferred embodiments and features, and that modifications and variations can be made by persons of ordinary skill in the art based on the spirit of the present invention. Such modifications and variations are considered to be within the scope of the present invention and the scope of the independent and appended claims.
Claims
1. A method for constructing a model for improving the accuracy of analyzing the transmission relationship of Mycobacterium tuberculosis using a feature combination, characterized in that: The feature set consists of the proportion of non-synonymous mutations, the total number of SNPs, the number of gene presence / absence variants, the number of rare mutations, the Euclidean distance of the patient's residence, the sampling time interval, whether the patients are in the same community, and the strain lineage classification; The total number of SNPs refers to the total number of sites of DNA sequence diversity caused by conversion, transversion, insertion or deletion of a single nucleotide in the genome; the strain lineage classification refers to the division of the MTB complex group into 8 main lineages and multiple sublineages according to the current internationally accepted classification system, each lineage has unique genomic characteristics, geographical distribution and clinical phenotypes; the Euclidean distance of the patient's residence refers to the Euclidean distance between the residence of the patient to be tested and the source of infection; the sampling time interval refers to the time difference between the dates of clinical sample collection of two tuberculosis patients; Analyzing the transmission relationship of Mycobacterium tuberculosis refers to calculating the probability value of whether Mycobacterium tuberculosis in patient B was transmitted from patient A, so as to analyze the transmission relationship between patient B and patient A. If the probability value is higher than a threshold, it proves that Mycobacterium tuberculosis in patient B was transmitted from patient A. The model was constructed using random forest, and the sensitivity of the model reached 100%.
2. A method for constructing a model for improving the accuracy of analyzing the transmission relationship of Mycobacterium tuberculosis using a feature combination, characterized in that: The feature set consists of the proportion of non-synonymous mutations, the total number of SNPs, the number of gene presence / absence variants, the number of rare mutations, the Euclidean distance of the patient's residence, the sampling time interval, whether the patients are in the same community, and the strain lineage classification; The total number of SNPs refers to the total number of sites of DNA sequence diversity caused by conversion, transversion, insertion or deletion of a single nucleotide in the genome; the strain lineage classification refers to the division of the MTB complex group into 8 main lineages and multiple sublineages according to the current internationally accepted classification system, each lineage has unique genomic characteristics, geographical distribution and clinical phenotypes; the Euclidean distance of the patient's residence refers to the Euclidean distance between the residence of the patient to be tested and the source of infection; the sampling time interval refers to the time difference between the dates of clinical sample collection of two tuberculosis patients; The analysis of the transmission relationship of Mycobacterium tuberculosis refers to calculating the probability value of whether the Mycobacterium tuberculosis in patient B was transmitted from patient A, so as to analyze the transmission relationship between patient B and patient A. If the probability value is higher than a threshold, it proves that the Mycobacterium tuberculosis in patient B was transmitted from patient A.
3. The use according to claim 2, characterized in that The model is constructed using a machine learning algorithm, which includes any one or more of a neural network, naive Bayes, random forest, generalized linear model, gradient boosting machine and support vector machine.
4. The use according to claim 3, characterized in that The model was constructed using random forest.
5. A Mycobacterium tuberculosis transmission relationship detection system, characterized in that: In the system, a model is constructed using a combination of the proportion of non-synonymous mutations, the total number of SNPs, the number of gene presence / absence variants, the number of rare mutations, the Euclidean distance of the patient's residence, the sampling time interval, whether the patients are in the same community, and the strain lineage classification; the total number of SNPs refers to the total number of sites of DNA sequence diversity caused by conversion, transversion, insertion or deletion of a single nucleotide in the genome; the strain lineage classification refers to the division of the MTB complex group into 8 main lineages and multiple sublineages according to the current internationally accepted classification system, each lineage has unique genomic characteristics, geographical distribution and clinical phenotypes; the Euclidean distance of the patient's residence refers to the Euclidean distance between the residence of the patient to be tested and the source of infection; the sampling time interval refers to the time difference between the clinical sample collection dates of two tuberculosis patients; The model analyzes the transmission relationship between patient B and patient A by calculating the probability value of whether the Mycobacterium tuberculosis in patient B was transmitted from patient A. If the probability value is higher than a threshold, it proves that the Mycobacterium tuberculosis in patient B was transmitted from patient A. The model is constructed using random forest.
6. A feature combination for evaluating the transmission relationship of Mycobacterium tuberculosis, characterized in that: It consists of the proportion of non-synonymous mutations, the total number of SNPs, the number of gene presence / absence variants, the number of rare mutations, the Euclidean distance of the patient's residence, the sampling time interval, whether the patients are in the same community and the strain lineage classification; the total number of SNPs refers to the total number of sites of DNA sequence diversity caused by conversion, transversion, insertion or deletion of a single nucleotide in the genome; the strain lineage classification refers to the division of the MTB complex group into 8 main lineages and multiple sub-lineages according to the current internationally accepted classification system, each lineage has unique genomic characteristics, geographical distribution and clinical phenotype; the Euclidean distance of the patient's residence refers to the Euclidean distance between the residence of the patient to be tested and the source of infection; the sampling time interval refers to the time difference between the clinical sample collection dates of two tuberculosis patients.
Citation Information
Patent Citations
Methods, systems and processes of determining transmission paths of infectious agents
CN108475297A