Mycobacterium tuberculosis propagation relation detection system and application thereof
By combining the system of pan-genomics and machine learning, and integrating multiple genetic differences and epidemiological information, the problems of inaccurate judgment of communication relationships and insufficient evaluation of transmission probability in the existing technology are solved, and higher accuracy in prediction of transmission modes and scientificity of prevention and control strategies are achieved.
Patent Information
- Application Number
- CN202510520125.5
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-04-24
- Publication Date
- 2025-05-27
- Estimated Expiration
- 2045-04-24
AI Technical Summary
The prior art has errors in the judgment of the transmission relationship between Mycobacterium tuberculosis, and failure to fully consider a variety of genetic variant types and epidemiological information, resulting in incomplete identification of transmission chains and insufficient assessment of transmission probability.
A system based on pan-genomics and machine learning is adopted to integrate multiple genetic differences characteristics and epidemiological information, and the optimal model is screened through model training and evaluation, overcome the fixed SNP threshold error, comprehensively consider a variety of genetic variant types, make full use of epidemiological information, and achieve quantitative assessment of the probability of transmission.
It significantly improves the accuracy of the prediction of Mycobacterium tuberculosis transmission mode, enhances the scientific nature of prevention and control measures, and can more accurately assess the transmission risk and identify high-risk groups, providing a scientific basis for preventing transmission.
Smart Images

Figure CN120048362A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the fields of pathogenic microbiology and epidemiological research. Specifically, it relates to a system for detecting the transmission relationship of Mycobacterium tuberculosis based on pan-genomics and machine learning and its application. Background Art
[0002] Tuberculosis is a serious infectious disease caused by Mycobacterium tuberculosis ( Mycobacterium tuberculosis , MTB), and still poses a major threat to global public health. As an infectious disease, understanding its transmission pattern is the key to prevention and control. During contact with patients, healthy people establish pulmonary infections by inhaling bacteria-carrying droplets through the respiratory tract. To achieve the ultimate goal of eliminating tuberculosis, it is necessary to strengthen the monitoring and prevention and control of interpersonal transmission in communities with a high burden. Since the 1990s, molecular biology methods have been applied to identify the genetic relationship between strains, and molecular epidemiology established by combining traditional epidemiological research has enabled people to better understand the transmission characteristics of tuberculosis. Traditional methods for detecting tuberculosis transmission mainly rely on techniques such as IS6110-RFLP and MIRU-VNTR, but the resolution of these methods is limited. In recent years, with the maturity of high-throughput sequencing technology and the reduction of costs, whole-genome sequencing (WGS) has been widely used in the transmission research of tuberculosis to trace the transmission direction and transmission chain by describing the substitution order of nucleotides. Currently, an international common practice is to use the difference threshold of 5 single nucleotide polymorphisms (SNPs) to define strains with a recent transmission relationship within 2 to 3 years, while in scientific research in China, it is customary to choose the difference of 12 SNPs as the threshold, and there is no unified standard.
[0003] However, this transmission detection technology based only on SNP differences faces the following key problems: 1. Fixed threshold leads to errors: Different strains have different mutation rates, and a unified SNP difference threshold may not accurately reflect the true transmission relationship; 2. Singularity of genetic variation types: Existing methods mainly focus on SNPs, ignoring other types of genetic variations such as insertions and deletions (InDel) and structural variants (SV) widely existing in MTB, resulting in incomplete identification of the transmission chain; 3. Insufficient utilization of epidemiological information: Epidemiological data such as the geographical location of cases, sampling time, social interaction network, and past medical history are not fully considered, making the identification of the transmission pattern less comprehensive; 4. Lack of quantitative assessment of transmission probability: Existing methods only judge whether it is related to transmission or not, and cannot accurately assess the transmission risk and formulate differential control strategies.
[0004] In a previously published paper ("TransFlow: a Snakemake workflow for transmission analysis of Mycobacterium tuberculosis whole-genome sequencing data"), previous research proposed a data analysis pipeline for MTB pan-genomics to analyze WGS data and detect recent transmission. However, this detection method still only uses SNP information and identifies recent transmission relationships based on a fixed SNP difference number threshold. Studies have shown that the mutation rates of different lineages of MTB vary significantly (e.g., the mutation rate of Lineage 2 Beijing strains is higher than that of Lineage 4), and the fixed threshold may lead to false positives / negatives.
[0005] Therefore, there is an urgent need to construct a new detection system for Mycobacterium tuberculosis transmission relationships, collect genomic and epidemiological information of more MTB isolates in community transmission to minimize the influence of different sampling locations and patients, and improve it by integrating pan-genomics and epidemiological information and optimizing machine learning models to solve the key technical problems in the existing technology, such as inaccurate judgment of transmission relationships, incomplete consideration of genetic variations, insufficient utilization of epidemiological information, and inability to quantitatively evaluate transmission probability, so as to significantly improve the accuracy, reliability and interpretability of MTB transmission detection, provide strong technical support for tuberculosis prevention and control, and contribute to the global tuberculosis prevention and control process. Summary of the Invention
[0006] Aiming at the problems existing in the prior art, the present invention provides a detection system for Mycobacterium tuberculosis transmission relationships and its application, which combines pan-genomics and machine learning algorithms, comprehensively considers various genetic differences between strains and epidemiological information between cases, and determines the transmission relationships between cases in a more comprehensive and accurate manner. Through sample collection, data preprocessing, extracting features such as genetic differences, epidemiological information, and strain lineages and integrating and normalizing them, and then through data partitioning, clustering analysis, multiple algorithm modeling training and evaluation to screen out the optimal model, it overcomes the error of the fixed SNP threshold, comprehensively considers various types of genetic variations, makes full use of epidemiological information, realizes quantitative evaluation of transmission probability, can accurately evaluate the transmission mode of Mycobacterium tuberculosis, realizes the traceability analysis of the transmission path of Mycobacterium tuberculosis, has wide applicability, provides a scientific basis for preventing transmission, improves the prevention and control efficiency and accuracy, and plays a key role in the global tuberculosis prevention and control process.
[0007] On the one hand, the present invention provides a system for detecting the transmission relationship of Mycobacterium tuberculosis. In this system, a combination of one or more of genetic difference features, epidemiological information features, strain lineage information features, and other features is used to construct a model for detecting recent transmission of Mycobacterium tuberculosis.
[0008] Further, the genetic difference features include any one or more selected from the following: non-synonymous mutation ratio, synonymous mutation ratio, number of mutations in promoter regions, total number of SNPs, total number of InDels, number of structural variations, number of gene presence / absence variations, number of rare mutations; The epidemiological information features include any one or more selected from the following: Euclidean distance of patient residence, sampling time interval, whether patients are in the same community; The strain lineage information features include any one or more selected from the following: strain lineage classification, number of lineage-specific mutations; The other features include any one or more selected from the following: number of drug resistance-related mutations, number of virulence-related mutations.
[0009] The transmission relationship of Mycobacterium tuberculosis described in the present invention refers to analyzing current Mycobacterium tuberculosis patients to find out where or from which patient they were infected with Mycobacterium tuberculosis. This is equivalent to a traceability analysis to explore the infection source of the current patient, so as to find the previous Mycobacterium tuberculosis patient who infected this patient. When the infection sources of all patients are detected in sequence, a complete MTB transmission network diagram can be constructed, from which it can be analyzed which aspects need to be noted to cut off the prevention and control path and reduce the transmission of Mycobacterium tuberculosis when preventing its transmission in the future, providing a scientific basis for preventing transmission.
[0010] Pan-genomics is the discipline that studies all the genomes of a species. It goes beyond the traditional concept of a reference genome and provides a more comprehensive view of genomic variations. In the field of pathogen transmission detection, the application of pan-genomics is particularly important because it can reveal the genetic composition and diversity of microbial populations, thus helping us understand the virulence, pathogenesis, and microevolution of pathogens. Pan-genomics research allows for the reclassification of species, thereby clarifying and improving the previously proposed traditional criteria. In the detection of the transmission of infectious diseases such as tuberculosis, pan-genomics can help identify and compare the genomic variations of different strains, including single nucleotide polymorphisms (SNPs), insertions and deletions (InDels), copy number variations (CNVs), and presence / absence variations (PAVs). These variations together constitute the variability of the pan-genome, which is crucial for understanding the transmission patterns and evolutionary paths of pathogens. The application of pan-genomics in pathogen transmission detection can not only improve the accuracy of variant detection but also significantly enhance the detection of large structural variations. When the relative differences among individuals in a population are relatively large, using PAV analysis can better reflect the differences within the population compared to using variant information such as SNPs, and it can also be regarded as a kind of marker. Based on PAVs, genetic evolutionary association analysis of species can also be carried out. Therefore, pan-genomics provides a relatively comprehensive framework to understand and track the transmission of MTB.
[0011] In the detection of the transmission of infectious diseases, the integration of epidemiological information is crucial because it provides direct clues to the risk factors of disease transmission. The core of epidemiological research lies in collecting and analyzing population-based health data, including the geographical distribution of cases, the time of onset, treatment and medication situations, exposure history, social behaviors, and environmental factors, etc. These data are crucial for understanding the transmission patterns of infectious diseases, identifying high-risk populations, and formulating effective preventive measures.
[0012] Due to the lack of all comprehensive and reliable information related to the recent transmission of MTB, more complex algorithms are required to process and analyze massive amounts of pan-genomics and epidemiological data. As an important branch of artificial intelligence, machine learning can identify complex non-linear relationships and reveal hidden patterns that are difficult to capture by traditional statistical methods. It has been widely applied in various fields, including genomics and epidemiology. Machine learning techniques are used to process and analyze large amounts of health data, including various genetic variation data in pan-genomics and traditional epidemiological characteristic data. Machine learning techniques, such as neural networks, support vector machines, and random forests, are powerful classification systems that have been used in health sciences, such as cancer genomics. In the field of machine learning, model training and algorithm optimization are the keys to improving prediction accuracy. In the diagnosis and treatment of tuberculosis, machine learning techniques have also been proven to have high potential. The conditional random forest model established based on the random forest and bagging ensemble algorithms of conditional inference decision trees performs best in distinguishing active tuberculosis from latent Mycobacterium tuberculosis infection.
[0013] The present invention combines pan-genomics and machine learning algorithms. First, through a rigorous sample collection process, samples with clear transmission chains and control samples are selected from tuberculosis transmission cases. After preprocessing such as data cleaning, genome alignment, and epidemiological information collation, multi-dimensional features including genetic differences, epidemiology, and strain lineages are screened, extracted, and normalized for model construction. The constructed model is used to analyze the transmission relationship of Mycobacterium tuberculosis between patients, so as to conduct traceability analysis and find out where the Mycobacterium tuberculosis is transmitted from.
[0014] Furthermore, the model is used to analyze the transmission relationship of Mycobacterium tuberculosis between patients. By analyzing the transmission relationship between patient B and patient A, it is determined whether the Mycobacterium tuberculosis on patient B is transmitted from patient A.
[0015] The evaluation and analysis of the model are carried out for a Mycobacterium tuberculosis and a patient carrying the bacterium (such as patient B). The infection route of this patient (patient B) is traced and analyzed to determine where the Mycobacterium tuberculosis of this patient (patient B) is transmitted from and whether it is obtained from another patient (patient A).
[0016] Furthermore, the model analyzes the transmission relationship between patient B and patient A by calculating the probability value of whether the Mycobacterium tuberculosis on patient B is transmitted from patient A. If the probability value is higher than the threshold, it proves that the Mycobacterium tuberculosis on patient B is transmitted from patient A.
[0017] By analyzing the transmission relationship between Patient B and Patient A, the probability value is high, thus determining that Patient B was infected with Mycobacterium tuberculosis from Patient A. Then, by analyzing Patient A, it is possible to find out where Patient A was infected, and in turn, the traceability analysis of Mycobacterium tuberculosis can be achieved.
[0018] The non-synonymous mutation ratio refers to the ratio of the number of sites with non-synonymous mutations to the number of all mutant sites among polymorphic sites. Among them, non-synonymous mutations refer to mutations that cause changes in the codons encoding amino acids, thereby affecting the amino acid sequence of proteins; synonymous mutations refer to mutations that occur in the third nucleotide of the codon. Due to the degeneracy of the genetic code, the mutated codon still encodes the same amino acid and does not change the amino acid sequence of the protein. The number of promoter region mutations refers to the number of sites with promoter mutations. Promoter mutations occur in the promoter region or other DNA sequences regulating genes, and are a type of mutation that can enhance or reduce the transcriptional initiation function of the promoter. The total number of SNPs refers to the total number of sites with DNA sequence diversity caused by single nucleotide transitions, transversions, insertions, or deletions in the genome. The total number of InDels refers to the total number of all insertion (Insertion, abbreviated as In) and deletion (Deletion, abbreviated as Del) variations in the genome. The number of structural variations refers to the number of variations in the genome that cause chromosomal structural abnormalities. The number of gene presence / absence variations refers to the number of variations in the presence (Presence) or absence (Absence) of genes detected in the Mycobacterium tuberculosis genome through pan-genomics analysis. Such variations reflect the plasticity of the genomes among different strains. The number of rare mutations refers to the number of genetic variations with a mutation frequency lower than a preset threshold (usually 5%) detected across the entire genome, including low-frequency SNPs / InDels: single nucleotide variations or insertions and deletions with a frequency <5% in the population, and private mutations: mutations that only appear within a single transmission cluster and can be used as molecular markers for the transmission chain.
[0019] The Euclidean distance of the patient's residence refers to the Euclidean distance between the residence of the patient to be tested (such as Patient B) and the residence of its source of infection (such as Patient A). The sampling time interval refers to the time difference between the clinical sample collection dates of two tuberculosis patients (such as Patient A and Patient B). The calculation formula is: , where and The sample sampling times for patients A and B respectively. Whether the patients are in the same community means whether the patient to be tested (such as patient B) and its source of infection (such as patient A) live in the same community (the same community refers to a community within a straight-line distance of less than 1000m and within the range of the same enclosure where there may be a probability of cross-infection).
[0020] The strain lineage classification refers to the phylogenetic typing based on single nucleotide polymorphisms (SNPs) of the whole genome of Mycobacterium tuberculosis, which divides the strains into genetic branches with common evolutionary origins. The currently internationally used classification system (such as the Global MTBC Lineage System) divides the MTB complex into 8 main lineages (Lineage 1 - 8) and multiple sub-lineages (such as Lineage 2.2.1 Beijing strain), and each lineage has unique genomic characteristics, geographical distributions, and clinical phenotypes.. The number of lineage-specific mutations refers to the total number of genetic variant sites (such as Rv3346c C913T in Lineage 4) that occur frequently (≥95% of the strains carry) in a certain MTB lineage and are rare (≤5%) in other lineages.
[0021] The number of drug resistance-related mutations in the other characteristics refers to the total number of genetic variant sites (such as the drug resistance mutations S450L, D435V, etc. in the rpoB gene of rifampicin) in the Mycobacterium tuberculosis genome that are experimentally verified to be significantly related to anti-tuberculosis drug resistance. TB-Profiler is used for mutation annotation based on the World Health Organization's "Catalogue of mutations in Mycobacterium tuberculosis complex and their association with drug resistance, 2nd ed". The number of virulence-related mutations refers to the total number of genetic variant sites in the Mycobacterium tuberculosis genome that are directly related to the ability of host immune escape, intracellular survival, or tissue destruction. Nonsynonymous mutations of known virulence genes are screened based on the Virulence Factor Database (VFDB).
[0022] After evaluating the importance of the characteristics, genetic difference characteristics such as the number of rare mutations, the proportion of nonsynonymous mutations, and presence / absence variation (PAV) of genes are finally determined as the core characteristics, and epidemiological information (such as geographical distance, sampling time interval) and strain lineage information (such as lineage classification, lineage-specific mutations) are used as auxiliary characteristics, which helps to deeply understand the transmission mechanism and provides key support for the precise prevention and control of tuberculosis, such as locking high-risk populations and optimizing prevention and control strategies.
[0023] Furthermore, a model is constructed by using a combination of genetic difference features, epidemiological information features, and strain lineage information features; the genetic difference features include a combination of the proportion of non-synonymous mutations, the total number of SNPs, the number of gene presence / absence variations, and the number of rare mutations; the epidemiological information features include a combination of the Euclidean distance of the patient's residence, the sampling time interval, and whether the patients are in the same community; the strain lineage information features include a combination of strain lineage classification and the number of lineage-specific mutations; the strain lineage information features include strain lineage classification.
[0024] The present invention examines the analysis efficacy of constructing models with different feature combinations, and screens out eight feature combinations with better analysis evaluation accuracy and specificity, higher AUC values, and fewer required features. The models constructed by using these eight feature combinations can more accurately analyze the transmission relationship of Mycobacterium tuberculosis among different patients, and help to more efficiently complete the traceability analysis of Mycobacterium tuberculosis.
[0025] Furthermore, the model is constructed by using a machine learning algorithm, and the machine learning algorithm includes any one or more of neural network, naive Bayes, random forest, generalized linear model, gradient boosting machine, and support vector machine.
[0026] Furthermore, the model is constructed by using a random forest.
[0027] In some ways, in the random forest algorithm, the number of trees is 500, the maximum depth of a single tree is 12, the minimum number of samples for node splitting is 5, and the minimum number of samples for leaf nodes is 2.
[0028] Due to the complexity of sample characteristics, it is impossible to estimate which machine learning algorithm is more suitable for the data of this study without prior knowledge. To further improve the accuracy, stability, and generalization ability of the model for predicting the transmission of Mycobacterium tuberculosis, the present invention has conducted in-depth screening and optimization of machine learning algorithms (neural network, naive Bayes, random forest, generalized linear model, gradient boosting machine, and support vector machine). The research results clearly show that different algorithms perform differently when dealing with this prediction problem. Among them, the model constructed by the ranger algorithm based on the random forest has the largest median ROC, specifically: AUC value: 0.97, significantly superior to other algorithms (such as support vector machine AUC 0.92, gradient boosting machine AUC 0.94); sensitivity: 100%, ensuring that all transmission relationships are accurately identified without false negatives; accuracy: 98.1%, indicating that the model prediction results are highly consistent with the true transmission relationships; specificity: 96.2%, effectively reducing the false positive rate and avoiding misjudging non-transmission clusters as transmission clusters; Kappa value: 0.962, further verifying the high reliability of the model prediction results. It shows that this model has the best performance in distinguishing transmission clusters and non-transmission clusters. In practical applications, it can more accurately judge the recent transmission situation of Mycobacterium tuberculosis, providing a solid algorithm guarantee for achieving accurate prediction of the recent transmission of Mycobacterium tuberculosis. Therefore, it is defined as the final prediction model.
[0029] For the eight feature combinations obtained through screening, the present invention divides the population to be evaluated into a training set and a validation set according to a certain proportion, uses the DBSCAN clustering algorithm to mine the data structure, trains the model with a variety of machine learning algorithms and screens according to performance, and determines that the random forest is the best algorithm to construct the model. The random forest models constructed based on the eight feature combinations have AUC values of 0.98 and 0.97 on the training set and the validation set respectively, the accuracy in the validation set is 0.981, the Kappa value is 0.962, the sensitivity is 100%, the specificity is 96.2%, the positive predictive rate is 96.4%, and the negative predictive rate is 100%, which can accurately identify the recent transmission relationships of Mycobacterium tuberculosis.
[0030] The prediction results output by the random forest model constructed by the present invention present the transmission relationships in the form of probability values, which not only enhances the operability of transmission monitoring but also provides a key basis for formulating the optimal tuberculosis prevention and control strategy. Moreover, this system can flexibly adapt to the environmental differences in different regions and the epidemic situation changes at different times, demonstrating broad applicability. It can also assist in efficiently identifying high-risk populations of Mycobacterium tuberculosis, effectively preventing the transmission of Mycobacterium tuberculosis, and comprehensively promoting the development of tuberculosis prevention and control work.
[0031] On the other hand, the present invention provides a combination of characteristics for evaluating the transmission relationship of Mycobacterium tuberculosis, and the combination of characteristics includes genetic difference characteristics, epidemiological information characteristics, and strain lineage information characteristics; the genetic difference characteristics include the proportion of non-synonymous mutations, the total number of SNPs, the number of gene presence / absence variations, and the number of rare mutations; the epidemiological information characteristics include the Euclidean distance of the patient's residence, the sampling time interval, and whether the patients are in the same community; the strain lineage information characteristics include strain lineage classification.
[0032] The combination of characteristics provided by the present invention comprehensively considers factors such as the specific mutation situation of the strain, geographical location, sampling time, etc., determines the relationship of transmission clusters, does not rely on a fixed SNP difference threshold, effectively avoids errors caused by different strain mutation rates, and greatly improves the accuracy of judging the transmission relationship. And it calculates and gives the probability of transmission between cases through the ranger algorithm based on random forest, and this probability value can provide an accurate quantitative reference for prevention and control work. Statistics on the results of transmission detection and drawing relevant statistical charts, these charts can visually present the transmission rules and trends, and provide clear visual information for researchers and prevention and control personnel. Speculating on the transmission risk factors can accurately identify high-risk strains, samples, and transmission scenarios, providing a key basis for taking targeted prevention and control measures in advance. Demonstrating the transmission hotspots and patterns helps to reasonably allocate prevention and control resources, concentrate more energy and resources on key areas and key transmission paths, and achieve efficient prevention and control. It significantly improves the depth and breadth of the understanding of the transmission law of Mycobacterium tuberculosis, provides a new perspective and strong data support for the basic research of tuberculosis; also provides a precise and operable guidance plan for the prevention and control work of tuberculosis, can effectively reduce the transmission risk of tuberculosis and reduce the number of infections; at the same time, by improving the accuracy and reliability of the traceability detection of the transmission relationship, it optimizes the process of formulating the prevention and control strategy of tuberculosis, saves a large amount of human, material and time costs, and strongly promotes the process of global tuberculosis prevention and control work, and has important innovation and application value in the field of tuberculosis prevention and control.
[0033] On the other hand, the present invention provides a use of a feature for analyzing the transmission relationship of Mycobacterium tuberculosis, wherein the feature includes any one or more of a genetic difference feature, an epidemiological information feature, a strain lineage information feature, and other features; the genetic difference feature includes any one or more selected from the following: non-synonymous mutation ratio, synonymous mutation ratio, number of mutations in the promoter region, total number of SNPs, total number of InDels, number of structural variations, number of gene presence / absence variations, number of rare mutations; the epidemiological information feature includes any one or more selected from the following: Euclidean distance of the patient's residence, sampling time interval, whether the patients are in the same community; the strain lineage information feature includes any one or more selected from the following: strain lineage classification, number of lineage-specific mutations; the other features include any one or more selected from the following: number of drug resistance-related mutations, number of virulence-related mutations.
[0034] Further, the feature includes a genetic difference feature, an epidemiological information feature, and a strain lineage information feature; the genetic difference feature includes non-synonymous mutation ratio, total number of SNPs, number of gene presence / absence variations, and number of rare mutations; the epidemiological information feature includes Euclidean distance of the patient's residence, sampling time interval, and whether the patients are in the same community; the strain lineage information feature includes strain lineage classification.
[0035] Further, the constructed model is used to evaluate the traceability of the transmission relationship of Mycobacterium tuberculosis, find out how Mycobacterium tuberculosis spreads to a large number of patients step by step, and provide a theoretical basis for preventing and controlling the transmission of Mycobacterium tuberculosis.
[0036] Further, a model is constructed by using a machine learning algorithm, and the machine learning algorithm includes any one or more combinations of neural network, naive Bayes, random forest, generalized linear model, gradient boosting machine, and support vector machine.
[0037] Further, the machine learning algorithm includes a random forest algorithm.
[0038] The beneficial effects of the present invention are as follows: 1. The present invention provides a brand-new detection system for the transmission relationship of Mycobacterium tuberculosis. By integrating a combination of various genetic difference features (including the proportion of non-synonymous mutations, the proportion of synonymous mutations, the number of mutations in promoter regions, the total number of SNPs, the total number of InDels, the number of structural variations, the number of gene presence / absence variations, the number of rare mutations), epidemiological information features (including the Euclidean distance of patients' residence, the sampling time interval, whether patients are in the same community), strain lineage information features (including strain lineage classification, the number of lineage-specific mutations), and other features (including strain lineage classification, the number of lineage-specific mutations), it breaks through the limitation of traditional research that only focuses on partial gene loci or single mutation types, deeply explores the genetic diversity of strains, and provides richer information for accurately understanding the genetic evolution and transmission basis of Mycobacterium tuberculosis. It improves the detection coverage of the genomes of different lineages of MTB and can comprehensively reflect the genetic relationship between strains.
[0039] 2. The present invention comprehensively utilizes a variety of genetic differences and epidemiological information, effectively reducing the error of detecting the threshold of the number of single SNP differences. Instead of relying on a fixed threshold, it combines various genetic variation types of strains and detailed epidemiological data, making the judgment of the transmission relationship more accurate, significantly improving the accuracy of predicting the transmission pattern of Mycobacterium tuberculosis, and avoiding incorrect conclusions caused by one-sided information.
[0040] 3. The present invention particularly focuses on the presence and absence of specific mutations, providing a more accurate assessment of the transmission possibility. By deeply analyzing the impact of each mutation on the transmission characteristics of strains, especially the key role of rare mutations in the adaptive evolution and transmission process of strains, it can more precisely judge the probability of strain transmission, providing a more targeted basis for the formulation of prevention and control measures and enhancing the scientific nature of prevention and control work.
[0041] 4. The present invention adopts machine learning clustering and classification algorithms, which can quickly process the complex data of pan-genomics and epidemiology. When facing a large amount of gene data and complex epidemiological information, this algorithm demonstrates powerful data processing capabilities, quickly extracts valuable information, greatly improves the research efficiency, ensures providing support for prevention and control decisions in a short time, and meets the requirements of timeliness for modern tuberculosis prevention and control.
[0042] 5. The present invention uses the prediction results of machine learning to represent the transmission possibility in probability values, which helps to improve the operability of transmission monitoring. Prevention and control personnel can conduct hierarchical management of different regions, populations, and strains according to the probability size, formulate personalized and efficient tuberculosis prevention and control strategies, and reasonably allocate prevention and control resources, such as focusing on monitoring and intervening in high-risk regions and populations, improving the prevention and control effect, and reducing the transmission risk of tuberculosis.
[0043] 6. The present invention can adapt to the transmission situations of Mycobacterium tuberculosis in different regions and at different times, and has wide applicability. Whether in high-incidence areas or low-incidence areas of tuberculosis, as well as in different epidemic stages, the detection system for the transmission relationship of Mycobacterium tuberculosis provided by the present invention can effectively play its role, providing a unified and reliable technical means for global tuberculosis prevention and control, strongly promoting the process of global tuberculosis prevention and control work, and having important application value in different prevention and control scenarios.
[0044] 1. Tuberculosis patients The tuberculosis patients refer to those who show symptoms such as cough, expectoration, low fever, night sweats, fatigue, weight loss, etc. due to infection with Mycobacterium tuberculosis and are diagnosed with tuberculosis through relevant examinations.
[0045] Tuberculosis is a chronic infectious disease caused by Mycobacterium tuberculosis and is mainly transmitted through the air. After infection, most people are in an asymptomatic latent infection state, and the bacteria are controlled by the immune system; however, when the immunity declines, about 5-10% of latent infected individuals will develop into active tuberculosis, mainly manifested as persistent cough, expectoration, hemoptysis, chest pain, fever, night sweats, weight loss, and fatigue, etc. Tuberculosis can affect the lungs (pulmonary tuberculosis) or other organs (extrapulmonary tuberculosis) such as the kidneys, brain, spine, and skin, etc. Diagnosis relies on imaging, microbiology, and immunology examinations, and treatment requires long-term combination drug therapy, and the problem of drug resistance is becoming increasingly severe. Preventive measures include BCG vaccination and infection control.
[0046] Tuberculosis diagnosis can be confirmed through laboratory examinations of Mycobacterium tuberculosis, mainly including sputum smear acid-fast staining, mycobacterium culture, and molecular biology detection (such as PCR). Among them, mycobacterium culture is the gold standard for tuberculosis laboratory diagnosis, including solid culture and liquid culture, which can confirm the presence of pathogens and provide drug sensitivity test results, but it takes a long time (usually 2-6 weeks). Molecular biology detection (such as Xpert MTB / RIF) can detect Mycobacterium tuberculosis and its drug resistance simultaneously in a shorter time, significantly improving the diagnostic efficiency. Combining clinical symptoms, imaging examinations, and immunological tests (such as tuberculin skin test or interferon-γ release assay) can further improve the diagnostic accuracy.
[0047] 2. Mycobacterium tuberculosis Mycobacterium tuberculosis, abbreviated as tubercle bacilli, is the pathogen of human tuberculosis. It is a type of obligate aerobic bacteria, positive for acid-fast staining. It has no flagella, has pili, has a microcapsule but does not form spores, and its cell wall has neither teichoic acid of Gram-positive bacteria nor lipopolysaccharide of Gram-negative bacteria. The German bacteriologist Robert Koch (1843 - 1910) discovered and proved it to be the pathogen of human tuberculosis in 1882. Tuberculosis caused by the infection of this bacterium is an infectious disease that seriously affects human life and health. After centuries of struggle with it, it has gradually been brought under control. However, in recent years, due to the influence of various factors, this disease has become increasingly serious.
[0048] Mycobacterium tuberculosis can undergo variations in morphology, colony, virulence, immunogenicity, and drug resistance. Bacillus Calmette - Guérin (BCG) is an attenuated live vaccine strain obtained by Calmette and Guérin (1908) through 230 subcultures of Mycobacterium bovis in a medium containing glycerol, bile, and potatoes over 13 years, and is now widely used for vaccination.
[0049] 3. Transmission relationship of Mycobacterium tuberculosis The study of the transmission relationship of Mycobacterium tuberculosis is to analyze the transmission relationship among patients, so as to conduct traceability analysis, find out where the Mycobacterium tuberculosis is transmitted from, where the source is, which patient infects which patient, and thus be able to construct a complete MTB transmission network diagram.
[0050] In the present invention, by constructing a detection system for the transmission relationship of Mycobacterium tuberculosis, the evaluation and analysis are carried out for a Mycobacterium tuberculosis and a patient carrying this bacterium (such as patient B). The transmission route of this patient (patient B) is traced and analyzed to determine where the Mycobacterium tuberculosis of this patient (patient B) is infected from, whether it is obtained from another patient (patient A). Through this system, a probability value can be calculated to judge whether the Mycobacterium tuberculosis on patient B is transmitted from patient A, so as to analyze the transmission relationship between patient B and patient A. When the probability value is higher than the threshold, it is judged that the Mycobacterium tuberculosis on patient B is transmitted from patient A.
[0051] In some ways, the threshold of the probability value is 0.7. When the probability value ≥ 0.7, it is related to transmission, and it is judged that patient B is transmitted from patient A; when the probability value < 0.7, it is judged to be irrelevant, that is, patient B is not transmitted from patient A.
[0052] The Mycobacterium tuberculosis transmission relationship detection system based on pan-genomics and machine learning provided by the present invention covers the whole process from sample collection to feature extraction and integration in terms of data processing. The sample collection module can accurately screen out cases with clear recent transmission chains and samples of their close contacts from known tuberculosis transmission cases, and at the same time include samples with no obvious epidemiological transmission associations as controls, and strictly screen according to specific criteria or supplemented and improved case diagnosis criteria to ensure the reliability and representativeness of the samples.
[0053] The data preprocessing module uses fastp and Kraken software to efficiently remove low-quality, repetitive, and non-MTB source sequences in the sequencing data, and uses bwa software to carefully align the qualified sequencing data with a specific reference genome to obtain the variant positions of the sample genome and the reference genome and the coverage information of the sequencing data for the genome, and to standardize the epidemiological data covering the comprehensive information of the patients.
[0054] Based on pan-genomics analysis, the feature engineering module uses GATK and SNPEff software packages to deeply mine the number, type (such as SNP, InDel, SV, etc.) and frequency information of gene mutations of strains, extracts drug resistance and virulence gene information to construct genomic mutation numerical features; uses the geographical information GIS system to calculate the Euclidean distance of the patients' residence, determine the sampling time interval and encode the epidemiological category variables; based on the TBprofiler software package, conducts strain lineage classification and specific mutation identification; finally integrates and normalizes the data to eliminate the dimension difference and provide a high-quality data basis for model training.
[0055] Through this system, the key problems in the study of Mycobacterium tuberculosis transmission in the prior art can be effectively solved. It overcomes the drawback of inaccurate judgment of transmission relationships caused by the traditional method based only on a fixed SNP difference threshold, comprehensively considers various genetic differences and epidemiological information, and greatly reduces the error; makes up for the defect of incomplete consideration of genetic variant types in the past, comprehensively includes various variant types; makes full use of the epidemiological information that has been ignored in the past to make the recognition of transmission patterns more comprehensive; realizes the quantitative assessment of transmission probability and changes the limitation that the existing methods can only make qualitative judgments.
[0056] Finally, it achieves the purpose of accurately detecting the transmission relationship of Mycobacterium tuberculosis, accurately assessing the transmission risk, effectively identifying high-risk populations, providing a scientific basis for preventing transmission, accurately calculating the transmission probability between strains, strongly promoting the prevention and control work of tuberculosis, improving the prevention and control efficiency and accuracy, playing a key role in the global tuberculosis prevention and control process, significantly reducing the transmission risk of tuberculosis, reducing the number of infections, saving a large amount of human, material and time costs, and bringing innovative solutions and important application values to the field of tuberculosis prevention and control. Description of the Drawings
[0057] Figure 1 It is the workflow of a Mycobacterium tuberculosis transmission detection system based on pan-genomics and machine learning.
[0058] Figure 2 It is the k-distance distribution graph for DBSCAN clustering to select the appropriate eps (neighborhood distance threshold) parameter.
[0059] Figure 3 It is the scatter plot of DBSCAN clustering of MTB samples based on pan-genomics and epidemiological data.
[0060] Figure 4 It is the performance evaluation of the optimal model constructed by six algorithms.
[0061] Figure 5 It is the performance evaluation of the optimal model constructed by six algorithms.
[0062] Figure 6 It is the MTB transmission network graph predicted by the model. Specific implementation manner
[0063] The present invention will be further described in detail below in conjunction with the accompanying drawings and embodiments. It should be noted that the following embodiments are intended to facilitate the understanding of the present invention and do not impose any limitations on it. The reagents used in this embodiment are all known products and are obtained by purchasing commercially available products.
[0064] Example 1 Collection and preliminary screening of characteristics of Mycobacterium tuberculosis transmission samples In the field of Mycobacterium tuberculosis transmission research, comprehensively and accurately obtaining sample information and extracting key characteristics are crucial for constructing an efficient transmission relationship detection model. This embodiment is dedicated to developing a computer analysis method for detecting recent transmissions of MTB based on pan-genomics and machine learning, aiming to accurately predict the transmission paths and associations of Mycobacterium tuberculosis among cases by deeply integrating multi-source data and advanced technical means.
[0065] The specific workflow of the Mycobacterium tuberculosis transmission relationship detection system based on pan-genomics and machine learning is as Figure 1 shown, and the steps are as follows: I. Sample collection From known tuberculosis transmission cases, a total of 150 samples of cases with clear recent transmission chains and their close contacts were screened. Through epidemiological investigation and tracking among these samples, direct or indirect transmission relationships were determined. Among them, there were 100 confirmed cases and 50 close contacts, all of whom were confirmed within the past 6 months.
[0066] Meanwhile, to ensure the rigor of the study, 60 samples with no obvious epidemiological transmission associations were also collected as the control group. These control samples were all patients who visited the same area during the same time period but had no contact with cases in the known transmission chains, in order to better evaluate the specificity and accuracy of the model.
[0067] The criteria for cases with a clear transmission chain and their close contacts in the samples are as follows: Referring to the Law of the People's Republic of China on the Prevention and Control of Infectious Diseases and the Guidelines for the Diagnosis and Treatment of Pulmonary Tuberculosis formulated by the Chinese Medical Association's Tuberculosis Branch, specifically including: Case diagnosis criteria: Meeting the clinical diagnosis criteria in "Diagnosis of Pulmonary Tuberculosis (WS 288-2017)", including: having typical tuberculosis symptoms (such as persistent cough, expectoration, low fever, night sweats, etc.); imaging examinations (such as chest X-ray or CT) showing pulmonary tuberculosis lesions; positive laboratory tests (such as sputum smear, culture or molecular detection).
[0068] Criteria for close contacts: Defined as those who had the following contacts with a confirmed case within 2 weeks before the case was confirmed: co-resident family members; those with long-term contact in the same workplace or learning environment; those participating in the same social activities.
[0069] II. Data preprocessing 1. Sequencing data cleaning: Use fastp and Kraken software to remove low-quality sequences, duplicate sequences, and non-MTB source sequences in the sequencing data to ensure the quality and consistency of the data; 2. Genome sequence alignment: Align the qualified sequencing data with the reference genome H37Rv (NCBI Reference Sequence: NC_000962.3) through bwa software to determine the variant positions between the sample genome and the reference genome and the coverage information of the sequencing data for the genome; 3. Epidemiological information collation: Standardize the epidemiological information collected through the questionnaire survey (covering patients' basic demographic information, details of living environment, social and travel activity trajectories, past medical history, and treatment details, etc.) to ensure the accuracy and integrity of the data.
[0070] III. Feature engineering 1. Genome feature extraction: Genome features are extracted from Mycobacterium tuberculosis samples based on sputum culture. The specific steps are as follows: Sample collection: Collect sputum samples from patients diagnosed with tuberculosis.
[0071] Bacterial culture: Culture the sputum samples on a selective medium to isolate Mycobacterium tuberculosis (such as using Lowenstein-Jensen medium).
[0072] DNA Extraction and Library Construction: Extract genomic DNA from the cultured bacteria.
[0073] Genome Sequencing: Perform whole-genome sequencing on the extracted DNA using high-throughput sequencing technology (such as the Illumina sequencing platform) to obtain the genomic information of Mycobacterium tuberculosis.
[0074] Data Analysis: (1) Based on pan-genomics analysis, calculate and annotate the number, type (SNP (single nucleotide variation), InDel (insertion or deletion fragment), SV (larger structural variation, etc.)) and frequency (>95% common mutations and <5% rare mutations) of gene mutations in each sample using the GATK and SNPEff software packages to comprehensively understand the genetic variation of the strains. (2) Construct numerical features of genomic mutations, convert each type of mutation into a numerical representation, with presence marked as 1 and absence marked as 0.
[0075] (3) Extract drug resistance gene mutation information in the genome based on the "Catalogue of Gene Mutations Associated with Drug Resistance in Mycobacterium tuberculosis (2nd Edition)" published by the World Health Organization (WHO). Extract mutation information of MTB virulence genes in the genome according to the VFDB database.
[0076] 2. Epidemiological Feature Extraction: Clinical features are from patients' medical records and epidemiological investigations.
[0077] (1) Use the Geographic Information GIS system (such as ArcGIS software) to convert the place of residence into a geographic area code, calculate the Euclidean distance between patients' places of residence, and construct a distance matrix between patients. (2) Calculate the interval of sampling time between samples as a time series feature (R language). (3) Encode categorical variables in epidemiology (such as gender, occupation, medical history, treatment plan, etc. and travel history and contact history (if any)), and convert each type into a binary vector representation (R language).
[0078] 3. Strain Lineage Information: Strain lineage information is annotated based on the data of whole-genome sequencing of Mycobacterium tuberculosis: (1) Strain lineage classification: Based on the TBprofiler software package (default parameters), classify the strains according to the whole-genome data (such as Lineage 2, Lineage 4, etc.), analyze the differences in transmission characteristics, pathogenicity, and drug resistance of different lineage strains, and provide an evolutionary perspective for transmission tracing.
[0079] (2)Lineage-specific mutations: Identify mutation sites unique to different lineages based on the TBprofiler lineage annotation database, analyze the key roles of these specific mutations in the process of strain adaptive evolution and the formation of transmission advantages, and explore their potential value as strain transmission markers.
[0080] 4. Data integration and normalization: Align relevant data columns of pan-genomics and epidemiological data according to row indices. Use the Scikit-learn framework to standardize the data, perform one-hot encoding on categorical features (such as mutation types, lineage classifications), and standardize or normalize numerical features (such as the number of mutations, geographical distances) data, mapping to a specific interval (range with a mean of 0 and a standard deviation of 1) to eliminate the impact of data dimension differences on the model and improve the stability and accuracy of the model.
[0081] IV. The impact of combinations of different features on evaluation performance Through univariate analysis and multivariate combination analysis, screen out features that have a significant contribution to transmission detection. The specific steps are as follows: Univariate feature analysis: Construct univariate classification models (such as logistic regression) for the 15 candidate features obtained by screening, and calculate the AUC, sensitivity, specificity, and SHAP values (see Table 1).
[0082] Multivariate combination analysis: According to the univariate results, group the features by contribution degree (genetic differences, epidemiological information, strain lineages, other features), and test the model performance of different combinations (see Table 2).
[0083] Table 1. Results of individual feature importance analysis
[0084] It can be seen from Table 1 that the five features of non-synonymous mutation ratio, total SNP number, presence / absence variation (PAV) of genes, number of rare mutations, and strain lineage classification all have a relatively high contribution degree (AUC > 0.75 and SHAP > 0.25), while the three features of synonymous mutation ratio, number of structural variations, and number of virulence-related mutations are low contribution features (AUC < 0.70 or SHAP < 0.15). In addition, although the SHAP values of the three features of the Euclidean distance of the patient's residence, sampling time interval, and whether the patients are in the same community are relatively low, the Euclidean distance of the patient's residence, sampling time interval, and community association enhance the model robustness through non-linear interaction and have irreplaceable epidemiological value in key transmission windows (such as short distance / short time) and prevention and control operability (such as delimiting the scope of precise screening), and can still be considered to have a relatively high contribution degree.
[0085] In this embodiment, different features of the same group and different features of different groups were examined, combined respectively, and a detection model for the transmission relationship of Mycobacterium tuberculosis was constructed respectively. The diagnostic performance of different models was evaluated respectively. A random forest was used to construct the model. During the model construction process, the above parameters were combined and tested through grid search (GridSearchCV), and the optimal parameter combination was finally selected, specifically including: 1) The number of trees (n_estimators) was set to 500 to improve the model stability and prediction accuracy by increasing the number of base learners and reduce the risk of overfitting; 2) The maximum depth of a single tree (max_depth) was set to 12 to limit the excessive growth of the tree to balance the model complexity and generalization ability; 3) The minimum number of samples for node splitting (min_samples_split) was set to 5 to prevent overfitting to a small number of samples; 4) The minimum number of samples in a leaf node (min_samples_leaf) was set to 2 to ensure that the leaf node has a statistically significant sample size; 5) The feature selection method (max_features) was set to'sqrt', that is, each tree randomly selects the square root of the number of features for splitting, enhancing the model diversity and reducing the impact of feature correlation on the results; 6) The splitting criterion (criterion) was set to 'gini', using the Gini coefficient as the node splitting criterion, which is suitable for classification tasks, has high computational efficiency and is robust to noisy data; 7) The sample sampling method (bootstrap) was set to True, and the bootstrap sampling method was used to generate the training set for each tree to further improve the model generalization ability; 8) The sample sampling ratio (max_samples) was set to 0.8, and each tree was trained using 80% of the samples, and 20% of the samples were reserved for verification to avoid the model's over-reliance on specific samples.
[0086] The evaluation results are shown in Table 2.
[0087] Table 2. Influence of Feature Combinations of Different Groups on Evaluation Performance
[0088] As can be seen from Table 2, when the high-contribution features (5) are combined with the epidemiological information (3), compared with other combination methods, it can significantly improve the performance of the model in assessing the recent transmission risk of Mycobacterium tuberculosis. The AUC value reaches 0.97, the sensitivity is 100%, and the specificity is 96.2%. However, the addition of low-contribution features (such as the proportion of synonymous mutations, the number of structural variations, and the number of virulence-related mutations) not only fails to improve the performance of the model but also decreases the performance (for example, the AUC of the combination of all 15 features is only 0.93). On the contrary, the performance of the model with all 15 features (AUC 0.93) is lower than that of the high-contribution feature subset (AUC 0.95), which proves that blindly increasing the number of features will introduce noise and reduce the generalization ability of the model. Therefore, it is preferred to construct an assessment model for the recent transmission risk of Mycobacterium tuberculosis using the feature combination of high-contribution features (proportion of non-synonymous mutations, total number of SNPs, presence / absence variation of genes, number of rare mutations, strain lineage classification) and epidemiological information (Euclidean distance of patient residence, sampling time interval, whether patients are in the same community). Reducing redundant features can improve the computational efficiency by 30% and avoid overfitting at the same time.
[0089] The final model did not use all 15 features. Instead, through the screening of Table 1 and Table 2, 8 key features were selected to build the model. The 8 features include: (1) Genetic differences: proportion of non-synonymous mutations, total number of SNPs, presence / absence variation of genes (PAV), number of rare mutations; (2) Epidemiology: Euclidean distance of patient residence, sampling time interval, whether patients are in the same community; (3) Strain lineage information: strain lineage classification.
[0090] The detection methods for the 8 key features are as follows: The detection methods for the proportion of non-synonymous mutations, total number of SNPs, presence / absence variation of genes (PAV), and number of rare mutations are: 1. Proportion of non-synonymous mutations Steps: 1) Use GATK HaplotypeCaller and BCFtools to detect SNPs and InDels in the whole-genome sequencing data; 2) Annotate the biological effects (non-synonymous, synonymous, and nonsense) of the variations through SNPEff; Calculate the proportion:
[0091] 2. Total number of SNPs Steps: 1) Based on the H37Rv reference genome (NC_000962.3), use BWA-MEM for sequence alignment; 2) Remove sequencing / alignment errors through GATK FilterMutectCalls (parameters: --min-reads 5, --min-allele-fraction 0.05); Count the total number of SNPs after filtering (including homologous recombination regions and non-coding regions).
[0092] 3. Number of presence / absence variations (PAVs) of genes Steps: 1) Construct the MTB pan-genome: Integrate multi-strain gene sets using Panaroo (parameters: --core_threshold 0.95); 2) Determine the presence / absence of genes in the target strain through BLASTn (E-value ≤ 1e-5, coverage ≥ 80%); 3) Count the number of PAVs relative to the reference genome (presence = 1, absence = 0).
[0093] 4. Number of rare mutations On the basis of the first step of detecting the total number of SNPs, further detect low-frequency mutations (frequency < 5%) using LoFreq (parameters: --min-cov 20, --min-alt 3); The detection methods for the Euclidean distance of the patient's residence, the sampling time interval, and whether the patients are in the same community are as follows: 1. Euclidean distance of the patient's residence Steps: 1) Data source: Extract the patient's address from the case registration system or electronic health record (EHR); 2) Geocoding: Use the map software API to convert the address into longitude and latitude coordinates; 3) Distance calculation: Based on the spherical Euclidean distance formula (Haversine formula), use the Pythongeopy.distance library to calculate the straight-line distance between two points (unit: meter): 5. Detection method for the sampling time interval: 1) Data extraction: Export the sample collection timestamp from the laboratory information management system (LIMS) (format: YYYY-MM-DD HH:MM:SS); 2) Time difference calculation: Use the Python datetime module to calculate the number of days between two samples; 3) Normalization: Convert the number of days into months (divide by 30) and standardize it to a Z-score (mean = 9.2, standard deviation = 4.1); 2. Detection method for whether the patients are in the same community: Geographical proximity (meeting either one is acceptable): 1) Straight-line distance ≤ 1000 meters (calculated by the above Euclidean distance); 2) Administrative boundary matching: Use GIS layers (such as ArcGIS Shapefile) to determine whether they belong to the same administrative village / street.
[0094] 3. The detection method for strain pedigree classification is: The results of total SNP detection were input into TB-profiler software (database version: v4.4.0) to annotate based on lineage-specific marker SNPs and output the lineage classification results of the strains.
[0095] Through the above systematic and rigorous sample collection and feature extraction process, this embodiment successfully obtained rich and high-quality characteristic data related to the transmission of Mycobacterium tuberculosis, laying a solid foundation for the construction of subsequent prediction models, and effectively promoting the development of the model in the direction of accuracy and efficiency. It is expected to significantly improve the ability to predict the transmission path and association of the disease, and provide a scientific and reliable basis for clinical decision-making and public health strategy formulation for tuberculosis prevention and control. It has broad application prospects in the field of Mycobacterium tuberculosis transmission research and prevention and control.
[0096] Example 2 Construction of the optimal detection model for the recent spread of Mycobacterium tuberculosis
[0097] In Example 1, sample collection and feature initial screening have been successfully completed, and feature data related to the spread of Mycobacterium tuberculosis have been obtained. In order to fully explore the information contained in these data and build a spread detection model with excellent performance, this example uses a series of model training and optimization strategies to build the best detection model for the recent spread of Mycobacterium tuberculosis. The specific process is as follows: 1. Construction of the best model 1. Divide the training set and test set The data is randomly divided into training set and validation set in a ratio of 7:3, that is, 70% of the samples are used as training set for model training, and 30% of the samples are used as validation set for model evaluation and hyperparameter adjustment.
[0098] 2. Clustering model implementation and evaluation The DBSCAN clustering algorithm in the Scikit-learn framework is used to perform data clustering analysis to discover the potential cluster structure in MTB propagation, reduce the computational complexity and data volume of subsequent classification tasks, make the classification process more focused on representative clusters, and effectively improve the efficiency and accuracy of classification. Figure 2As shown, the optimization of DBSCAN parameters mainly includes determining the neighborhood radius (Epsilon value) and the minimum number of samples. Among them, the reference value of the neighborhood radius is 0.5, and the minimum number of samples is 5. Figure 2 By observing the change in the distribution density of data points at different distances and combining the elbow method, when the appropriate neighborhood radius is determined to be 0.5, the density of data points. When the k-distance increases, the density of data points gradually decreases. At a certain inflection point, the downward trend of the density significantly slows down, and the distance corresponding to this inflection point can be used as the preliminary reference value of the neighborhood radius. Within this distance range, the distribution of data points begins to show a relatively stable state, which can better reflect the local density characteristics of the data, thus initially determining a suitable neighborhood range so that the data points within this range can form a relatively tight clustering structure. At the same time, according to the scale and distribution characteristics of the data, as well as the requirements for clustering compactness and separation, the minimum number of samples is determined through multiple experiments and verifications. For example, in the case of a small-scale dataset with relatively concentrated data distribution, the minimum number of samples can be appropriately reduced; while for large-scale and complex data, in order to avoid generating too many noise clusters, the minimum number of samples needs to be increased.
[0099] From Figure 3 it can be seen that the clustering results show an obvious cluster-like distribution. There is a certain degree of separation between different clusters, and the samples within the clusters have high similarity in genetic characteristics and epidemiological characteristics. Specifically manifested as: 1) The cluster-like distribution corresponds to the transmission relationship: 6 main clusters (Cluster 0 - 5) can be seen in the figure, and each cluster corresponds to an independent transmission chain (for example, Cluster 0 contains 10 samples, all of which are transmission cases within the same community); 2) The separation between clusters is significant: The center distance between different clusters is greater than 2.0 (Euclidean distance), indicating no cross-infection between transmission chains; 3) The consistency of intra-cluster characteristics: The genetic difference characteristics (such as the average number of SNPs ≤ 5, the rare mutation ratio ≥ 80%) and epidemiological information (such as the geographical distance ≤ 1km, the sampling time interval ≤ 30 days) of the samples within the same cluster are highly consistent, conforming to the biological laws of recent transmission.
[0100] Store the clustering labels into the corresponding variables. These labels will serve as important auxiliary information for the subsequent classification model, helping the model better understand the potential relationships between samples and further improving the performance of the model. The specific functions include: 1) Reduce data complexity: Divide the samples into subsets through cluster labels, reducing the amount of noise data that the classification model needs to process; 2) Enhance feature expression: Use the cluster labels as one of the input features of the classification model, providing the population membership information of the samples and assisting the model in capturing the spatio-temporal patterns of transmission relationships; 3) Optimize training efficiency: Design differentiated training strategies for different clusters (such as increasing sample weights for clusters with high transmission risks) to enhance the model's ability to identify key transmission chains.
[0101] Therefore, it can be seen that the DBSCAN clustering results not only verify the clustering characteristics of Mycobacterium tuberculosis transmission but also provide key structured prior knowledge for machine learning models, enabling the models to more efficiently and accurately identify transmission relationships and solving the performance bottleneck problem caused by data redundancy and noise in traditional methods.
[0102] 3. Screening of the best machine learning algorithm and construction of the best prediction model Due to the complexity of sample features, it is impossible to estimate which machine learning algorithm is more suitable for the data in this study without prior knowledge. Therefore, in this embodiment, six machine learning algorithms, namely neural network, naive Bayes, random forest, generalized linear model, gradient boosting machine, and support vector machine, are used to construct prediction models respectively. The construction methods of the six prediction models are as follows: First, train the six machine learning models with the best hyperparameters, and the hyperparameter configuration is determined through grid search and 10-fold cross-validation, as shown in Table 3 specifically. Subsequently, complete model training, validation, and optimization according to the following process.
[0103] 1) Model training: According to the hyperparameter configuration in Table 3, accelerate the training of the neural network with GPU and parallelize the training of other models with multi-core CPU to optimize the computing efficiency.
[0104] 2) Dynamic evaluation: Automatically switch evaluation metrics according to the characteristics of data distribution (such as whether the categories are balanced), and additionally test the prediction stability for time series data.
[0105] 3) Performance comparison: Comprehensively compare the six models from three dimensions: prediction accuracy (such as AUC), operation speed, and feature importance, and screen out random forest and neural network as the dominant models.
[0106] 4) Ensemble optimization: Combine the prediction results of the selected models with the original features and perform secondary training through logistic regression. Finally, the false positive rate of the ensemble model is reduced by 40%.
[0107] Table 3. Hyperparameter selection of machine learning models
[0108] Determine whether strains belong to the same transmission cluster according to the transmission probability threshold (≥0.7) (that is, a transmission probability ≥0.7 is determined to be transmission-related, and <0.7 is determined to be unrelated). Use the 10-fold cross-validation method to screen out the optimal models under 6 algorithms. The specific steps are as follows: Determine whether strains belong to the same transmission cluster according to the transmission probability threshold (≥0.7) (that is, a transmission probability ≥0.7 is determined to be transmission-related, and <0.7 is determined to be unrelated). Use the 10-fold cross-validation method to screen out the optimal models under 6 algorithms. The specific steps are as follows: 0107. Cross-validation design: Randomly divide the training set into 10 subsets. Each time, take 9 subsets to train the model, and the remaining 1 subset is used to verify the performance. Repeat 10 times and take the average result; Performance evaluation indicators: Use the area under the ROC curve (AUC) as the core indicator, supplemented by sensitivity, specificity, and accuracy. Optimal model screening: Select the model with the highest median AUC in cross-validation as the optimal model for each algorithm.
[0109] Use the ROC method to evaluate the performance of the optimal model. The results are as Figure 4 shown. Among them, the model constructed by the ranger algorithm based on random forest has the largest median ROC, specifically 0.97 (95% confidence interval: 0.95 - 0.99), which is significantly better than other algorithms (such as gradient boosting machine AUC = 0.94, support vector machine AUC = 0.92). This model shows the following advantages in cross-validation: Sensitivity 100%: Ensure that all true transmission clusters are accurately identified; Specificity 96.2%: Effectively avoid misjudging non-transmission clusters as transmission clusters; Kappa value 0.962: The model prediction results are highly consistent with the true transmission relationship.
[0110] The above results show (Table 4) that the random forest model has the best performance in distinguishing transmission clusters and non-transmission clusters. Its core advantages are: Robustness to high-dimensional data: It can effectively handle the non-linear relationship between genetic difference features (such as rare mutations) and epidemiological information (such as geographical distance); Anti-overfitting ability: By randomly selecting features (max_features = sqrt) and sample bootstrap sampling (bootstrap = True), enhance the generalization ability of the model; Interpretability: The feature importance ranking based on SHAP values (such as the contribution degree of the number of rare mutations is the highest) provides a basis for the analysis of the transmission mechanism.
[0111] Therefore, define the ranger algorithm based on random forest as the final prediction model, providing a high-precision and high-reliability technical solution for the detection of recent transmission of Mycobacterium tuberculosis.
[0112] Table 4. Performance comparison of the optimal models of 6 algorithms
[0113] 4. Performance Verification of the Optimal Prediction Model The performance of the final prediction model was verified in the validation set. The ROC analysis method was used for model performance evaluation, and the performance metrics included: accuracy, Kappa value, sensitivity, specificity, positive predictive rate, negative predictive rate, and feature importance. The results are as Figure 5 shown. The AUC of the final prediction model in the validation set was 0.97, indicating that the model had extremely high accuracy and reliability in distinguishing positive and negative samples and could effectively identify sample pairs with and without transmission relationships. The accuracy was 0.981, indicating that the proportion of samples correctly predicted by the model was as high as 98.1%, further demonstrating the high precision of the model. The Kappa value was 0.962, reflecting a very high degree of consistency between the model's prediction results and the actual situation and ruling out the high accuracy caused by accidental factors. The sensitivity was 100%, meaning that the model could accurately identify all samples with actual transmission relationships without false negatives. The specificity was 96.2%, indicating that the model also had high accuracy in identifying samples without transmission relationships with only a small number of misjudgments. The positive predictive rate was 96.4%, indicating that among the samples predicted by the model as positive (with transmission relationships), the proportion of actual positives was relatively high. The negative predictive rate was 100%, that is, all samples predicted by the model as negative (without transmission relationships) were actually negative, ensuring the reliability of the model in negative prediction.
[0114] Therefore, the above performance metrics comprehensively indicate that the optimal prediction model constructed in the present invention has excellent performance, can accurately detect the recent transmission of Mycobacterium tuberculosis, and provides strong technical support for tuberculosis prevention and control.
[0115] 5. Feature Importance Evaluation In this embodiment, the random forest model was used to calculate the importance score of each feature Calculation process of feature importance score: 1) Data preparation: Collect and clean the dataset to ensure it contains all relevant features (such as non-synonymous mutation ratio, total SNP count, presence / absence variation of genes, number of rare mutations, Euclidean distance of patient residence, sampling time interval, strain lineage classification, and number of drug resistance-related mutations). Divide the dataset into feature variables (X) and target variable (y), where the target variable is the label related to Mycobacterium tuberculosis transmission. 2) Train the random forest model: Use the training set data to train the random forest model. During training, the model will learn how to predict the target variable based on the input features. 3) Calculate feature importance: The random forest model will automatically calculate the importance score for each feature during training. The importance score reflects the contribution of the feature to the model's decision-making. Specifically, the model will analyze the role of each feature in the node splitting process, calculate the reduction in impurity when using the feature, and thus obtain the importance score for each feature. 4) Result collation: Organize the importance scores of each feature into a table for easy comparison and analysis. Sort the features according to the importance scores, usually from high to low, to identify the features that have the greatest impact on the model prediction.
[0116] The specific importance scores of the features of the MTB transmission detection model are shown in Table 5 as follows.
[0117] Table 5. Results of the importance scores of the features of the MTB transmission detection model
[0118]
[0119] Result analysis and interpretation Features with high importance scores (the higher the score, the more important): Non-synonymous mutation ratio (0.30): This is the feature with the greatest impact in the model, indicating that this feature plays a significant role in judging the transmission risk.
[0120] Total SNP count (0.25): This feature also has an important impact on the model's prediction and may be related to the construction of the transmission chain.
[0121] Features with medium importance scores: Presence / absence variation of genes (PAV) (0.22), number of rare mutations (0.20), and whether patients are from the same community (0.18): These three features also provide some support for the model's judgment and offer additional information.
[0122] Features with low importance scores: The scores of features such as the Euclidean distance of patient residence (0.12), sampling time interval (0.05), and strain lineage classification (0.1) are relatively low, indicating that their contributions to the model are limited.
[0123] Through the above rigorous model construction, performance verification, and feature importance evaluation processes, this embodiment successfully constructs a Mycobacterium tuberculosis near-transmission relationship detection model with high accuracy and reliability, laying a solid foundation for subsequent result visualization and analysis as well as practical applications. It has important technological innovation and application value in the field of Mycobacterium tuberculosis transmission research and is expected to bring significant improvements to tuberculosis prevention and control work.
[0124] After successfully constructing a Mycobacterium tuberculosis transmission relationship detection model with high accuracy and reliability according to Example 2 and comprehensively evaluating the model performance and feature importance, this embodiment focuses on the in-depth mining and intuitive display of the model output results, aiming to transform complex data information into knowledge with practical guiding significance and provide direct support for tuberculosis prevention and control work.
[0125] In the tuberculosis outbreak investigation of a certain city in this embodiment, a transmission cluster composed of 1 strain was systematically detected, and 152 initially infected patients were collected. For each patient sample, the detection results of eight features were obtained, including the non-synonymous mutation ratio, total SNP count, presence / absence variation (PAV) of genes, rare mutation count, Euclidean distance of patient residence, sampling time interval, and whether the patients are in the same community, for the remaining 151 patients.
[0126] According to the random forest model constructed in Example 2, the detection results are input, the judgment threshold is 0.7. If the output value is higher than or equal to 0.7, it is judged as positive (that is, patient B is likely to be infected from patient A), and if the output value is lower than 0.7, it is judged as negative (that is, patient B is not infected from patient A).
[0127] Retrospective verification was carried out on the culture-positive strains of 152 pairs of patients with a known transmission chain determined by on-site epidemiological investigation in combination with a common exposure environment (such as family, school, workplace, etc.) or close contact history. The results are shown in Table 6: Table 6. Verification of Mycobacterium tuberculosis transmission relationship detection results for 152 patients
[0128] Since there is currently no gold standard for judging tuberculosis transmission relationships, it can only be limited by very strict screening conditions to verify the evaluation and analysis effect of the model. For example, there should be a transmission relationship between patients in the same family, which is judged as a true positive; to show a non-transmission relationship, samples are taken from different regions. For example, 4 transmission cases in a family are selected in region A, and 3 transmission cases in a family are selected in region B. Through epidemiological investigation, it is judged that there is no intersection in the life and travel of these two family members, so there is no transmission relationship between the two family members. If it is judged to have a transmission relationship, it is a false positive.
[0129] Performance indicators: Positive predictive value (PPV): 112 / 115 = 97.4%; Negative predictive value (NPV): 31 / 37 = 83.8%; Youden index: 0.94.
[0130] According to the detection results of the Mycobacterium tuberculosis transmission relationship detection model, the iGraph software was used to construct a visualization network for the transmission clustering results (see Figure 6 ), and a visualization relationship network was constructed for strains with a transmission probability greater than 0.7. In the network, nodes represent samples, and samples predicted to have a transmission relationship are connected by edges. The numbers on the edges represent the predicted transmission probability between samples. It can be seen from Figure 6 that the diffusion process of tuberculosis starting from the source case C: As an asymptomatic carrier, C infects B (probability 0.82) and D (0.78) through Route 1, and infects F (0.85) through Route 2. Subsequently, B infects A (0.89) through Route 3; D infects E (0.72) through Route 4. As a super spreader, F infects G (0.75) and H (0.81) through Route 5. G then infects I (0.76) through Route 6, while H infects J (0.83) through Route 7, and finally I also forms a transmission with J (0.71) through Route 8, fully demonstrating the hidden diffusion trajectory of tuberculosis through the family-community-medical institution chain.
[0131] All patents and publications mentioned in the specification of the present invention indicate that these are publicly available technologies in the field and can be used in the present invention. All patents and publications cited herein are equally listed in the references, as if each publication was specifically referenced individually. The present invention described herein can be implemented in the absence of any one or more elements, one or more limitations, where such limitations are not specifically stated. For example, in each example herein, the terms "comprising", "consisting essentially of", and "consisting of" can be replaced by the other two of the three terms. The so-called "a" herein only means "one", and does not exclude including only one, nor does it exclude including more than two. The terms and expressions adopted herein are for descriptive purposes and are not limiting, and there is no intention to indicate that the terms and interpretations described in this book exclude any equivalent features, but it can be understood that any suitable changes or modifications can be made within the scope of the present invention and the claims. It can be understood that the embodiments described in the present invention are all preferred embodiments and features, and any person of ordinary skill in the art can make some changes and variations based on the essence described in the present invention, and these changes and variations are also considered to be within the scope of the present invention and the scope limited by the independent claims and the dependent claims.
Claims
1. A Mycobacterium tuberculosis transmission relationship detection system, characterized in that: In the system, a model for detecting the recent spread of Mycobacterium tuberculosis is constructed by using a combination of one or more of genetic difference features, epidemiological information features, strain lineage information features, and other features.
2. The system according to claim 1, characterized in that The genetic difference characteristics include any one or more selected from the following: non-synonymous mutation ratio, synonymous mutation ratio, number of promoter region mutations, total number of SNPs, total number of InDels, number of structural variations, number of gene presence / absence variations, number of rare mutations; The epidemiological information features include any one or more selected from the following: Euclidean distance of the patient's residence, sampling time interval, whether the patients are in the same community; The strain pedigree information features include any one or more selected from the following: strain pedigree classification, number of pedigree-specific mutations; The other characteristics include any one or more selected from the following: the number of drug resistance-related mutations and the number of virulence-related mutations.
3. The system according to claim 2, characterized in that The model is used to analyze the transmission relationship of Mycobacterium tuberculosis between patients. By analyzing the transmission relationship between patient B and patient A, it is determined whether the Mycobacterium tuberculosis in patient B is acquired from patient A.
4. The system according to claim 3, characterized in that The model analyzes the transmission relationship between patient B and patient A by calculating the probability value of whether the Mycobacterium tuberculosis in patient B is transmitted from patient A. If the probability value is higher than the threshold, it proves that the Mycobacterium tuberculosis in patient B is transmitted from patient A.
5. The system according to claim 4, characterized in that The model is constructed by combining genetic difference characteristics, epidemiological information characteristics, and strain pedigree information characteristics; the genetic difference characteristics include a combination of the proportion of non-synonymous mutations, the total number of SNPs, the number of gene presence / absence variations, and the number of rare mutations; the epidemiological information characteristics include a combination of the Euclidean distance of the patient's residence, the sampling time interval, and whether the patients are in the same community; the strain pedigree information characteristics include a combination of strain pedigree classification and the number of pedigree-specific mutations; the strain pedigree information characteristics include strain pedigree classification.
6. The system according to claim 5, characterized in that The model is constructed using a machine learning algorithm, which includes any one or more of a neural network, naive Bayes, random forest, generalized linear model, gradient boosting machine, and support vector machine.
7. The system according to claim 6, characterized in that The model was constructed using random forest.
8. A feature combination for evaluating the transmission relationship of Mycobacterium tuberculosis, characterized in that: It includes genetic difference characteristics, epidemiological information characteristics and strain pedigree information characteristics; the genetic difference characteristics include the proportion of non-synonymous mutations, the total number of SNPs, the number of gene presence / absence variations, and the number of rare mutations; the epidemiological information characteristics include the Euclidean distance of the patient's residence, the sampling time interval, and whether the patients are in the same community; the strain pedigree information characteristics include the strain pedigree classification.
9. A method for analyzing the transmission relationship of Mycobacterium tuberculosis, characterized in that: The characteristics include any one or more of genetic difference characteristics, epidemiological information characteristics, strain pedigree information characteristics, and other characteristics; the genetic difference characteristics include any one or more selected from the following: non-synonymous mutation ratio, synonymous mutation ratio, number of promoter region mutations, total number of SNPs, total number of InDels, number of structural variations, number of gene presence / absence variations, number of rare mutations; the epidemiological information characteristics include any one or more selected from the following: Euclidean distance of the patient's residence, sampling time interval, whether the patient is in the same community; the strain pedigree information characteristics include any one or more selected from the following: strain pedigree classification, number of pedigree-specific mutations; The other characteristics include any one or more selected from the following: the number of drug resistance-related mutations and the number of virulence-related mutations.
10. The use according to claim 9, characterized in that The characteristics include genetic difference characteristics, epidemiological information characteristics and strain lineage information characteristics; the genetic difference characteristics include the proportion of non-synonymous mutations, the total number of SNPs, the number of gene presence / absence variations and the number of rare mutations; the epidemiological information characteristics include the Euclidean distance of the patient's residence, the sampling time interval and whether the patients are in the same community; the strain lineage information characteristics include strain lineage classification.
Citation Information
Patent Citations
Methods, systems and processes of determining transmission paths of infectious agents
CN108475297A
Tuberculosis report case recent propagation prediction method based on machine learning
CN119446564A
Infectious disease prediction and transmission link speculation method and system based on random forest
CN119626576A
Construction of a Comparative Database and Identification of Virulence Factors Comparison of Polymorphic Regions in Clinical Isolates of Infectious Organisms
US20080085284A1