A method for constructing a tuberculosis incidence risk prediction model and application thereof

CN122531749APending Publication Date: 2026-08-07SHANGHAI MUNICIPAL CENT FOR DISEASE CONTROL & PREVENTION
View PDF 0 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
SHANGHAI MUNICIPAL CENT FOR DISEASE CONTROL & PREVENTION
Filing Date
2026-06-25
Publication Date
2026-08-07

AI Technical Summary

Technical Problem

[0004]本申请提供一种结核病发病风险预测模型的构建方法及其应用,旨在解决现有技术存在标志物组合存在数量冗余、验证样本量不足以及预测结果可靠性差的技术问题

Benefits of technology

1.本申请所提供的GBP6、SEPTIN4、FCGR1A、SERPING1、ETV7、BATF2、ANKRD22共 7个基因构成的基因标志物组合,是经过多维度严格筛选、多独立队列验证、与结核病发病风险高度关联的核心特征集合,具备高关联性、高稳定性、高特异性、精简高效的突出优势,能够从分子层面精准反映潜伏性结核感染者向活动性结核病进展的内在生物学特征,有效解决现有标志物组合冗余、效能不足、普适性差的技术缺陷。

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN122531749A_ABST
    Figure CN122531749A_ABST
Patent Text Reader

Abstract

The application discloses a construction method of a tuberculosis onset risk prediction model and application thereof, and belongs to the technical field of biological medicine. The method comprises the following steps: collecting multiple groups of independent latent tuberculosis infection cohort samples, obtaining transcriptome data of 7 genes, namely, GBP6, SEPTIN4, FCGR1A, SERPING1, ETV7, BATF2 and ANKRD22, and taking whether developing into active tuberculosis within 2 years as an outcome variable; obtaining a standard expression matrix through data preprocessing of quality control, normalization and batch correction; verifying a marker through differential expression analysis and WGCNA screening; constructing a tuberculosis onset risk prediction model by using a machine learning algorithm combined with hierarchical K-fold cross-validation, and verifying the model by using an independent cohort. The tuberculosis onset risk prediction model constructed by the application has excellent prediction performance and strong stability, can accurately quantify risks and stratify, is suitable for screening of high-risk groups of tuberculosis, can be used for preparing a detection kit, and has high clinical conversion value.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application belongs to the field of biomedical technology, specifically relating to a method for constructing a tuberculosis incidence risk prediction model and its application. Background Technology

[0002] About a quarter of the world's population is infected with Mycobacterium tuberculosis, but only 5%-15% of those with latent tuberculosis infection will progress to active tuberculosis within the next two years. Accurate identification of high-risk individuals is a core aspect of tuberculosis prevention and control, facilitating early intervention and reducing the incidence rate.

[0003] In recent years, the development of transcriptomics technology has provided new directions for tuberculosis risk prediction, and some tuberculosis-related gene biomarkers have been reported. However, the biomarker combinations used in existing tuberculosis infection detection methods are not concise enough, and some biomarkers have low association with the risk of disease, increasing the cost and complexity of detection and model construction; most biomarker combinations have not been validated by many independent cohorts, and the reliability and universality of the prediction results are insufficient, making it difficult to extend to different populations and clinical scenarios. Summary of the Invention

[0004] This application provides a method for constructing a tuberculosis incidence risk prediction model and its application, aiming to solve the technical problems of redundant biomarker combinations, insufficient validation sample size, and poor reliability of prediction results in existing technologies.

[0005] To achieve the above objectives, this application provides a combination of gene markers, which includes the GBP6, SEPTIN4, FCGR1A, SERPING1, ETV7, BATF2, and ANKRD22 genes.

[0006] This application also provides a method for constructing a tuberculosis incidence risk prediction model, comprising the following steps: (1) collecting multiple independent cohorts of latent tuberculosis infected individuals, obtaining transcriptome expression data of the above-mentioned gene marker combination in the samples, and recording whether the latent tuberculosis infected individuals progress to active tuberculosis within 2 years as an outcome variable; (2) performing quality control, normalization, and batch correction data preprocessing on the transcriptome expression data to obtain a standard expression matrix; (3) screening the standard expression matrix through differential expression analysis and weighted gene co-expression network analysis to verify the significant association between each gene in the gene marker combination and the risk of tuberculosis incidence; (4) using a machine learning algorithm, with the expression level of each gene in the gene marker combination as the independent variable and the incidence outcome of active tuberculosis within 2 years as the dependent variable to construct a tuberculosis incidence risk prediction model, and optimizing the model parameters through hierarchical K-fold cross-validation; (5) using an independent external cohort to validate the tuberculosis incidence risk prediction model and evaluate the predictive efficacy of the tuberculosis incidence risk prediction model.

[0007] Optionally, in step (1), the sample is a peripheral blood sample of a person with latent tuberculosis infection, covering people of different ages, genders, regions and HIV infection status, the total number of the sample is not less than 1200 cases, and the proportion of active tuberculosis progression in the sample is not less than 18%.

[0008] Optionally, in step (2), the data preprocessing for quality control, normalization and batch correction of the transcriptome expression data specifically includes: removing low-quality sequencing data from the transcriptome expression data, normalizing the expression level using the TPM method, and removing batch effects using the ComBat method.

[0009] Optionally, in step (3), the standard expression matrix is ​​screened by differential expression analysis and weighted gene co-expression network analysis. The specific screening criteria are: differential expression fold ≥ 1.5, P value < 0.01, and consistent expression trend in at least 5 independent datasets.

[0010] Optionally, in step (4), the machine learning algorithm is one or more combinations of LightGBM, logistic regression, and random forest; The K-value for the stratified K-fold cross-validation is 5-10, and the samples are stratified according to their disease status.

[0011] Optionally, in step (5), the indicators for validating the tuberculosis incidence risk prediction model include AUC, sensitivity, and specificity. For external validation, AUC ≥ 0.85, and sensitivity and specificity ≥ 75%.

[0012] Optionally, the tuberculosis incidence risk prediction model outputs the probability value of incidence within 2 years, and completes the three-level risk classification according to probability < 0.2 as low risk, 0.2 ≤ probability < 0.6 as medium risk, and probability ≥ 0.6 as high risk.

[0013] This application also provides an application of the above-described construction method in the preparation of a tuberculosis incidence risk detection kit.

[0014] Optionally, the tuberculosis incidence risk detection kit is used to detect the expression levels of GBP6, SEPTIN4, FCGR1A, SERPING1, ETV7, BATF2, and ANKRD22 genes.

[0015] Compared with the prior art, this application has the following beneficial effects: 1. The gene biomarker combination consisting of seven genes—GBP6, SEPTIN4, FCGR1A, SERPING1, ETV7, BATF2, and ANKRD22—provided in this application is a set of core features that have undergone rigorous multi-dimensional screening and multi-independent cohort validation and are highly associated with the risk of tuberculosis. It has outstanding advantages of high correlation, high stability, high specificity, and conciseness and efficiency. It can accurately reflect the intrinsic biological characteristics of the progression from latent tuberculosis infection to active tuberculosis at the molecular level, effectively solving the technical defects of existing biomarker combinations such as redundancy, insufficient efficacy, and poor universality.

[0016] 2. The method for constructing the tuberculosis incidence risk prediction model in this application solves the technical problems of redundant biomarkers, insufficient predictive efficacy, and poor universality in existing technologies through a complete process of sample collection, data preprocessing, biomarker validation, model construction, and multi-cohort validation. The tuberculosis incidence risk prediction model constructed in this application achieves an internal validation AUC of over 0.958 and an external validation AUC of over 0.886. The number of biomarkers is reduced and there is no redundancy, significantly reducing detection and analysis costs. Multi-cohort validation shows no significant population heterogeneity. It can accurately quantify the risk of latent tuberculosis infection within 2 years, with a sensitivity and specificity of ≥75% for predicting the risk of incidence within 2 years. This fills the gap in clinically practical prediction models and provides standardized technical support for the precise prevention and control of tuberculosis. Compared with existing methods, it has the advantages of high efficiency, low cost, and high universality, meeting the needs of clinical translation and large-scale application. Attached Figure Description

[0017] Figure 1 This is a flowchart illustrating the method for constructing a tuberculosis incidence risk prediction model in this application embodiment. Detailed Implementation

[0018] The technical solutions of the embodiments of this application will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only a part of the embodiments of this application, and not all of the embodiments. Based on the embodiments of this application, all other embodiments obtained by those of ordinary skill in the art without creative effort are within the scope of protection of this application.

[0019] This application provides a combination of gene markers, which includes seven genes: GBP6, SEPTIN4, FCGR1A, SERPING1, ETV7, BATF2, and ANKRD22.

[0020] Specifically, the gene biomarker combination consisting of seven genes—GBP6, SEPTIN4, FCGR1A, SERPING1, ETV7, BATF2, and ANKRD22—provided in this application is a set of core features that have undergone rigorous multi-dimensional screening and multi-independent cohort validation and are highly associated with the risk of tuberculosis. It has outstanding advantages of high correlation, high stability, high specificity, and conciseness and efficiency. It can accurately reflect the intrinsic biological characteristics of the progression from latent tuberculosis infection to active tuberculosis at the molecular level, effectively solving the technical defects of existing biomarker combinations such as redundancy, insufficient efficacy, and poor universality.

[0021] From a biological function perspective, all seven genes mentioned above are involved in the host immune response, inflammation regulation, macrophage activation, and post-infection cell signaling processes induced by Mycobacterium tuberculosis infection. Changes in their expression levels directly correspond to the host's resistance to Mycobacterium tuberculosis and the progression of the disease, making them ideal molecular indicators reflecting the risk of tuberculosis. Among them, GBP6, ETV7, and BATF2 are closely related to interferon-mediated anti-mycobacterial immunity; SEPTIN4 and ANKRD22 are involved in cytoskeleton remodeling and regulation of the microenvironment of infected cells; and FCGR1A and SERPING1 play key roles in immune cell phagocytosis and inflammatory cascade reactions. Together, they constitute a complete biomarker panel covering immune recognition, signal transduction, and inflammation regulation, which can comprehensively capture molecular characteristics related to the risk of disease.

[0022] like Figure 1 As shown, this application also provides a method for constructing a tuberculosis incidence risk prediction model, including the following steps: S1, collecting multiple independent cohorts of latent tuberculosis infected individuals, obtaining transcriptome expression data of the above-mentioned gene marker combinations in the samples, and recording whether latent tuberculosis infected individuals progress to active tuberculosis within 2 years as an outcome variable; S2, performing quality control, normalization, and batch correction on the transcriptome expression data for data preprocessing to obtain a standard expression matrix; S3, screening the standard expression matrix through differential expression analysis and weighted gene co-expression network analysis to verify the significant association between each gene in the gene marker combination and the risk of tuberculosis incidence; S4, using a machine learning algorithm, constructing a tuberculosis incidence risk prediction model with the expression level of each gene in the gene marker combination as the independent variable and the incidence outcome of active tuberculosis within 2 years as the dependent variable, and optimizing the model parameters through hierarchical K-fold cross-validation; S5, validating the tuberculosis incidence risk prediction model using an independent external cohort to evaluate the predictive efficacy of the tuberculosis incidence risk prediction model.

[0023] Specifically, this application protects a method for constructing a tuberculosis incidence risk prediction model. This method uses the aforementioned seven genes as core biomarkers and, through a complete process of sample collection, data preprocessing, biomarker validation, model construction, and multi-cohort validation, solves the technical problems of biomarker redundancy, insufficient predictive efficacy, and poor universality in existing technologies. The tuberculosis incidence risk prediction model constructed in this application achieves an internal validation AUC of over 0.958 and an external validation AUC of over 0.886. The number of biomarkers is streamlined and free of redundancy, significantly reducing detection and analysis costs. Multi-cohort validation shows no significant population heterogeneity. It can accurately quantify the incidence risk within two years in individuals with latent tuberculosis infection, filling a gap in clinically practical prediction models and providing standardized technical support for precise tuberculosis prevention and control. Compared to existing methods, it has the advantages of high efficiency, low cost, and high universality, meeting the needs of clinical translation and large-scale application.

[0024] In other embodiments, in step S1, the sample is a peripheral blood sample from a person with latent tuberculosis infection, covering people of different ages, genders, regions and HIV infection statuses, with a total number of samples of not less than 1200, and the proportion of active tuberculosis progression in the sample is not less than 18%.

[0025] Specifically, this embodiment further defines the sample collection steps. The samples are peripheral blood samples from individuals with latent tuberculosis infection, covering people of different ages, genders, regions, and HIV infection statuses, with a sample size of no less than 1200 cases. Among them, the proportion of those with active tuberculosis progression is no less than 18%. Peripheral blood sample acquisition is non-invasive and convenient, complies with clinical testing ethics and operating standards, and is suitable for large-scale population screening scenarios. Including multi-dimensional subgroups can comprehensively simulate real clinical scenarios, avoid model bias caused by a single sample, and improve the model's generalization ability. Sufficient sample size and a reasonable proportion of those with disease progression ensure the statistical power of differential analysis and model training, avoiding the randomness of results caused by small samples.

[0026] In other embodiments, step S2 involves data preprocessing for quality control, normalization, and batch correction of the transcriptome expression data. Specifically, this includes removing low-quality sequencing data from the transcriptome expression data, normalizing the expression level using the TPM method, and removing batch effects using the ComBat method.

[0027] Specifically, this embodiment further defines the data preprocessing in step S3. Data preprocessing includes removing low-quality sequencing data, normalizing expression levels using the TPM method, and removing batch effects using the ComBat method. Filtering low-quality data eliminates noise interference, ensuring the accuracy of the original data; normalizing expression levels using the TPM method eliminates differences in sequencing depth and gene length between samples, making gene expression levels comparable across different samples; removing batch effects using the ComBat method effectively removes technical biases caused by different platforms and experimental procedures, allowing for the combined analysis of data from multiple cohorts. This embodiment solves common technical error problems in transcriptome data analysis, ensuring uniform and reliable data quality, avoiding biomarker selection bias and model distortion caused by improper preprocessing, and making the expression differences of the seven core genes more realistically reflect the association with disease risk, thus improving the accuracy of subsequent analysis. Data processed by this preprocessing workflow shows higher consistency in multi-cohort integrated analysis, providing data assurance for biomarker stability validation and efficient model construction, and improving the accuracy and reproducibility of prediction results.

[0028] In other embodiments, in step S3, the standard expression matrix is ​​screened through differential expression analysis and weighted gene co-expression network analysis. The specific screening criteria are: differential expression fold ≥ 1.5, P value < 0.01, and consistent expression trend in at least 5 independent datasets.

[0029] Specifically, the standard expression matrix obtained through steps S1-S2 in this application is characterized by being noise-free, comparable, batch-bias-free, and highly stable. It can provide a high-quality, standardized, and reproducible data foundation for subsequent differential expression analysis, WGCNA screening, and machine learning model construction, directly ensuring accurate biomarker screening, efficient model training, and robust prediction results.

[0030] This embodiment uses differential expression analysis combined with weighted co-expression network analysis (WGCNA) to screen genes from the preprocessed standard expression matrix. Strict screening criteria are set: differential expression fold increase ≥ 1.5, p-value < 0.01, and consistent expression trends across at least five independent datasets. This is a key technology design for accurately identifying stable, strongly correlated, and highly reliable core biomarkers from massive transcriptome data. It fundamentally solves the problems of insufficient biomarker screening, poor stability, high redundancy, and clinical unreliability in existing technologies, providing the highest quality feature foundation for subsequent model construction.

[0031] In other embodiments, in step S4, the machine learning algorithm is one or more combinations of LightGBM, logistic regression, and random forest. The K value for stratified K-fold cross-validation is 5-10, and the samples are stratified according to their disease status.

[0032] Specifically, this embodiment selects one or more combinations of LightGBM, logistic regression, and random forest as the modeling algorithm, and employs K-fold cross-validation with a K value of 5-10 and stratified by disease status, which is a key optimization design to improve the model's predictive accuracy, generalization ability, and clinical robustness. LightGBM has high fitting efficiency for high-dimensional transcriptome data, strong ability to capture nonlinear associations, and excellent predictive performance; logistic regression can output clear risk coefficients, which are easy to interpret clinically; random forest can reduce the risk of overfitting and has stronger robustness. The flexible combination of multiple algorithms can adapt to different data distributions and analysis needs, balancing predictive performance and clinical applicability.

[0033] Hierarchical K-fold cross-validation with a K-value of 5-10 ensures validation effectiveness while maintaining reasonable computational costs, avoiding insufficient validation or resource waste. Stratifying samples by disease status ensures that the proportion of diseased to non-diseased samples in each fold matches the total set, eliminating model bias caused by uneven data distribution and preventing the tuberculosis incidence risk prediction model from favoring the majority class while ignoring minority diseased samples. This setting significantly improves the stability and reliability of the tuberculosis incidence risk prediction model on independent external data, avoiding overfitting and underfitting, and ensuring that the model maintains high AUC, high sensitivity, and high specificity in real clinical scenarios, providing algorithmic support for accurate quantification and reliable prediction of tuberculosis incidence risk.

[0034] In other embodiments, in step S5, the indicators for validating the tuberculosis incidence risk prediction model include AUC, sensitivity, and specificity. External validation AUC ≥ 0.85, sensitivity and specificity are both ≥ 75%.

[0035] Specifically, this embodiment further defines the validation indicators for the tuberculosis incidence risk prediction model. These indicators include AUC, sensitivity, and specificity. External validation requires an AUC ≥ 0.85, and both sensitivity and specificity ≥ 75%. AUC reflects the overall discriminative ability of the tuberculosis incidence risk prediction model, while sensitivity and specificity balance the risks of missed and misdiagnosed diagnoses in clinical practice. Clear numerical thresholds define the clinically applicable standards for the tuberculosis incidence risk prediction model. These limitations provide a quantitative evaluation standard for the effectiveness of the tuberculosis incidence risk prediction model, avoiding vague descriptions and ensuring that the model meets clinical testing requirements. This addresses the problem of existing models lacking clear standards for predictive effectiveness and being difficult to apply clinically.

[0036] In other embodiments, the tuberculosis incidence risk prediction model outputs the probability value of incidence within 2 years, and completes the three-level risk classification according to probability < 0.2 as low risk, 0.2 ≤ probability < 0.6 as medium risk, and probability ≥ 0.6 as high risk.

[0037] Specifically, this embodiment further limits the model output. The model outputs the probability value of disease incidence within 2 years, dividing it into three strata: low, medium, and high risk. Low risk is defined as probability < 0.2, medium risk as 0.2 ≤ probability < 0.6, and high risk as probability ≥ 0.6. Quantifying the probability value directly reflects an individual's disease risk. The three-stratification aligns with the needs of clinical stratified management, and the risk thresholds have been validated as scientifically reasonable through external cohort analysis. This limitation achieves precise stratification of disease risk, facilitating rapid clinical identification of high-risk individuals, targeted intervention for high-risk individuals, and reduced overtreatment for low-risk individuals, thus optimizing the allocation of medical resources. After stratification, the disease risk of each group shows a significant gradient difference, and the survival curves are clearly separated. This addresses the deficiency of existing technologies in accurately quantifying stratification, providing an operational standard for personalized tuberculosis prevention and control, improving prevention and control efficiency, and reducing medical costs.

[0038] This application also provides an application of the above-described construction method in the preparation of a tuberculosis incidence risk detection kit. The tuberculosis incidence risk detection kit is used to detect the expression levels of seven genes: GBP6, SEPTIN4, FCGR1A, SERPING1, ETV7, BATF2, and ANKRD22.

[0039] Specifically, this embodiment protects the application of the above-described method in the preparation of a tuberculosis incidence risk detection kit. The kit contains reagents for detecting the expression levels of seven genes: GBP6, SEPTIN4, FCGR1A, SERPING1, ETV7, BATF2, and ANKRD22. Based on a simplified gene combination design, the kit can be used with technologies such as qPCR and transcriptome sequencing. It is highly standardized, easy to operate, and suitable for large-scale production and clinical application. This application transforms the tuberculosis incidence risk prediction model into a tangible clinical product, addressing the lack of accurate prediction capabilities in existing testing products. It provides a standardized testing tool for clinical use, facilitates risk screening in primary healthcare institutions, promotes the widespread adoption of precision tuberculosis control technologies, and possesses both economic and social value, significantly improving the early prevention and control of tuberculosis.

[0040] Example 1: Screening and Validation of Gene Markers Data source: Transcriptome datasets of the above 7 genes from 6 independent studies obtained from public databases such as GEO and ArrayExpress, including a total of 1230 samples of latent tuberculosis infected individuals, of which 229 were those who progressed to active tuberculosis within 2 years, and 1001 were those who did not develop the disease.

[0041] Data analysis workflow: All transcriptome datasets underwent uniform quality control, expression level normalization, and batch correction; differential expression analysis was performed using R software to screen for genes with significant expression differences between patients with and without tuberculosis progression; weighted gene co-expression network analysis was also performed to identify gene modules highly associated with the risk of tuberculosis incidence.

[0042] Screening results: Through multidimensional analysis and multi-cohort validation, GBP6, SEPTIN4, FCGR1A, SERPING1, ETV7, BATF2, and ANKRD22 were finally identified as the core gene biomarker combination. The expression of these 7 genes differed significantly between the progressive and non-progressive groups, and showed a consistent expression trend in multiple independent datasets, with good stability and reproducibility.

[0043] This embodiment uses a large sample size and multiple independent cohorts to screen biomarkers. Through standardized preprocessing and combined analysis using two methods, technical noise and random interference are eliminated at the source, ensuring the reliability of the biomarker screening results. The seven genes obtained possess both statistical significance and biological relevance, are concise in number, and have high correlation, effectively overcoming the redundancy and poor stability of existing biomarkers. This provides a precise and robust core feature foundation for the subsequent construction of a tuberculosis incidence risk prediction model.

[0044] Example 2: Construction and Validation of a Tuberculosis Incidence Risk Prediction Model Model construction: Using the expression levels of the 7 genes screened in Example 1 as independent variables and whether the disease progresses to active tuberculosis within 2 years as the dependent variable, a tuberculosis incidence risk prediction model was constructed using the LightGBM algorithm; the model parameters were optimized through 10-fold stratified K-fold cross-validation to determine the optimal model structure.

[0045] Internal validation: In a total sample of 1230 cases, the tuberculosis incidence risk prediction model had an AUC of 0.958, a sensitivity of 89.1%, and a specificity of 87.6%, demonstrating excellent internal predictive power.

[0046] External validation: Two independent external cohorts of latent tuberculosis-infected individuals were selected for validation, with a total of 420 samples included, of which 68 were progressive and 352 were non-progressive. The validation results showed that the AUC of the tuberculosis incidence risk prediction model was 0.886, the sensitivity was 78.2%, and the specificity was 82.1%. The predictive effect remained stable in different age, gender, and HIV infection status subgroups.

[0047] This embodiment employs a high-performance machine learning algorithm combined with hierarchical cross-validation to construct a tuberculosis incidence risk prediction model, effectively avoiding overfitting and improving the generalization ability of the model. Both internal and external validation results have reached clinical applicability levels, demonstrating that the tuberculosis incidence risk prediction model can accurately distinguish between high- and low-risk individuals and possesses stable predictive capabilities across multiple populations and scenarios. This addresses the key issues of insufficient efficiency, poor universality, and inability to be clinically translated into practical applications found in existing models.

[0048] Example 3: Risk Stratification Application Based on Predictive Models Risk classification criteria: Based on external validation results, the incidence risk probability output by the tuberculosis incidence risk prediction model is divided into three intervals: low risk group (probability < 0.2), medium risk group (0.2 ≤ probability < 0.6), and high risk group (probability ≥ 0.6).

[0049] Stratification results: The 420 external validation cohort samples were stratified as follows: low-risk group (n=285, 12 cases progressed, 4.2% progression rate); intermediate-risk group (n=103, 27 cases progressed, 26.2% progression rate); and high-risk group (n=32, 29 cases progressed, 90.6% progression rate). The risk of disease showed a significant gradient among the three groups, and Kaplan-Meier survival curves showed a clear separation in event rates among the groups.

[0050] This embodiment achieves precise stratification of the risk of developing tuberculosis in latently infected individuals by quantifying probability and classifying risk into three tiers. The high-risk group has an extremely high incidence rate and can be considered as a key intervention target; the low-risk group has an extremely low incidence rate, reducing the need for overtreatment. This stratification standard is scientific, intuitive, and operable, directly guiding personalized clinical intervention and optimal resource allocation, significantly improving the efficiency of precise tuberculosis prevention and control.

[0051] Comparative Example The existing tuberculosis incidence risk prediction biomarker combination RISK6 and CORsignature were selected as controls and parallel tests were conducted in the same external validation cohort as in Examples 2 and 3. The results showed that the above-mentioned 7-gene biomarker combination provided in this application had a higher AUC and better predictive efficacy in external validation; fewer genes were required, resulting in lower detection and analysis costs; and stronger predictive stability and lower heterogeneity were observed in different age groups, regions, and HIV infection status subgroups, making it more suitable for clinical multi-scenario applications and large-scale promotion.

[0052] The above are merely preferred embodiments of this application and do not limit the patent scope of this application. Any equivalent structural or procedural transformations made using the content of this application's specification and drawings, or direct or indirect applications in other related technical fields, are similarly included within the patent protection scope of this application.

Claims

1. A combination of gene markers, characterized in that, The genetic marker combination includes the GBP6, SEPTIN4, FCGR1A, SERPING1, ETV7, BATF2, and ANKRD22 genes.

2. A method for constructing a tuberculosis incidence risk prediction model, characterized in that, Includes the following steps: (1) Collect multiple independent cohorts of latent tuberculosis infected individuals, obtain transcriptome expression data of the gene marker combination as described in claim 1 in the samples, and record whether the latent tuberculosis infected individuals progress to active tuberculosis within 2 years as the outcome variable; (2) Perform quality control, normalization and batch correction on the transcriptome expression data to obtain a standard expression matrix; (3) Screen the standard expression matrix through differential expression analysis and weighted gene co-expression network analysis to verify the significant association between each gene in the gene marker combination and the risk of tuberculosis incidence; (4) Use machine learning algorithms to construct a tuberculosis incidence risk prediction model with the expression level of each gene in the gene marker combination as the independent variable and the incidence outcome of active tuberculosis within 2 years as the dependent variable, and optimize the model parameters through hierarchical K-fold cross-validation; (5) Validate the tuberculosis incidence risk prediction model using an independent external cohort to evaluate the predictive efficacy of the tuberculosis incidence risk prediction model.

3. The construction method according to claim 2, characterized in that, In step (1), the sample is a peripheral blood sample from a person with latent tuberculosis infection, covering people of different ages, genders, regions and HIV infection statuses. The total number of samples is not less than 1200, and the proportion of active tuberculosis progression in the sample is not less than 18%.

4. The construction method according to claim 2, characterized in that, In step (2), the data preprocessing for quality control, normalization and batch correction of the transcriptome expression data specifically includes: removing low-quality sequencing data from the transcriptome expression data, normalizing the expression level using the TPM method, and removing batch effects using the ComBat method.

5. The construction method according to claim 2, characterized in that, In step (3), the standard expression matrix is ​​screened through differential expression analysis and weighted gene co-expression network analysis. The specific screening criteria are: differential expression fold ≥ 1.5, P value < 0.01, and consistent expression trend in at least 5 independent datasets.

6. The construction method according to claim 2, characterized in that, In step (4), the machine learning algorithm is one or more combinations of LightGBM, logistic regression, and random forest; The K-value for the stratified K-fold cross-validation is 5-10, and the samples are stratified according to their disease status.

7. The construction method according to claim 2, characterized in that, In step (5), the indicators for validating the tuberculosis incidence risk prediction model include AUC, sensitivity, and specificity. The external validation AUC is ≥0.85, and both sensitivity and specificity are ≥75%.

8. The construction method according to any one of claims 2-8, characterized in that, The tuberculosis incidence risk prediction model outputs the probability value of incidence within 2 years, and completes the three-level risk classification according to probability < 0.2 as low risk, 0.2 ≤ probability < 0.6 as medium risk, and probability ≥ 0.6 as high risk.

9. The application of the construction method as described in any one of claims 2-8 in the preparation of a tuberculosis incidence risk detection kit.

10. The application according to claim 9, characterized in that, The tuberculosis incidence risk detection kit is used to detect the expression levels of GBP6, SEPTIN4, FCGR1A, SERPING1, ETV7, BATF2, and ANKRD22 genes.