Biomarker for early risk early warning of latent tuberculosis infection, early risk early warning model and application thereof

By detecting the expression levels of SCO2, MT1G, CREB5, PARP9, ATF3, MUC1 and MGST1 genes, and combining statistical analysis to establish an early risk warning model, solving the problem of inability to distinguish between LTBI and ATB in the existing technology, and achieving early risk warning and personalized treatment of latent tuberculosis infection.

CN120400330APending Publication Date: 2025-08-01中国人民解放军总医院第八医学中心
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510563085.2
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-04-30
Publication Date
2025-08-01

AI Technical Summary

Technical Problem

The prior art cannot effectively distinguish between latent tuberculosis infection (LTBI) and active tuberculosis (ATB), and lacks fast and efficient early-stage risk warning methods, making it impossible to predict whether LTBI patients will progress to ATB.

Method used

The SCO2, MT1G, CREB5, PARP9, ATF3, MUC1 and MGST1 genes were used as biomarkers, and their expression levels were detected by RT-PCR and RT-qPCR. The early risk warning model was established in combination with binary logistic regression analysis, and the early risk warning model for latent tuberculosis infection was constructed using statistical analysis methods.

Benefits of technology

Accurate, specific and sensitive early risk warnings for latent tuberculosis infections are achieved, personalized treatment suggestions are provided, and the incidence and mortality of active tuberculosis are reduced, and the accuracy and efficiency of diagnosis are improved.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure SMS_1
    Figure SMS_1
  • Figure SMS_2
    Figure SMS_2
  • Figure SMS_3
    Figure SMS_3
Patent Text Reader

Abstract

The invention discloses a biomarker for early risk early warning of latent tuberculosis infection, an early risk early warning model and application thereof. Specifically, the invention discloses a biomarker (an SCO2 gene, an MT1G gene, a CREB5 gene, a PARP9 gene, an ATF3 gene, an MUC1 gene and / or an MGST1 gene) and / or an application of a substance for detecting the biomarker in preparation of a product for early risk warning of latent tuberculosis infection or an application in construction of an early risk warning model of latent tuberculosis infection. The biomarker and the model constructed by using the biomarker are accurate and reliable through ROC curve verification, and are suitable for wide popularization. By performing early risk early warning on latent tuberculosis infection on the to-be-detected subject, decision support can be provided for clinicians, personalized or further prevention or treatment suggestions can be provided for patients, or early dry treatment and the like can be performed on high-risk patients, so that optimal disease management is realized.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the fields of bioinformatics and biomedicine, and specifically to biomarkers for early risk warning of latent tuberculosis infection, an early risk warning model, and their applications. Background Art

[0002] Tuberculosis (TB) is mainly caused by infection with Mycobacterium tuberculosis (MTB). According to statistics, approximately one-fourth of the global population is infected with MTB, and 5%-15% of these infected individuals may develop active tuberculosis (ATB). Therefore, early identification and treatment of latent tuberculosis infection (LTBI) are particularly important.

[0003] Latent tuberculosis infection refers to a state in which the body produces a continuous immune response to MTB antigen stimulation but shows no obvious clinical symptoms or imaging evidence. Currently, there is no gold standard for the diagnosis of LTBI, and it mainly relies on the tuberculin skin test (TST) and interferon-gamma release assay (IGRA). However, these methods all have certain limitations. Although a variety of tuberculosis diagnostic kits have been launched globally in recent years, these methods still have limitations and cannot completely distinguish ATB from LTBI, nor can they predict whether LTBI patients will progress to ATB. Therefore, finding biomarkers that can effectively identify LTBI and ATB is of great significance for the early detection and preventive treatment of tuberculosis.

[0004] Therefore, early warning and diagnosis of latent tuberculosis infection and early intervention treatment are important strategies for reducing the incidence and mortality of active tuberculosis. Currently, there is still a lack of rapid and efficient early risk warning methods for latent tuberculosis infection, and the booming development of artificial intelligence technology and multi-omics technology provides new means to solve this problem. Summary of the Invention

[0005] The object of the present invention is to provide an accurate, specific, and sensitive biomarker for early risk warning of latent tuberculosis infection, an early risk warning model, and their applications. The technical problems to be solved are not limited to the described technical topics, and those skilled in the art can clearly understand other technical topics not mentioned herein through the following description.

[0006] To achieve the above objectives, the present invention first provides the use of a biomarker and / or a substance for detecting the biomarker in any of the following:

[0007] A1) Use in the preparation of a product for early warning of the risk of latent tuberculosis infection;

[0008] A2) Application in building an early warning model for the risk of latent tuberculosis infection;

[0009] The biomarker may be SCO2 gene, MT1G gene, CREB5 gene, PARP9 gene, ATF3 gene, MUC1 gene and / or MGST1 gene.

[0010] In the above application, the substance may include a reagent and / or an instrument for detecting the expression level of SCO2 gene, MT1G gene, CREB5 gene, PARP9 gene, ATF3 gene, MUC1 gene and / or MGST1 gene.

[0011] Furthermore, the substance may include reagents and / or instruments for detecting the expression levels of SCO2 gene, MT1G gene, CREB5 gene, PARP9 gene, ATF3 gene, MUC1 gene and / or MGST1 gene through reverse transcription-polymerase chain reaction (RT-PCR), reverse transcription real-time quantitative polymerase chain reaction (RT-qPCR), and transcriptome sequencing technology (RNA-seq).

[0012] In the above application, the reagents may include primers that specifically amplify the SCO2 gene, MT1G gene, CREB5 gene, PARP9 gene, ATF3 gene, MUC1 gene and / or MGST1 gene and / or probes that specifically recognize the SCO2 gene, MT1G gene, CREB5 gene, PARP9 gene, ATF3 gene, MUC1 gene and / or MGST1 gene.

[0013] Furthermore, the primers include primer pair 1 for detecting or specifically amplifying the SCO2 gene, primer pair 2 for detecting or specifically amplifying the MT1G gene, primer pair 3 for detecting or specifically amplifying the CREB5 gene, primer pair 4 for detecting or specifically amplifying the PARP9 gene, primer pair 5 for detecting or specifically amplifying the ATF3 gene, primer pair 6 for detecting or specifically amplifying the MUC1 gene and / or primer pair 7 for detecting or specifically amplifying the MGST1 gene.

[0014] Furthermore, the primer pair 1 may consist of a forward primer having a nucleotide sequence as shown in SEQ ID NO: 1 and a reverse primer having a nucleotide sequence as shown in SEQ ID NO: 2;

[0015] The primer pair 2 may consist of a forward primer with a nucleotide sequence as shown in SEQ ID NO:3 and a reverse primer with a nucleotide sequence as shown in SEQ ID NO:4;

[0016] The primer pair 3 may consist of a forward primer with a nucleotide sequence as shown in SEQ ID NO:5 and a reverse primer with a nucleotide sequence as shown in SEQ ID NO:6;

[0017] The primer pair 4 may consist of a forward primer with a nucleotide sequence as shown in SEQ ID NO:7 and a reverse primer with a nucleotide sequence as shown in SEQ ID NO:8;

[0018] The primer pair 5 may consist of a forward primer with a nucleotide sequence as shown in SEQ ID NO:9 and a reverse primer with a nucleotide sequence as shown in SEQ ID NO:10;

[0019] The primer pair 6 may consist of a forward primer with a nucleotide sequence as shown in SEQ ID NO:11 and a reverse primer with a nucleotide sequence as shown in SEQ ID NO:12;

[0020] The primer pair 7 may consist of a forward primer with a nucleotide sequence as shown in SEQ ID NO:13 and a reverse primer with a nucleotide sequence as shown in SEQ ID NO:14.

[0021] In this article, the product may include, but is not limited to, reagents, kits, detection chips (such as gene chips), test strips, test cards, immunosensors, and devices.

[0022] Furthermore, the application may include using the data of the expression levels of the SCO2 gene, MT1G gene, CREB5 gene, PARP9 gene, ATF3 gene, MUC1 gene, and / or MGST1 gene in the samples of known latent tuberculosis infection patients and active tuberculosis patients as training samples, and adopting statistical analysis methods (such as binary logistic regression analysis method) to establish an early risk warning model for latent tuberculosis infection, and conducting early risk warning of latent tuberculosis infection on the subjects according to this model.

[0023] The sample described in this article may be a blood sample (including whole blood, plasma, and serum).

[0024] Furthermore, the sample may be whole blood.

[0025] The present invention also provides a composition or kit for early risk warning of latent tuberculosis infection, and the composition or kit may include reagents for detecting the expression levels of the SCO2 gene, MT1G gene, CREB5 gene, PARP9 gene, ATF3 gene, MUC1 gene, and / or MGST1 gene.

[0026] Furthermore, the reagent can be a primer for specifically amplifying the SCO2 gene, MT1G gene, CREB5 gene, PARP9 gene, ATF3 gene, MUC1 gene, and / or MGST1 gene described herein (such as primer pair 1, primer pair 2, primer pair 3, primer pair 4, primer pair 5, primer pair 6, and / or primer pair 7 described herein) and / or a probe for specifically recognizing the SCO2 gene, MT1G gene, CREB5 gene, PARP9 gene, ATF3 gene, MUC1 gene, and / or MGST1 gene.

[0027] Furthermore, the composition or kit may further include reagents required for PCR detection, such as DNA polymerases (such as Taq DNA polymerase, Tth DNA polymerase, Vent DNA polymerase, and Pfu DNA polymerase, etc.), dNTPs, Mg 2+ solution (such as MgSO4 or MgCl2 solution) and PCR buffer (such as Tris-HCl).

[0028] Furthermore, the composition or kit may further include reagents required for reverse transcription reaction, transcriptome sequencing, Northern blot, in situ hybridization, gene chip detection, Nanopore sequencing, and / or PacBio sequencing, etc.

[0029] Furthermore, the composition or kit may further include total RNA extraction reagents.

[0030] The kit described herein can be a detection kit based on reverse transcription quantitative real-time polymerase chain reaction (RT-qPCR).

[0031] The various reagent components of the kit can be present in separate containers, or can be pre-combined into a reagent mixture in whole or in part.

[0032] The components of the kit can be provided in the form of a solution, such as in the form of an aqueous solution. In the case of being in an aqueous solution state, the concentration or content of these components can be conveniently determined by those skilled in the art according to different needs. For example, for storage purposes, the concentration of the components can exist in a higher form, and when in a working state or in use, the concentration can be reduced to the working concentration by diluting the above-mentioned higher concentration solution.

[0033] The present invention also provides an early risk warning model for latent tuberculosis infection, and the model can be constructed using the biomarkers described herein.

[0034] Furthermore, the formula of the model can be:

[0035] P = 1 / [1 + e-(31.875-0.795*MT1G+0.193*SCO2-1.102*MUC1-0.775*PARP9-0.624*CREB5-1.157*MGST1-0.284*ATF3)

[0036] Among them, P represents the probability value of latent tuberculosis infection; MT1G, SCO2, MUC1, PARP9, CREB5, MGST1, and ATF3 respectively represent the expression levels of the MT1G, SCO2, MUC1, PARP9, CREB5, MGST1, and ATF3 genes; e represents the base of the natural logarithm.

[0037] The present invention also provides a method for constructing an early risk warning model for latent tuberculosis infection. The method includes using the data of the expression levels of the SCO2 gene, MT1G gene, CREB5 gene, PARP9 gene, ATF3 gene, MUC1 gene, and / or MGST1 gene in the samples of known latent tuberculosis infection patients and active tuberculosis patients as training samples, and adopting statistical analysis methods to establish an early risk warning model for latent tuberculosis infection.

[0038] In the above method, the statistical analysis method may include, but is not limited to, binary logistic regression analysis, support vector machine analysis, and LASSO regression analysis methods.

[0039] Further, the statistical analysis method may be a binary logistic regression analysis method.

[0040] The application method of any of the early risk warning models for latent tuberculosis infection in the present invention may include the following steps:

[0041] (1) Detect the data of the expression levels of the SCO2 gene, MT1G gene, CREB5 gene, PARP9 gene, ATF3 gene, MUC1 gene, and / or MGST1 gene in the sample of the subject to be tested;

[0042] (2) Input the data into any of the early risk warning models for latent tuberculosis infection in the present invention;

[0043] (3) Obtain the early risk warning result of latent tuberculosis infection of the subject to be tested through the model.

[0044] Further, the formula of the model in step (2) may be:

[0045] P = 1 / [1 + e -(31.875-0.795*MT1G+0.193*SCO2-1.102*MUC1-0.775*PARP9-0.624*CREB5-1.157*MGST1-0.284*ATF3)

[0046] Among them, P represents the probability value of latent tuberculosis infection; MT1G, SCO2, MUC1, PARP9, CREB5, MGST1, and ATF3 respectively represent the expression levels of the MT1G, SCO2, MUC1, PARP9, CREB5, MGST1, and ATF3 genes; e represents the base of the natural logarithm. ​​

[0047] Furthermore, based on the early risk warning results of latent tuberculosis infection, personalized or further prevention or treatment recommendations can be provided to the tested subjects, or early intervention treatment can be carried out for high-risk patients.

[0048] The present invention also provides a system for early warning of the risk of latent tuberculosis infection, which may include:

[0049] A data receiving module is used to: receive data on the expression level of the biomarker in a sample of a test subject from at least one terminal;

[0050] A data processing module is used to: input the data into the model described herein, or into the model constructed by the method described herein, to obtain the probability value of latent tuberculosis infection;

[0051] A data output module is used to output the probability value to at least one client.

[0052] The present invention also provides a method for early warning of the risk of latent tuberculosis infection, which may include the following steps:

[0053] S1. Data reception: receiving data on the expression level of the biomarker in a sample of a test subject from at least one terminal;

[0054] S2. Data processing: inputting the data into the model described herein, or into the model constructed by the method described herein, to obtain the probability value of latent tuberculosis infection;

[0055] S3. Data output: output the probability value to at least one client.

[0056] The sequences of the SCO2, MT1G, CREB5, PARP9, ATF3, MUC1, and MGST1 genes described herein are all known in the art and are publicly available from the NCBI website. Exemplary sequences of these genes are shown below:

[0057] The nucleotide sequence of the coding sequence (CDS) of the SCO2 gene described herein may be positions 211-1011 of GenBank Accession No. NM_001169109.2 (Update Date 13-FEB-2025).

[0058] The nucleotide sequence of the coding sequence (CDS) of the MT1G gene described herein may be positions 73-261 of GenBank Accession No. NM_001301267.2 (Update Date 01-OCT-2024).

[0059] The nucleotide sequence of the coding sequence (CDS) of the CREB5 gene described in this article may be positions 19 - 1128 of GenBank Accession No. NM_001011666.3 (Update Date 23 - JUL - 2024).

[0060] The nucleotide sequence of the coding sequence (CDS) of the PARP9 gene described in this article may be positions 146 - 2710 of GenBank Accession No. NM_001146102.2 (Update Date 06 - APR - 2024).

[0061] The nucleotide sequence of the coding sequence (CDS) of the ATF3 gene described in this article may be positions 82 - 627 of GenBank Accession No. NM_001030287.4 (Update Date 03 - SEP - 2024).

[0062] The nucleotide sequence of the coding sequence (CDS) of the MUC1 gene described in this article may be positions 67 - 861 of GenBank Accession No. NM_001018016.3 (Update Date 12 - FEB - 2025).

[0063] The nucleotide sequence of the coding sequence (CDS) of the MGST1 gene described in this article may be positions 82 - 549 of GenBank Accession No. NM_001260511.2 (Update Date 07 - APR - 2024).

[0064] The expression levels of the biomarkers (SCO2 gene, MT1G gene, CREB5 gene, PARP9 gene, ATF3 gene, MUC1 gene, and / or MGST1 gene) described in this article may be applicable to conventional reverse transcription quantitative real - time polymerase chain reaction (RT - qPCR) detection experiments. Although the examples provided in the present invention exemplarily use the RT - qPCR method to detect the expression levels of biomarkers in whole - blood samples, the present invention is not limited to this specific detection method. Those skilled in the art can use any other suitable well - known detection methods (such as RNA - seq, Northern blot, in situ hybridization, gene chip detection, Nanopore sequencing, and PacBio sequencing, etc.) to detect the expression levels of biomarkers in whole - blood samples. These alternative methods do not deviate from the scope of the present invention, and the present invention should include these alternative methods.

[0065] The present invention first proposes to use the SCO2 gene, MT1G gene, CREB5 gene, PARP9 gene, ATF3 gene, MUC1 gene and / or MGST1 gene as biomarkers for early risk warning of latent tuberculosis infection and the construction of an early risk warning model for latent tuberculosis infection. By deeply mining the transcriptome data in the GEO database, the present invention has identified a total of 925 differentially expressed genes related to LTBI. Through GO and KEGG enrichment analysis, combined with artificial intelligence technologies such as weighted gene co-expression network analysis (WGCNA) and two machine learning algorithms, and integrating bioinformatics methods, copper death / ferroptosis biomarkers related to latent tuberculosis infection are identified. Finally, SCO2, MT1G, CREB5, PARP9, ATF3, MUC1 and MGST1 are selected as copper death / ferroptosis-related biomarkers for latent tuberculosis infection patients. Based on these seven genes, an early risk warning model is constructed. And the model is verified through the GEO dataset, prospective cohort study and RT-qPCR. Using GSE28623 as an external dataset for verification, the results show that AUC = 0.930 (CI: 0.865 - 0.981). The expression levels of these 7 genes are verified through population cohort study and RT-qPCR, and the results show that they are consistent with the trend in the training set. The verification results of clinical samples show that the AUC of the LTBI early risk warning model constructed by the present invention is 0.778. When the Youden index is 0.43243, the sensitivity of the model is 0.81081, the specificity is 0.62162, and the accuracy is 0.71622.

[0066] In addition, considering the key role of LTBI-related differential genes in the occurrence and development of LTBI, we have not only identified biomarkers related to the disease, but also discovered important immune signaling pathways through KEGG enrichment analysis. For example, cytokine-cytokine receptor interaction, complement and coagulation cascades, and NOD-like receptor signaling pathways are mainly enriched and play key roles in the pathogenic mechanism of latent tuberculosis infection. These new findings not only provide potential biological targets for the early detection and screening of LTBI patients, but also offer new ideas and directions for the treatment and prevention of related diseases. The present invention has developed 7 reliable copper death / ferroptosis-related LTBI biomarkers, and the constructed risk assessment model shows significant prediction efficiency, providing an early screening strategy for LTBI.

[0067] The biomarker of the present invention and the model constructed using this biomarker have been verified to be accurate and reliable through the ROC curve, and are suitable for wide promotion. By providing early risk warnings for latent tuberculosis infection in subjects to be tested, it can provide decision-making support for clinicians, guide the selection and adjustment of treatment plans, provide personalized or further preventive or treatment suggestions for patients, or conduct early intervention treatment for high-risk patients, etc., to achieve optimal disease management, and has broad clinical application value. Description of the Drawings

[0068] Figure 1 It is a flow chart for the construction and verification of an early risk warning model for latent tuberculosis infection (LTBI) patients.

[0069] Figure 2 It is differential analysis. A. Grouping and sample size of the GSE37250 dataset. B. Volcano plot showing 925 differentially expressed genes between the LTBI group and the ATB group. Up-regulated genes are shown in red, and down-regulated genes are shown in blue. C. Heat map of differentially expressed genes in GSE37250. There are two vertical dotted lines in the figure, indicating log2 FC at -1 and 1. Genes up-regulated are represented by red and genes down-regulated are represented by blue on both sides of the vertical dotted lines, and the names of the top five up-regulated and down-regulated genes are marked; the horizontal dotted line represents an adjusted p-value of 0.05.

[0070] Figure 3 It is GO and KEGG enrichment analysis of differentially expressed genes. A. GO enrichment results of GSE37250. Orange represents the enrichment results of GO biological processes; green represents the enrichment results of GO cellular components; blue represents the enrichment results of GO molecular functions. An adjusted p < 0.05 was determined as a significant change in GO. The x-axis represents the adjusted P-value and (or) the number of gene enrichments for each term. The y-axis represents the entries of GO-BP, GO-CC, and GO-MF. B. KEGG pathway enrichment results. An adjusted p < 0.05 was determined as a significant change in KEGG. The x-axis represents the adjusted P-value and (or) the number of gene enrichments for each term. The entries of KEGG in the figure are further classified into five categories. The length of the bar chart represents the number of genes related to each entry.

[0071] Figure 4 Weighted gene co-expression network analysis (WGCNA) identifies key LTBI modules in GSE37250. A. Clustering tree of genes, with different colors representing different modules. B. WGCNA identifies key LTBI modules in GSE37250.

[0072] Figure 5Screening LTBI biomarkers for machine learning. A. Venn diagram of DEGs, genes in the green module of WGCNA, cuproptosis-related genes, and Ferroptosis-related genes. B. Screening of feature genes using the Support Vector Recursive Feature Elimination (SVM-RFE) algorithm. C. Identification of feature genes using the Least Absolute Shrinkage and Selection Operator (LASSO) logistic regression algorithm. The regularization parameter λ is used to select covariates. D. Venn diagram showing overlapping key feature genes screened by LASSO and SVM-RFE.

[0073] Figure 6 Verification of the expression levels and evaluation of the diagnostic efficacy of the core genes of the HeptaTB Dx Model between LTBI and ATB populations. A-G. Expression of 7 core genes in GSE28623. H. ROC curve of 7 core genes in the training set GSE37250. I. ROC curve of the core genes of the HeptaTB Dx Model in the validation set GSE28623.

[0074] Figure 7 Identification of cuproposis / ferroptosis-related molecular subtypes in LTBI. A. CDF curve showing consistent distribution from k = 2 to k = 5. B. Area fraction under the CDF curve for k = 2-9. The horizontal axis represents the number of classes (k), while the vertical axis represents the relative change in the area under the CDF curve. C-E. Consensus clustering matrices were generated for k values from 2 to 4. F. Principal component analysis (PCA) showing the distribution of two clusters. G. Expression of 7 core genes in the two groups of ATB and LTBI. H. Expression of 7 core genes in cuproposis / ferroptosis-related molecular subtypes.

[0075] Figure 8 Nomogram of 7 core genes. A. Nomogram of 7 genes. B. Calibration curve. C. Decision curve.

[0076] Figure 9 Expression analysis of biomarkers and risk warning models for LTBI in a prospective cohort (sequencing data) study. A-G. mRNA expression analysis of 7 DE-FRG / DE-CRG in the prospective cohort.

[0077] Figure 10 Verification results of biomarkers and risk warning models for LTBI in RT-qPCR. A-G. mRNA expression analysis of 7 DE-FRG / DE-CRG in the RT-qPCR cohort.

[0078] Figure 11 ROC curve analysis of the early risk warning model for LTBI. Detailed implementation manners

[0079] The present invention will be further described in detail below in conjunction with the specific implementation manners. The provided embodiments are only for clarifying the present invention, rather than limiting the scope of the present invention. The following provided embodiments can be used as a guide for those of ordinary skill in the art to make further improvements, and do not constitute any limitation to the present invention in any way.

[0080] The experimental methods in the following embodiments are all conventional methods unless otherwise specified, and are carried out according to the techniques or conditions described in the literature in this field or according to the product instructions. The materials, reagents, etc. used in the following embodiments can be obtained from commercial channels unless otherwise specified.

[0081] The following embodiments use GraphPad Prism 10.1.2 software (San Diego, CA, USA) statistical software to process the data. According to the characteristics of the data distribution, different statistical description methods are adopted: for continuous variables that conform to the normal distribution, they are expressed as mean ± standard deviation (Mean ± SD); for continuous variables that do not conform to the normal distribution, they are expressed as quartiles [50% (25% - 75%)]; the count data are directly presented as numerical values. In the comparison of two groups of data, if the data both conform to the normal distribution, an unpaired t-test is used for analysis; if they do not conform to the normal distribution, a Mann-Whitney U test is used. For the experimental data of multiple groups, if the data of each group meet the normality and homogeneity of variance, a one-way ANOVA is used for statistics; if any group of data does not meet the normality and / or homogeneity of variance, a Kruskal-Wallis test is used. In the present invention, P < 0.05 is used as the threshold for statistical significance of differences, and the following significance marks are used: *P < 0.05; **P < 0.01; ***P < 0.001; ****P < 0.0001. In the following embodiments, quantitative experiments are set with three biological replicate experiments and the results are averaged unless otherwise specified.

[0082] The acquisition of clinical samples in the following embodiments all obtained the informed consent of the patients and was approved by the Ethics Committee of the Eighth Medical Center of the Chinese People's Liberation Army General Hospital (Approval No.: 30920230825701232).

[0083] Example 1, Data processing and identification of DEGs in LTBI patients

[0084] 1. Data source and preprocessing

[0085] First, we designed the flow chart of this study ( Figure 1) Then, we searched the GEO database (https: / / www.ncbi.nlm.nih.gov / geo / ) using the keywords "latent tuberculosis infection" and "tuberculosis". On this basis, we carefully screened and selected based on the following inclusion criteria. Inclusion criteria: (1) Samples were from ATB patients or LTBI individuals; (2) The sample size of each group was ≥6; (3) Each sample had complete gene expression data; (4) The subjects were >15 years old; (5) The subjects had not received tuberculosis-related treatment; (6) The samples were blood samples not affected by HIV infection, autoimmune diseases, or malignancies; (7) Samples used for gene expression profiling must be RNA expressed at the whole human genome level. Finally, we identified two eligible datasets, namely GSE37250 and GSE28623. The former will be used as the training set, and the latter will be used as the validation set (Table 1).

[0086] Table 1. Summary of the included GEO data

[0087]

[0088] Abbreviations: ATB, active tuberculosis; LTBI, latent tuberculosis infection; GEO, Gene Expression Omnibus

[0089] 2. Differential analysis

[0090] Based on the GEO datasets that met the screening criteria, we used the interactive web tool GEO2R to preliminarily analyze the microarray gene expression data to obtain basic information on the gene expression profile. Subsequently, we used RStudio (version 2023.12.0) for statistical analysis and used the R package limma to screen for DEGs in the ATB and LTBI samples in the training set. The screening criteria were set as: adjusted p-value (adj.p-value) ≤ 0.05 and |log2 fold change (log2FC)| ≥ 1. Next, we used the R package ClusterProfiler to perform gene ontology (GO) and Kyoto Encyclopedia of Genes and Genomes (KEGG) functional enrichment analysis on the DEGs between ATB and LTBI. The significance criterion for the enrichment results was adj.p-value ≤ 0.05. Finally, we used the R package ggplot2 to perform visual analysis of the differentially expressed genes, drew a volcano plot to show the distribution of the differential genes, and showed the gene expression pattern through a heatmap.

[0091] 3. Results

[0092] Two datasets, GSE37250 and GSE28623, were screened from the GEO database by keyword search and based on strict inclusion and exclusion criteria. Among them, GSE37250 was used as the training set, and GSE28623 was used as an independent validation set. The R package "limma" was used to analyze the differentially expressed genes (DEGs) between the LTBI and ATB groups in the training set (GSE37250) ( Figure 2 in A). The analysis results showed that a total of 925 DEGs were identified in the GSE37250 dataset, among which 405 genes were significantly up-regulated and 520 genes were significantly down-regulated. To more intuitively display the distribution and expression patterns of differentially expressed genes, we plotted a volcano plot ( Figure 2 in B) and a heatmap ( Figure 2 in C). The volcano plot shows the relationship between the significance (-log10p value) and the fold change in expression (log2FC) of DEGs, while the heatmap intuitively presents the expression patterns of DEGs in LTBI and ATB samples.

[0093] Example 2. GO and KEGG enrichment analysis of differentially expressed genes

[0094] We used the R package ClusterProfiler to perform gene ontology (GO) and Kyoto Encyclopedia of Genes and Genomes (KEGG) functional enrichment analysis on the DEGs between ATB and LTBI. The significance criterion for the enrichment results was adj.p-value ≤ 0.05.

[0095] The GO results showed ( Figure 3 in A) that in terms of biological process (BP), the DEGs were mainly enriched in terms such as bacterial defense response, activation of T cells, regulation of type II interferon production, regulation of natural killer cell activation, and classical pathways of immune response regulatory signaling pathways; in terms of cellular component (CC), the DEGs were mainly enriched in components such as specific granules, granule secretory lumen, cytoplasmic vesicle lumen, and specific granule lumen; in terms of molecular function, the enriched functions were cytokine activity, protein binding, and cytokine binding. KEGG analysis found that the DEGs were mainly enriched in pathways such as complement and coagulation cascades, cytokine receptor interactions, NOD-like receptor signaling pathways, and efferocytosis ( Figure 3 in B).

[0096] Example 3. Weighted co-expression network analysis

[0097] After completing the DEGs analysis, to further screen for key genes related to LTBI, we performed Weighted Gene Co-expression Network Analysis (WGCNA). WGCNA is a well-established method for biological network analysis and is widely used in the identification of disease-related gene modules. We used the R package WGCNA to analyze the training dataset to identify key modules related to LTBI. First, a scale-free network was constructed by evaluating different β parameters, and the pickSoftThreshold function was used to select the optimal soft threshold (β) to ensure that the network conformed to the scale-free distribution characteristics. An adjacency matrix was constructed based on gene expression data, and hierarchical clustering was performed according to the Topological Overlap Matrix (TOM) to identify co-expression modules. Then, the module eigengene (ME) of each module was calculated, and similar modules were merged based on the ME results to generate a hierarchical clustering dendrogram. To identify functional modules related to LTBI, we calculated the correlation between the module eigengene and the LTBI clinical phenotype. Modules with high correlation coefficients were considered candidate modules related to LTBI characteristics and were used for subsequent analysis.

[0098] The present invention constructs a gene co-expression network based on gene expression data. Through weighted gene co-expression network analysis, all genes are divided into different modules according to the similarity of their expression patterns, and each module is identified by a unique color ( Figure 4 in A). To further explore the association between these modules and clinical traits, a module-trait relationship heatmap was drawn, and the correlation between each module and two clinical characteristics was calculated. The results showed that in the GSE37250 dataset, the magenta module and the green module had the most significant correlations with the LTBI group, where the magenta module showed a significant positive correlation (r = 0.66, p = 4e-24), while the green module showed a significant negative correlation (r = -0.7, p = 1e-27) ( Figure 4 in B). These results indicate that the genes in the magenta and green modules may play important roles in the pathophysiological process of LTBI, providing important clues for further functional research and biomarker screening.

[0099] Example 4: Screening and verification of key genes related to copper death / ferroptosis in LTBI

[0100] 1. Screening of key genes related to copper death / ferroptosis in LTBI

[0101] First, DEGs, key module genes associated with LTBI from WGCNA analysis, and Cuproptosis-related genes (CRGs) and Ferroptosis-related genes (FRGs) were uploaded to an online analysis platform (https: / / hiplot.cn / ?lang=zh_cn). Intersection analysis identified common intersection genes. These genes were considered potential key genes associated with LTBI. To improve the accuracy of the screening results and prevent overfitting, two machine learning algorithms—least absolute shrinkage and selection operator (LASSO) regression and support vector machine recursive feature elimination (SVM-RFE)—were used to further screen candidate genes. LASSO regression analysis was performed using the R package "glmnet." The optimal penalty parameter λ was adjusted through 10-fold cross-validation to identify the eigengenes that contributed most to LTBI classification. SVM-RFE was implemented using the R packages "e1071" and "caret." The most discriminatory eigengenes were identified by calculating the point at which the cross-validation error was minimized. The characteristic genes screened by LASSO regression and SVM-RFE algorithm were uploaded to the online analysis platform (https: / / hiplot.cn / ?lang=zh_cn) and the intersection was taken. The overlapping genes obtained were used as the final core DEGs. These genes were considered to be the best candidate genes for constructing the diagnostic model.

[0102] Through WGCNA analysis, we screened out the two modules with the highest correlation with LTBI. To further explore potential key genes, we performed intersection analysis on the green module with stronger correlation (containing 874 genes) with DEGs, CRGs and FRGs, and finally obtained 14 intersection genes. Among them, 11 genes were related to ferroptosis, including MUC1, MIR9-3, PARP9, CREB5, MGST1, LCN2, GALNT14, MAPK14, ATF3, MEG3 and SLC2A14; 2 genes were related to copper death, namely SCO2 and CD274; in addition, MT1G was identified as a co-regulatory gene of copper death and ferroptosis ( Figure 5 Middle A).

[0103] To avoid overfitting of the diagnostic model and take into account the economic costs of practical applications, we used two machine learning methods, LASSO regression and SVM, to perform dimensionality reduction screening on the core genes. Based on the SVM algorithm, when N = 13, the classification error reached the minimum, and finally 13 characteristic genes were determined, including PARP9, GALNT14, MGST1, SCO2, CREB5, MT1G, MAPK14, ATF3, MUC1, CD274, LCN2, SLC2A14, and MEG3( Figure 5 in B). Meanwhile, based on the LASSO algorithm and through ten-fold cross-validation, we screened out 7 characteristic genes, namely MT1G, SCO2, MUC1, PARP9, CREB5, MGST1, and ATF3( Figure 5 in C). Finally, we obtained 7 core genes screened by the two machine learning algorithms through a Venn diagram: MT1G, SCO2, MUC1, PARP9, CREB5, MGST1, and ATF3( Figure 5 in D). These genes may serve as potential biomarkers for LTBI, providing an important basis for subsequent functional verification and clinical applications.

[0104] 2. Verification of key genes related to cuproptosis / ferroptosis in LTBI

[0105] To evaluate the ability of the screened core DEGs to distinguish between ATB and LTBI, first, the "limma" and "ggpubr" packages in R software were used to verify the expression levels of the core genes screened by machine learning. The screening criterion was a p-value < 0.05 to ensure the statistical significance of gene expression differences. Subsequently, the R software package "pROC" (version 1.15.0) was used to perform receiver operating characteristic curve (ROC) analysis, and the area under the curve (AUC) was calculated to evaluate the accuracy of the diagnostic model. The ROC curve is a statistical tool for evaluating the discriminatory ability of a binary diagnostic test. The AUC value is usually between 0.5 and 1, and the closer this value is to 1, the stronger the discriminatory ability of the diagnostic model. Subsequently, the differential expression of the core genes in LTBI and ATB was further verified in the validation dataset, and a diagnostic model based on these genes was constructed. Finally, the diagnostic model was verified using the training dataset and the validation dataset respectively to ensure the stability and reliability of the model.

[0106] After screening out 7 core genes, we first analyzed the expression of these genes in the validation set GSE28623. The results showed that there were significant differences in the expression of SCO2 (P<0.0001), PARP9 (P<0.0001), CREB5 (P<0.0001), MGST1 (P<0.0001), ATF3 (P<0.0001), MT1G (P = 0.00021) and MUC1 (P = 0.027) between the two groups of ATB and LTBI ( Figure 6 A-G in), indicating that these genes met the modeling requirements.

[0107] Based on CRGs / FRGs, we constructed an LTBI diagnostic model by binary logistic regression analysis method: P = 1 / [1+e -(31.875-0.795*MT1G+0.193*SCO2-1.102*MUC1-0.775*PARP9-0.624*CREB5-1.157*MGST1-0.284*ATF3) , and named it HeptaTB Dx Model. Among them, P represents the probability value of latent tuberculosis infection; the gene names in the model represent the expression levels of the genes; e represents the base of the natural logarithm.

[0108] We further evaluated the diagnostic efficacy of the above 7 genes in differentiating LTBI from ATB through the ROC curve. In the training set GSE37250, the combined diagnostic model composed of MT1G, SCO2, MUC1, PARP9, CREB5, MGST1 and ATF3 showed excellent discrimination ability, and its AUC value was 0.963 ( Figure 6 H in). Similarly, in the validation set GSE28623, the AUC value of this model reached 0.930 ( Figure 6 I in), further verifying its stability and reliability. These results indicate that the diagnostic model based on CRGs / FRGs has high clinical application potential in differentiating LTBI from ATB.

[0109] Example 5, Identification of the copper death / ferroptosis subtypes in LTBI

[0110] Consensus Clustering is a method for identifying molecular subtypes based on the approximate number of clusters, which can effectively reveal gene expression patterns related to specific biological processes (such as cuproptosis and ferroptosis). In this study, the k-means method was used to subgroup LTBI samples based on the expression levels of CRGs / FRGs to explore the potential molecular subtypes of LTBI. We used the R package "ConsensusClusterPlus" for consensus clustering analysis and calculated the number of clusters and their stability for the samples through an unsupervised clustering method. Based on the expression levels of CRGs / FRGs, hierarchical clustering was performed on LTBI samples. 50 iterations were set to ensure the stability of the clustering results. The optimal number of clusters was determined through the cumulative distribution function (CDF) curve of the consensus score and the heatmap of the consensus matrix. To verify the effectiveness of the consensus clustering, principal components analysis (PCA) was performed using GraphPad Prism (version 10.1.2) to visualize the distribution of samples in the reduced-dimensional space and further confirm the reliability and biological significance of the clustering results.

[0111] To deeply explore the roles of CRGs and FRGs in the pathogenesis of LTBI and their molecular associations, we performed subtype analysis on the LTBI population using an unsupervised learning method. In the present invention, based on the expression differences of CRGs and FRGs, a clustering analysis method was used to subgroup LTBI samples. The clustering analysis results showed that when k = 2, the samples within the two clusters had the smallest difference, while the difference between the clusters was maximized, thus successfully dividing the LTBI samples into two different subtypes ( Figure 7 in A - E). PCA further verified the reliability of the clustering results, and the PCA plot showed good discrimination efficiency between the two clusters ( Figure 7 in F).

[0112] In addition, we compared the expression of the 7 core genes of the HeptaTB Dx Model between the ATB and LTBI groups and found that there were significant differences in the expression of these genes between the two groups. Notably, among the two subtypes of LTBI, the four genes MT1G, SCO2, PARP9, and ATF3 showed high expression in cluster 2, but their expression levels were still significantly lower than those in the ATB group; while the three genes MUC1, CREB5, and MGST1 showed no significant difference in expression between the two clusters, but their expression trends were consistent with the LTBI group and were lower than those in cluster 1 ( Figure 7 in G - H). These results indicate that CRGs and FRGs may have different regulatory roles in different subtypes of LTBI, providing a new research direction for further revealing the molecular mechanism of LTBI.

[0113] Example 6: Construction and Evaluation of Nomogram

[0114] To evaluate the predictive diagnostic ability and clinical value of the model, a nomogram based on the above-mentioned signature genes was constructed using the R package "rms" in R software (version 4.3.2). Nomograms are widely used in the medical field, especially in oncology and other fields that require accurate prediction. A nomogram can include multiple different influencing factors that may affect the prognosis simultaneously to predict the survival or occurrence of a study cohort. To verify the predictive accuracy of the nomogram, the present invention plotted a calibration curve to evaluate the degree of consistency between the expected probability and the actual results. The closer the calibration curve is, the stronger the predictive ability of the model. In addition, the present invention also used decision curve analysis (DCA) to evaluate the clinical usefulness and advantages of the nomogram. DCA evaluates the practical application value of the model in clinical decision-making by quantifying the net benefit of the model at different threshold probabilities.

[0115] To visualize the HeptaTB Dx model to improve its convenience in actual clinical applications, we plotted a nomogram of the HeptaTB Dx model ( Figure 8 as shown in A). In the nomogram, the contribution value corresponding to each gene can be intuitively shown by its position on the scale. We can calculate the total score by locating the score of each gene on the scale and projecting these scores onto the total point scale at the top. The total score can then be mapped to the probability scale at the bottom to evaluate the likelihood that a patient has LTBI.

[0116] Through calibration curve analysis, we found a high degree of consistency between the predicted probability of the HeptaTB Dx model and the actual observations, indicating that the HeptaTB Dx model has a good fitting effect ( Figure 8 as shown in B). In addition, DCA showed that when the threshold probability of the patient was between 50% and 99%, using the HeptaTB Dx model could significantly improve the net benefit of clinical decision-making, further verifying its clinical application value ( Figure 8 as shown in C).

[0117] In summary, the 7-gene diagnostic model HeptaTB Dx constructed based on CRGs and FRGs showed extremely high credibility and clinical application potential in differentiating LTBI from ATB. This model not only provides a new tool for the accurate diagnosis of tuberculosis but also lays an important foundation for subsequent clinical translational research.

[0118] Example 7: Verification of the Early Risk Warning Model for Latent Tuberculosis Infection

[0119] 1. Validation of the LTBI Early Risk Warning Model through a Cohort Study

[0120] To further validate the expression of the above 7 core genes in the real-world population, we recruited 10 patients with active tuberculosis (ATB), 10 patients with latent tuberculosis infection (LTBI), and 10 healthy controls (HC) for a prospective cohort study. This study has been approved by the Medical Ethics Committee of the Eighth Medical Center of the Chinese PLA General Hospital (Ethical Approval Number: 3092023122013297234), and all participants have signed informed consent forms. We collected 5 ml of peripheral blood from each participant to isolate peripheral blood mononuclear cells (PBMCs), and then extracted total RNA. After ensuring that the quality of the RNA samples met the standards, we constructed small RNA libraries using the Illumina sequencing platform and performed sequencing.

[0121] Inclusion criteria for HCs: A. No history of contact with tuberculosis patients; B. Negative enzyme-linked immunospot assay; C. No clinical manifestations of tuberculosis, normal chest X-ray, excluding the diagnosis of active tuberculosis; D. HIV negative.

[0122] Exclusion criteria for HCs: A. History of travel or residence in high-risk tuberculosis areas; B. Staff in tuberculosis specialized hospitals or laboratories; C. Children under 12 years old; D. Previous history of tuberculosis or old lesions in pulmonary imaging; E. Unable to perform CE antigen detection or allergic; F. HIV positive; G. Unable to perform CE antigen detection or allergic.

[0123] Inclusion criteria for LTBI: A. Close contact history with pulmonary tuberculosis patients; B. Staff in specialized tuberculosis hospitals or laboratories; C. No clinical symptoms of pulmonary tuberculosis; D. Normal chest X-ray; E. Positive interferon-gamma release assay (IGRA); F. HIV negative; G. 12 years old and above.

[0124] Exclusion criteria for LBTI: A. Pulmonary tuberculosis patients; B. Pregnant or lactating women; C. HIV positive; D. Anti-tuberculosis treatment for one month or more; F. Children under 12 years old.

[0125] Inclusion criteria for ATB: The TB patients referred to in the present invention are tuberculosis patients of the lung tissue, trachea, bronchus, and pleura diagnosed according to the "Diagnostic Criteria for Pulmonary Tuberculosis WS288-2017". The diagnosis of tuberculosis is based on etiological and pathological results as the confirmation basis, and is comprehensively analyzed and diagnosed in combination with the epidemiological history, clinical manifestations, chest imaging, relevant auxiliary examinations, and differential diagnosis.

[0126] Exclusion criteria for ATB: A. Those using hormones; B. Diseases affecting immune function such as HIV infection, post-transplantation, and autoimmune diseases; C. Malnourished; D. Children under 12 years old.

[0127] The results showed that there were significant differences in the expression of genes SCO2 and PARP9 between the ATB and LTBI groups (P < 0.05). Specifically, compared with the ATB group, genes SCO2 and PARP9 were significantly downregulated in the LTBI group (P < 0.05). Although the other 5 CRGs / FRGs showed a certain trend of difference between the ATB and LTBI groups, this difference did not reach statistical significance ( Figure 9 A-G).

[0128] 2. RT-qPCR verification of the early risk warning model for LTBI

[0129] Finally, we verified the reliability of the bioinformatics prediction results by RT-qPCR. A total of 111 research subjects from the Eighth Medical Center of the Chinese PLA General Hospital were recruited, including HCs individuals (n = 37), LTBI individuals (n = 37), and ATB patients (n = 37). The inclusion and exclusion criteria for each group of people were the same as described above. The diagnosis of TB patients was based on the "Diagnostic Criteria for Pulmonary Tuberculosis WS288-2017", and LTBI patients were diagnosed according to the results of IGRAs. Immunodeficiency, the use of glucocorticoids, pregnant women, and children under 12 years old were excluded.

[0130] 5 ml of peripheral venous blood was collected, and total RNA was extracted from whole blood samples using the IVD Pure Total Blood RNA Extraction Kit (IVD, China) according to the manufacturer's instructions. Then, reverse transcription was performed using the FastKing gDNA Dispelling RT SuperMix Kit (Tiangen, China), incubated at 42 °C for 15 minutes, and then incubated at 95 °C for 3 minutes. After that, the samples were diluted with RNase-free water. RT-qPCR was performed using the FastReal qPCR PreMix (SYBR Green) (Tiangen, China) kit and the Roche 480 system. The reaction conditions were: pre-denaturation (95 °C, 2 minutes), 40 denaturation cycles (95 °C, 5 seconds), annealing and extension (60 °C, 30 seconds). Glyceraldehyde-3-phosphate dehydrogenase (GAPDH) was used as an internal control for amplification. The relative expression level was measured using the 2 -ΔΔCt method. The primer sequences used in this experiment are shown in Table 2 in detail.

[0131] Table 2. Detailed information on the primer sequences of each gene symbol and its internal reference gene

[0132]

[0133]

[0134] The results showed that compared with the HCs group, there were significant differences in the expression of genes CREB5, ATF3, MT1G, PARP9, and MGST1 between the ATB and LTBI groups (P<0.05). Specifically, compared with the ATB group, genes CREB5, ATF3, and PARP9 were significantly downregulated in the LTBI group (P<0.05), while genes MT1G (P<0.01) and MGST1 (P<0.05) were significantly upregulated in the LTBI group. Although the differences in the expression of genes SCO2 and MUC1 between the ATB and LTBI groups did not reach statistical significance, their expression trends were consistent with those of the above genes( Figure 10 as shown in A-G). Subsequently, we plotted the ROC curve, and the results of ROC analysis showed that the overall AUC of this model was 0.778 (CI: 0.673-0.883). When the Youden index was 0.43243, the sensitivity was 0.81081, the specificity was 0.62162, and the accuracy was 0.71622( Figure 11 ).

[0135] The above results indicate that the seven core genes (SCO2, MT1G, CREB5, PARP9, ATF3, MUC1, MGST1) of the present invention can be used as biomarkers for early risk warning of LTBI. The LTBI risk early warning model constructed based on these seven core genes has good accuracy, providing a new method for early risk warning of LTBI.

[0136] The present invention has been described in detail above. For those skilled in the art, without departing from the purpose and scope of the present invention and without unnecessary experiments, the present invention can be implemented within a wide range under equivalent parameters, concentrations, and conditions. Although specific embodiments of the present invention are given, it should be understood that the present invention can be further improved. In short, according to the principle of the present invention, this application intends to include any changes, uses, or improvements to the present invention, including changes made using conventional techniques known in the art that are outside the scope disclosed in this application.

Claims

1. Use of a biomarker and / or a substance for detecting the biomarker in any of the following: A1) Use in the preparation of a product for early risk warning of latent tuberculosis infection; A2) Use in the construction of an early risk warning model for latent tuberculosis infection; The biomarker is the SCO2 gene, MT1G gene, CREB5 gene, PARP9 gene, ATF3 gene, MUC1 gene, and / or MGST1 gene.

2. The application according to claim 1, characterized in that, The substance includes reagents and / or instruments for detecting the expression levels of the SCO2 gene, MT1G gene, CREB5 gene, PARP9 gene, ATF3 gene, MUC1 gene, and / or MGST1 gene.

3. The application according to claim 2, characterized in that, The reagent includes primers for specifically amplifying the SCO2 gene, MT1G gene, CREB5 gene, PARP9 gene, ATF3 gene, MUC1 gene, and / or MGST1 gene and / or probes for specifically recognizing the SCO2 gene, MT1G gene, CREB5 gene, PARP9 gene, ATF3 gene, MUC1 gene, and / or MGST1 gene.

4. A composition or kit for early risk warning of latent tuberculosis infection, characterized in that, The composition or kit includes reagents for detecting the expression levels of the SCO2 gene, MT1G gene, CREB5 gene, PARP9 gene, ATF3 gene, MUC1 gene, and / or MGST1 gene.

5. A tuberculosis latent infection early risk warning model, characterized in that, The model is constructed using the biomarker described in claim 1.

6. The model according to claim 5, characterized in that, The formula of the model is: P = 1 / [1 + e -(31.875-0.795*MT1G+0.193*SCO2-1.102*MUC1-0.775*PARP9-0.624*CREB5-1.157*MGST1-0.284*ATF3) ​ Wherein, P represents the probability value of latent tuberculosis infection; MT1G, SCO2, MUC1, PARP9, CREB5, MGST1, ATF3 respectively represent the expression levels of the MT1G, SCO2, MUC1, PARP9, CREB5, MGST1, ATF3 genes; e represents the base of the natural logarithm.

7. A method for constructing an early risk warning model for latent tuberculosis infection, characterized in that, The method includes using the data of the expression levels of the SCO2 gene, MT1G gene, CREB5 gene, PARP9 gene, ATF3 gene, MUC1 gene, and / or MGST1 gene in samples of known latent tuberculosis infection patients and active tuberculosis patients as training samples, and adopting statistical analysis methods to establish an early risk warning model for latent tuberculosis infection.

8. The method according to claim 7, wherein The statistical analysis methods include binary logistic regression analysis, support vector machine analysis, and LASSO regression analysis methods.

9. A system for early risk warning of latent tuberculosis infection, characterized in that, The system includes: A data receiving module, configured to: receive data of the expression level of the biomarker described in claim 1 in a sample of a subject to be tested from at least one terminal; A data processing module, configured to: input the data into the model described in claim 5 or 6, or a model constructed by the method described in claim 7 or 8, to obtain the probability value of latent tuberculosis infection; A data output module, configured to: output the probability value to at least one client.

10. A method for early risk warning of latent tuberculosis infection, characterized in that, The method includes the following steps: S1. Data reception: Receive data of the expression level of the biomarker described in claim 1 in a sample of a subject to be tested from at least one terminal; S2. Data processing: Input the data into the model described in claim 5 or 6, or the model constructed by the method described in claim 7 or 8, to obtain the probability value of latent tuberculosis infection; S3. Data output: Output the probability value to at least one client.