Biomarker for early risk early warning of lung cancer complicated with pulmonary tuberculosis, early risk early warning model and application thereof
Through the expression detection of ANK2, CA1, EPB42, HBB, HBD and MYL4 genes, an early risk warning model for lung cancer combined with tuberculosis was established, which solved the problem of early diagnosis of lung cancer combined with tuberculosis, and improved diagnostic accuracy and treatment effect.
Patent Information
- Application Number
- CN202510562552.X
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-04-30
- Publication Date
- 2025-07-29
AI Technical Summary
The prior art is difficult to effectively distinguish the early symptoms and imaging characteristics of lung cancer and tuberculosis, which leads to difficulty in early diagnosis of lung cancer combined with tuberculosis, increasing the difficulty of treatment and mortality.
The expression level of ANK2, CA1, EPB42, HBB, HBD and MYL4 genes were used as biomarkers, and their expression levels were detected by RT-qPCR and other methods, combined with binary logistic regression analysis, and early risk warning model was established to provide personalized treatment suggestions.
It improves the accuracy and sensitivity of early diagnosis of lung cancer combined with tuberculosis, reduces mortality, provides personalized prevention and treatment strategies, and improves patients' quality of life.
Smart Images

Figure SMS_1 
Figure SMS_3 
Figure SMS_4
Abstract
Description
Technical Field
[0001] The present invention relates to the fields of bioinformatics and biomedicine, and in particular to biomarkers for early risk warning of lung cancer combined with pulmonary tuberculosis, an early risk warning model and applications thereof. Background Art
[0002] Lung cancer (LC) is the most common type of cancer in the world and the leading cause of cancer-related deaths. Its morbidity and mortality rates rank first among cancers, posing a serious threat to human health and life. The pathogenesis of lung cancer is complex, with smoking, occupational exposure, and air pollution being common causes. The early symptoms of lung cancer are often not obvious and are easily overlooked. As the disease progresses, patients may experience symptoms such as coughing, sputum, and chest pain. In the late stage, it may be accompanied by distant metastasis, leading to multiple organ failure. Tuberculosis (TB) is a chronic infectious disease caused by Mycobacterium tuberculosis (MTB).
[0003] Previous studies have found that the risk of MTB infection in lung cancer patients is nine times higher than in the general population, and anti-tumor treatment further increases the risk of MTB infection in cancer patients. Furthermore, studies have shown that although the pathological features of lung cancer and pulmonary tuberculosis (PTB) differ, their clinical manifestations overlap to a certain extent. In particular, the clinical manifestations of PTB in patients with lung cancer are more subtle and often overlap with those of lung cancer, leading to delayed early diagnosis. This delay not only affects the effective treatment of PTB but also may aggravate the progression of lung cancer, making treatment more difficult, the prognosis worse, and the mortality rate higher in patients with lung cancer and pulmonary tuberculosis (LC-PTB).
[0004] With the increasing prevalence of LC-PTB, its early prevention, diagnosis, and treatment have become key issues that need to be addressed urgently. However, the difficulty of early diagnosis is greatly increased due to the overlap of early symptoms of lung cancer and pulmonary tuberculosis (such as cough, sputum, chest pain, etc.), overlapping imaging features (pulmonary tuberculosis lesions and lung cancer nodules may be difficult to distinguish on X-ray or CT images), and the possibility that traditional lung cancer markers may be interfered with by tuberculosis infection, resulting in false positives or false negatives. Therefore, early warning and diagnosis of TB in LC patients and early interventional treatment are important strategies to reduce the morbidity and mortality of lung cancer combined with pulmonary tuberculosis (LC-PTB). Currently, there is still a lack of rapid and efficient early risk warning methods for LC-PTB. The vigorous development of artificial intelligence technology and multi-omics technology has provided new means to solve this problem. Summary of the invention
[0005] The object of the present invention is to provide a biomarker, an early risk warning model and their applications for early risk warning of lung cancer complicated with pulmonary tuberculosis, which are accurate, specific and sensitive. The technical problems to be solved are not limited to the described technical topics, and those skilled in the art can clearly understand other technical topics not mentioned herein through the following description.
[0006] To achieve the above object, the present invention first provides the application of a biomarker and / or a substance for detecting the biomarker in any of the following:
[0007] A1) Application in the preparation of a product for early risk warning of lung cancer complicated with pulmonary tuberculosis;
[0008] A2) Application in the construction of an early risk warning model for lung cancer complicated with pulmonary tuberculosis;
[0009] The biomarker can be any of the following:
[0010] B1) ANK2 gene, CA1 gene, EPB42 gene, HBB gene, HBD gene and / or MYL4 gene;
[0011] B2) ANK2 gene, CA1 gene, HBD gene and MYL4 gene.
[0012] In the above applications, the substance may include reagents and / or instruments for detecting the expression levels of ANK2 gene, CA1 gene, EPB42 gene, HBB gene, HBD gene and / or MYL4 gene.
[0013] Furthermore, the substance may include reagents and / or instruments for detecting the expression levels of ANK2 gene, CA1 gene, EPB42 gene, HBB gene, HBD gene and / or MYL4 gene by reverse transcription-polymerase chain reaction (RT-PCR), reverse transcription real-time quantitative polymerase chain reaction (RT-qPCR), transcriptome sequencing technology (RNA-seq), Northern blot, in situ hybridization technology (ISH), fluorescence in situ hybridization (FISH), gene chip technology, Nanopore sequencing technology and / or PacBio sequencing technology.
[0014] In the above applications, the reagent may include primers for specifically amplifying ANK2 gene, CA1 gene, EPB42 gene, HBB gene, HBD gene and / or MYL4 gene and / or probes for specifically recognizing ANK2 gene, CA1 gene, EPB42 gene, HBB gene, HBD gene and / or MYL4 gene.
[0015] Further, the primers include primer pair 1 for detecting or specifically amplifying the ANK2 gene, primer pair 2 for detecting or specifically amplifying the CA1 gene, primer pair 3 for detecting or specifically amplifying the EPB42 gene, primer pair 4 for detecting or specifically amplifying the HBB gene, primer pair 5 for detecting or specifically amplifying the HBD gene, and primer pair 6 for detecting or specifically amplifying the MYL4 gene.
[0016] Further, primer pair 1 may be composed of a forward primer with a nucleotide sequence as shown in SEQ ID NO:1 and a reverse primer with a nucleotide sequence as shown in SEQ ID NO:2;
[0017] Primer pair 2 may be composed of a forward primer with a nucleotide sequence as shown in SEQ ID NO:5 and a reverse primer with a nucleotide sequence as shown in SEQ ID NO:6;
[0018] Primer pair 3 may be composed of a forward primer with a nucleotide sequence as shown in SEQ ID NO:3 and a reverse primer with a nucleotide sequence as shown in SEQ ID NO:4;
[0019] Primer pair 4 may be composed of a forward primer with a nucleotide sequence as shown in SEQ ID NO:7 and a reverse primer with a nucleotide sequence as shown in SEQ ID NO:8;
[0020] Primer pair 5 may be composed of a forward primer with a nucleotide sequence as shown in SEQ ID NO:9 and a reverse primer with a nucleotide sequence as shown in SEQ ID NO:10;
[0021] Primer pair VI may be composed of a forward primer with a nucleotide sequence as shown in SEQ ID NO:11 and a reverse primer with a nucleotide sequence as shown in SEQ ID NO:12.
[0022] In this paper, the products may include reagents, reagent kits, detection chips (such as gene chips), test strips, test cards, immunosensors, and devices, but are not limited thereto.
[0023] Further, the application may include using the data of the expression levels of the ANK2 gene, CA1 gene, EPB42 gene, HBB gene, HBD gene, and / or MYL4 gene in the samples of known patients with lung cancer complicated with pulmonary tuberculosis and lung cancer patients as training samples, and adopting statistical analysis methods (such as binary logistic regression analysis method) to establish an early risk warning model for lung cancer complicated with pulmonary tuberculosis, and conducting early risk warning for lung cancer complicated with pulmonary tuberculosis on the subjects according to this model.
[0024] The samples described in this paper may be blood samples (including whole blood, plasma, and serum).
[0025] Further, the sample may be whole blood.
[0026] The present invention also provides a composition or kit for early risk warning of lung cancer complicated with pulmonary tuberculosis, and the composition or kit may include reagents for detecting the expression levels of ANK2 gene, CA1 gene, EPB42 gene, HBB gene, HBD gene and / or MYL4 gene.
[0027] Further, the reagent may be a primer for specifically amplifying ANK2 gene, CA1 gene, EPB42 gene, HBB gene, HBD gene and / or MYL4 gene (such as primer pair 1, primer pair 2, primer pair 3, primer pair 4, primer pair 5 and / or primer pair 6 described herein) and / or a probe for specifically recognizing ANK2 gene, CA1 gene, EPB42 gene, HBB gene, HBD gene and / or MYL4 gene.
[0028] Further, the composition or kit may also include reagents required for PCR detection, such as DNA polymerase (such as Taq DNA polymerase, Tth DNA polymerase, Vent DNA polymerase and Pfu DNA polymerase, etc.), dNTPs, Mg 2+ solution (such as MgSO4 or MgCl2 solution) and PCR buffer (such as Tris-HCl).
[0029] Further, the composition or kit may also include reagents required for reverse transcription reaction, transcriptome sequencing, Northern blot, in situ hybridization, gene chip detection, Nanopore sequencing and / or PacBio sequencing, etc.
[0030] Further, the composition or kit may also include total RNA extraction reagents.
[0031] The kit described herein may be a detection kit based on reverse transcription real-time quantitative polymerase chain reaction (RT-qPCR).
[0032] The various reagent components of the kit may be present in separate containers, or may be pre-combined into a reagent mixture in whole or in part.
[0033] The components of the kit may be provided in the form of a solution, for example, in the form of an aqueous solution. In the case of being in an aqueous solution state, the concentration or content of these components can be conveniently determined by those skilled in the art according to different requirements. For example, for storage purposes, the concentration of the components can be in a higher form, and when in a working state or in use, the concentration can be reduced to the working concentration by diluting the above higher concentration solution.
[0034] The present invention also provides an early risk warning model for lung cancer complicated with pulmonary tuberculosis, and the model can be constructed by using the biomarkers described herein.
[0035] Further, the formula of the model can be:
[0036] logit(p) = 2.022 + 416.907 * ANK2 - 358 * CA1 - 626 * HBD - 884 * MYL4
[0037] Wherein, logit(p) represents the risk score; ANK2 represents the expression level of the ANK2 gene; CA1 represents the expression level of the CA1 gene; HBD represents the expression level of the HBD gene; MYL4 represents the expression level of the MYL4 gene.
[0038] The present invention also provides the application of the model in the preparation of a product for early risk warning of lung cancer complicated with pulmonary tuberculosis.
[0039] The present invention also provides a method for constructing an early risk warning model for lung cancer complicated with pulmonary tuberculosis, and the method includes using the data of the expression levels of the ANK2 gene, CA1 gene, EPB42 gene, HBB gene, HBD gene and / or MYL4 gene in the samples of known patients with lung cancer complicated with pulmonary tuberculosis and lung cancer patients as training samples, and adopting a statistical analysis method to establish an early risk warning model for lung cancer complicated with pulmonary tuberculosis.
[0040] In the above method, the statistical analysis method includes binary logistic regression analysis.
[0041] Further, the statistical analysis method may also include, but is not limited to, random forest analysis, decision tree analysis, support vector machine analysis and neural network analysis methods.
[0042] The application method of any of the early risk warning models for lung cancer complicated with pulmonary tuberculosis in the present invention may include the following steps:
[0043] (1) Detect the data of the expression levels of the ANK2 gene, CA1 gene, EPB42 gene, HBB gene, HBD gene and / or MYL4 gene in the sample of the subject to be tested;
[0044] (2) Input the data into any of the early risk warning models for lung cancer complicated with pulmonary tuberculosis in the present invention;
[0045] (3) Obtain the early risk warning result of lung cancer complicated with pulmonary tuberculosis of the subject to be tested through the model.
[0046] Further, the formula of the model in step (2) can be:
[0047] logit(p) = 2.022 + 416.907 * ANK2 - 358 * CA1 - 626 * HBD - 884 * MYL4
[0048] Among them, logit(p) represents the risk score; ANK2 represents the expression level of the ANK2 gene; CA1 represents the expression level of the CA1 gene; HBD represents the expression level of the HBD gene; MYL4 represents the expression level of the MYL4 gene.
[0049]
[0050] The combine formula is used to convert the linear prediction value into a probability value. In this way, a probability value between 0 and 1 can be obtained, representing the probability of the occurrence of lung cancer complicated with pulmonary tuberculosis.
[0051] Furthermore, personalized or further prevention or treatment suggestions can be provided for the subjects to be tested according to the early risk warning results of lung cancer complicated with pulmonary tuberculosis, or early intervention treatment can be carried out for high-risk patients.
[0052] The present invention also provides a system for early risk warning of lung cancer complicated with pulmonary tuberculosis, and the system may include:
[0053] A data receiving module, configured to: receive data on the expression levels of the biomarkers described in the present invention in the samples of the subjects to be tested from at least one terminal;
[0054] A data processing module, configured to: input the data into the model described in the present invention, or a model constructed by the method described in the present invention, to obtain a risk score for lung cancer complicated with pulmonary tuberculosis;
[0055] A data output module, configured to: output the risk score to at least one client; or convert the risk score into a probability value of the occurrence of lung cancer complicated with pulmonary tuberculosis and then output it to at least one client.
[0056] Furthermore, the risk score can be converted into a probability value of the occurrence of lung cancer complicated with pulmonary tuberculosis through the combine formula described herein.
[0057] The subjects to be tested described herein may be selected from healthy subjects, lung cancer patients, and suspected patients with lung cancer complicated with pulmonary tuberculosis.
[0058] The gene sequences of the ANK2 gene, CA1 gene, EPB42 gene, HBB gene, HBD gene, and MYL4 gene described in the present invention are all well-known in the art, and their gene sequences can be publicly obtained from the NCBI website. Exemplary sequences of these genes are shown as follows:
[0059] The nucleotide sequence of the coding sequence (CDS) of the ANK2 gene described in this article may be positions 200 - 5791 of GenBank Accession No. NM_001127493.3 (Update Date 12 - OCT - 2024).
[0060] The nucleotide sequence of the coding sequence (CDS) of the CA1 gene described in this article may be positions 168 - 953 of GenBank Accession No. NM_001128829.4 (Update Date 06 - APR - 2024).
[0061] The nucleotide sequence of the coding sequence (CDS) of the EPB42 gene described in this article may be positions 194 - 2359 of GenBank Accession No. NM_000119.3 (Update Date 27 - APR - 2025).
[0062] The nucleotide sequence of the coding sequence (CDS) of the HBB gene described in this article may be positions 51 - 494 of GenBank Accession No. NM_000518.5 (Update Date 27 - APR - 2025).
[0063] The nucleotide sequence of the coding sequence (CDS) of the HBD gene described in this article may be positions 51 - 494 of GenBank Accession No. NM_000519.4 (Update Date 27 - APR - 2025).
[0064] The nucleotide sequence of the coding sequence (CDS) of the MYL4 gene described in this article may be positions 163 - 756 of GenBank Accession No. NM_001002841.2 (Update Date 03 - APR - 2024).
[0065] The expression levels of the biomarkers (ANK2 gene, CA1 gene, EPB42 gene, HBB gene, HBD gene, and / or MYL4 gene) described in this article are applicable to conventional reverse transcription quantitative real-time polymerase chain reaction (RT-qPCR) detection experiments. Although the examples provided in the present invention exemplarily use the RT-qPCR method to detect the expression levels of biomarkers in whole blood samples, the present invention is not limited to this specific detection method. Those skilled in the art can use any other suitable detection methods well-known in the art (such as RNA-seq, Northern blot, in situ hybridization, gene chip detection, Nanopore sequencing, and PacBio sequencing, etc.) to detect the expression levels of biomarkers in whole blood samples. These alternative methods do not deviate from the scope of the present invention, and the present invention should include these alternative methods.
[0066] The present invention first proposed the use of ANK2 gene, CA1 gene, EPB42 gene, HBB gene, HBD gene, and / or MYL4 gene as biomarkers for early risk warning of lung cancer complicated with pulmonary tuberculosis and the construction of an early risk warning model for lung cancer complicated with pulmonary tuberculosis. By analyzing the blood transcriptome data of clinical patients with lung cancer complicated with pulmonary tuberculosis and lung cancer patients, a total of 604 differentially expressed genes related to lung cancer complicated with pulmonary tuberculosis were identified. Through GSEA, GO, and KEGG enrichment analysis, combined with weighted gene co-expression network analysis (WGCNA) and PPI protein interaction network and other technologies, and integrating bioinformatics methods to identify immune biomarkers related to lung cancer complicated with pulmonary tuberculosis. Finally, ANK2, EPB42, CA1, HBB, HBD, and MYL4 were selected as biomarkers related to patients with lung cancer complicated with pulmonary tuberculosis. The expression levels of these six genes were verified by RT-qPCR and immunohistochemistry, and the results showed that the expression levels of these six genes were consistent with the trend in the training set. Based on these six genes, an early risk warning model can be constructed. The present invention constructed an early risk warning model for lung cancer complicated with pulmonary tuberculosis using four of these genes (ANK2 gene, CA1 gene, HBD gene, and MYL4 gene). The results showed that the overall AUC of this model was 0.94. When the Youden index was 0.7, the sensitivity of this model was 0.9 and the specificity was 0.8. The verification results of clinical samples showed that the AUC of the model constructed by the present invention was 0.715. When the Youden index was 0.38, the sensitivity of this model was 0.69 and the specificity was 0.69.
[0067] In addition, considering the crucial role of the immune system in the occurrence and development of lung cancer complicated with pulmonary tuberculosis, we not only identified immune biomarkers related to the disease, but also found important immune signaling pathways through KEGG enrichment analysis. For example, oxidative phosphorylation and nitrogen metabolism play key roles in the pathogenic mechanism of lung cancer complicated with pulmonary tuberculosis. At the same time, through immune infiltration analysis, we revealed that some adaptive immune cells in patients with lung cancer complicated with pulmonary tuberculosis were significantly inhibited, while inflammatory cells were significantly activated. These new findings not only provide potential immunological targets for the early detection and screening of patients with lung cancer complicated with pulmonary tuberculosis, but also offer new ideas and directions for the treatment and prevention of related diseases. The present invention can provide strong support for the tuberculosis risk management of lung cancer patients, thereby improving the quality of life and survival rate of patients. The present invention has developed 6 reliable biomarkers related to lung cancer complicated with pulmonary tuberculosis, and the constructed risk assessment model shows significant prediction efficiency, providing an early screening strategy for lung cancer complicated with pulmonary tuberculosis.
[0068] The biomarkers of the present invention and the model constructed using the biomarkers are verified by the ROC curve to be accurate, reliable, with good sensitivity and specificity, and are suitable for wide promotion. By providing early risk warnings for lung cancer complicated with pulmonary tuberculosis in the subjects to be tested, it can provide decision-making support for clinicians, guide the selection and adjustment of treatment plans, provide personalized or further preventive or treatment suggestions for patients, or conduct early intervention treatment for high-risk patients, etc., to achieve the best disease management and has broad clinical application value. Brief Description of the Drawings
[0069] Figure 1 Flow chart for the construction and verification of the early risk warning model for patients with lung cancer complicated with pulmonary tuberculosis (LC-PTB).
[0070] Figure 2 For differential analysis. a. The volcano plot shows that compared with the Control group, there are 305 up-regulated genes and 435 down-regulated genes in the LC group; while in the LC-PTB group, there are 144 up-regulated genes and 435 down-regulated genes. In addition, compared with the LC group, there are 216 up-regulated genes and 388 down-regulated genes in the LC-PTB group; b. The heat map shows the clustering of differential genes in the three groups of people.
[0071] Figure 3GSEA enrichment analysis of differentially expressed genes. a. Compared with the Control group, the top five pathways of DEGs in LC group patients were neutrophil degranulation, oxidative phosphorylation, MHC class II antigen presentation, DNA replication, and cellular response to chemical stress, etc.; b. Compared with the Control group, the DEGs of LC-PTB group were mainly enriched in five pathways: antigen processing and presentation, translation, Huntington's disease, oxidative phosphorylation, and red blood cell uptake of carbon dioxide and release of oxygen; c. Compared with the LC group, the related pathways enriched by the DEGs of LC-PTB group patients (including histone deacetylation, RUNX1-regulated megakaryocyte differentiation and platelet function-related genes, and DNA methylation, etc.) were all down-regulated in the analyzed samples and also showed significant negative enrichment scores in the dataset.
[0072] Figure 4 Overlapping regions of differential genes among three groups of people. We used Venn Diagram to display and compare the co-existing differential genes among the three groups.
[0073] Figure 5 GO and KEGG functional enrichment analysis of intersection differentially expressed genes. (a, d). GO and KEGG enrichment analysis of intersection DEGs between LC group and LC-PTB group compared with the Control group; (b, e). GO and KEGG enrichment analysis of intersection DEGs of LC-PTB group compared with the Control group and LC group; (c, f). GO and KEGG enrichment analysis of intersection DEGs of LC group compared with the Control group and LC-PTB group.
[0074] Figure 6 Protein-protein interaction network (PPI) and core genes of intersection differential genes. (a, b, c). This network intuitively presented the interaction relationship between the intersection differential genes and proteins of the three groups; (d, e, f). The relevant hub genes of the intersection differential genes of the three groups were identified using the MCODE plugin.
[0075] Figure 7 Screening of LC-PTB immune biomarkers. (a, b). Six core genes were screened using the PPI network; c. The optimal soft threshold was determined to be 12 by scale independence and average connectivity. d. Dendrogram of clustering of highly connected genes in key modules. e. Relationship between modules and lymphocytes. (f, g). Six core genes screened using the PPI network were used as a reference, and cross-comparison was performed with 1793 immune genes in the ImmPort database and key module genes screened by the WGCNA network, and six core genes were screened out.
[0076] Figure 8For the construction of the early risk warning model for LC-PTB. a. Visualize the risk assessment model using a nomogram; b. ROC curves of the ANK2, HBD, and MYL4 genes; c. ROC curve of the combined diagnostic model.
[0077] Figure 9 For the verification results of the biomarkers of LC-PTB in RT-qPCR. a. Expression of 6 genes in a prospective cohort study. b. Show the expression of 6 genes in a prospective cohort study and RT-qPCR verification.
[0078] Figure 10 For the verification of the early risk warning model for LC-PTB.
[0079] Figure 11 For the verification results of the biomarkers of LC-PTB in immunohistochemistry. This figure shows the expression of 6 genes in pathological tissues. Specific implementation manners
[0080] The present invention will be further described in detail below in conjunction with specific implementation manners. The provided embodiments are only for clarifying the present invention, rather than limiting the scope of the present invention. The following provided embodiments can be used as a guide for those of ordinary skill in the art to make further improvements, and do not limit the present invention in any way.
[0081] The experimental methods in the following embodiments, unless otherwise specified, are all conventional methods, carried out according to the techniques or conditions described in the literature in this field or according to the product instructions. The materials, reagents, etc. used in the following embodiments, unless otherwise specified, can be obtained from commercial channels.
[0082] The acquisition of clinical samples in the following embodiments has obtained the informed consent of the patients and has been approved by the Ethics Committee of the Eighth Medical Center of the Chinese People's Liberation Army General Hospital.
[0083] Example 1. Collection and sequencing of blood samples
[0084] The screened samples were 10 confirmed LC-PTB patients, 10 confirmed LC patients who visited the Eighth Medical Center of the Chinese People's Liberation Army General Hospital, and 10 healthy subjects as controls.
[0085] Inclusion criteria for lung cancer patients (LC patients): ① Age and medical history: Age ≥ 18 years old, with a smoking history (≥ 20 pack-years) or high-risk factors (such as occupational carcinogen exposure, family history of lung cancer, etc.); ② Pathological diagnosis: Definitively diagnosed as LC by tissue biopsy or cytological examination, including non-small cell lung cancer (NSCLC) and small cell lung cancer (SCLC); ③ Imaging features: Chest CT shows a space-occupying lesion in the lung, with or without mediastinal lymph node enlargement, pleural invasion, or distant metastasis; ④ Clear staging: Divided into stages I-IV according to the TNM staging system (9th edition).
[0086] Exclusion criteria for lung cancer patients (LC patients): ① Other malignancies: Complicated with other primary malignancies (such as breast cancer, colorectal cancer, etc.); ② Active infections: Complicated with active bacterial, fungal, or viral infections (such as pneumonia, HIV infection); ③ Severe complications: Presence of severe heart, liver, and kidney insufficiency or coagulation dysfunction; ④ Treatment contraindications: Unable to tolerate standard treatments such as surgery, radiotherapy, or chemotherapy.
[0087] Inclusion criteria for lung cancer patients with pulmonary tuberculosis (LC-PTB patients): ① Dual diagnostic basis: LC needs to meet the above inclusion criteria, and PTB needs to meet the diagnostic basis in the "Uniform Diagnostic Criteria for PTB" (WS288-2017), including positive sputum smear or culture, typical imaging features (such as miliary nodules, cavity formation), or pathological confirmation of tuberculous granuloma; ② Coexistence evidence: The PTB and LC lesions are located in the same lung lobe or there is a pathological association (such as carcinoma in situ within a tuberculous scar); ③ Temporal association: Active PTB or a previous history of TB (≥ 1 year), and the time interval from the diagnosis of LC is ≤ 5 years; ④ Symptom overlap: Simultaneously have PTB symptoms (such as low fever, night sweats, hemoptysis) and LC symptoms (such as irritating dry cough, weight loss).
[0088] Exclusion criteria for lung cancer patients with pulmonary tuberculosis (LC-PTB patients): ① Single disease: Only diagnosed as LC or PTB, without coexistence evidence; ② Non-tuberculous mycobacterial infection: Sputum culture or gene detection confirms non-tuberculous mycobacterial (NTM) infection; ③ Other lung diseases: Complicated with pulmonary embolism, pulmonary fibrosis, or other chronic lung diseases; ④ Immunosuppressed state: Long-term use of immunosuppressants or presence of primary immunodeficiency diseases.
[0089] Inclusion criteria for healthy subjects (Control): ① Age: Age ≥ 18 years old; ② Medical history: No history of malignancy, TB, HIV positivity, or autoimmune diseases, and no history of immunotherapy or glucocorticoid therapy in the past 3 months.
[0090] Exclusion criteria for healthy subjects (Control): ① Malignant tumors: primary or secondary malignant tumors (especially LC); ② Active infections: combined with active bacterial, fungal or viral infections (especially MTB infection); ③ Underlying diseases: chronic obstructive pulmonary disease, diabetes and other underlying diseases, etc.
[0091] 1. Sample collection and processing
[0092] According to the aforementioned inclusion and exclusion criteria, professional TB and oncology doctors and research personnel are jointly responsible for enrolling each group of subjects, and 5 mL of peripheral venous blood is collected by intravenous puncture from the anterior side of the arm (the specific enrollment process is as Figure 1 ). PBMCs are extracted from 5 mL of peripheral venous blood according to the following method: ① Transfer the collected blood sample to a 15 mL centrifuge tube, and gently pipette an equal volume of lymphocyte separation medium (Ficoll-Paque PLUS, restored from low temperature to room temperature) into the centrifuge (temperature adjusted to room temperature 20 - 23 °C, time 20 min, 2500 r, acceleration 9 and deceleration 1); ② Pipette the PBMCs layer into a new centrifuge tube, add 5 mL of 1640 medium to each tube, invert and shake well, and then place it in the centrifuge (temperature adjusted to room temperature 20 - 23 °C, time 10 min, 1500 r, acceleration 9 and deceleration 9); ③ After centrifugation, discard the supernatant, add 5 mL of 1640 medium again, gently pipette the cell suspension at the bottom 3 - 5 times to resuspend the cells, and then centrifuge; ④ If there is red blood cell precipitation at the bottom of the reagent tube after centrifugation, discard the supernatant, add 1 - 3 mL of hemolysin (depending on the amount of RBC) and pipette evenly, place it on ice or in the refrigerator for 3 - 5 min and then centrifuge; ⑤ After centrifugation, discard the supernatant, add 5 mL of 1640 medium to resuspend the cells and then centrifuge (if cell counting is to be measured at this time, 200 - 300 μL of PBS can be added to resuspend, then placed in a hemocytometer for counting, and after counting, centrifuge again); ⑥ After centrifugation, discard the supernatant, place it in a 1.5 μL centrifuge tube, add 1 mL of Trizol and gently pipette and mix well, and store it in a -80 °C refrigerator.
[0093] 2. Transcriptome sequencing
[0094] First, use the MJzol Animal RNA Isolation Kit (Majorivd) to extract total RNA according to the standard operating procedures provided by the manufacturer, and use the RNA Clean XP Kit (Beckman Coulter) and RNase-Free DNase Set (QIAGEN) for purification. Detect the integrity of RNA with an Agilent 2100 Bioanalyzer / Agilent 4200 TapeStation (Agilent technologies), and use a Qubit 2.0 fluorometer (Thermo Fisher Scientific) and a NanoDrop ND-2000 spectrophotometer (Thermo Fisher Scientific) to measure the total amount and purity of RNA. Secondly, perform steps such as separation, fragmentation, first-strand cDNA synthesis, second-strand cDNA synthesis, end repair, 3'-end adenylation, adapter ligation, and enrichment on the purified total RNA to complete the construction of the mRNA sequencing library. Use a Qubit 2.0 fluorometer (Thermo Fisher Scientific) to detect the library concentration, and an Agilent 4200 TapeStation (Agilent technologies) to detect the library fragment distribution. Finally, perform sequencing according to the effective concentration of the library and the data output requirements. The sequencing platform uses an Illumina NovaSeq 6000, and the sequencing mode is PE150 (Pair-end 150bp), that is, 150bp is sequenced for each end in paired-end sequencing.
[0095] Example 2: Identification of DEGs and GSEA enrichment analysis in three groups of people
[0096] 1. Differential analysis
[0097] To identify differentially expressed genes (DEGs) among the 50,869 genes obtained from sequencing, we used the R software package "DESeq2" for DEG identification, along with background correction and normalization to ensure data comparability and reliability. Based on the set statistical criteria (Fold Change ≥ 2 and P-value < 0.05), "DESeq2" calculated the expression differences of each gene between LC-PTB or LC patients and control samples and determined which genes met the conditions for significant differences. To visually display the analysis results of DEGs, we used the "ggplot2" package to draw volcano plots, the GraphPad Prism software (version: 10.0.0, San Diego, California, USA) to draw gene expression level plots, and performed correlation analysis on gene expression profiles for visualization.
[0098] 2. Results
[0099] Our sequencing data included 10 LC-PTB patients, 10 LC patients, and 10 control subjects, and a total of 50,869 genes were detected. Through R language analysis, it was found that compared with the Control group, there were 305 upregulated genes and 435 downregulated genes in the LC group; while in the LC-PTB group, there were 144 upregulated genes and 435 downregulated genes. In addition, compared with the LC group, there were 216 upregulated genes and 388 downregulated genes in the LC-PTB group. These changes in gene expression profiles were shown by volcano plots and heatmaps ( Figure 2 ).
[0100] In the volcano plot ( Figure 2 a)), green represents downregulated differential genes, and red represents upregulated differential genes. Among them, compared with the Control group, the top six significantly different genes in the LC group were: AGAP9, AC022392.1, TRAV23DV6, PPIL6, AC021087.1, and TP53I3; the top five significantly different genes in the LC-PTB group were: SDR16C5, ADTRP, MAP1S, AC011462.1, and HLA-DRB; while compared with the LC group, the top six significantly different genes in the LC-PTB group were: HBA1, INKA2-AS1, RBM38, TMCC2, CCDC144B, and RGS13. The heatmap ( Figure 2 b)) visually demonstrated the clustering of DEGs in Control, LC, and LC-PTB group patients.
[0101] In addition, we also systematically analyzed the functional characteristics of DEGs in each group and their potential roles in the pathological process through GSEA ( Figure 3) The results showed that compared with the Control group, the top five pathways of DEGs in the LC group were neutrophil degranulation, oxidative phosphorylation, MHC class II antigen presentation, DNA replication, and cellular response to chemical stress, etc.; compared with the Control group, the DEGs of the LC-PTB group were mainly enriched in five pathways including antigen processing and presentation, translation, Huntington's disease, oxidative phosphorylation, and red blood cell uptake of carbon dioxide and release of oxygen; compared with the LC group, the related pathways enriched by the DEGs of the LC-PTB group patients (including histone deacetylation, RUNX1-regulated megakaryocyte differentiation and platelet function-related genes, and DNA methylation, etc.) were all down-regulated in the analyzed samples and also showed significant negative enrichment scores in the dataset. In the context of LC-PTB, the significant negative enrichment of RUNX1-related genes indicates that this pathway may be inhibited during the occurrence and development of the disease, thus affecting platelet production and function, and further affecting the immune response and inflammatory process.
[0102] Through the comparative analysis of the three GSEA plots, we found that compared with Control, a significant positive enrichment of the oxidative phosphorylation pathway was observed in both the LC group and the LC-PTB group. These analyses enabled us to understand the molecular-level complexity of LC-PTB disease and also revealed the different adaptation mechanisms of tumor cells and host cells in aspects such as immune response, energy metabolism, epigenetic regulation, and viral infection under different pathological conditions. These findings provided important clues for understanding the molecular mechanism of LC-PTB.
[0103] Example 3. DEGs of LC-PTB patients and GO and KEGG enrichment analysis of DEGs
[0104] To further explore the LC-PTB-related characteristic genes in LC patients, we first used Venn Diagram to display and compare the co-existing differential genes among the three groups ( Figure 4)。To deeply understand the roles of these intersecting DEGs in the pathogenesis of LC-PTB, we performed Gene Ontology (GO) enrichment analysis and Kyoto Encyclopedia of Genes and Genomes (KEGG) pathway analysis on the three groups of intersecting DEGs. GO and KEGG analyses are two indispensable gene function analysis tools, which together provide a systematic and comprehensive perspective for revealing the gene functions and key biological regulatory networks closely related to LC-PTB. GO enrichment analysis aims to identify the significant enrichment of these genes in three dimensions: Biological Process (BP), Cellular Component (CC), and Molecular Function (MF) by comparing the gene list in experimental data with the entries in the GO database. KEGG pathway analysis examines the molecular mechanism of LC-PTB from another angle, focusing on analyzing how these genes participate in complex biological metabolic pathways and signal transduction networks.
[0105] We used GO enrichment analysis to deeply explore the biological functions and potential pathogenesis of these DEGs. After strictly screening GO terms with a P-value less than 0.05, we sorted and analyzed them in order according to the number of genes involved ( Figure 5 ).
[0106] Compared with the Control group, there were 1052 GO entries in the enriched Biological Process (BP) category of the intersecting DEGs in the LC group and the LC-PTB group. These items widely covered many key areas such as glomerular basement membrane development, glomerular development, positive regulation of neurotransmitter secretion, response to glucose, intermediate filament organization, response to monosaccharides, establishment of mitochondrial localization, and regulation of ventricular cardiomyocyte membrane repolarization; in the Cellular Component (CC) category, 150 enriched entries were found, mainly concentrated on structures such as basement membrane, collagen trimer complex, intermediate filaments, keratin fibers, hippocampal mossy fiber to CA3 synapse, cell-cell junction, membrane raft, and membrane microdomain; in the Molecular Function (MF) category, we found 160 significantly enriched GO entries, including bioactive lipid receptor activity, extracellular matrix structural components conferring tensile strength, Ras-nucleotide exchange factor activity, cell-cell adhesion mediator activity, guanyl-nucleotide exchange factor activity, α-catenin binding, 1-acyl-2-lyso-phosphatidylserine acylhydrolase activity, and structural components of the cytoskeleton ( Figure 5 a). In addition, through KEGG enrichment analysis ( Figure 5In d), we further confirmed that these intersecting DEGs were closely associated with extracellular matrix-receptor interaction, cancer pathways, cell adhesion molecules, and the PI3K-Akt signaling pathway. These results suggest that these differential genes play important roles in the occurrence and development of LC-PTB. In particular, the significant enrichment of the extracellular matrix-receptor interaction and cell adhesion molecule pathways suggests that these genes may promote the invasion and metastasis of tumor cells by affecting the interaction between cells and the extracellular environment.
[0107] Among the intersecting DEGs of the LC-PTB group compared with the Control group and the LC group, we found 395 significantly enriched BP categories, 94 significantly enriched CC categories, and 98 significantly enriched MF category GO terms ( Figure 5 In b). Among them, in terms of BP, there was significant enrichment in biological functions such as response to chemical substances, cell differentiation, nuclear part, cell developmental process, regulation of gene expression, and nuclear protein metabolic process; in terms of CC, there was significant enrichment in cytoskeleton, regulation of biological quality, cell macromolecule biosynthesis process, phylogeny, and DNA metabolic process; in terms of MF, it included cation binding, transport, catalytic activity, cell nitrogen compound biosynthesis process, and protein-containing complexes. Meanwhile, KEGG enrichment analysis showed that the intersecting DEGs of this group were related to cytoskeleton regulation, phosphatidylinositol signaling system, primary immunodeficiency, shigellosis, AGE-RAGE signaling pathway, inositol phosphate metabolism, glycerophospholipid metabolism, and complement and coagulation cascades ( Figure 5 In e).
[0108] Finally, among the intersecting DEGs of the LC group compared with the Control group and the LC-PTB group, platelet aggregation, blood coagulation, regulation of immune response, cytokine production, and food transport were significantly enriched in the BP category; ribosome, ribosome structural component, autophagosome, and cell signaling were significantly enriched in the CC category; the MF category included molecular functions such as RNA binding, sulfation, oxygen transport, cell-cell adhesion, glycosyltransferase activity, and transferase activity ( Figure 5 In c). In addition, in the KEGG enrichment analysis, metabolic pathways such as nitrogen metabolism, alanine, aspartate, and glutamate metabolism, glycosphingolipid biosynthesis, and choline metabolism were significantly enriched ( Figure 5 In f).
[0109] Example 4, PPI network, WGCNA, and immune gene set screening of LC-PTB immune-related genes
[0110] 1. Screening of LC-PTB core genes by PPI network
[0111] To further screen for the core genes among these intersecting DEGs, we constructed PPI networks for the three groups of intersecting DEGs. This network visually presented the interaction relationships between these genes and proteins, thereby deepening our understanding of their potential roles and biological significance in the occurrence and development of LC and LC-PTB ( Figure 6 as shown in a, b, and c in Figure 6 Figure 4d). To further screen for characteristic biomarkers, we also used the MCODE plugin to identify relevant hub genes. Among them, compared with the Control group ( Figure 6 as shown in Figure 4e), six hub genes (RAPGEF4, NEFM, ANK2, SDR16C5, KRT7, and RET) were screened out from the intersecting DEGs of the LC group and the LC-PTB group; in the visualization graph of the hub genes of the intersecting DEGs of the LC-PTB group with the Control group and the LC group ( Figure 6 as shown in Figure 4f), we found 15 hub genes including PFN4, TMOD1, PRKY, OR2W3, USP9Y, OR2L13, RPS4Y1, ZFY, OR2A7, TMSB4Y, EIF1AY, KDM5D, UTY, and NLGN4Y, as well as DDX3Y; while Figure 7 as shown in Figure 5a, Figure 7 as shown in Figure 5b).
[0112] 2. Screening for LC-PTB core genes by WGCNA
[0113] To gain an in-depth understanding of the close relationship between DEGs and diseases and further screen for core genes, we employed WGCNA. WGCNA is a method for studying gene correlations in different samples. It can identify genes that are highly coordinated in expression and help discover core genes related to phenotypes. In this invention, we used the R package "WGCNA" to identify gene modules related to LC-PTB immunity. To construct a suitable weighted gene co-expression network, we adopted the soft-threshold conversion method. By evaluating the scale-free characteristics of the network at different β values, we finally determined the optimal soft-threshold parameter β = 12. In the module detection stage, we combined the hierarchical clustering algorithm and the dynamic tree-cutting method. Using the topological overlap degree between genes as a metric, we clustered genes with similar expression profiles into different modules. To balance the resolution and stability of the modules, we set the minimum genome size of the gene dendrogram to 30 and adjusted the sensitivity parameter to 3. To optimize the module structure, we merged modules with a distance less than 0.25 according to the cutting line of the module dendrogram to reduce redundancy and enhance the biological significance of the modules.
[0114] We used WGCNA to construct a co-expression network. Through network topology analysis, we determined the optimal soft-threshold parameter β = 12. At this time, the scale-free R 2 value of the network reached 0.84, indicating that the network has a good scale-free topological structure and is suitable for subsequent module detection analysis. At the same time, the average connectivity of the network also shows the stability of the network ([ Figure 7 in c and Figure 7 in d). Finally, the WGCNA model divided 11,994 genes into 11 modules ([ Figure 7 in d). Among them, the grey module (Grey) is a set of genes that cannot be assigned to any module and has no reference significance. However, when we explored the relationship between the modules and clinical phenotype data in detail, we found that grey60 was significantly negatively correlated with lymphocytes (LYM) (R = -0.6, P < 0.0001); black was positively correlated with LYM (R = 0.37, P < 0.05) ([ Figure 7 in e). This suggests that the genes in these two modules may affect the immune response process of LC-PTB by regulating the function of lymphocytes.
[0115] 3. Screening for LC-PTB Immunity-Related Genes by ImmPort
[0116] Given that the occurrence and development of LC-PTB are inseparable from the functional status of the immune system. Therefore, it is crucial to deeply understand the underlying immune regulatory mechanisms. Thus, we obtained a comprehensive dataset of immune-related genes through the ImmPort database (https: / / www.immport.org / shared). ImmPort is an immunology database maintained by authoritative institutions and widely recognized. It aggregates high-quality immunology data from research institutions around the world, including gene expression profiles, proteomics, clinical information, etc., providing valuable resources for scientific researchers. Subsequently, in order to precisely lock the target genes directly related to the immune regulation of LC-PTB, we performed an intersection analysis of the immune-related gene dataset obtained from ImmPort with the key gene lists identified by WGCNA and PPI network.
[0117] We used the 6 core genes screened by the PPI network as a reference and cross-compared them with 1793 immune genes in the ImmPort database and the key module genes screened by the WGCNA network ( Figure 7 in f, Figure 7 in g). Finally, 6 key genes (ANK2, EPB42, CA1, HBB, HBD, and MYL4) were identified. These genes may play a key role in the pathogenesis of LC-PTB. They not only have the potential to be biomarkers but also provide new perspectives and clues for our in-depth understanding of the immune pathogenesis of LC-PTB.
[0118] Example 5: Construction and verification of the early risk warning model for LC-PTB
[0119] 1. Construction of the early risk warning model for LC-PTB
[0120] To more precisely conduct early risk warning for the development of LC into LC-PTB, we constructed an LC-PTB risk warning model based on four of the six identified core genes (ANK2, EPB42, CA1, HBB, HBD, and MYL4). The model constructed by ANK2, CA1, HBD, and MYL4 was developed using binary logistic regression analysis. In logistic regression, we used logit(p) to model the relationship between the dependent variable and the independent variables.
[0121] logit(p) represents the log odds of the event occurring. Its definition is: where p is the probability of the event occurring.
[0122] The form of the logistic regression model is: logit(p) = β0 + β1X1 + β2X2 + … + β n X n,where β is the regression coefficient and X is the independent variable. The constructed LC-PTB early risk warning model is as follows:
[0123] logit(p) = 2.022 + 416.907*ANK2 - 358*CA1 - 626*HBD - 884*MYL4
[0124] where logit(p) represents the risk score; ANK2 represents the ANK2 gene expression level; CA1 represents the CA1 gene expression level; HBD represents the HBD gene expression level; MYL4 represents the MYL4 gene expression level.
[0125]
[0126] The combine formula is used to convert the linear predicted value into a probability value. In this way, a probability value between 0 and 1 can be obtained, representing the probability of the occurrence of lung cancer complicated with pulmonary tuberculosis.
[0127] The results showed that ANK2 (OR: NA, CI: 0.000–inf; P = 0.7419), CA1 (OR: 0.001, CI: 0.000–1.816; P = 0.0718), HBD (OR: 0.394, CI: 0.121–1.286; P = 0.1226), MYL4 (OR: 0.009, CI: 0.010–0.798; P = 0.0306) were important differential genes between LC-PTB and lung cancer (Table 1). To further visually display the prediction efficacy of this model, we visualized the model through a nomogram ( Figure 8 in a).
[0128] Table 1. Logistic regression analysis of four hub differentially expressed genes
[0129]
[0130] Next, to evaluate the accuracy of this risk warning model in predicting LC-PTB, we verified its performance using the ROC curve. First, the ROC analysis results of the four genes showed that the AUC of ANK2 was 0.540 (CI: 0.000-Inf), the AUC of CA1 was 0.940 (CI: 0.000-1.186), the AUC of HBD was 0.930 (CI: 0.121-1.286), and the AUC of MYL4 was 0.880 (CI: 0.010-0.798)( Figure 8 in b). The overall AUC of this model was 0.94 (CI: 0.846-1.000). When the Youden index was 0.7, the sensitivity was 0.9 and the specificity was 0.8( Figure 8In c). These data indicate that the early warning ability of the four-gene model is significantly better than that of the single-gene model.
[0131] 2. RT-qPCR verification of the LC-PTB early risk warning model
[0132] To verify the expression of the above six core genes, we recruited 88 subjects, including 28 LC-PTB patients, 30 LC patients, and 30 healthy controls, for a prospective cohort study. This study has been approved by the Medical Ethics Committee of the Eighth Medical Center of the Chinese PLA General Hospital (Ethical Approval Number: 3092023122013297233), and all participants have signed informed consent forms.
[0133] Control inclusion criteria: ① Age: Age ≥ 18 years old; ② Medical history: No history of malignant tumors, TB, HIV positivity, or autoimmune diseases, and no history of immunotherapy or glucocorticoid therapy in the past 3 months.
[0134] Control exclusion criteria: ① Malignant tumors: Primary or secondary malignant tumors (especially LC); ② Active infections: Complicated with active bacterial, fungal, or viral infections (especially MTB infection); ③ Underlying diseases: Chronic obstructive pulmonary disease, diabetes, and other underlying diseases, etc.
[0135] LC inclusion criteria: ① Age and medical history: Age ≥ 18 years old, with a smoking history (≥ 20 pack-years) or high-risk factors (such as occupational carcinogen exposure, family history of lung cancer, etc.); ② Pathological diagnosis: Definitively diagnosed as LC by tissue biopsy or cytological examination, including non-small cell LC (Non-small cell lung cancer, NSCLC) and small cell LC (Small cell lung cancer, SCLC); ③ Imaging features: Chest CT shows a pulmonary space-occupying lesion, with or without mediastinal lymph node enlargement, pleural invasion, or distant metastasis; ④ Clear staging: Divided into stages Ⅰ - Ⅳ according to the TNM staging system (9th edition).
[0136] LC exclusion criteria: ① Other malignant tumors: Complicated with other primary malignant tumors (such as breast cancer, colorectal cancer, etc.); ② Active infections: Complicated with active bacterial, fungal, or viral infections (such as pneumonia, HIV infection); ③ Severe complications: Presence of severe heart, liver, and kidney insufficiency or coagulation dysfunction; ④ Treatment contraindications: Unable to tolerate standard treatments such as surgery, radiotherapy, or chemotherapy.
[0137] Inclusion criteria for LC-PTB: ① Dual diagnostic basis: LC should meet the above inclusion criteria, and PTB should meet the diagnostic basis in the "Uniform Diagnostic Criteria for PTB" (WS288-2017), including positive sputum smear or culture, typical imaging features (such as miliary nodules, cavity formation) or pathological confirmation of tuberculous granuloma; ② Coexistence evidence: The foci of PTB and LC are located in the same lung lobe or there is a pathological association (such as canceration within a tuberculous scar); ③ Temporal association: Active PTB or a previous history of TB (≥1 year), and the time interval from the diagnosis of LC is ≤5 years; ④ Symptom overlap: Simultaneously have PTB symptoms (such as low fever, night sweats, hemoptysis) and LC symptoms (such as irritating dry cough, weight loss).
[0138] Exclusion criteria for LC-PTB: ① Single disease: Only diagnosed with LC or PTB, without coexistence evidence; ② Non-tuberculous mycobacterial infection: Sputum culture or gene detection confirms non-tuberculous mycobacterial (NTM) infection; ③ Other pulmonary diseases: Complicated with pulmonary embolism, pulmonary fibrosis or other chronic pulmonary diseases; ④ Immunosuppressed state: Long-term use of immunosuppressants or presence of primary immunodeficiency diseases.
[0139] Detection of gene expression levels: Collect 5 ml of peripheral venous blood, and use the IVD Pure Total Blood RNA Extraction Kit (IVD, China) to extract total RNA from whole blood samples according to the manufacturer's instructions. Then, use the FastKing gDNA Dispelling RT SuperMix Kit (Tiangen, China) for reverse transcription, incubate at 42 °C for 15 minutes, and then incubate at 95 °C for 3 minutes. After that, dilute the sample with RNase-free water. Use the FastReal qPCR PreMix (SYBR Green) (Tiangen, China) kit and the Roche 480 system for RT-qPCR. The reaction conditions are: pre-denaturation (95 °C, 2 minutes), 40 denaturation cycles (95 °C, 5 seconds), annealing and extension (60 °C, 30 seconds). Glyceraldehyde-3-phosphate dehydrogenase (GAPDH) is used as an internal control for amplification. The relative expression level is measured by the 2 -ΔΔCt method. The primer sequences used in this experiment are shown in Table 2 in detail.
[0140] Table 2. Details of the primer sequences of each gene symbol and its internal reference gene
[0141]
[0142]
[0143] The results showed that six genes presented a changing trend consistent with the sequencing data ( Figure 9 ).
[0144] 3. Validation of the LC-PTB Early Risk Warning Model through a Cohort Study
[0145] To further validate the efficacy of the early risk warning model, we used the RT-qPCR experimental data of 88 patients we collected for model validation.
[0146] The results showed that the efficacy of the four genes in warning LC-PTB individually was as follows: the AUC of ANK2 was 0.555 (CI: 0.468 - 0.641), the AUC of CA1 was 0.674 (CI: 0.592 - 0.757), the AUC of HBD was 0.659 (CI: 0.576 - 0.741), and the AUC of MYL4 was 0.703 (CI: 0.623 - 0.782)( Figure 10 ). The bar chart showed that there were no significant differences in ANK2, CA1, and HBD among the Control group, LC group, and LC-PTB group, but they showed a changing trend consistent with the sequencing results( Figure 9 in b)). And the AUC of the four genes in warning LC-PTB as a whole was 0.715 (CI: 0.637 - 0.793)( Figure 10 ). When the Youden index was 0.38, the sensitivity of this model was 0.69 and the specificity was 0.69.
[0147] 4. Immunohistochemical Validation of the LC-PTB Early Risk Warning Model
[0148] Finally, we verified the reliability of the bioinformatics prediction results through immunohistochemical detection. A total of 30 subjects from the Eighth Medical Center of the Chinese PLA General Hospital were recruited (different from the blood samples in Example 1, they were pathological tissue samples), including 10 LC-PTB patients, 10 LC patients, and 10 healthy controls. The inclusion and exclusion criteria for each group of people were as described above. The specific operations were as follows: ① Cut 15 consecutive sections from the pathological tissue wax block of each patient and place them in water at 42°C for spreading; ② Dewaxing and hydration: Before dewaxing, bake the tissue sections at 65°C for 2 hours or overnight in an incubator at 56°C, and soak the sections in xylene, absolute ethanol, 95% ethanol, and 75% ethanol in sequence; ③ Remove endogenous peroxidase: Place the sections in freshly prepared 3% H2O2 and incubate at room temperature for 10 min, then wash the sections thoroughly with water and wash them 2-3 times with 0.01M PBS for 3 min each; ④ Antigen repair: Place the tissue sections in 2L of 0.01M sodium citrate buffer solution for high-pressure repair, and then wash the sections thoroughly with water; ⑤ Block non-specific staining with goat serum and incubate at room temperature for 30 min; ⑥ Dropwise add primary antibody (CD3, CD4, CD8, CD20, CD56, CD68, ANK2, CA1, EPB42, HBB, HBD, and MYL4), incubate overnight at 4°C or for 1 hour at 37°C, and then wash thoroughly with water; ⑦ Dropwise add secondary antibody (enzyme-labeled goat anti-mouse / rabbit antibody), incubate at room temperature for 30 min, and then wash thoroughly with water; ⑧ DAB color development for 1-10 min; ⑨ Hematoxylin staining for 5 min and then wash thoroughly with water; ⑩ Dehydration, clearing, and mounting. Finally, place the prepared sections in a panoramic slide scanner for panoramic scanning, and use Qupath software (v0.5.1, https: / / qupath.github.io / ) to analyze the number of positive cells and the positive rate.
[0149] The results showed that there were significant differences in ANK2, CA1, and EPB42 between the LC group and the LC-PTB group (P = 0.0481, P = 0.0412, P = 0.0105), while there were significant differences in HBB between the Control group and the LC-PTB group (P = 0.0082)( Figure 11 ).
[0150] The above results indicate that the six core genes of the present invention (ANK2, EPB42, CA1, HBB, HBD, and MYL4) can be used as biomarkers for early risk warning of LC-PTB. The early risk warning model of LC-PTB constructed based on four of these core genes has good accuracy, providing a new method for early risk warning of LC-PTB.
[0151] The present invention has been described in detail above. For those skilled in the art, without departing from the spirit and scope of the present invention and without unnecessary experiments, the present invention can be implemented within a relatively wide range under equivalent parameters, concentrations and conditions. Although specific embodiments of the present invention are given, it should be understood that the present invention can be further improved. In short, according to the principle of the present invention, this application intends to cover any modifications, uses or improvements of the present invention, including those that depart from the scope disclosed in this application and are made by conventional techniques known in the art.
Claims
1. Use of a biomarker and / or a substance for detecting the biomarker in any of the following: A1) Use in the preparation of a product for early risk warning of lung cancer complicated with pulmonary tuberculosis; A2) Use in the construction of an early risk warning model for lung cancer complicated with pulmonary tuberculosis; The biomarker is any one of the following: B1) ANK2 gene, CA1 gene, EPB42 gene, HBB gene, HBD gene and / or MYL4 gene; B2) ANK2 gene, CA1 gene, HBD gene and MYL4 gene.
2. The application according to claim 1, wherein The substance includes reagents and / or instruments for detecting the expression levels of ANK2 gene, CA1 gene, EPB42 gene, HBB gene, HBD gene and / or MYL4 gene.
3. The application according to claim 2, wherein The reagent includes primers for specifically amplifying ANK2 gene, CA1 gene, EPB42 gene, HBB gene, HBD gene and / or MYL4 gene and / or probes for specifically recognizing ANK2 gene, CA1 gene, EPB42 gene, HBB gene, HBD gene and / or MYL4 gene.
4. A composition or kit for early risk warning of lung cancer complicated with pulmonary tuberculosis, characterized in that, The composition or kit includes reagents for detecting the expression levels of ANK2 gene, CA1 gene, EPB42 gene, HBB gene, HBD gene and / or MYL4 gene.
5. Early risk warning model for lung cancer complicated with pulmonary tuberculosis, characterized in that, The model is constructed using the biomarker described in claim 1.
6. The model according to claim 5, wherein The formula of the model is: logit(p) = 2.022 + 416.907*ANK2 - 358*CA1 - 626*HBD - 884*MYL4 where logit(p) represents the risk score; ANK2 represents the expression level of ANK2 gene; CA1 represents the expression level of CA1 gene; HBD represents the expression level of HBD gene; MYL4 represents the expression level of MYL4 gene.
7. Use of the model according to claim 5 or 6 in the preparation of a product for early risk warning of lung cancer complicated with pulmonary tuberculosis.
8. A method for constructing an early risk warning model for lung cancer complicated with pulmonary tuberculosis, characterized in that, The method includes using the data of the expression levels of ANK2 gene, CA1 gene, EPB42 gene, HBB gene, HBD gene and / or MYL4 gene in the samples of known patients with lung cancer complicated with pulmonary tuberculosis and lung cancer patients as training samples, and adopting statistical analysis methods to establish an early risk warning model for lung cancer complicated with pulmonary tuberculosis.
9. The method according to claim 8, characterized in that, The statistical analysis method includes binary logistic regression analysis.
10. A system for early risk warning of lung cancer complicated with pulmonary tuberculosis, characterized in that, The system includes: A data receiving module for receiving, from at least one terminal, data on the expression level of the biomarker described in claim 1 in a sample of a subject to be tested; A data processing module for inputting the data into the model according to claim 5 or 6, or into the model constructed by the method according to claim 8 or 9, to obtain a risk score for lung cancer complicated with pulmonary tuberculosis; A data output module for outputting the risk score to at least one client; or for converting the risk score into a probability value of the occurrence of lung cancer complicated with pulmonary tuberculosis and then outputting it to at least one client.