Visualized online prediction system for immune-related pneumonia and construction method thereof
Patent Information
- Application Number
- CN202311214027.6
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2023-09-19
- Publication Date
- 2026-09-18
- Estimated Expiration
- 2043-09-19
AI Technical Summary
[0007]但是,目前尚无任何基于蛋白组学质谱分析技术获得高特异性的预测CIP发生的生物标志物的研究报道
[0042] This invention provides a visualized online prediction system for CIP (Chronic Illness Infection) based on a combination of five specific predictive proteins obtained from serum proteomics mass spectrometry analysis. The five specific predictive proteins are: SERPINA5, AKAP6, TUBA4A, COL1A2, and Pro-SHAP. The expression levels of these differentially expressed proteins in the experimental and control groups were quantitatively detected using ELISA. The results showed significant differences in the expression of SERPINA5, AKAP6, TUBA4A, COL1A2, and Pro-SHAP between the experimental and control groups, with all five proteins highly expressed in CIP patients. The predictive accuracy of most of these five proteins exceeded 85%, making them suitable as specific biomarkers for predicting CIP. Based on these results, this invention also utilizes R language, with SERPINA5, AKAP6, TUBA4A, COL1A2, and Pro-SHAP as modeling elements, and employs machine learning algorithms to develop a visualized online prediction system for immune-associated pneumonia. The system achieved an AUC of 0.996, with specificity and sensitivity of 97.6% and 95.5%, respectively. The model validation set was then formed using 21 newly included patients (control group: 12 patients; experimental group: 9 patients), and the AUC of the ROC curve of the validation set reached 0.968.
Smart Images

Figure CN117275588B_ABST
Abstract
Description
Technical Field
[0001] This invention belongs to the field of biomedical technology, specifically relating to a visual online prediction system for immune-related pneumonia and its construction method. Background Technology
[0002] Malignant tumors have always been an intractable problem in the history of human medicine. With rapid population growth and increasing aging, the global incidence and mortality rates of malignant tumors are rising year by year, becoming a major disease burden on humanity. With the development of human medicine, various new anti-tumor treatments have rapidly emerged, such as targeted therapy and immunotherapy, significantly improving humanity's ability to fight tumors. Among them, immunotherapy has ushered in a new era in tumor treatment. Since the first CTLA-4 antibody—ipilimumab—was approved for the treatment of malignant melanoma in 2011, the field of tumor immunotherapy has experienced one great renaissance after another (Nat Rev Drug Discov, 2016, 15:235-247.). In recent years, numerous clinical trials of drugs have been widely conducted globally, demonstrating significant anti-tumor efficacy. Immune checkpoint inhibitors (ICIs), represented by PD-1 / PD-L1 inhibitors, have been widely used in various tumors, achieving satisfactory therapeutic effects and ushering in a new era of anti-tumor therapy, such as malignant melanoma (N Engl J Med, 2017, 377: 1345-1356.), renal cell carcinoma (N Engl J Med, 2019, 380: 1116-1127.), and non-small cell lung cancer (N Engl J Med, 2018, 379: 2040-2051.).
[0003] However, the enthusiasm for immunotherapy is largely based on its long-term clinical benefits, which only occur in a small number of patients—approximately 20%-40% of patients are sensitive to immunotherapy (J ClinOncol, 2019, 37:2518-2527). Furthermore, the clinical benefits of immunotherapy depend not only on its efficacy but also on its unique adverse reactions related to its mechanism of action. Severe immune-related adverse reactions significantly hinder the clinical application of immunotherapy, among which checkpoint immune pneumonitis (CIP) is one of the most common and deadliest adverse reactions (JAMA Oncol, 2016, 2:1346-1353). Previous reports and preliminary research by the inventors' research group indicate that the incidence of CIP can reach 10%-20% (J Thorac Oncol, 2018, 13:1930-1939.), and the mortality rate of severe CIP is as high as 14%-35% (JAMA Oncol, 2018, 4:1721-1728.). Severe CIP not only causes irreversible damage to patients and shortens their survival time, but also incurs huge costs and consumes a large amount of medical resources. Imaging examination is considered the gold standard for diagnosing CIP, but the imaging changes in CIP are delayed, complex, and have poor specificity. In addition, the clinical symptoms of early CIP are not obvious, making it difficult to differentiate from other respiratory diseases, which easily leads to missed diagnosis or progression to severe CIP and endangering life. Therefore, early screening of high-risk CIP patients and actively exploring new approaches with higher specificity and stronger predictive efficacy are urgent and crucial for reducing the incidence and mortality of CIP.
[0004] Multiple studies have attempted to explore the risk factors and biomarkers of chronic lung injury (CIP). Previous studies have shown that male sex, older age, smoking, a history of high-dose chemotherapy, a history of combination therapy, a pre-existing interstitial lung disease, and emphysema are more likely to develop CIP (JAMA Oncol, 2018, 4:1721-1728.; J Thorac Oncol, 2018, 13:1930-1939.; J Thorac Oncol, 2018, 13:1138-1145.; Cancer Immunol Immunother, 2020, 69:15-22.). Other studies have shown that an increase in baseline absolute eosinophil count and neutrophil-lymphocyte ratio during treatment also indicates a higher risk of CIP (Lung Cancer, 2020, 150:76-82.; Sci Rep, 2021, 11:1324.). While the aforementioned factors can predict CIP occurrence to some extent, each single-dimensional prediction has its own bias, and current research on predictive biomarkers for CIP is limited to clinical characteristics and non-specific hematological indicators, lacking research on specific proteomics. Furthermore, the area under the ROC curve for identifying CIP is only about 0.60-0.70 (Lung Cancer, 2020, 150:76-82.; Clin Lung Cancer, 2019, 20:442-450.e4.), exhibiting drawbacks such as poor accuracy, specificity, and stability, failing to achieve the goal of "precise prediction," and lacking strong evidence-based medicine support. To overcome the shortcomings of these macroscopic characteristics, developing more accurate and effective biomarkers and predictive models based on microomics is an effective method for screening high-risk populations and providing early warning of CIP.
[0005] The mechanism of CIP (collateral pulmonary inflammatory response) is currently unclear, but it is the result of multiple factors. An immunological mechanism primarily based on biochemical signals is considered a classic mechanistic hypothesis. Studies have shown that CIP may be due to the presence of common antigens between tumors and inflammatory organs, leading to collateral damage to inflammatory organs in the lungs during the antitumor process of ICIs (N Engl J Med, 2016, 375:1749-1755; JAMA Oncol, 2019, 5:1043-1047.). Furthermore, the infiltration of inflammatory cells in the lungs by autoreactive T cells, autoantibodies, and various inflammatory cytokines generated by antitumor responses is also a potential mechanism for CIP (NEngl J Med, 2018, 378:158-168.). Suresh et al. found a significant increase in CD4+ T cells in the bronchoalveolar lavage fluid of CIP patients (J Clin Invest, 2019, 129:4305-4315.). Suzuki et al. reported a significant increase in the proportion of PD-1, Tim-3, and Tighit-positive CD8+ T cells in the bronchoalveolar lavage fluid (BALF) of CIP patients (Int Immunol, 2020, 32:547-557.). Similarly, Franken et al. also found enrichment of CD4+ and CD8+ T cells in BALF, especially an increase in pathogenic T helper 17.1 cells, with high expression of TBX21 and RORC, IFN-G, IL-17A, CSF2 (GM-CSF), and cytotoxicity-related genes (J Immunother Cancer, 2022, 10:undefined.). In summary, immune dysregulation and excessive immune responses caused by ICIs damage normal tissues, leading to the occurrence and development of CIP. However, biochemical mechanisms cannot fully explain the occurrence of CIP; as a typical fibrotic disease, mechanobiological mechanisms may also play a crucial role. Fibroblast activation, ECM remodeling, and endothelial cell dysfunction are common pathological changes in the formation of fibrosis across different tissue types (Cells, 2021, 10:undefined.). The alveolar extracellular matrix (ECM) is a key mediator of mechanical responses. Pathological changes in the ECM disrupt mechanical homeostasis; increased collagen secretion leads to increased tissue stiffness, which can mediate pulmonary fibrosis (Elife, 2018, 7:undefined.). Furthermore, studies have indicated that increased pulmonary mechanical tension activates the TGF-β signaling cycle in alveolar stem cells, thereby driving the progression of pulmonary fibrosis (Cell, 2021, 184:845-846.). In conclusion, mechanobiology may be involved in the development of CIP, but further research is needed.
[0006] Mass spectrometry, a rapidly developing analytical technique in recent years, is considered one of the most important analytical technologies in protein and biomolecular research. It boasts significant advantages such as high sensitivity, accuracy, and resolution. It can detect not only the molecular weight and amino acid sequence of proteins and peptides but also protein binding sites and post-translational modifications. This allows for large-scale protein identification and assessment of protein expression levels, contributing to a deeper understanding of cellular and molecular mechanisms and providing new biomarkers for disease and drug treatment. In clinical applications, its more reliable, sensitive, and specific serum biomarker screening capabilities have propelled proteomics-based clinical serum biomarker screening to the forefront, becoming a new trend in precision medicine. This suggests that proteomics-based mass spectrometry analysis has the potential to become a novel detection method for predicting CIP (Chronic Inflammatory Disease). Utilizing proteomics-based mass spectrometry to screen for quantifiable biological indicators and biomarkers of CIP is not only beneficial for early screening and diagnosis but also crucial for exploring the pathogenesis of the disease, establishing diagnostic criteria, and developing therapeutic drugs.
[0007] However, there are currently no reports of studies that have obtained highly specific biomarkers for predicting CIP based on proteomics mass spectrometry analysis. Summary of the Invention
[0008] In order to overcome the shortcomings of the prior art, the present invention aims to provide a visual online prediction system for immune-related pneumonia and a method for constructing the system.
[0009] To achieve the above objectives, the present invention employs the following technical solution:
[0010] This invention discloses a visualized online prediction system for immune-related pneumonia, comprising:
[0011] The serum protein biomarker expression level detection module is used to detect the expression level of protein biomarkers in the serum of the test subject;
[0012] The intelligent prediction module is used to calculate the probability of a subject developing immune-related pneumonia based on the expression level of protein biomarkers using an intelligent prediction model.
[0013] The results output module is used to output results through a visual online interactive interface.
[0014] Preferably, the protein markers in the serum include SERPINA5, AKAP6, TUBA4A, COL1A2, and Pro-SHAP.
[0015] More preferably, the intelligent prediction module further includes:
[0016] The differential protein acquisition unit is used to acquire the expression levels of differentially expressed proteins in serum samples.
[0017] The intelligent prediction model construction unit plots the ROC curve of differentially expressed proteins based on their expression levels and calculates the optimal cutoff value for the differentially expressed proteins based on the maximum Yoden index. Then, it classifies the expression levels of serum samples based on the optimal cutoff value, constructs a Logistic regression equation, and obtains the intelligent prediction model.
[0018] Preferably, the collected serum samples are obtained by mass spectrometry to obtain differential peaks using primary mass spectrometry, the differential peaks are identified by secondary mass spectrometry, and the expression level of differentially expressed proteins is verified by ELISA.
[0019] Preferably, the optimal cutoff value for the serum protein marker SERPINA5 in predicting immune-related pneumonia is 642.27 pg / mL, with an AUC of 0.908 for the ROC curve and p < 0.0001.
[0020] The optimal cutoff value for the serum protein marker AKAP6 in predicting immune-related pneumonia was 1050.88 pg / mL, with an AUC of 0.857 for the ROC curve and p < 0.0001.
[0021] The optimal cutoff value for the serum protein marker TUBA4A to predict immune-related pneumonia was 226.32 ng / L, with an AUC of 0.874 for the ROC curve and p < 0.0001.
[0022] The optimal cutoff value for the serum protein marker COL1A2 in predicting immune-related pneumonia was 894.83 pg / mL, with an AUC of 0.825 for the ROC curve and p < 0.0001.
[0023] The optimal cutoff value for the serum protein marker Pro-SHAP to predict immune-related pneumonia was 198.01 pg / mL, with an AUC of 0.865 for the ROC curve and p < 0.0001.
[0024] Preferably, the constructed Logistic regression equation is as follows:
[0025] P = 1 / (1+e) -(-25.4216+8.5143*A+1.9459*B+9.1706*C+15.7390*D+7.7367*E) );
[0026] Where P represents the probability of having CIP; A, B, C, D and E represent the expression levels of SERPINA5, AKAP6, TUBA4A, COL1A2 and Pro-SHAP, respectively.
[0027] This invention also discloses a method for constructing a visual online prediction system for immune-related pneumonia, comprising:
[0028] Serum samples were collected, and differential peaks were obtained by mass spectrometry using a primary mass spectrometer, and the differential peaks were identified by a secondary mass spectrometer. The expression levels of differentially expressed proteins were verified by ELISA.
[0029] Plot the ROC curve of the differentially expressed protein and determine the optimal cutoff value of the differentially expressed protein based on the maximum Oden index.
[0030] The expression levels of serum samples were classified using the optimal cutoff value, and a Logistic regression equation was constructed.
[0031] An online interactive visualization prediction model is constructed using the Logistic regression equation.
[0032] Preferably, the differentially expressed proteins include serum protein markers: SERPINA5, AKAP6, TUBA4A, COL1A2, and Pro-SHAP.
[0033] More preferably, the optimal cutoff value for the serum protein marker SERPINA5 in predicting immune-related pneumonia was 642.27 pg / mL, with an AUC of 0.908 for the ROC curve and p < 0.0001.
[0034] The optimal cutoff value for the serum protein marker AKAP6 in predicting immune-related pneumonia was 1050.88 pg / mL, with an AUC of 0.857 for the ROC curve and p < 0.0001.
[0035] The optimal cutoff value for the serum protein marker TUBA4A to predict immune-related pneumonia was 226.32 ng / L, with an AUC of 0.874 for the ROC curve and p < 0.0001.
[0036] The optimal cutoff value for the serum protein marker COL1A2 in predicting immune-related pneumonia was 894.83 pg / mL, with an AUC of 0.825 for the ROC curve and p < 0.0001.
[0037] The optimal cutoff value for the serum protein marker Pro-SHAP to predict immune-related pneumonia was 198.01 pg / mL, with an AUC of 0.865 for the ROC curve and p < 0.0001.
[0038] More preferably, the constructed Logistic regression equation is:
[0039] P = 1 / (1+e) -(-25.4216+8.5143*A+1.9459*B+9.1706*C+15.7390*D+7.7367*E) );
[0040] Where P represents the probability of having CIP; A, B, C, D and E represent the expression levels of SERPINA5, AKAP6, TUBA4A, COL1A2 and Pro-SHAP, respectively.
[0041] Compared with the prior art, the present invention has the following beneficial effects:
[0042] This invention provides a visualized online prediction system for CIP (Chronic Illness Infection) based on a combination of five specific predictive proteins obtained from serum proteomics mass spectrometry analysis. The five specific predictive proteins are: SERPINA5, AKAP6, TUBA4A, COL1A2, and Pro-SHAP. The expression levels of these differentially expressed proteins in the experimental and control groups were quantitatively detected using ELISA. The results showed significant differences in the expression of SERPINA5, AKAP6, TUBA4A, COL1A2, and Pro-SHAP between the experimental and control groups, with all five proteins highly expressed in CIP patients. The predictive accuracy of most of these five proteins exceeded 85%, making them suitable as specific biomarkers for predicting CIP. Based on these results, this invention also utilizes R language, with SERPINA5, AKAP6, TUBA4A, COL1A2, and Pro-SHAP as modeling elements, and employs machine learning algorithms to develop a visualized online prediction system for immune-associated pneumonia. The system achieved an AUC of 0.996, with specificity and sensitivity of 97.6% and 95.5%, respectively. The model validation set was then formed using 21 newly included patients (control group: 12 patients; experimental group: 9 patients), and the AUC of the ROC curve of the validation set reached 0.968.
[0043] This invention utilizes machine learning algorithms to construct a visual online prediction system composed of five highly specific proteins: SERPINA5, AKAP6, TUBA4A, COL1A2, and Pro-SHAP. This system enables highly efficient and specific prediction of CIP. The online interactive format developed using R language has significant advantages in terms of portability, visualization, and ease of operation, facilitating the clinical translation of the model. It provides a new strategy for screening high-risk individuals for CIP and increasing the safety of immunotherapy. Attached Figure Description
[0044] Figure 1 : Research roadmap of this invention;
[0045] Figure 2 : A spectrum of differentially expressed protein and peptide peaks in CIP patients compared to non-CIP patients, identified using MALDI-TOF-MS mass spectrometry analysis.
[0046] Figure 3 Differences in the expression of SERPINA5 (A), AKAP6 (B), TUBA4A (C), COL1A2 (D), and Pro-SHAP (E) between CIP patients (red) and non-CIP patients (green);
[0047] Figure 4Gel chromatographic separation results of SERPINA5 (A), AKAP6 (B), TUBA4A (C), COL1A2 (D) and Pro-SHAP (E);
[0048] Figure 5 ELISA quantitative analysis results of SERPINA5 (A), AKAP6 (B), TUBA4A (C), COL1A2 (D) and Pro-SHAP (E);
[0049] Figure 6 ROC curves for identifying CIP occurrence using SERPINA5, AKAP6, TUBA4A, COL1A2(D), and Pro-SHAP(E);
[0050] Figure 7A A nomogram-based comprehensive prediction model for predicting CIP occurrence; Figure 7B To illustrate the application of the online interactive visualization comprehensive prediction model; Figure 7C The ROC curve for the prediction model; Figure 7D The calibration curve for the prediction model; Figure 7E The decision curve for the prediction model; Figure 7F The ROC curve for the validation set of the prediction model; Figure 7G Calibration curves for the validation set of the prediction model; Figure 7H The decision curve is the validation set of the prediction model. Detailed Implementation
[0051] To enable those skilled in the art to better understand the present invention, the technical solutions of the present invention will be clearly and completely described below with reference to the accompanying drawings of the embodiments of the present invention. Obviously, the described embodiments are only some embodiments of the present invention, and not all embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those skilled in the art without creative effort should fall within the scope of protection of the present invention.
[0052] It should be noted that the terms "first," "second," etc., in the specification, claims, and accompanying drawings of this invention are used to distinguish similar objects and are not necessarily used to describe a specific order or sequence. It should be understood that such data can be interchanged where appropriate so that the embodiments of the invention described herein can be implemented in orders other than those illustrated or described herein. Furthermore, the terms "comprising" and "having," and any variations thereof, are intended to cover a non-exclusive inclusion; for example, a process, method, system, product, or apparatus that comprises a series of steps or units is not necessarily limited to those steps or units explicitly listed, but may include other steps or units not explicitly listed or inherent to such processes, methods, products, or apparatus.
[0053] The present invention will now be described in further detail with reference to the accompanying drawings:
[0054] This invention provides a visualized online prediction system for CIP occurrence based on a combination of five specific predictive proteins obtained from serum proteomics mass spectrometry analysis. The five specific predictive proteins are: SERPINA5 (human plasma serine protease inhibitor), AKAP6 (kinase ankyrin 6), TUBA4A (tubulin α-4a chain), COL1A2 (type I collagen α-2(I) chain), and Pro-SHAP (α-trypsin inhibitor heavy chain 2 precursor, also known as human serum-derived hyaluronic acid-associated protein precursor). Wherein:
[0055] SERPINA5, its amino acid sequence is as follows:
[0056] R.SARLNSQRLVFNRPFLMFIVDNNILFLGKVNRP.-(shown in SEQ.ID.NO.1);
[0057] AKAP6, its amino acid sequence is as follows:
[0058] S.DVNVSMIVNVSCTSACTDDEDDSDLLSSSTLT.L (shown in SEQ.ID.NO.2);
[0059] TUBA4A, with the amino acid sequence: R.LISQIVSSITASLR.F (as shown in SEQ.ID.NO.3);
[0060] COL1A2, its amino acid sequence is: GHAGLAGAR (SEQ.ID.NO.4);
[0061] Pro-SHAP, with the amino acid sequence: FLHVPDTFEGHFDGVPVISKGQQK (as shown in SEQ.ID.NO.5).
[0062] 1. Clinical serum sample collection and processing
[0063] The overall research roadmap of this invention is as follows: Figure 1As shown. This invention includes 139 patients with malignant tumors who underwent ICIs for the first time at the First Affiliated Hospital of Xi'an Jiaotong University and Tangdu Hospital of Air Force Medical University from January 1, 2019 to December 31, 2021. Baseline whole blood samples were collected from the patients before ICIs treatment. Factors such as age, gender, collection time, consistency of storage conditions, and presence of underlying diseases were fully considered to ensure that baseline conditions were basically consistent. Blood was collected from the subjects in the morning on an empty stomach using vacuum blood collection tubes (red cap, with insulating gel, and no added coagulants, anticoagulants, or other additives). The blood was incubated at 4°C for 4 hours, and centrifuged within 8 hours to obtain serum. Centrifugation conditions were: 4°C, 3.0 rpm, 20 minutes. The supernatant serum was aliquoted into 100 μL / tube and immediately stored at -80°C, avoiding repeated freeze-thaw cycles. All patients were followed up for a minimum of 6 months. Based on the "Expert Consensus on the Management of Immune Checkpoint Inhibitor-Related Adverse Reactions" published by Chinese experts and the international NCCN guidelines for the management of immunotherapy-related toxicities (2021 edition), three radiologists with extensive clinical experience (at the associate chief physician level or above) assessed the included patients for CIP (Critical Illness Inhibition). Patients were ultimately divided into an experimental group (i.e., patients who developed CIP) and a control group (i.e., patients who did not develop CIP). Exclusion criteria were: (1) prior immune-related or immune-mediated therapy for the target lesion; (2) history of more than one primary malignant tumor; (3) lack of baseline characteristics and imaging evidence preventing dynamic follow-up; and (4) poor blood sample quality, such as severe hemolysis or low blood concentration. Ultimately, this invention included 117 patients with malignant tumors (87 males and 30 females). Fifty-four patients were randomly selected as the mass spectrometry training set (control group: 36 patients; experimental group: 18 patients), and 63 patients were randomly selected as the ELISA validation set (control group: 41 patients; experimental group: 22 patients).
[0064] This study was conducted in accordance with the Declaration of Helsinki (2013 revision). All patients or their legal representatives voluntarily signed written informed consent before participating in the study. This study was approved by the Ethics Committee of the First Affiliated Hospital of Xi'an Jiaotong University (Approval No.: XJTU1AF2021LSK-001).
[0065] 2. Proteomics mass spectrometry analysis and identification
[0066] 2.1 Reagents and Instruments
[0067] Serum proteins were extracted using the magnetic bead kit "copper ion type" (IMAC-Cu) from Xiamen Primagene Biotechnology Co., Ltd., along with spectroscopically pure (HPLC grade) acetonitrile, trifluoroacetic acid (Merck, Germany), and α-cyano-4-hydroxycinnamic acid (HCCA) (Sigma, USA).
[0068] Magnetic bead separator, 600 / 384 AnchorChip target plate, and AutoFlex III matrix-assisted laser desorption / ionization time-of-flight mass spectrometry (MALDI-TOF-MS) (Bruker Daltonics, Germany).
[0069] The preparation of serum protein samples involved capturing serum protein peptides using copper ion-modified (IMAC-Cu) magnetic beads. The specific operational steps are as follows:
[0070] ①Use a mixer to thoroughly mix the magnetic bead suspension for 1 minute;
[0071] ② Add 10 μL of IMAC-Cu binding solution and 10 μL of IMAC-Cu magnetic beads to the PCR tube, mix well, then add 5 μL of serum, mix at least 5 times, and let stand for 5 min.
[0072] ③ Place the PCR tube into the magnetic column separator, allow the magnetic beads to adhere to the wall for 1 minute, and discard the supernatant after the liquid becomes clear;
[0073] ④ Add 100 μL of IMAC-Cu washing buffer, move the PCR tube back and forth 10 times on the magnetic column separator, and discard the supernatant after the magnetic beads adhere to the wall.
[0074] ⑤ Repeat steps ③ and ④ twice;
[0075] ⑥ Add 5 μL of IMAC-Cu elution buffer to wash the adhered magnetic beads and repeatedly blow them 10 times. The magnetic beads adhered to the wall for 2 minutes. Transfer the supernatant into a clean centrifuge tube.
[0076] ⑦ Add 5 μL of IMAC-Cu stabilizing solution to the centrifuge tube and mix well. The extracted protein peptides can be used for direct MALDI-TOF-MS detection or frozen in a -20℃ freezer for up to 24 hours for mass spectrometry analysis.
[0077] 2.2 Mass Spectrometry Analysis
[0078] 1 μL of the isolated protein sample was mixed with 10 μL of the matrix α-cyano-4-hydroxycinnamic acid. 1 μL of the mixture was then spotted onto an Anchorchip target plate (Bruker, Germany), with three spots spotted for each sample for three replicates. After drying at room temperature, the target plate was placed in a mass spectrometer for analysis. Standard calibration was performed using FlexControl 2.0 software (Bruker, Germany) before sample detection began. Each sample underwent a total of 300 laser targeting cycles (5 spotting cycles, 2 × 30 cycles per cycle) to generate a mass spectrum, obtaining protein-peptide spectra composed of different mass-to-nucleus ratios (M / Z). ClinProTools 2.1 software (Bruker, Germany) combined with biostatistical and bioinformatics methods, including genetic algorithms, was used to analyze the protein-peptide spectra of the two serum samples. The total ion chromatogram was normalized and smoothed to eliminate chemical and electrophysical noise; differentially expressed proteins between groups were analyzed and the magnitude of the differences was calculated. The proteins were sorted from largest to smallest difference to identify the peaks of proteins with significant differences in expression between groups (p<0.05).
[0079] Serum samples from the CIP experimental group and the control group without CIP were processed using a magnetic bead separation system and then analyzed by MALDI-TOF-MS. Protein and peptide profiles were plotted for each sample from both groups. A total of 129 protein and peptide peaks were detected within the molecular weight range of 1000 Da to 10000 Da. The stability of the three replicates for each sample was high. Figure 2 As shown.
[0080] Serum protein and peptide profiles of CIP patients and non-CIP control groups captured by mass spectrometry were analyzed using ClinProTools 2.1 software. The serum peptide profiles of CIP patients were compared with those of non-CIP patients. Significantly high expression of protein and peptide peaks with molecular weights of 3901.24 Daltons (SERPINA5), 3351.81 Daltons (AKAP6), 1488.72 Daltons (TUBA4A), 808.46 Daltons (COL1A2), and 2682.71 Daltons (Pro-SHAP) was detected in the serum of CIP patients (CIP patients vs. non-CIP patients: SERPINA5: 1.66±0.63 vs 1.37±0.40, p=0.044; AKAP6: 1.82±0.56 vs 2.81±1.38, p<0.001; TUBA4A: 2.99±1.09 vs). 2.32±0.67, p<0.001; COL1A2: 957.15±74.54 vs 847.45±94.54, p<0.001; Pro-SHAP: 216.23±12.60 vs 193.09±16.63, p<0.001. Results are as follows... Figure 3 As shown, the expression of M / Z: 3901.24, 3351.81, 1488.72, 808.46, and 2682.71 in CIP patients (red, curves with peaks at the top) and non-CIP patients (green, curves with peaks at the bottom) was compared, revealing that their protein-peptide peaks were significantly highly expressed in the serum of CIP patients. Therefore, these five protein-peptides were identified by secondary mass spectrometry.
[0081] Specifically, liquid chromatography coupled with mass spectrometry (LC-MS) was used to identify the serum peptide markers M / Z: 1488.72, 3351.81, 3901.24, 808.46, and 2682.71 in CIP patients. Two-dimensional gel chromatography (GLC) was used to separate the remaining serum protein peptides collected after magnetic bead separation and mass spectrometry loading. 15–30 peptide fractions were collected, and the target proteins were detected in the collected fractions. The sequences of the upregulated protein peptides M / Z: 1488.72, 3351.81, 3901.24, 808.46, and 2682.71 in the serum of CIP patients were then identified using a Thermo Fisher LTQ Orbitrap XL mass spectrometry system.
[0082] The specific operating steps are as follows:
[0083] 2.3.1 Sample Pretreatment
[0084] Combine the extracted protein samples, reflux at 1300 rpm for 10 minutes, collect the supernatant, and freeze-dry to a final volume of 50 μL to obtain liquid A. Concentrate liquid A using an Agilent Ziptip extraction column. The treatment method is as follows: ① Activate the Ziptip column by blowing and aspirating it 5 times with 100% acetonitrile; ② Repeat the blowing and aspirating process 10 times with the activated Ziptip in liquid 1, minimizing bubble formation; ③ Wash the Ziptip column 3 times with a 50% ACN and 0.1% TFA aqueous solution; ④ Elute the sample by repeatedly blowing and aspirating the Ziptip column in 0.1% TFA to obtain eluent 2; ⑤ Repeat steps 1-4 above 30 times; ⑥ Combine the 30 eluents 2, freeze-dry to 10 μL, and use for mass spectrometry identification.
[0085] 2.3.2 Chromatographic separation
[0086] Add 10 μL of mobile phase A to the original sample and transfer it to a vial, for a total of 20 μL.
[0087] One-dimensional ultra-high performance liquid chromatography system: Waters Nano Aquity UPLC (Waters Corporation, Milford, USA). Column:
[0088] Trapping column: C18,5μm,180μm×20mm,nanoAcquity TM Column
[0089] Analysis column: C18,3.5μm,75μm×150mm,nanoAcquity TM Column
[0090] Mobile phase A: an aqueous solution of 5% acetonitrile and 0.1% formic acid.
[0091] Mobile phase B: 95% acetonitrile, 0.1% formic acid aqueous solution; all solutions were HPLC grade.
[0092] The capture flow rate was 15 μl / min, the capture time was 3 min, the analysis flow rate was 400 nl / min, the analysis time was 60 min, the column temperature was 35 ℃, and the injection was performed in Partial Loop mode with an injection volume of 18 μL.
[0093] The gradient elution procedure is shown in Table 1 below.
[0094] Table 1
[0095]
[0096]
[0097] Gel chromatography separation results as follows Figure 4 As shown in the figure. The horizontal axis of the chromatogram represents the sample elution time, and the vertical axis represents the relative abundance of peptides. The chromatographic setting time is 60 min, and the fractions are collected starting from 10 min. The peptide components are mainly separated after 15 min and gradient elution is used to improve the elution efficiency. The set capture time for collecting fractions is 15 to 30 peptide fractions.
[0098] 2.3.3 LTQ-Orbitrap XL Mass Spectrometry Analysis
[0099] A Thermo Fisher LTQ Obitrap XL mass spectrometry system was used. A nano ion source (Michrom Bioresources, Auburn, USA) was used with a spray voltage of 1.8 kV. The mass spectrometry scan time was 60 min. The experimental modes were data-dependent and dynamic exclusion. The precursor ion was cascaded twice within 10 seconds and then added to the exclusion list for 90 seconds. The scan range was 400-2000 M / Z. The primary scan (MS) used the Obitrap with a resolution of 100,000. The CID and secondary scans used the LTQ. The 10 strongest ions from the MS spectrum were selected as single isotopes as precursor ions for MS / MS (single charge exclusion, not considered as precursor ions). The detection results are as follows: Figure 4 As shown.
[0100] Data Analysis: Using Bioworks Browser 3.3.1SP1 data analysis software for Sequest TM Search was performed. The parent ion error was set to 100 ppm, the fragment ion error to 1 Da, the digestion method was non-enzymatic digestion, and the variable modification was M(Methionine) oxidation. The search result parameter was set to deltacn >= 0.10. The search results were:
[0101] (1) M / Z: 3901.24; Protein ID: P05154; Gene Symbol = SERPINA5; Sequence: R.SARLNSQRLVFNRPFLMFIVDNNILFLGKVNRP.-;
[0102] (2) M / Z: 3351.81; Protein ID: Q13023; Gene Symbol = AKAP6; Sequence: S.DVNVSMIVNVSCTSACTDDEDDSDLLSSSTLT.L;
[0103] (3) M / Z: 1488.72; Protein ID: P68366; Gene Symbol = TUBA4A; Sequence: R.LISQIVSSITASLR.F;
[0104] (4) M / Z: 808.46; Protein ID: P08123; Gene Symbol = COL1A2; Sequence: GHAGLAGAR;
[0105] (5) M / Z: 2682.71; Protein ID: P19823; Gene Symbol = Pro-SHAP; Sequence: FLHVPDTFEGHFDGVPVISKGQQK.
[0106] The isolated M / Z: 3901.24 is named human plasma serine protease inhibitor (SERPINA5), with an exact molecular weight of 3901.24 Daltons and the sequence R.SARLNSQRLVFNRPFLMFIVDNNILFLGKVNRP.-; the isolated M / Z: 3351.81 is named kinase ankylosing protein 6 (AKAP6), with an exact molecular weight of 3351.81 Daltons and the sequence S.DVNVSMIVNVSCTSACTDDEDDSDLLSSSTLT.L; the isolated M / Z: 1488.72 is named tubulin α-4a chain (TUB). The exact molecular weight of the isolated component (A4A) is 1488.72 Daltons, and its sequence is R.LISQIVSSITASLR.F. The isolated component with a molecular weight of 808.46 is called type I collagen α-2(I) chain (COL1A2), with an exact molecular weight of 808.46 Daltons and its sequence is GHAGLAGAR. The isolated component with a molecular weight of 2682.71 is called α-trypsin inhibitor heavy chain 2 precursor, also known as human serum-derived hyaluronic acid-related protein precursor (Pro-SHAP), with an exact molecular weight of 2681.72 Daltons and its sequence is FLHVPDTFEGHFDGVPVISKGQQK. The functions of these potential biomarkers are mainly concentrated in: (1) immune-inflammatory mechanisms related to complement and coagulation cascade reactions, fibrin-related pathways, and inflammatory responses, such as SERPINA5; (2) physical microenvironment mechanisms related to extracellular matrix secretion and matrix stiffness, such as COL1A2 and Pro-SHAP; (3) cell functional pathways related to cell structural stability and energy supply, such as AKAP6 and TUBA4A.
[0107] The results above suggest that SERPINA5, AKAP6, TUBA4A, COL1A2, and Pro-SHAP are proteins specifically associated with CIP and can serve as potential predictive biomarkers for CIP. Further quantitative detection using ELISA can be used to verify their clinical value as predictive biomarkers.
[0108] 3. Quantitative Validation Analysis Using Enzyme-Linked Immunosorbent Assay (ELISA)
[0109] 1) Serum samples: Serum samples were collected from 22 patients with CIP (14 males and 8 females) and 41 patients without CIP (27 males and 14 females) for ELISA quantitative validation analysis. Serum samples were obtained from the First Affiliated Hospital of Xi'an Jiaotong University and Tangdu Hospital of Air Force Medical University, and were collected from January 2019 to December 2021.
[0110] 2) Detection Method: The expression levels of serum SERPINA5, AKAP6, TUBA4A, COL1A2, and Pro-SHAP in patients with and without CIP were detected using an enzyme-linked immunosorbent assay (ELISA). The kit was purchased from R&D Corporation, USA. The kit used a one-step sandwich ELISA with double antibodies: Samples, standards, and HRP-labeled detection antibodies were added sequentially to microwells pre-coated with antibodies against human SERPINA5, TUBA4A, AKAP6, COL1A2, and Pro-SHAP proteins. After incubation and thorough washing, the substrate TMB was used for color development. TMB was converted to blue under the catalysis of peroxidase, and then to yellow under acidic conditions. The color intensity was positively correlated with the levels of SERPINA5, AKAP6, TUBA4A, COL1A2, and Pro-SHAP proteins in the sample. The absorbance (OD value) was measured at 450 nm using a microplate reader, and the sample concentration was calculated. Refer to the kit instructions for specific experimental procedures, and the criteria for determining a positive result should be defined in accordance with the kit instructions.
[0111] 3) Statistical methods: One-way ANOVA and independent samples t-tests were performed using GraphPad.Prism.v8 (GraphPad Software, La Jolla, CA, USA); ROC analysis and optimal cutoff value analysis were performed using IBM SPSS Statistics 25 (SPSS, Inc., Chicago, IL, USA).
[0112] 4) Quantitative Results Analysis: ELISA quantitative analysis showed that the expression levels of SERPINA5, AKAP6, TUBA4A, COL1A2, and Pro-SHAP in CIP patients and non-CIP patients were as follows: CIP patients vs. non-CIP patients: SERPINA5: 738.25±58.2 pg / ml (625.36~807.75 pg / ml) vs. 609.24±27.19 pg / ml (558.67~675.35 pg / ml), p<0.001; AKAP6: 1126.90±64.04 pg / ml (1028.15~1232.52 pg / ml) vs. 1017.13±55.22pg / ml (914.41~1162.23pg / ml), p<0.001; TUBA4A: 228.23±18.48ng / l (190.33~257.03ng / l) vs 201.93±10.08ng / l (184.53~233.67ng / l), p<0.001; COL1A2: 957.15±74.54pg / ml (827.681~1084.242pg / ml) vs The effective values were 847.45±94.54 pg / ml (715.151~1079.28 pg / ml), p<0.001; Pro-SHAP: 216.23±12.60 pg / ml (191.219~240.751 pg / ml) vs 193.09±16.63 pg / ml (170.117~234.175 pg / ml), p<0.001. Specific results are shown in Table 2. Figure 5 As shown above, SERPINA5, AKAP6, TUBA4A, COL1A2, and Pro-SHAP are proteins closely related to CIP.
[0113] Table 2. Expression levels of serum protein biomarkers in different groups
[0114]
[0115] 5) Analysis of optimal cutoff values: ROC curves of SERPINA5, AKAP6, TUBA4A, COL1A2, and Pro-SHAP were plotted using SPSS software. The optimal cutoff values of SERPINA5, AKAP6, TUBA4A, COL1A2, and Pro-SHAP were determined based on the maximum Youden index to obtain the ability of SERPINA5, AKAP6, TUBA4A, COL1A2, and Pro-SHAP to distinguish between CIP patients and non-CIP patients (Youden index = sensitivity + specificity - 1). The optimal cutoff values for SERPINA5, AKAP6, TUBA4A, COL1A2, and Pro-SHAP were 642.27 pg / ml, 1050.88 pg / ml, 226.32 ng / L, 894.83 pg / ml, and 198.01 pg / ml, respectively, with corresponding AUCs of 0.908 (p < 0.0001), 0.857 (p < 0.0001), 0.874 (p < 0.0001), 0.825 (p < 0.0001), and 0.865 (p < 0.0001). Specific results are as follows... Figure 6 As shown above, SERPINA5, AKAP6, TUBA4A, COL1A2, and Pro-SHAP are reliable biomarkers for predicting CIP and have good disease identification capabilities.
[0116] 4. Construction and Validation of Online Interactive Visualized Comprehensive Prediction Model
[0117] The prediction model was constructed using R language V4.1.2 (Vienna, Austria). To obtain better prediction performance, we modeled the model based on four machine learning artificial intelligence methods: Random Forest, Support Vector Machine (SVM), LASSO regression, and Logistic regression (Table 3). After comparing the ROC values of the models, the final results show that LASSO regression and Logistic regression algorithms have the greatest ability to distinguish CIP, both reaching 0.996. However, considering the visualization of the model and its ease of clinical application, we ultimately chose to construct a nomogram prediction model based on the immune-physical microenvironment using the Logistic regression algorithm. An online interactive visual nomogram prediction model was built, accessible at https: / / mass-cip.shinyapps.io / DynNomapp / . The prediction model is shown in Figure 7 (AB). Its ROC curve AUC reaches 0.996 (Figure 7 (C), indicating excellent disease identification ability for CIP. The calibration curve shows good consistency between predicted and actual values (Figure 7 (D)). The clinical decision curve shows that the comprehensive model has the highest clinical net benefit rate compared to each individual factor prediction (Figure 7 (E)). The model was validated using a validation set of 21 patients. The validation set's ROC curve AUC was 0.968 (Figure 7 (F)). The calibration curve shows good consistency between predicted and actual values in the validation set (Figure 7 (G)). The decision curve supports the conclusion that the comprehensive model has the highest clinical net benefit rate (Figure 7 (H)). After validation, the prediction model showed good clinical predictive value for CIP. This online interactive and visualized comprehensive prediction model has significant advantages such as low detection cost, ease of use, visualization, high clinical application value, and ease of promotion. It can be used as an efficient prediction model for the occurrence of CIP to screen high-risk groups of malignant tumor patients before receiving ICIs treatment, thereby increasing the safety of immunotherapy.
[0118] Table 3 Comparison of four machine learning artificial intelligence modeling methods
[0119] Random Forest 0.944(0.880-1.000) SVM 0.965(0.915-1.000) LASSO 0.996(0.986-1.000) Logistics Returns 0.996(0.976-1.000)
[0120] Examples illustrating the clinical application of this online prediction system:
[0121] Taking two patients, A and B, with malignant tumors admitted to the clinic as examples, according to guidelines, before considering ICIs treatment, it is necessary to assess the patient's risk of developing CIP in the future. Specifically, fasting serum was collected from patient A, and the expression levels of SERPINA5, AKAP6, TUBA4A, COL1A2, and Pro-SHAP in the serum were detected. Based on the optimal cutoff values for the five modeling proteins (SERPINA5 642.27 pg / ml, AKAP6 1050.88 pg / ml, TUBA4A 226.32 ng / L, COL1A2 894.83 pg / ml, and Pro-SHAP 198.01 pg / ml), patient A's protein expression levels were classified as high or low (i.e., above the optimal cutoff value is high, represented by "1"; otherwise, it is low, represented by "0"). The final expression levels of the five modeling proteins for patient A were: SERPINA5 "1", AKAP6 "0", TUBA4A "1", COL1A2 "0", and Pro-SHAP "1". After inputting these results into the prediction website https: / / mass-cip.shinyapps.io / DynNomapp / , the risk of patient A developing CIP after ICI treatment was calculated to be 50.1%. In other words, patient A has a 50.1% chance of developing CIP in the future, and ICI treatment can be considered on a case-by-case basis. Similarly, the expression levels of the five modeling proteins for patient B were: SERPINA5 "1", AKAP6 "0", TUBA4A "1", COL1A2 "1", and Pro-SHAP "1". Therefore, patient B's risk of developing CIP was as high as 99.9%, meaning patient B has a 99.9% chance of developing CIP in the future, and clinical use of immune checkpoint inhibitors is not recommended.
[0122] To further evaluate the predictive ability of this visualized online prediction system, this invention obtained the critical value of the scoring system based on the scores of each patient in the training set, and divided patients into high-risk and low-risk groups accordingly. The analysis results showed that the optimal cutoff value for the prediction system score was 155 points; that is, patients with scores above 155 points were considered high-risk, otherwise low-risk. Univariate regression analysis showed that the risk of CIP was significantly higher in the high-risk group than in the low-risk group (training set: OR = 0.024, 95% CI = 0.004–0.169, p < 0.001; validation set: OR = 0.125, 95% CI = 0.018–0.863, p = 0.005). These results demonstrate that the visualized online prediction system for CIP constructed in this invention can effectively screen out high-risk individuals for CIP.
[0123] In summary, this invention discloses an online interactive and visual comprehensive prediction model for predicting the risk of CIP in patients with malignant tumors before immunotherapy and its clinical application. The modeling elements are SERPINA5 (optimal cutoff value: 642.27 pg / ml), AKAP6 (optimal cutoff value: 1050.88 pg / ml), TUBA4A (optimal cutoff value: 226.32 ng / L), COL1A2 (optimal cutoff value: 894.83 pg / ml) and Pro-SHAP (optimal cutoff value: 198.01 pg / ml), all of which are highly expressed in CIP patients.
[0124] The above content is only for illustrating the technical concept of the present invention and should not be construed as limiting the scope of protection of the present invention. Any modifications made to the technical solution based on the technical concept proposed in this invention shall fall within the scope of protection of the claims of this invention.
Claims
1. A visualized online prediction system for immune-related pneumonia, characterized in that, include: The serum protein biomarker expression level detection module is used to detect the expression levels of protein biomarkers in the serum of the test subjects; the protein biomarkers in the serum include SERPINA5, AKAP6, TUBA4A, COL1A2 and Pro-SHAP; The intelligent prediction module is used to calculate the probability of a subject developing immune-related pneumonia based on the expression level of protein biomarkers using an intelligent prediction model. The results output module is used to output results through a visual online interactive interface. The intelligent prediction module includes: The differential protein acquisition unit is used to acquire the expression levels of differentially expressed proteins in serum samples. The intelligent prediction model construction unit plots the ROC curve of differentially expressed proteins based on their expression levels and calculates the optimal cutoff value of the differentially expressed proteins based on the maximum Yoden index. Then, it classifies the expression levels of serum samples based on the optimal cutoff value, constructs a Logistic regression equation, and obtains the intelligent prediction model. The collected serum samples were analyzed by mass spectrometry to obtain differential peaks using primary mass spectrometry, and the differential peaks were identified by secondary mass spectrometry. The expression levels of differentially expressed proteins were then verified by ELISA. The optimal cutoff value for the serum protein marker SERPINA5 in predicting immune-related pneumonia was 642.27 pg / mL, with an AUC of 0.908 for the ROC curve. p <0.0001; The optimal cutoff value for the serum protein marker AKAP6 in predicting immune-related pneumonia was 1050.88 pg / mL, with an AUC of 0.857 for the ROC curve. p <0.0001; The optimal cutoff value for the serum protein marker TUBA4A in predicting immune-related pneumonia was 226.32 ng / L, with an AUC of 0.874 for the ROC curve. p <0.0001; The optimal cutoff value for the serum protein marker COL1A2 in predicting immune-related pneumonia was 894.83 pg / mL, with an AUC of 0.825 for the ROC curve. p <0.0001; The optimal cutoff value for the serum protein biomarker Pro-SHAP in predicting immune-related pneumonia was 198.01 pg / mL, with an AUC of 0.865 for the ROC curve. p <0.0001; The constructed Logistic regression equation is as follows: P =1 / (1+e -(-25.4216+8.5143 A+1.9459 B+9.1706 C+15.7390 D+7.7367 E) ); in, P The probability of having CIP is represented by A, B, C, D, and E, which represent the expression levels of SERPINA5, AKAP6, TUBA4A, COL1A2, and Pro-SHAP, respectively.
2. A method for constructing a visual online prediction system for immune-related pneumonia, characterized in that, include: Serum samples were collected, and differential peaks were obtained by mass spectrometry using a primary mass spectrometer and identified by a secondary mass spectrometer. The expression levels of differentially expressed proteins were verified by ELISA. The differentially expressed proteins included serum protein markers: SERPINA5, AKAP6, TUBA4A, COL1A2, and Pro-SHAP. Plot the ROC curve of the differentially expressed protein and determine the optimal cutoff value of the differentially expressed protein based on the maximum Oden index. The expression levels of serum samples were classified using the optimal cutoff value, and a Logistic regression equation was constructed. An online interactive and visual prediction model is constructed using the Logistic regression equation; The optimal cutoff value for the serum protein marker SERPINA5 in predicting immune-related pneumonia was 642.27 pg / mL, with an AUC of 0.908 for the ROC curve. p <0.0001; The optimal cutoff value for the serum protein marker AKAP6 in predicting immune-related pneumonia was 1050.88 pg / mL, with an AUC of 0.857 for the ROC curve. p <0.0001; The optimal cutoff value for the serum protein marker TUBA4A in predicting immune-related pneumonia was 226.32 ng / L, with an AUC of 0.874 for the ROC curve. p <0.0001; The optimal cutoff value for the serum protein marker COL1A2 in predicting immune-related pneumonia was 894.83 pg / mL, with an AUC of 0.825 for the ROC curve. p <0.0001; The optimal cutoff value for the serum protein biomarker Pro-SHAP in predicting immune-related pneumonia was 198.01 pg / mL, with an AUC of 0.865 for the ROC curve. p <0.0001; The constructed Logistic regression equation is as follows: P =1 / (1+e -(-25.4216+8.5143 A+1.9459 B+9.1706 C+15.7390 D+7.7367 E) ); in, P The probability of having CIP is represented by A, B, C, D, and E, which represent the expression levels of SERPINA5, AKAP6, TUBA4A, COL1A2, and Pro-SHAP, respectively.
Citation Information
Patent Citations
Biomarker for predicting benefit degree and prognosis of non-small cell lung cancer immunotherapy and application of biomarker
CN116042832A
Interleukin 37-based early warning of (critical) severe respiratory virus infection
WO2022056896A1