Combination of four serum protein markers for predicting recurrence of non-small cell lung cancer after surgery and application thereof
Patent Information
- Application Number
- CN202611017221.9
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2026-07-09
- Publication Date
- 2026-09-29
AI Technical Summary
由于实践存在困难,许多被提出的预测复发转移的标志物未能广泛应用于临床
[0023]一、 预测质量与诊断效能的显著提高
Smart Images

Figure CN122836334A_ABST
Abstract
Description
Technical Field
[0001] This invention belongs to the field of biomedical technology and relates to a combination of four serum protein biomarkers for predicting postoperative recurrence of non-small cell lung cancer and their applications. Background Technology
[0002] Non-small cell lung cancer (NSCLC) is the most common type of lung cancer, mainly including adenocarcinoma, squamous cell carcinoma and large cell carcinoma. Its treatment and prognosis are closely related to the stage.
[0003] Currently, the following methods are mainly used to predict postoperative recurrence or prognosis in patients with early-stage NSCLC:
[0004] i. Based on clinicopathological features and imaging. The TNM staging system consists of three dimensions: tumor size (T), lymph node metastasis (N), and distant metastasis (M). Although routinely used, it has significant limitations in predicting outcomes for early-stage NSCLC patients and guiding postoperative adjuvant therapy. Furthermore, imaging techniques are often used for recurrence monitoring, primarily relying on microscopic observation of postoperative pathological sections and CT scans. However, these methods have a high false-positive rate and cannot provide early, accurate warnings of recurrence at the molecular level, often leading patients to miss the optimal intervention window.
[0005] ii. Blood-based single or generalized biomarkers. Previous biomarker studies for early recurrent NSCLC have largely focused on the impact of clinical and pathological features. These biomarkers are primarily used for the prognosis of immunotherapy in advanced patients or for assessing cancer risk in the general population, rather than specifically for monitoring recurrence and metastasis after early radical resection. Due to practical difficulties, many proposed biomarkers for predicting recurrence and metastasis have not been widely applied in clinical practice. Summary of the Invention
[0006] To address the problems existing in the prior art, the present invention aims to provide a combination of four serum protein biomarkers for predicting postoperative recurrence of non-small cell lung cancer and their applications. This combination can accurately distinguish between non-recurrent lung cancer (BN) and recurrent lung cancer (RR), filling the gap in high-precision postoperative recurrence prediction.
[0007] To achieve the above objectives, the technical solution adopted by the present invention is as follows:
[0008] This invention provides a combination of four serum protein biomarkers for predicting postoperative recurrence of non-small cell lung cancer, including four proteins: C9 (complement component 9), MMP12 (matrix metalloproteinase 12), MMP15 (matrix metalloproteinase 15), and FLT4 (Fms-associated tyrosine kinase 4).
[0009] This invention also provides the application of the above four serum protein biomarkers in the preparation of products for predicting postoperative recurrence and metastasis of non-small cell lung cancer.
[0010] The present invention also provides in vitro diagnostic reagents or kits for predicting postoperative recurrence and metastasis of non-small cell lung cancer, including a combination of the above four serum protein markers.
[0011] Furthermore, the in vitro diagnostic reagent or kit is a targeted mass spectrometry detection reagent or kit, or an immunoassay reagent or kit.
[0012] This invention also provides a system for predicting the postoperative recurrence probability of non-small cell lung cancer based on the above-mentioned combination of four serum protein markers, comprising:
[0013] Data input module: used to input feature data of the combination of four serum protein biomarkers into the prediction model;
[0014] Prediction model: Used to receive feature data from the data input module, perform calculations, and output prediction results.
[0015] Furthermore, the characteristic data are the expression levels of four proteins C9, MMP12, MMP15, and FLT4 obtained by panoramic qualitative and quantitative detection of serum samples using mass spectrometry or by specific detection using Western blotting / enzyme-linked immunosorbent assay (ELISA).
[0016] Furthermore, the abundance of the four proteins was calculated using intensity-based absolute quantification (iBAQ) and normalized to the total protein fraction (FOT).
[0017] Furthermore, the prediction model is constructed using a logistic regression algorithm to build a classifier, dividing the total samples into a training set of 60% and an independent test set of 40%, using the FOT-normalized expression levels of four proteins as input features, and training it using a 10-fold cross-validation method.
[0018] Furthermore, when the output probability in the prediction result is greater than the set cutoff value, it is judged as having a high risk of recurrence.
[0019] Furthermore, the calculation formula for the prediction model is as follows:
[0020] Logit(P) = 0.4591 + 2.5858 C9-0.5185 MMP12+2.0634 MMP15 + 0.6096 FLT4
[0021] The protein concentration in the serum of non-small cell lung cancer patients detected by mass spectrometry was substituted into the prediction model formula. When the P value was ≥0.767, the patients were identified as those who were prone to recurrence after surgery, and when the Logit(P) value was <0.767, the patients were identified as those who were not prone to recurrence after surgery.
[0022] The beneficial effects of this invention are as follows:
[0023] I. Significant improvement in predictive quality and diagnostic efficacy
[0024] This application proposes a specific combination of four serum proteins—C9, MMP12, MMP15, and FLT4—for predicting postoperative recurrence in non-small cell lung cancer (NSCLC). This novel biomarker combination was first discovered and validated based on in-depth proteomics and lipidomics analysis of serum samples from 58 discovery cohorts and 40 validation cohorts, and after rigorous statistical screening (expression increase ≥1.5-fold, FDR <0.01). Compared to the limitations of conventional TNM staging in predicting the prognosis of early-stage NSCLC patients and the high false-positive rate of imaging monitoring, the protein combination of this invention achieves precise early warning at the molecular level, significantly improving prediction quality.
[0025] II. Savings in procedures and ease of operation
[0026] The technical solution of this invention significantly simplifies the diagnostic process and improves clinical operability during implementation.
[0027] 1. Minimally invasive and simple sample acquisition: This invention only requires drawing peripheral venous blood from the patient to separate serum for testing. Compared with pathological examinations that rely on surgical removal of tissue samples or invasive puncture biopsies, the operation is extremely simple, patient compliance is high, and samples can be repeatedly collected for dynamic monitoring.
[0028] 2. Avoids the cumbersome procedures of multi-omics joint detection: Although integrating proteomics and lipidomics data can provide a holistic understanding, performing both omics detection procedures simultaneously in clinical practice is extremely cumbersome and time-consuming. This invention, through rigorous statistical screening, narrows down the range from 7,050 proteins to only 4 core proteins, eliminating the need for panoramic omics scanning in clinical practice. Only quantitative detection (such as ELISA or targeted mass spectrometry) of these 4 targets is required, significantly reducing the detection procedures and operational complexity.
[0029] III. Significant savings in testing costs
[0030] 1. Cost savings in reagents and instruments: Panoramic non-targeted proteomics detection (e.g., identifying 7,050 proteins) requires expensive mass spectrometry equipment and a large amount of reagents. This invention simplifies the detection targets to C9, MMP12, MMP15, and FLT4, allowing for the development of small-scale targeted mass spectrometry detection kits or immunoassay kits (e.g., chemiluminescent immunoassay) for clinical applications. This will result in an exponential decrease in reagent and instrument wear and tear costs.
[0031] 2. Cost savings in healthcare economics: Due to its extremely high sensitivity (97%) and specificity (96%), this invention can accurately identify high-risk individuals for recurrence. This allows clinicians to avoid unnecessary overtreatment and frequent imaging re-examinations for patients with low recurrence risk, thereby significantly reducing patients' medical expenses and the consumption of public medical resources.
[0032] IV. The Precision of Mechanism Guidance
[0033] The four-protein combination of this invention is not randomly assembled. MMP12 and MMP15 are matrix metalloproteinases closely related to extracellular matrix degradation, tumor invasion, and angiogenesis in metastatic lesions. C9, as a component of the complement system, participates in immune responses and inflammatory recruitment in the tumor microenvironment. FLT4 (vascular endothelial growth factor receptor 3) mediates lymphangiogenesis. The synergistic high expression of these four proteins suggests active immune inflammatory recruitment, matrix remodeling, and angiogenesis / lymphangiogenesis in the tumor microenvironment, thereby driving recurrence and metastasis of minimal residual disease after surgery. This mechanistic clarity addresses the lack of specificity for postoperative recurrence in existing blood biomarkers and the difficulty in clinical translation. It provides precise companion diagnostic evidence for subsequent interventions targeting these pathways (such as anti-angiogenic targeted therapy), and has extremely high clinical translational value.
[0034] The combination of C9, MMP12, MMP15, and FLT4 in this application constructs a prediction model with a sensitivity of up to 97%, a specificity of up to 96%, an independent test set accuracy of 92%, and an AUC of 0.91. This performance significantly outperforms existing models of the same type, filling the gap in high-precision postoperative recurrence protein classifiers. Attached Figure Description
[0035] Figure 1. Serum proteomics and lipidomics analysis of non-recurrent lung cancer (BN) and recurrent lung cancer (RR) cases. A. Overview of the serum proteomics and lipidomics research workflow, including cohort construction (serum discovery cohort: BN=30 cases, RR=28 cases; serum validation cohort: BN=20 cases, RR=20 cases), data collection, and data analysis (proteomics data, lipidomics data, and clinical information). B. Display of all baseline clinical characteristics of individuals included in the discovery cohort. C. Proteins identified in the serum of each BN and RR group case in the discovery cohort. D. Arranged in descending order of protein abundance in each BN and RR group case in the discovery cohort, presenting the dynamic range of protein identification for each sample, reaching 10. 7 Order of magnitude.
[0036] Figure 2. Serum proteomic profiles differ between non-recurrent lung cancer (BN) and recurrent lung cancer (RR) cases. A. Principal component analysis (PCA) of proteomic data from BN and RR cases. Red dots represent RR cases, and blue dots represent BN cases. B. Differences in protein abundance between BN and RR cases. C, D. Differential proteomic pathways between BN and RR cases.
[0037] Figure 3. Serum protein biomarkers used for the diagnosis of NSCLC patients. A. Screening criteria for lung cancer diagnostic protein biomarkers applicable to the discovery and validation cohorts. B. Heatmap showing the relative abundance (Z-score) of the four proteins in the discovery cohort. C. Classification error matrix for distinguishing recurrent lung cancer (RR) from non-recurrent (BN) using a logistic regression classifier (60% training set, 40% test set) in the discovery cohort, with each square indicating the number of identified samples. D. Receiver operating characteristic (ROC) curves for predicting the combination of the four proteins in the discovery cohort.
[0038] Figure 4 The ROC curves of subjects were observed when four protein biomarkers were detected individually and in combination in the cohort.
[0039] Figure 5 ROC curves of subjects when four protein biomarkers were detected individually and in combination in the validation cohort. Detailed Implementation
[0040] The present invention will now be described in detail with reference to specific embodiments. The following specific embodiments will help those skilled in the art to further understand the present invention, but do not limit the present invention in any way.
[0041] Example 1: Construction of a prediction model for a combination of four protein biomarkers based on a strict screening threshold (Fold Change ≥ 1.5)
[0042] The specific steps are as follows:
[0043] 1. Cohort construction and serum sample collection
[0044] Peripheral venous blood was collected from patients with non-small cell lung cancer (NSCLC) postoperatively, and serum samples were obtained. A discovery cohort (58 patients: 30 with non-recurrent bronchopulmonary nocturnal (BN) and 28 with recurrent recurrent (RR)) and an independent validation cohort (40 patients: 20 with BN and 20 with RR) were constructed. Figure 1 As shown in Figure A. Record the baseline clinical characteristics of all patients, including age, sex, tumor differentiation, BMI, smoking index, Ki-67, and various biochemical indicators (such as white blood cell count, neutrophil count, lymphocyte count, monocyte count, eosinophil count, and platelet count). Figure 1 As shown in B. All patients signed informed consent forms.
[0045] 2. Serum proteomics data acquisition
[0046] A comprehensive qualitative and quantitative analysis of proteins in serum samples was performed using the next-generation UltraOmic™ mass spectrometry technology (or other existing mass spectrometry technologies). Protein abundance was first calculated using intensity-based absolute quantification (iBAQ) and then normalized to a total protein fraction (FOT) for comparison across different experiments.
[0047] according to Figure 2 As shown in Figure A, principal component analysis (PCA) of proteomic data from non-recurrent lung cancer (BN) cases and recurrent lung cancer (RR) cases revealed significant differences between patients with recurrent and non-recurrent lung cancer; significant differences were also observed in protein abundance (Wilcoxon test, P < 0.05). Figure 2 As shown in B; the differential proteomic pathways between BN cases and RR cases are as follows: Figure 2 As shown in C and D, the high-density lipoprotein pathway and glycolysis / gluconeogenesis pathway were enriched in the non-relapse group; while in the relapse group, the complement pathway / acute immune pathway, MAPK pathway and cell adhesion pathway were significantly enriched.
[0048] In proteomics analysis, a total of 7,050 proteins were identified, with false discovery rates (FDR) at both peptide and protein levels controlled at 1%, averaging 1,800 proteins identified per sample, and a dynamic range reaching [value missing]. ,like Figure 1 As shown in C and D.
[0049] 3. Rigorous screening of differentially expressed proteins
[0050] In serum discovery and validation samples, the following three-step rigorous screening strategy was used to identify candidate proteins:
[0051] • Condition 1: The candidate protein must be expressed in at least 50% of the samples;
[0052] • Condition 2: The expression level of the candidate protein in the relapsed sample (RR) is at least 1.5 times higher than that in the non-relapsed sample (BN) (Fold Change ≥ 1.5).
[0053] Condition 3: The expression level of the candidate protein in the tumor sample is significantly increased compared with that in the normal sample (FDR <0.01).
[0054] Through the above screening, 66 proteins that were significantly and stably overexpressed in the discovery and validation samples were first obtained.
[0055] 4. Determination of the core marker combination
[0056] Combining the cross-validation results of the discovery and validation cohorts, four proteins—C9, MMP12, MMP15, and FLT4—were ultimately identified from 66 candidate proteins as significantly increased in samples from recurrent lung cancer patients. Figure 3 As shown in A and B.
[0057] The UniProt protein numbers for the four proteins are: C9: P02748, MMP12: P39900, MMP15: P51511, and FLT4: P359164.
[0058] 5. Construction and validation of prediction models (classifiers)
[0059] A logistic regression algorithm is used to construct a classifier. The total samples are divided into two sets: 60% for training and 40% for independent test sets. Figure 3 As shown in C. The model was trained using the FOT-normalized expression levels of the four proteins as input features and a 10-fold cross-validation method.
[0060] • Training set performance: The model exhibits high sensitivity (true positive rate, 97%) and high specificity (true negative rate, 96%).
[0061] • Independent test set performance: accuracy reaches 92%.
[0062] Overall efficacy: The receiver operating characteristic (ROC) curves for predicting the four protein combinations in the discovery cohort showed good performance, with an area under the curve (AUC) of 0.918. Figure 3 As shown in D, the AUC value is significantly higher than the AUC value predicted using four protein biomarkers individually, as shown in Figure D. Figure 4 As shown.
[0063] Meanwhile, the AUC value predicted by the four protein combinations in the validation cohort reached 0.924, significantly higher than the AUC value predicted by using the four protein biomarkers individually. Figure 5 As shown.
[0064] In this embodiment, a prediction model for the postoperative recurrence probability of non-small cell lung cancer is preferably constructed by fitting these four protein indicators using a Logit regression equation. The calculation formula for the prediction model is as follows:
[0065] Logit(P) = 0.4591 + 2.5858 C9-0.5185 MMP12+2.0634 MMP15 + 0.6096 FLT4
[0066] The model has a sensitivity of 0.733, a specificity of 0.967, and an optimal cutoff value of 0.767. By substituting the protein concentration in the serum of non-small cell lung cancer patients detected by mass spectrometry into the prediction model formula, the Logit(P) value is obtained. When the Logit(P) value is ≥0.767, the patient is diagnosed as a person with a high risk of recurrence after lung cancer surgery; when the Logit(P) value is <0.767, the patient is diagnosed as a person with a low risk of recurrence after lung cancer surgery.
[0067] Example 2: Construction of a protein biomarker combinatorial prediction model based on a moderate screening threshold (Fold Change ≥ 1.2)
[0068] To verify the predictive power of biomarker combinations at lower difference fold thresholds, the screening criteria were adjusted as follows in this embodiment:
[0069] 1. Queue construction and data acquisition are the same as in Example 1.
[0070] 2. Screening and adjustment of differentially expressed proteins
[0071] • Condition 1: The candidate protein is expressed in at least 50% of the samples;
[0072] • Condition 2: The expression level of the candidate protein in the relapsed sample (RR) is at least 1.2-fold higher than that in the non-relapsed sample (BN) (Fold Change ≥ 1.2).
[0073] Condition 3: FDR < 0.01.
[0074] With this relaxed threshold, the number of candidate proteins screened will be significantly greater than 66 (covering more proteins with low fold differences). Similarly, through cross-validation between two independent cohorts, proteins that are stably overexpressed in both cohorts are selected to determine an expanded biomarker combination including C9, MMP12, MMP15, FLT4, and other cooperating proteins (such as S100A8, S100A9, etc.).
[0075] 3. The prediction model was constructed and validated using the same logistic regression algorithm and a 6:4 training / test set split ratio as in Example 1, to build a multi-parameter prediction model. Because this example incorporates more markers with low fold differences, the model may capture more marginal recurrence features, potentially improving its sensitivity. However, its specificity may fluctuate due to the increased number of included variables, requiring cross-validation to determine the optimal cutoff value.
[0076] Example 3: Construction of a protein biomarker combinatorial prediction model based on a high-specificity screening threshold (Fold Change ≥ 2.0)
[0077] To achieve extremely high specificity and minimize false positive interference, this embodiment employs the most stringent difference fold threshold:
[0078] 1. Queue construction and data acquisition are the same as in Example 1.
[0079] 2. Screening and adjustment of differentially expressed proteins
[0080] • Condition 1: The candidate protein is expressed in at least 50% of the samples;
[0081] • Condition 2: The expression level of the candidate protein in the relapsed sample (RR) is at least 2.0 times higher than that in the non-relapsed sample (BN) (Fold Change ≥ 2.0).
[0082] Condition 3: FDR < 0.01.
[0083] Under this extremely stringent threshold, the selected candidate proteins will be highly concentrated on targets that are extremely highly expressed in the relapse group. C9, MMP12, MMP15, and FLT4, due to their significant pro-angiogenic and matrix remodeling effects in the relapse group, can still be stably selected.
[0084] 3. The predictive model construction and validation adopted the same algorithm and data partitioning strategy as in Example 1, using only core biomarkers (including C9, MMP12, MMP15, and FLT4) that satisfy Fold Change ≥ 2.0 to construct a simplified classifier. Because this example excludes all interfering proteins with low fold differences, the specificity (true negative rate) of the model prediction is expected to reach close to 100%, minimizing the possibility of non-relapsed patients being misdiagnosed as relapsed and subjected to overtreatment; however, the sensitivity may be correspondingly reduced.
[0085] Obviously, the above embodiments of the present invention are merely examples to illustrate the present invention more clearly, and are not intended to limit the implementation of the present invention. For those skilled in the art, other variations or modifications can be made based on the above description. It is impossible to exhaustively list all implementation methods here. Any obvious variations or modifications derived from the technical solutions of the present invention are still within the protection scope of the present invention.
Claims
1. A combination of four serum protein biomarkers for predicting postoperative recurrence of non-small cell lung cancer, including C9, MMP12, MMP15 and FLT4.
2. The use of the combination of four serum protein biomarkers as described in claim 1 in the preparation of a product for predicting postoperative recurrence and metastasis of non-small cell lung cancer.
3. An in vitro diagnostic reagent or kit for predicting postoperative recurrence and metastasis of non-small cell lung cancer, characterized in that, It includes the combination of the four serum protein markers described in claim 1.
4. The in vitro diagnostic reagent or kit according to claim 3, characterized in that, The in vitro diagnostic reagent or kit is a targeted mass spectrometry detection reagent or kit, or an immunoassay reagent or kit.
5. A system for predicting the probability of postoperative recurrence of non-small cell lung cancer based on the combination of four serum protein biomarkers as described in claim 1, comprising: Data input module: used to input feature data of the combination of four serum protein biomarkers into the prediction model; Prediction model: Used to receive feature data from the data input module, perform calculations, and output prediction results.
6. The system according to claim 5, characterized in that, The characteristic data refers to the expression levels of four proteins C9, MMP12, MMP15, and FLT4 obtained by using mass spectrometry to perform panoramic qualitative and quantitative detection of proteins in serum samples, or by specific detection using Western blotting / enzyme-linked immunosorbent assay.
7. The system according to claim 6, characterized in that, The abundance of the four proteins was calculated using intensity-based absolute quantification (iBAQ) and normalized to the total protein fraction (FOT).
8. The system according to any one of claims 5 to 7, characterized in that, The prediction model is constructed using a logistic regression algorithm to build a classifier. The total samples are divided into a training set of 60% and an independent test set of 40%. The FOT-normalized expression levels of four proteins are used as input features, and the model is trained using a 10-fold cross-validation method.
9. The system according to claim 5, characterized in that, When the output probability in the prediction result is greater than the set cutoff value, it is judged as having a high risk of recurrence.
10. The system according to claim 5, characterized in that, The calculation formula for the prediction model is as follows: Logit(P)=0.4591+2.5858 C9-0.5185 MMP12+2.0634 MMP15+0.6096 FLT4 The protein concentration in the serum of non-small cell lung cancer patients detected by mass spectrometry was substituted into the prediction model formula. When the Logit(P) value was ≥0.767, the patient was identified as a person who was prone to recurrence after surgery, and when the Logit(P) value was <0.767, the patient was identified as a person who was not prone to recurrence after surgery.