Marker for diagnosing and / or predicting invasive lung cancer, prediction model and establishment method
By constructing a random forest and logistic regression model based on blood test data, markers such as FIB, TBIL, CEA and CYFRA21-1 were screened out, and a comprehensive prediction model for predicting microvascular invasion of invasive lung cancer was established. This model addresses the shortcomings of existing technologies in the diagnosis and prediction of early invasive lung cancer and achieves higher prediction accuracy.
Patent Information
- Application Number
- CN202311608981.3
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2023-11-29
- Publication Date
- 2025-09-12
AI Technical Summary
Existing clinical prediction models for lung cancer are insufficient in the diagnosis and prediction of early invasive lung cancer, especially in the lack of effective blood test data combined models for identifying microvascular infiltration.
By collecting blood laboratory test data from patients with lung nodules diagnosed with lung cancer, markers such as FIB, TBIL, CEA and CYFRA21-1 were screened, and random forest models and binary logistic regression models were constructed to establish a comprehensive prediction model for predicting microvascular invasion of invasive lung cancer.
This model can effectively predict microvascular invasion of invasive lung cancer, with an area under the ROC curve of 0.768, showing good predictive accuracy and clinical application value.
Smart Images

Figure CN120636784A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the field of biotechnology, and in particular to a marker for diagnosing and / or predicting invasive lung cancer, a prediction model and an establishment method thereof. Background Art
[0002] Lung cancer is currently one of the most common and fatal malignancies worldwide. With an aging population and improved medical care, the incidence and burden of lung cancer in China are steadily increasing. For surgical treatment of lung cancer, clear and appropriate early diagnosis, early treatment, and surgical scheduling can significantly improve prognosis. High-resolution CT is currently the most commonly used technique for lung cancer diagnosis. The primary imaging manifestation of early-stage lung cancer is ground-glass opacity nodules (GGNs). However, chest CT is not very effective in distinguishing tumor status and early-stage subtypes, especially in small cell lung cancer (SCLC). Most lung cancers are often accompanied by microvascular invasion (MVI). Early-stage adenocarcinomas can be pathologically divided into three subtypes based on the degree of lesion invasion: carcinoma in situ (AIS), minimally invasive adenocarcinoma (MIA), and invasive adenocarcinoma (IAC). Treatment and prognosis vary significantly depending on the pathological type and stage of lung cancer.
[0003] Surgical resection is the treatment of choice for lung cancer, but current technology makes it difficult to determine the extent of tumor tissue resection in advance, resulting in a high postoperative recurrence rate. With the advancement of precision and personalized cancer treatment, the treatment of early-stage lung cancer is gradually moving towards minimally invasive and precise methods. Improving patients' postoperative quality of life without increasing their burden of diagnosis and treatment is a hot topic in current clinical research. Currently, diagnostic prediction models based on imaging features are still inaccurate for invasive lung cancer. Clinical diagnosis still relies on postoperative histopathological examination results as the gold standard, and there is still a certain lag in the accurate understanding and precise treatment of lung cancer. The precise pathological mechanisms underlying tumor progression and metastasis are currently unclear, and our understanding of preinvasive and invasive lesions is limited to postoperative pathological diagnosis. Therefore, combining common preoperative laboratory indicators with the patient's ability to predict the nature of the nodule is crucial for planning surgical treatment for lung cancer and improving patient outcomes.
[0004] Existing clinical prediction models for lung cancer include the Mayo Clinic model, the VA model, and the Peking University People's Hospital (PKUPH) model. Their primary algorithms rely on CT image morphology to identify and predict lung cancer (see Zhu Peng. Analysis of Risk Factors for Malignant Solitary Pulmonary Nodules and Comparison of Related Lung Cancer Prediction Models [D]; Nanchang University, 2019). Their ability to distinguish and predict early-stage lung cancer invasion is limited. No model exists that accurately predicts MVI of lung cancer by combining routine preoperative blood test data.
[0005] In this study, based on routine clinical blood test data, we aimed to establish a comprehensive prediction model for whether lung nodules are combined with MVI diagnosis, conduct predictive analysis on different pathological subgroups, and establish a nomogram clinical application tool. Summary of the Invention
[0006] The primary purpose of the present invention is to provide a marker for diagnosing and / or predicting invasive lung cancer, wherein the marker comprises FIB, TBIL, CEA and CYFRA21-1.
[0007] A second object of the present invention is to provide a method for establishing a prediction model for invasive lung cancer, wherein the method comprises:
[0008] (1) To collect variables of patients with lung nodules diagnosed with lung cancer, including baseline data, blood laboratory test data, all routine blood biochemical function test parameters, lung cancer-related serum tumor markers, and pathological data;
[0009] (2) Patients in step (1) were divided into experimental group and control group according to whether the pathological examination results were combined with cytological microvascular infiltration;
[0010] (3) Based on the baseline data obtained in step (1), 1:1 propensity score matching was performed to screen patients in the experimental and control groups;
[0011] (4) constructing a random forest model to identify the importance of the variables described in step (1) and screen out the variables;
[0012] (5) Ten-fold cross validation and internal test set validation were used to verify the reliability of the variables screened in step (4);
[0013] (6) Construct a binary logistic regression model, incorporate the variables obtained in step (4) into the binary logistic regression model in univariate and multivariate forms, and select variables with p < 0.05;
[0014] (7) constructing a nomogram using the variables described in step (6);
[0015] (8) After the nomogram model described in step (7) is established, the model is validated using an external validation group.
[0016] Preferably, the baseline clinical characteristics of the patients are compared using chi-square test and t-test before and after propensity score matching in step (3).
[0017] Preferably, the random forest in step (4) optimizes each variable around its default value.
[0018] Preferably, the variables screened in step (6) include FIB, TBIL, CEA and CYF-211.
[0019] The third object of the present invention is to provide a prediction model for invasive lung cancer established by the method.
[0020] A fourth object of the present invention is to provide a system for constructing a prediction model for invasive lung cancer, which is applied to the construction method, and comprises:
[0021] A data acquisition module is at least used for data acquisition and obtaining a sample data set;
[0022] A data processing module, at least used to extract valid samples from the sample data set that can be used to build an evaluation model;
[0023] A model building module is at least used to randomly divide the incomplete data set of the valid samples into a training set and a validation set, fit the training set using a random forest method, and record the optimal model parameters based on the out-of-bag error;
[0024] The threshold calculation module is at least used to calculate the model classification threshold using the validation set according to the ROC curve.
[0025] A fourth object of the present invention is to provide a prediction model diagnostic system for invasive lung cancer, comprising:
[0026] It is a pre-input module of the evaluation model, at least used to input the data to be diagnosed;
[0027] The invasive lung cancer diagnostic model constructed by the method is at least used to evaluate the data to be evaluated;
[0028] The display module is at least used to display the diagnosis result.
[0029] A fifth object of the present invention is to provide a computer device comprising a memory, a processor, and a computer program stored in the memory and executable on the processor, wherein when the processor executes the program, the system implements the following steps, including:
[0030] (1) Collect and / or input variable data of patients with lung nodules diagnosed with lung cancer;
[0031] (2) Substitute the variable data into the model for calculation to obtain the conclusion of whether lung cancer has pathological microvascular invasion;
[0032] (3) Output the conclusion of whether lung cancer is invasive.
[0033] A sixth object of the present invention is to provide a computer-readable storage medium on which the computer device is stored.
[0034] The present invention has the following beneficial effects: The present invention provides a method for establishing a prediction model for invasive lung cancer. The prediction model established by the method is used to predict invasive lung cancer. The results show that FIB, TBIL, CEA, and CYF-211 are independent risk factors for MVI infiltration of lung cancer nodules. The area under the receiver operating characteristic (ROC) curve (AUC) of the model is 0.768 (95% CI, 0.706-0.830). The calibration curve and clinical decision and clinical impact curves demonstrate that the model has good predictive accuracy and predictive differential diagnostic value. The independent validation results show that the AUC curve for independent validation prediction is 0.851 (95% CI, 0.724-0.934). The calibration curve shows that the model's predicted probability and actual probability are highly consistent. DCA and CIC curve tests show that the model has a high clinical net benefit and the accuracy is within the fitting range. The model meets clinical application requirements. BRIEF DESCRIPTION OF THE DRAWINGS
[0035] Figure 1 RF variable screening results
[0036] Figure 2 Model prediction discrimination ROC curve
[0037] Figure 3 Model prediction accuracy calibration curve
[0038] Figure 4 Model prediction accuracy DCA curve
[0039] Figure 5 Model prediction accuracy CIC curve
[0040] Figure 6 Model Nomogram Prediction Chart
[0041] Figure 7 ROC curve of model external validation
[0042] Figure 8 Calibration curve for external validation of the model
[0043] Figure 9 DCA curve of model external validation
[0044] Figure 10 CIC curve of model external validation DETAILED DESCRIPTION
[0045] The present invention is described in detail below with reference to the accompanying drawings and specific embodiments. This embodiment is implemented based on the technical solution of the present invention, and provides a detailed implementation method and specific operation process, but the protection scope of the present invention is not limited to the following embodiments.
[0046] Example 1: A prediction model for invasive lung cancer
[0047] 1. Materials and Methods
[0048] 1.1 General Information
[0049] This study used PSM and a retrospective cohort study approach. Patients with lung nodules confirmed by surgical pathology for lung cancer admitted to the Department of Thoracic Surgery, Lanzhou University First Hospital, between January 2018 and December 2021 were included in the study. Patients who met the inclusion and exclusion criteria were also included from January to June 2022 to validate the predictive value of the model. The study adhered to the Declaration of Helsinki and was approved by the Ethics Committee of Lanzhou University First Hospital (LDYYLL2021-257). All patients provided written informed consent before surgery. The clinical prediction model for this study adhered to the TRIPOD principles.
[0050] 1.2 Inclusion and exclusion standards
[0051] Inclusion criteria: (1) patients with lung nodules who underwent surgical treatment in our hospital and whose postoperative pathological examination confirmed lung cancer; (2) patients who underwent blood tests before surgery;
[0052] Exclusion criteria: (1) history of other tumors and metastasis; (2) preoperative chemotherapy and immunotherapy; (3) other major systemic diseases; (4) incomplete clinical data that cannot meet the diagnostic requirements.
[0053] 1.3 Research Methods
[0054] Upon admission, all patients underwent a detailed medical history and a comprehensive physical examination. The recorded data were collected and collated independently by two researchers. 80 parameters were collected and assessed before modeling: (1) Baseline data of patients included age, gender, history of hypertension, diabetes, smoking history, alcohol consumption history, and family history. (2) Blood laboratory tests included all coagulation parameters and routine blood test data. (3) All routine blood biochemical function test parameters. (4) Lung cancer-related serum tumor markers included carcinoembryonic protein (CEA), cytokeratin 19 fragment antigen 21-1 (CYFRA21-1), carbohydrate antigen 19-9 (CA19-9), neuron-specific enolase (NSE), and ferritin (FER). (5) Pathological data included lung cancer pathological subtype, tumor differentiation degree, tumor MVI infiltration, TNM stage, etc. Patients were divided into combined MVI and non-MVI control groups according to whether the postoperative pathological diagnosis of infiltration was present.
[0055] 1.4 Observation indicators and evaluation criteria
[0056] Observation indicators included: (1) the presence of MVI, lung cancer pathological classification and stage by tumor pathological microscopy and immunohistochemistry; and (2) comparison of patients' baseline data and differences in PSM.
[0057] 2 Prediction model establishment
[0058] 2.1 Grouping
[0059] Patients were divided into an experimental group and a control group based on whether the pathological examination results included cytological microvascular invasion. A 1:1 propensity score matching was performed based on the available baseline data. The chi-square test and t-test were used to compare the baseline clinical characteristics of the patients before and after PSM to exclude collinearity and further screen the patients in the experimental and control groups.
[0060] 2.2 Variable screening
[0061] Machine learning algorithms were implemented in R (version 4.2.1). Random forests were used to identify variable importance features and to screen variables based on their impact on outcome predictions. Random forests were optimized around their default values (using a maximum of 500 trees and a random space with a dimension equal to the square of the number of features, rounded off). Ten-fold cross-validation and internal test set validation were used to verify the reliability of the variable screening results.
[0062] 2.3Logistic Regression Modeling
[0063] A binary logistic regression model was constructed. The variables obtained in 2.2 were analyzed in univariate and multivariate form using stepwise regression to identify variables and test for association (backward method, p < 0.05). A P < 0.05 was considered statistically significant, and the adjusted odds ratio (OR) and corresponding 95% confidence interval (95% CI) were calculated. The Hosmer-Lemeshow test was used to assess model variability. The receiver operating characteristic (ROC) curve and area under the curve (AUC) were used to test the predictive discrimination of the model. The predictive accuracy of the model was tested by plotting decision and clinical impact curves using predicted and actual probabilities.
[0064] 2.4 External Validation
[0065] Patients who were admitted to the First Hospital of Lanzhou University between January and June 2022 according to the same inclusion and exclusion criteria were included as an independent external validation group. Receiver operating characteristic (ROC) curves, direct current analysis (DCA) curves, and circumstantial influence curve (CIC) curves were used for analysis. The model accuracy was verified by applying a modeling nomogram and calculating the prediction value.
[0066] 2.5 Subgroup Analysis
[0067] After the model was constructed, subgroups were set based on histocytopathological classification. The model's predictive accuracy for different lung cancer subtypes was evaluated by exploring the patient's pathological diagnosis type, including squamous cell carcinoma, adenocarcinoma, and other tumor types.
[0068] 3. Results
[0069] 3.1 Patient characteristics
[0070] Comparison of PSM status and general characteristics between the two groups before and after matching: Of the 428 patients included in the modeling study, 285 were diagnosed with MVI by pathological examination, and 143 were diagnosed without MVI. The two groups did not differ in gender, hypertension, diabetes, family history, alcohol consumption, or history of preoperative chronic lung disease. However, patients with a history of smoking and preoperative age had a significantly higher incidence of cancer (Table 1). After a 1:1 matching of the two groups based on these factors, 108 patients were obtained in each group. No significant differences in baseline characteristics were found between the two groups (all P > 0.05).
[0071] Table 1 Analysis of baseline data of all patients
[0072]
[0073]
[0074] 3.2 Variable screening
[0075] After 1:1 PSM, 80 variables of the included patients were screened, collinearity was excluded before modeling, and random forest modeling was used for screening. The minimum feature set was obtained, and the correlation of each variable was calculated and ranked in turn; 30 variables were identified by the RF method, such as Figure 1 These steps are shown in Figure 1 The "randomForest" package implementation is shown.
[0076] 3.3Logistic Predictive Modeling
[0077] The variables initially screened by the random forest algorithm were included in univariate and multivariate binary logistic regression analysis. For patients with MVI of lung cancer, serum CEA level higher than 2.3 μg / L (OR = 1.201; 95% CI, 1.016-1.419) had a sensitivity of 60.6% and a specificity of 72.1%. Serum CYF-211 level higher than 2.7 μg / L (OR = 1.259; 95% CI, 1.021-1.552) had a sensitivity of 58.33% and a specificity of 72.22%. Serum total bilirubin (Total A bilirubin (TBIL) level higher than 15.3 U / L (OR = 1.042; 95% CI, 1.011-1.074), with a sensitivity of 60.20% and a specificity of 66.70%, and a fibrinogen (FIB) level higher than 3.05 (OR = 0.1.651; 95% CI, 1.067-2.555), with a sensitivity of 52.80% and a specificity of 77.80%, were high-risk predictors of lung cancer nodule invasion, as shown in Table 2.
[0078] Table 2 Univariate and multivariate logistic regression analysis
[0079]
[0080]
[0081] 3.4 Model prediction performance evaluation
[0082] After binary logistic regression modeling, the predictive performance of the model was evaluated, mainly including the model prediction discrimination and accuracy evaluation. In terms of prediction discrimination, the AUC of the model ROC curve was 0.768 (95% CI, 0.706-0.830), and the Hosmer-Lemeshow test p = 0.215 ( Figure 2 The prediction accuracy test is mainly carried out through the calibration curve ( Figure 3 ), the model prediction accuracy is basically in line with the expected ideal state. For clinical application, the clinical decision curve ( Figure 4) and clinical impact curves ( Figure 5 ) showed that the model is consistent with clinical reality. For the clinical application of the model, a nomogram prediction tool for MVI invasion of lung cancer was established by calculation. The risk of invasion of patients can be judged by four risk factors ( Figure 6 ), the results will help to preliminarily identify whether the patient's nodules have MVI before surgery and assist in clinical diagnosis and treatment.
[0083] 3.5 External validation of the model
[0084] A total of 52 patients with 1:1 PSM matching (26 in the MVI group and 26 in the control group) were included. The nomogram was validated by incorporating the predictive factors of the training model, and ROC, DCA, and CIC curve analysis were performed. The predicted AUC curve was 0.851 (95% CI, 0.724-0.934), the sensitivity was 73.08%, the specificity was 84.62%, and the Hosmer-Lemeshow test p = 0.166. The calibration curve, DCA, and CIC curve tests showed that the accuracy was within the fitting range. The results of the external cohort study showed that the model met the requirements for clinical application ( Figure 7-10 ).
[0085] 3.6 Subgroup Analysis
[0086] To test the model's predictive power for lung cancer of different cytopathological types, different subgroups of patients were included in the modeling and prediction analysis was performed according to pathological type. For 118 patients with adenocarcinoma, the model showed good predictive performance, with an AUC of 0.776 (95% CI 0.731-0.816), significantly improving clinical diagnosis. For 84 patients with squamous cell carcinoma, the model's predictive power was poor, with an AUC of 0.632 (95% CI 0.582-0.679). For patients with other types of lung tumors, the model's prediction AUC was 0.643 (95% CI 0.593-0.690).
[0087] In summary, the present invention provides a method for establishing a prediction model for invasive lung cancer. The prediction model established by the method is used to predict invasive lung cancer. The results show that FIB, TBIL, CEA, and CYF-211 are independent risk factors for MVI infiltration of lung cancer nodules. The area under the receiver operating characteristic (ROC) curve (AUC) of the model is 0.768 (95% CI, 0.706-0.830). The calibration curve and clinical decision and clinical impact curves demonstrate that the model has good predictive accuracy and predictive differential diagnostic value. The external validation results show that the AUC curve for independent external validation prediction is 0.851 (95% CI, 0.724-0.934). The calibration curve shows that the model's predicted probability and actual probability are highly consistent. DCA and CIC curve tests show that the model has a high clinical net benefit and the accuracy is within the fitting range. The model meets the requirements for clinical application.
[0088] Those skilled in the art will appreciate that the logic and / or steps represented in the flowcharts or otherwise described herein, for example, can be considered as an ordered list of executable instructions for implementing the logical functions, and can be embodied in any computer-readable medium for use by, or in conjunction with, an instruction execution system, apparatus, or device (e.g., a computer-based system, a system including a processor, or other system that can fetch and execute instructions from an instruction execution system, apparatus, or device). For purposes of this specification, a "computer-readable medium" can be any device that can contain, store, communicate, propagate, or transport a program for use by, or in conjunction with, an instruction execution system, apparatus, or device.
[0089] More specific examples (a non-exhaustive list) of computer-readable media include the following: an electrical connection with one or more wires (electronic devices), a portable computer disk cartridge (magnetic devices), a random access memory (RAM), a read-only memory (ROM), an erasable and programmable read-only memory (EPROM or flash memory), a fiber optic device, and a portable compact disc read-only memory (CDROM). In addition, the computer-readable medium may even be paper or other suitable medium on which the program is printed, since the program may be obtained electronically, for example, by optically scanning the paper or other medium, followed by editing, deciphering, or processing in another suitable manner as necessary, and then stored in a computer memory.
[0090] It should be understood that various parts of the present invention can be implemented using hardware, software, firmware, or a combination thereof. In the above-described embodiments, multiple steps or methods can be implemented using software or firmware stored in a memory and executed by a suitable instruction execution system. For example, if implemented using hardware, as in another embodiment, any one of the following technologies known in the art or a combination thereof can be used to implement the hardware: a discrete logic circuit having a logic gate circuit for implementing a logic function on a data signal, an application-specific integrated circuit having a suitable combination of logic gate circuits, a programmable gate array (PGA), a field programmable gate array (FPGA), etc.
[0091] Throughout this specification, reference to terms such as "one embodiment," "some embodiments," "examples," "specific examples," or "some examples" means that a specific feature, structure, material, or characteristic described in conjunction with that embodiment or example is included in at least one embodiment or example of the present invention. In this specification, schematic representations of the above terms do not necessarily refer to the same embodiment or example. Furthermore, the specific features, structures, materials, or characteristics described may be combined in any suitable manner in any one or more embodiments or examples.
[0092] The above-described embodiments merely illustrate several implementations of the present invention, and while their descriptions are relatively specific and detailed, they should not be construed as limiting the scope of the present invention. It should be noted that a person skilled in the art would be able to make numerous variations and improvements without departing from the spirit of the present invention, all of which fall within the scope of protection of the present invention. Therefore, the scope of protection of the present invention shall be determined by the appended claims.
Claims
1. A marker for diagnosing and / or predicting invasive lung cancer, characterized in that: The markers include FIB, TBIL, CEA and CYFRA21-1.
2. A method for establishing a prediction model for invasive lung cancer, characterized in that: The prediction model establishment method is: (1) To collect variables of patients with lung nodules diagnosed with lung cancer, including baseline data, blood laboratory test data, all routine blood biochemical function test parameters, lung cancer-related serum tumor markers, and pathological data; (2) Patients in step (1) were divided into experimental group and control group according to whether the pathological examination results were combined with cytological microvascular infiltration; (3) Based on the baseline data obtained in step (1), 1:1 propensity score matching was performed to screen patients in the experimental and control groups; (4) constructing a random forest model to identify the importance of the variables described in step (1) and screen out the variables; (5) Ten-fold cross validation and internal test set validation were used to verify the reliability of the variables screened in step (4); (6) Construct a binary logistic regression model, incorporate the variables obtained in step (4) into the binary logistic regression model in univariate and multivariate forms, and select variables with p < 0.05; (7) constructing a nomogram using the variables described in step (6); (8) After the nomogram model described in step (7) is established, the model is validated using an external validation group.
3. The method for establishing a prediction model for invasive lung cancer according to claim 2, wherein: The chi-square test and t-test were used to compare the baseline clinical characteristics of the patients before and after propensity score matching in step (3).
4. The method for establishing a prediction model for invasive lung cancer according to claim 2, wherein: The random forest described in step (4) optimizes each variable around its default value.
5. The method for establishing a prediction model for invasive lung cancer according to claim 2, wherein: The variables screened in step (6) include FIB, TBIL, CEA and CYFRA21-1.
6. A prediction model for invasive lung cancer established by the method according to any one of claims 2 to 5.
7. A system for constructing a prediction model for invasive lung cancer, applied to the construction method according to any one of claims 2 to 5, comprising: A data acquisition module is at least used for data acquisition and obtaining a sample data set; A data processing module, at least used to extract valid samples from the sample data set that can be used to build an evaluation model; A model building module is at least used to randomly divide the incomplete data set of the valid samples into a training set and a validation set, fit the training set using a random forest method, and record the optimal model parameters based on the out-of-bag error; The threshold calculation module is at least used to calculate the model classification threshold using the validation set according to the ROC curve.
8. A prediction model diagnostic system for invasive lung cancer, characterized in that: include: It is a pre-input module of the evaluation model, at least used to input the data to be diagnosed; The invasive lung cancer diagnostic model constructed by the method according to any one of claims 2 to 5 is used at least to evaluate the data to be evaluated; The display module is at least used to display the diagnosis result.
9. A computer device comprising a memory, a processor, and a computer program stored in the memory and executable on the processor, wherein: When the processor executes the program, the system according to claim 7 or 8 implements the following steps, including: (1) Collect and / or input variable data of patients with lung nodules diagnosed with lung cancer; (2) Substitute the variable data into the model for calculation to obtain the conclusion of whether lung cancer has pathological microvascular invasion; (3) Output the conclusion of whether lung cancer is invasive.
10. A computer-readable storage medium having stored thereon the computer device according to claim 9.