Methods and systems for predicting prognosis of lung adenocarcinoma patients receiving EGFR-TKI targeted therapy
By constructing a combined prediction model based on pathological genomics and clinical data, the problem of difficult to accurately predict the prognosis of patients with lung adenocarcinoma treated with EGFR-TKI was solved, achieving more accurate prognostic assessment and personalized treatment guidance.
Patent Information
- Application Number
- CN202510721984.0
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-05-30
- Publication Date
- 2025-10-17
AI Technical Summary
Existing technologies make it difficult to accurately predict the prognosis of lung adenocarcinoma patients receiving EGFR-TKI targeted therapy, especially for patients with EGFR mutations, where the efficacy varies significantly, resulting in poor treatment outcomes.
A prediction model based on pathological genomic features was constructed. Pathological genomic features were extracted from the area with the highest tumor cell density, and the pathomicScore was obtained by training the Cox proportional hazard regression model. A combined prediction model was constructed in combination with clinical data to evaluate patient prognosis.
It improves the accuracy of prognostic prediction for patients with lung adenocarcinoma treated with EGFR-TKI, can guide personalized treatment strategies, and improve the ability to predict progression-free survival.
Smart Images

Figure BDA0005429685430000031 
Figure BDA0005429685430000161 
Figure BDA0005429685430000191
Abstract
Description
TECHNICAL FIELD
[0001] The present application relates to the field of computational pathology, cancer prognosis prediction, and in particular, to a method and system for predicting the prognosis of lung adenocarcinoma patients receiving EGFR-TKI targeted therapy. BACKGROUND
[0002] Lung cancer is the most common cancer worldwide and remains a leading cause of cancer-related deaths. Non-small cell lung cancer (NSCLC) is the most common pathological subtype, accounting for about 85% of all lung cancer cases. In NSCLC patients, epidermal growth factor receptor (EGFR) mutations play a key role in tumor progression, and about 12.8%-49.1% of cases will have mutations according to different populations. These mutations are more common in lung adenocarcinoma (LUAD) than in other NSCLC subtypes. Epidermal growth factor receptor tyrosine kinase inhibitors (EGFR-TKI) have revolutionized the treatment of EGFR-mutant NSCLC, with significantly prolonged progression-free survival (PFS) and overall survival (OS) compared with traditional chemotherapy.
[0003] However, despite these advances, relapse and drug resistance of EGFR-TKI remain a major challenge, severely limiting its long-term efficacy. NSCLC shows significant inter-tumor heterogeneity at both pathological and molecular levels, leading to significant differences in the efficacy of EGFR-TKI among patients. Even patients with the same pathological stage and EGFR mutation site may have significantly different efficacy when treated with the same EGFR-targeted drugs. This difference highlights the complexity of NSCLC and the ongoing challenge of accurately predicting treatment efficacy.
[0004] Histopathological analysis of tumor biopsy samples remains the basis for cancer diagnosis and prognosis. While traditional histological grading methods provide important insights, the advent of computational pathology has revolutionized the field, enabling the extraction of quantitative features from digital histopathology images, a method now referred to as "pathomics". This rapidly evolving discipline combines pathology with high-throughput image analysis, computational modeling, and machine learning, providing a multidisciplinary framework for extracting and interpreting complex pathological data. By capturing tumor heterogeneity and microenvironment features that human observers often fail to detect, pathomics holds promise for discovering key prognostic markers and improving precision oncology. Pathomics has shown great potential in predicting treatment response, survival outcomes, and molecular subtypes of various cancers.
[0005] Therefore, there is an urgent need in the art for a method and system for predicting the prognosis of LUAD patients, especially EGFR-mutant LUAD patients, receiving EGFR-TKI therapy using pathomics-based biomarkers. SUMMARY
[0006] The present application aims to provide a method and system for predicting the prognosis of a lung adenocarcinoma patient with EGFR mutation receiving EGFR-TKI targeted therapy.
[0007] In a first aspect, the present application provides a method for constructing a prediction model for the prognosis of a lung adenocarcinoma patient receiving EGFR-TKI targeted therapy, the method comprising the steps of:
[0008] (A1) using pathological omics data from lung adenocarcinoma patients receiving EGFR-TKI targeted therapy as a training set;
[0009] (A2) preprocessing the pathological omics data, comprising the steps of:
[0010] (a2.1) selecting a region with the highest tumor cell density from the pathological omics data; the region with the highest tumor cell density is a region with a tumor cell density of ≥100 per unit area;
[0011] (a2.2) extracting pathological omics features from the region with the highest tumor cell density;
[0012] (a2.3) normalizing the pathological omics features;
[0013] (a2.4) screening the pathological omics features to obtain preferred pathological omics features;
[0014] (A3) using the preferred pathological omics features for Cox proportional hazards regression model training to obtain a pathological omics feature model (pathomicScore).
[0015] In another preferred embodiment, the lung adenocarcinoma patient comprises a human.
[0016] In another preferred embodiment, the lung adenocarcinoma patient contains an EGFR mutation.
[0017] In another preferred embodiment, the EGFR mutation is selected from the group consisting of 19-del, 21-L858R.
[0018] In another preferred embodiment, the EGFR-TKI is selected from the group consisting of first-generation EGFR-TKI, second-generation EGFR-TKI, and third-generation EGFR-TKI.
[0019] In another preferred embodiment, the first-generation EGFR-TKI comprises gefitinib and icotinib.
[0020] In another preferred embodiment, the second-generation EGFR-TKI comprises afatinib and dacomitinib.
[0021] In another preferred embodiment, the third-generation EGFR-TKI comprises: osimertinib, amatinib, furmetinib.
[0022] In another preferred embodiment, the pathologyomic data comprises: a pathology tissue section, a digital pathology section.
[0023] In another preferred embodiment, the digital pathology section is obtained from the pathology tissue section using a digital scanner.
[0024] In another preferred embodiment, the pathologyomic data is a digital pathology section.
[0025] In another preferred embodiment, the unit area is per square millimeter.
[0026] In another preferred embodiment, the unit area is 512x512 pixels.
[0027] In another preferred embodiment, in step (a2.2), the pathologyomic features are extracted using tools comprising: CellProfiler.
[0028] In another preferred embodiment, the pathologyomic features comprise: image quality features, image co-localization features, image granularity features, or a combination thereof.
[0029] In another preferred embodiment, the pathologyomic features are image quality features, image co-localization features, and image granularity features.
[0030] In another preferred embodiment, in step (a2.2), 90 of the pathologyomic features are extracted.
[0031] In another preferred embodiment, in step (a2.2), 33 of the image quality features, 9 of the image co-localization features, and 48 of the image granularity features are extracted.
[0032] In another preferred embodiment, the formula for normalization is as follows:
[0033]
[0034] where x is the original intensity, f(x) is the normalized intensity, μ and δ are the mean and variance, respectively, and s is an optional scaling factor (default setting is 1).
[0035] In another preferred embodiment, in step (a2.4), it further comprises a step of:
[0036] (s1) using the maximum relevance minimum redundancy method (mRMR) to remove redundancy, retaining pathologyomic features with a correlation coefficient <0.8;
[0037] (s2) identifying pathological features most closely related to the disease progression by using LASSO-Cox regression method (LASSO-Cox regression model), thereby obtaining preferred pathological features.
[0038] In another preferred embodiment, in step (s2), further comprising the step of fine-tuning the LASSO-Cox regression model.
[0039] In another preferred embodiment, the LASSO-Cox regression model is fine-tuned by using five-fold cross-validation to evaluate the LASSO-Cox regression model.
[0040] In another preferred embodiment, the preferred pathological features include:
[0041] (1) Granularity_5_Eosin;
[0042] (2) Correlation_RWC_Eosin_Hematoxylin;
[0043] (3) Granularity_2_Hematoxylin;
[0044] (4) Granularity_11_Eosin;
[0045] (5) Granularity_4_Hematoxylin;
[0046] (6) a combination of (1)-(5) above.
[0047] In another preferred embodiment, the preferred pathological features are:
[0048] (1) Granularity_5_Eosin;
[0049] (2) Correlation_RWC_Eosin_Hematoxylin;
[0050] (3) Granularity_2_Hematoxylin;
[0051] (4) Granularity_11_Eosin; and
[0052] (5) Granularity_4_Hematoxylin.
[0053] In a second aspect of the present application, a system for predicting prognosis of lung adenocarcinoma patients is provided, comprising:
[0054] an input module configured to input data, the data being pathomic data of a subject to be tested;
[0055] a prediction module configured to execute a prediction model for prognosis of a lung adenocarcinoma patient receiving EGFR-TKI targeted therapy, to obtain a pathomic score according to the pathomic data, to determine the prognosis of the subject to be tested according to the pathomic score, and to obtain a prediction result; the prediction model is constructed by the method of the first aspect of the application;
[0056] an output module configured to output the prediction result of the prediction module.
[0057] In a third aspect of the application, a method for constructing a combined prediction model for prognosis of a lung adenocarcinoma patient receiving EGFR-TKI targeted therapy is provided, the method comprising the steps of:
[0058] (B1) constructing a pathomic score model by the method of the first aspect of the application;
[0059] (B2) constructing a clinical data score model, comprising the steps of:
[0060] (B2.1) providing clinical data of the lung adenocarcinoma patient;
[0061] (B2.2) screening the clinical data by a regression method to obtain preferred clinical data;
[0062] (B2.3) converting the preferred clinical data into numerical values to obtain a clinical data score model;
[0063] (B3) integrating the models in step (B1) and step (B2) to obtain a combined model.
[0064] In another preferred embodiment, the lung adenocarcinoma patient comprises a human.
[0065] In another preferred embodiment, the regression method is selected from the group consisting of Cox regression, stepwise regression, or a combination thereof.
[0066] In another preferred embodiment, the Cox regression is selected from the group consisting of univariate Cox regression, multivariate Cox regression, or a combination thereof.
[0067] In another preferred embodiment, the clinical data is screened by univariate Cox regression and multivariate Cox regression in sequence.
[0068] In another preferred embodiment, after screening by univariate Cox regression, the clinical data with P value < 0.05 is selected for screening by multivariate Cox regression.
[0069] In another preferred embodiment, the stepwise regression is selected from the group consisting of forward selection, backward selection, bidirectional stepwise selection, or a combination thereof.
[0070] In another preferred embodiment, the stepwise regression is forward selection and backward selection.
[0071] In another preferred embodiment, the stepwise regression is performed when performing multivariate Cox regression screening.
[0072] In another preferred embodiment, the clinical data comprises pathological stage, smoking history, histological subtype, cytokeratin 19 fragment (CYFRA21-1) level, age, gender, tumor volume, T stage, N stage, M stage, pathological stage, smoking history, ECOG score, visceral pleural invasion (VPI), EGFR mutation type, EGFR-TKI type, carbohydrate antigen 19-9 (CA19-9) level, carcinoembryonic antigen (CEA), squamous cell carcinoma antigen (SCC), neuron-specific enolase (NSE) level, pro-gastrin-releasing peptide (ProGRP) level.
[0073] In another preferred embodiment, the preferred clinical data comprises pathological stage, smoking history, histological subtype, cytokeratin 19 fragment (CYFRA21-1) level.
[0074] In another preferred embodiment, the pathological stage is selected from the group consisting of stage III, stage IV.
[0075] In another preferred embodiment, the histological subtype is selected from the group consisting of acinar type, papillary type, solid type, micro-papillary type.
[0076] In another preferred embodiment, the preferred clinical data is pathological stage.
[0077] In another preferred embodiment, the combined model is nomogram.
[0078] In a fourth aspect, the present application provides a combined prediction system for prognosis of a lung adenocarcinoma patient receiving EGFR-TKI targeted therapy, the system comprising:
[0079] an input module configured to input data, the data comprising pathological omics data and clinical data of a subject to be tested;
[0080] a prediction module configured to process the input data and convert the input data into a score; the prediction module comprising:
[0081] (M1) a prediction unit based on pathomic features, the (M1) prediction unit is configured to execute a prediction model of lung adenocarcinoma patient prognosis based on pathomics, thereby obtaining a pathomic score of the test subject; the prediction model of lung adenocarcinoma patient prognosis based on pathomics is constructed by the method of the first aspect of the present application;
[0082] (M2) a prediction unit based on clinical data, the (M2) prediction unit is configured to convert the clinical data factors into a score;
[0083] The scores obtained by integrating the (M1) and (M2) prediction units are combined to obtain a combined score, and the prognosis of the test subject receiving EGFR-TKI targeted therapy is determined according to the combined score, thereby obtaining a prediction result;
[0084] An output module, the output module is configured to output the prediction result of the prediction module.
[0085] In a fifth aspect of the present application, a method for typing and / or evaluating the prognosis of lung adenocarcinoma patients receiving EGFR-TKI targeted therapy is provided, the method comprising the steps of:
[0086] (C1) providing pathomic data and / or clinical data of a test subject;
[0087] (C2) inputting the pathomic data into a pathomicScore model to obtain a pathomic score, and / or inputting the pathomic data and clinical data into a combined model to obtain a combined score; the pathomicScore model is constructed by the method of the first aspect of the present application, and the combined model is constructed by the method of the third aspect of the present application;
[0088] (C3) comparing the pathomic score and / or the combined score with a reference value or a standard value, thereby typing and / or evaluating the prognosis of lung adenocarcinoma patients receiving EGFR-TKI targeted therapy.
[0089] In another preferred embodiment, the typing divides the test subject into a subject with good prognosis of receiving EGFR-TKI targeted therapy and a subject with poor prognosis of receiving EGFR-TKI targeted therapy.
[0090] In another preferred embodiment, the test subject comprises a human.
[0091] In another preferred embodiment, the method comprises the step of:
[0092] (C1) providing pathomic data of a test subject;
[0093] (C3) comparing the pathologyomic score with a reference value or a standard value, thereby stratifying and / or evaluating the prognosis of the lung adenocarcinoma patient receiving EGFR-TKI targeted therapy;
[0094] (C3) comparing the pathologyomic score with a reference value or a standard value, thereby stratifying and / or evaluating the prognosis of the lung adenocarcinoma patient receiving EGFR-TKI targeted therapy;
[0095] wherein if the pathologyomic score is lower than the reference value or the standard value, it indicates that the detection object is an object with good prognosis receiving EGFR-TKI targeted therapy; if the pathologyomic score is higher than the reference value or the standard value, it indicates that the detection object is an object with poor prognosis receiving EGFR-TKI targeted therapy.
[0096] In another preferred embodiment, the reference value or the standard value is determined according to a training set.
[0097] In another preferred embodiment, the reference value or the standard value is determined by using a tool comprising X-tile.
[0098] In another preferred embodiment, the reference value or the standard value is 1.25.
[0099] In another preferred embodiment, the method comprises the following steps:
[0100] (C1) providing pathologyomic data and clinical data of a detection object;
[0101] (C2) inputting the pathologyomic data and the clinical data into a combined model to obtain a combined score;
[0102] (C3) comparing the combined score with a reference value or a standard value, thereby stratifying and / or evaluating the prognosis of the lung adenocarcinoma patient receiving EGFR-TKI targeted therapy;
[0103] wherein if the combined score is lower than the reference value or the standard value, it indicates that the detection object is an object with good prognosis receiving EGFR-TKI targeted therapy; if the combined score is higher than the reference value or the standard value, it indicates that the detection object is an object with poor prognosis receiving EGFR-TKI targeted therapy.
[0104] In another preferred embodiment, the evaluation comprises predicting the progression-free survival (PFS) of the detection object.
[0105] In another preferred embodiment, the evaluation comprises predicting the progression-free survival of the detection object in a specific year.
[0106] In another preferred embodiment, the year is selected from the group consisting of 0.5 years, 1 year, 1.5 years, or a combination thereof.
[0107] In a sixth aspect, the present application provides a method for adjusting a medication / treatment regimen, comprising:
[0108] (I) providing pathological omics data and / or clinical data from a subject;
[0109] (II) inputting the pathological omics data into a pathomicScore model to obtain a pathological omics score, and / or inputting the pathological omics data and clinical data into a combination model to obtain a combination score; the pathomicScore model is constructed by the method of the first aspect of the present application, and the combination model is constructed by the method of the third aspect of the present application;
[0110] (III) adjusting the medication / treatment regimen according to the pathological omics score and / or the combination score.
[0111] In another preferred embodiment, the subject is a lung adenocarcinoma patient receiving EGFR-TKI targeted therapy.
[0112] In another preferred embodiment, when the pathological omics score and / or the combination score is higher than the reference value or standard value, it indicates that the subject is a patient with poor prognosis after receiving EGFR-TKI targeted therapy, and the treatment regimen comprises: changing to other EGFR-TKI targeted therapy, or combining EGFR-TKI targeted therapy with other treatment methods.
[0113] In another preferred embodiment, the other treatment methods comprise: radiotherapy, chemotherapy, immunotherapy, or other targeted therapy.
[0114] It should be understood that, within the scope of the present application, each of the technical features described above and each of the technical features specifically described below (e.g., in the examples) can be combined with each other to form new or preferred technical solutions. Due to the limited space, they will not be listed one by one here. BRIEF DESCRIPTION OF DRAWINGS
[0115] Figure 1The flow chart of the study cohort selection is shown. The baseline clinical data of the patients were collected, including age, gender, Eastern Cooperative Oncology Group performance status (ECOG PS), chronic obstructive pulmonary disease (COPD); the tumor size, T stage, N stage, M stage, pathological stage, EGFR mutation type, EGFR-TKI drug, visceral pleural invasion (VPI), smoking history, histological subtype, carcinoembryonic antigen (CEA), carbohydrate antigen 19-9 (CA19-9), squamous cell carcinoma antigen (SCC), cytokeratin 19 fragment (CYFRA21-1), neuron-specific enolase (NSE), progastrin-releasing peptide (ProGRP) and survival (progression or death) of the patients were evaluated within 2 weeks before the first TKI treatment.
[0116] Figure 2 The flow chart of the pathomics analysis is shown.
[0117] Figure 3 The selection process of features associated with disease progression is shown. (A) As the parameter λ increases, the binomial bias gradually decreases to a minimum point. (B) As the regularization parameter λ increases, the bias likelihood bias gradually decreases until it reaches a minimum value, corresponding to the optimal value of λ. (C) Pathomics features associated with disease progression selected by LASSO-cox regression and the corresponding coefficients. LASSO, least absolute shrinkage and selection operator.
[0118] Figure 4 The results of the analysis of the association between pathomicsScore and disease progression are shown. (A-B) The distribution level of pathomicsScore between the progression group and the non-progression group. The results show that the median value in the progression group is higher than that in the non-progression group, and the difference is statistically significant. (C-D) According to the pathomicsScore threshold, the patients are divided into high-risk group and low-risk group, and Kaplan-Meier analysis is performed on the high-risk group and low-risk group. The results show that the PFS of the patients in the low-risk group is significantly better than that in the high-risk group. (E-F) Correlation analysis between the total score (Total points) obtained when predicting disease progression according to pathomicsScore and PFS. The results show that the higher the total score obtained, the shorter the PFS. (G-H) The visualization form of the prediction results of the prognosis of LUAD patients receiving EGFR-TKI treatment using pathomicsScore.
[0119] Figure 5 The Kaplan-Meier curves of the high-risk group and the low-risk group with different smoking history (A-B), different CYFRA21-1 levels (C-D), different epidermal growth factor receptor mutation types (E-F) and pathological stages (G) in the training cohort are shown.
[0120] Figure 6 Kaplan-Meier curves of high-risk and low-risk groups in the validation cohort with different smoking history (A-B), different CYFRA21-1 levels (C-D), different epidermal growth factor receptor mutation types (E-F) and pathological stages (G) are shown.
[0121] Figure 7 The results of the establishment and validation of the individualized disease progression prediction model are shown. (A) The results of univariate and multivariate Cox regression analysis. (B) The prediction model. (C-D) The validation results of the sensitivity and accuracy of the model. (E-F) The validation results of the model on representative cases outside the two study cohorts.
[0122] Figure 8 The validation results of the prediction model on the performance of progression-free survival prediction are shown. DETAILED DESCRIPTION
[0123] The inventors, through extensive and in-depth research, have for the first time developed a prediction model based on the pathological features of conventional biopsy samples for evaluating the recurrence risk of EGFR-mutated LUAD patients receiving EGFR-TKI treatment. According to the pathological data, the inventors have for the first time discovered five pathological features, and based on the five features, a patient scoring model pathomicScore has been constructed. The results show that the pathomicScore model is negatively correlated with the progression-free survival of the disease. In addition, by combining the pathomicScore model with a model based on clinical data, a comprehensive prediction model of disease progression risk has been developed, which has good stratification ability for LUAD patients, thereby guiding individualized treatment strategies. On this basis, the present application is completed.
[0124] It should be understood that the following description of specific methods and experimental conditions of the present application in various details is provided to give an understanding of the essence of the present application. The definitions of certain terms used in the present specification are provided below. Unless otherwise defined, all technical and scientific terms used herein have the same meaning as commonly understood by one of ordinary skill in the art to which the present application belongs.
[0125] Terminology
[0126] Where a numerical range is provided, unless otherwise stated, the limits of the range are included in the disclosure. These limits can independently be included in the disclosure, and can be included in the disclosure as recited in any suitable wrapper. For example, "1 to 50" includes "2 to 25," "5 to 20," "25 to 50," "1 to 10," etc.
[0127] As used herein, the terms "containing" or "including" can be open, semi-closed and closed. In other words, the terms also include "consisting essentially of" or "consisting of."
[0128] As used herein, the term "and / or" relates to and encompasses any and all possible combinations of one or more of the associated listed items.
[0129] As used herein, the term "progression-free survival (PFS)" refers to the time from the start of EGFR-TKI treatment to objective disease progression or death from any cause without disease progression, whichever occurs first. Common methods for assessing PFS include: median PFS, i.e., the time at which 50% of patients have not progressed, which is usually represented by a Kaplan-Meier curve; hazard ratio (HR), with HR < 1 indicating a reduced risk of disease progression.
[0130] As used herein, the term "digital pathology slide" refers to a Whole Slide Image (WSI), which is a form of slide formed by scanning a traditional pathological tissue section into a high-resolution digital image using a digital scanner.
[0131] As used herein, the terms "Eastern Cooperative Oncology Group Performance Status Score" and "ECOG Score" can be used interchangeably, and are a scoring table used to quantify a patient's ability to perform daily activities and treatment tolerance, ranging from 0 (completely normal) to 5 (death).
[0132] As used herein, the terms "disease progression" and "disease progression" can be used interchangeably, and refer to the process of gradual deterioration of a disease from its initial state, usually accompanied by aggravation of the disease, exacerbation of symptoms, and impact on the patient's quality of life.
[0133] Method for constructing a prediction model of the prognosis of a lung adenocarcinoma patient receiving EGFR-TKI targeted therapy
[0134] The present application provides a method for constructing a prediction model of the prognosis of a lung adenocarcinoma patient receiving EGFR-TKI targeted therapy, and a method for constructing a combination prediction model.
[0135] The method for constructing the prediction model comprises the steps of:
[0136] (A1) Pathological omics data from lung adenocarcinoma patients receiving EGFR-TKI targeted therapy were used as training sets;
[0137] (A2) Preprocessing the pathological genomics data, including the following steps:
[0138] (a2.1) selecting an area with the highest tumor cell density from the pathological data; the area with the highest tumor cell density is an area with a tumor cell density of ≥100 per unit area;
[0139] (a2.2) extracting pathomic features from the area with the highest tumor cell density;
[0140] (a2.3) normalizing the pathomic features;
[0141] (a2.4) screening the pathological features to obtain preferred pathological features;
[0142] (A3) Using the preferred pathomic features for Cox proportional hazards regression model training to obtain a pathomic feature model (pathomicScore).
[0143] Preferably, the lung adenocarcinoma patient harbors an EGFR mutation. Preferably, the EGFR mutation is selected from the group consisting of 19-del and 21-L858R. 19-del and 21-L858R are two common EGFR mutations, where "19-del" is a deletion of EGFR exon 19, and "21-L858R" is a mutation of the L at position 858 in EGFR exon 21 to an R.
[0144] The lung adenocarcinoma patient is receiving EGFR-TKI targeted therapy. Preferably, the EGFR-TKI is selected from the group consisting of first-generation EGFR-TKI, second-generation EGFR-TKI, and third-generation EGFR-TKI. Preferably, the first-generation EGFR-TKI includes gefitinib and icotinib. Preferably, the second-generation EGFR-TKI includes afatinib and dacomitinib. Preferably, the third-generation EGFR-TKI includes osimertinib, ametinib, and furimetinib.
[0145] The pathology data include: pathology tissue sections, digital pathology sections. Generally, the pathology data obtained directly from the patient can be pathology tissue sections. The method for obtaining pathology tissue sections is well known to those skilled in the art. The digital pathology sections are obtained by scanning the obtained pathology tissue sections. Preferably, the digital pathology sections are obtained from the pathology tissue sections by using a digital scanner. The method for obtaining digital pathology sections from pathology tissue sections is also well known to those skilled in the art.
[0146] The region with the highest tumor cell density is selected from the pathology data for subsequent processing. The method and criteria for selecting the region with the highest tumor cell density from the pathology data (such as pathology tissue sections) are well known or recognized by those skilled in the art (such as pathologists). Generally, the region with tumor cell density ≥ 100 per unit area can be selected as the region with the highest tumor cell density. In the pathology tissue sections, the unit area can be per square millimeter; in the digital pathology sections, the unit area can be a unit area composed of pixels, such as 512 x 512 pixels. The method for selecting the region with the highest tumor cell density can be manual judgment, such as direct judgment by pathologists according to the image, or automatic recognition by systems or software according to the set threshold. It can be understood that any method for selecting the region with the highest tumor cell density from the pathology data is within the scope of the present application.
[0147] The pathology features are extracted from the region with the highest tumor cell density. The commonly used feature extraction method is well known to those skilled in the art. Generally, the pathology features are extracted from the pathology data containing the region with the highest tumor cell density by systems or software. Preferably, the pathology features are extracted by using CellProfiler. It can be understood that any method for extracting the pathology features from the pathology data is within the scope of the present application.
[0148] The present application extracts 90 pathology features, including 33 image quality features, 9 image co-localization features and 48 image granularity features. Among them, the "image quality features" also include image intensity features.
[0149] The pathological omics features are screened after normalization, including steps of: (s1) removing redundancy by using the maximum relevance minimum redundancy method (mRMR), and retaining pathological omics features with a correlation coefficient <0.8; (s2) identifying pathological omics features most closely related to disease progression by using LASSO-Cox regression method (LASSO-Cox regression model), thereby obtaining preferred pathological omics features. Preferably, the LASSO-Cox regression model is fine-tuned. Preferably, the LASSO-Cox regression model is evaluated by using five-fold cross-validation, thereby fine-tuning the LASSO-Cox regression model.
[0150] Five pathological omics features involved in the construction of the prediction model are finally obtained: Granularity_5_Eosin; Correlation_RWC_Eosin_Hematoxylin; Granularity_2_Hematoxylin; Granularity_11_Eosin and Granularity_4_Hematoxylin. The pathological omics feature model pathomicScore is constructed.
[0151] The construction method of the combined prediction model includes steps of:
[0152] (B1) constructing a pathological omics feature model (pathomicScore) by using the method described above;
[0153] (B2) constructing a clinical data score model, including steps of:
[0154] (B2.1) providing clinical data of the lung adenocarcinoma patient;
[0155] (B2.2) screening the clinical data by using a regression method, thereby obtaining preferred clinical data;
[0156] (B2.3) converting the preferred clinical data into numerical values, thereby obtaining a clinical data score model;
[0157] (B3) integrating the models in step (B1) and step (B2), thereby obtaining a combined model.
[0158] Preferably, the clinical data comprises: pathological stage, smoking history, histological subtype, cytokeratin 19 fragment (CYFRA21-1) level, age, gender, tumor volume, T stage, N stage, M stage, pathological stage, smoking history, ECOG score, visceral pleural invasion (VPI), EGFR mutation type, EGFR-TKI type, carbohydrate antigen 19-9 (CA19-9) level, carcinoembryonic antigen (CEA), squamous cell carcinoma antigen (SCC), neuron-specific enolase (NSE) level, pro-gastrin-releasing peptide (ProGRP) level.
[0159] Preferably, the regression method is selected from the group consisting of Cox regression, stepwise regression, or a combination thereof. Preferably, the Cox regression is univariate Cox regression and multivariate Cox regression. Preferably, the clinical data is screened by univariate Cox regression and multivariate Cox regression in sequence. Preferably, after screening by univariate Cox regression, the clinical data with P value < 0.05 is selected for screening by multivariate Cox regression. Preferably, the multivariate Cox regression is combined with stepwise regression for screening. Generally, multivariate Cox regression combined with stepwise regression can avoid including redundant variables in the constructed model, thereby improving the stability and interpretability of the Cox model. Through regression screening, the preferred clinical data obtained finally comprises: pathological stage, smoking history, histological subtype, cytokeratin 19 fragment (CYFRA21-1) level. Preferably, the preferred clinical data is pathological stage.
[0160] Methods for integrating two models are well known to those skilled in the art. Preferably, after integration, the combined model obtained is nomogram.
[0161] Patient stratification
[0162] The present application provides a method for stratifying lung adenocarcinoma patients receiving EGFR-TKI targeted therapy. The method comprises the steps of:
[0163] (C1) providing pathomic data and / or clinical data of a test subject;
[0164] (C2) inputting the pathomic data into a pathomicScore model to obtain a pathomic score, and / or inputting the pathomic data and clinical data into a combined model to obtain a combined score;
[0165] (C3) comparing the pathomic score and / or combined score with a reference value or standard value, thereby stratifying lung adenocarcinoma patients receiving EGFR-TKI targeted therapy and / or evaluating the prognosis of lung adenocarcinoma patients receiving EGFR-TKI targeted therapy.
[0166] As used herein, the terms "typing" and "stratification" can be used interchangeably, which refers to dividing the detection object into good prognosis objects / objects with longer progression-free survival and poor prognosis objects / objects with shorter progression-free survival based on the prediction results. Preferably, if the pathomic score and / or the combined score is lower than the reference value or standard value, the detection object is a good prognosis object for EGFR-TKI targeted therapy; if the pathomic score and / or the combined score is higher than the reference value or standard value, the detection object is a poor prognosis object for EGFR-TKI targeted therapy.
[0167] As used herein, the terms "high risk" and "high risk" can be used interchangeably, and "low risk" and "low risk" can be used interchangeably. Among them, high risk refers to patients with poor prognosis / shorter progression-free survival; low risk refers to patients with good prognosis / longer progression-free survival. Preferably, if the pathomic score (i.e. pathomicsScore score) / combined score is higher than the threshold value, the reference value or standard value, the detection object is a high risk object; if the pathomic score / combined score is lower than the reference value or standard value, the detection object is a low risk object.
[0168] The main advantages of the present application include:
[0169] (1) The present application first discovers 5 features closely related to disease progression after EGFR-TKI treatment according to the pathomic data of EGFR mutant lung adenocarcinoma patients receiving EGFR-TKI treatment.
[0170] (2) The present application constructs a lung adenocarcinoma pathology score model pathomicsScore based on 5 pathomic features. The model is negatively correlated with disease progression, and the higher the pathomicsScore score, the worse the prognosis of the disease.
[0171] (3) The present application combines the pathomicsScore prediction factor and the preoperative pathological stage prediction factor to construct a combined prediction model. Compared with the pathological stage and the pathomicsScore model alone, the prediction ability of the model for the prognosis of EGFR mutant lung adenocarcinoma patients receiving EGFR-TKI treatment is significantly improved.
[0172] The application will be further described in conjunction with specific examples. It should be understood that these examples are only used to illustrate but not limit the scope of the application. The experimental methods in the following examples, if not specified, are generally carried out according to the conventional conditions, for example, the conditions described in Sambrook et al., Molecular Cloning: A Laboratory Manual (New York: Cold Spring Harbor Laboratory Press, 1989), or the conditions recommended by the manufacturer. Unless otherwise specified, percentages and parts are weight percentages and weight parts.
[0173] Materials and methods
[0174] 1. Patient selection
[0175] This study was approved by the Institutional Review Board and the Human Ethics Committee of the Fifth Affiliated Hospital of Wenzhou Medical University. Since this study is a retrospective study, informed consent is not required. This study retrospectively screened patients with LUAD who were pathologically diagnosed and received EGFR-TKI monotherapy in the Fifth Affiliated Hospital of Wenzhou Medical University from January 2017 to April 2023.
[0176] Patient selection was based on the following criteria: (1) diagnosed with LUAD with EGFR mutation (19-del or 21-L858R) by histopathological examination; (2) age greater than 18 years at the time of diagnosis; (3) received EGFR-TKI monotherapy as first-line treatment; (4) had complete clinical and follow-up data; (5) had specimen slides available for analysis.
[0177] Exclusion criteria were as follows: (1) other concurrent malignancies; (2) insufficient or poor quality tissue samples; (3) had received surgical resection or local regional therapy before starting EGFR-TKI treatment; (4) permanently stopped EGFR-TKI treatment due to intolerable toxicity; (5) died within 6 months after starting EGFR-TKI treatment due to serious complications, which could not be evaluated for disease progression.
[0178] Finally, this study retrospectively identified 122 patients. These patients were randomly divided into a training group (84 patients; 40 had experienced disease progression, and 44 had not experienced disease progression) and a validation group (38 patients; 18 had experienced disease progression, and 20 had not experienced disease progression) at a ratio of 7:3. The case identification process is shown in Figure 1 .
[0179] 2. Treatment regimen and follow-up
[0180] Patients received epidermal growth factor receptor tyrosine kinase inhibitors (EGFR-TKI) as first-line treatment, including first-generation EGFR-TKI (gefitinib, icotinib), second-generation EGFR-TKI (afatinib, dacomitinib) and third-generation EGFR-TKI (osimertinib, amivantib, fostamatinib). Treatment continued until disease progression. The dosing regimen of EGFR-TKI followed the guidelines established by the Chinese Society of Clinical Oncology and the National Comprehensive Cancer Network (NCCN) for non-small cell lung cancer. The primary endpoint of this study was progression-free survival (PFS), defined as the time from the start of EGFR-TKI treatment to objective disease progression or death from any cause without disease progression, whichever occurred first. Disease progression was assessed according to the Response Evaluation Criteria in Solid Tumors (RECIST) version 1.1.
[0181] 3. CT-guided lung biopsy and tissue processing:
[0182] All patients included in this study underwent CT-guided percutaneous lung biopsy at the time of initial diagnosis as part of standard clinical practice for histopathological confirmation. Biopsies were performed by experienced interventional radiologists. Patient position (supine, prone or lateral) was determined based on lesion location to ensure optimal puncture path while avoiding critical structures such as ribs, pulmonary fissures and large blood vessels. Using coaxial technique, a guide needle was passed through the pleura, followed by insertion of a core needle to obtain tissue samples. Multiple low-dose CT scans were performed intraoperatively to guide needle placement and adjust the path, thereby achieving precise lesion localization and minimizing surgical complications. One to three core needles were collected per patient and immediately fixed with 10% neutral buffered formalin solution. All specimens were processed in the pathology department using standard protocols, including dehydration, paraffin embedding, sectioning and hematoxylin-eosin (H&E) staining, to prepare formalin-fixed paraffin-embedded (FFPE) slides for subsequent analysis. All biopsy specimens were from radiologically diagnosed primary lung adenocarcinoma lesions. After biopsy, the puncture needle was removed and manual compression was performed at the puncture site. Sterile dressings were applied and patients were instructed to lie down for observation. Postoperative chest CT scans were routinely performed to assess potential complications such as pneumothorax or hemorrhage.
[0183] 4. Digital pathology image acquisition and region of interest selection:
[0184] H&E-stained sections were prepared from formalin-fixed, paraffin-embedded tissue samples from all enrolled patients. Pathologist A (JJ.C) with more than 10 years of experience independently evaluated each section and selected the most representative areas with the highest tumor cell density as the Region of Interest (ROI) for the histopathological analysis. Subsequently, confirmation was performed by pathologist B (H.Z) with more than 15 years of experience. In case of disagreement, the final decision was made by the department head. Subsequently, formalin-fixed, paraffin-embedded samples corresponding to the selected areas were cut into 5 pm-thick sections and stained with H&E. All selected sections were scanned using the Kibfo ScanScope scanner system (KF-PRO-005) with a x40 objective.
[0185] To reduce the computation time, five non-overlapping representative tiles were randomly selected for each patient, each containing the area with the highest tumor cell density within a 512 x 512 pixel field of view. These tiles were initially selected by pathologist A and subsequently confirmed by pathologist B. If disagreement between the two pathologists occurred, consensus was reached by consultation with the department head. The selected five representative tiles were then saved as.jpg files and color normalized using the Macenko method for subsequent analysis
[0186] 5. Extraction of histopathological features from images:
[0187] Quantitative histopathological features were extracted from the selected tiles using CellProfiler (version 4.1.3), an open-source and free image analysis platform. The H&E-stained images were separated into hematoxylin-stained and eosin-stained grayscale images using the “UnmixColors” module. In addition, the digital H&E-stained images were converted into grayscale images using the “ColorToGray” module based on the “Combine” method for further analysis. First, the “MeasureImageQuality” and “MeasureImageIntensity” modules were used to assess image quality-related features in the grayscale H&E, individual hematoxylin, and eosin images. Subsequently, the “MeasureColocalization” module was used to quantify the colocalization and intensity correlation between the hematoxylin and eosin images on a pixel-by-pixel basis across the entire image. In addition, the “MeasureGranularity” module was used to analyze the granularity features of each image, which generated a spectrum that captured the texture size measurements within the image, covering 16 granularity spectral ranges.
[0188] Specifically, 90 pathology features were extracted, including 33 image quality features, 9 image co-localization features, and 48 image granularity features.
[0189] wherein the image quality features include: (1) MeanIntensity: the mean of pixel intensity within an object; (2) StdIntensity: the standard deviation of pixel intensity within an object; (3) LowerQuartileIntensity: the lower quartile of pixel intensity values within an object; (4) MedianIntensity: the median of intensity within an object; (5) MAD: the median absolute deviation (MAD) of intensity within an object. MAD is defined as median(|xi-median(x)|); (6) UpperQuartileIntensity: the upper quartile of pixel intensity values within an object; (7) FocusScore: a measure of image intensity variance. The score is computed using normalized variance, which is the best ranking algorithm for brightfield, phase contrast, and differential interference contrast images. The higher the focus score, the lower the blurring degree; (8) LocalFocusScore: a measure of intensity variance between sub-regions of an image. The LocalFocusScore is computed by dividing the image into non-overlapping tiles, computing the normalized variance for each tile, and taking the average of these values as the final measure; (9) PowerLogLogSlope: the slope of the log-log power spectrum of an image. The power spectrum contains the frequency information of an image, and the slope measures the blurring degree of the image. The higher the slope, the lower the frequency components, and thus the more blurred the image; (10) Correlation: a measure of correlation of an image at a given spatial scale, which is set to 20 in this study. This is a measure of spatial intensity distribution across sub-regions of an image at a given spatial scale. If an image is blurred, the correlation between neighboring pixels will be higher, resulting in a higher correlation value; (11) ThresholdOtsu: the threshold of each image computed automatically using the Otsu algorithm, which is used to identify the tissue foreground from the unstained background.
[0190] Image co-localization features include: (1) Correlation: the correlation between a pair of hematoxylin (H) and eosin (E) stained images, calculated as Pearson's correlation coefficient. The formula is covariance(H, E) / [std(H) x std(E)]; (2) Slope: the slope of the least squares regression between a pair of H and E images. The formula used is a x H + b = E, where a is the slope; (3) Overlap coefficient: the overlap coefficient is a modification of the Pearson's correlation coefficient, where the average intensity value of a pixel is not subtracted from the original intensity value. For a pair of H and E images, the formula for the overlap coefficient is sum(Hi x Ei) / sqrt(sum(Hi x Hi) x sum(Ei x Ei)); (4) Manders' coefficients: the formula for Manders' coefficients for a pair of H and E images is Ml = sum(Hi_coloc) / sum(Hi) and M2 = sum(Ei_coloc) / sum(Ei), where Hi_coloc = Hi when Ei > 0; otherwise 0; Ei_coloc = Ei when Hi > 0; (5) Manders' coefficients (Costes auto threshold): the Costes auto threshold estimates a maximum intensity threshold for each image based on the correlation value. Manders' coefficients are applied to the thresholded images, where Hi_coloc = Hi when Ei > Ethr; Ei_coloc = Ei when Hi > Hthr. Ethr and Hthr are the thresholds calculated using the Costes auto threshold method; (6) Rank-weighted co-localization coefficients: the formula for rank-weighted co-localization (RWC) coefficients for a pair of H and E images is RWC1 = sum(Hi_coloc*Wi) / sum(Hi) and RWC2 = sum(Ei_coloc*Wi) / sum(Ei); in this formula, Wi is the weight, defined as Wi = (Hmax - Di) / Rmax, where Rmax is the maximum rank between H and E based on maximum intensity, Di = abs(rank(Hi) - rank(Ei)) (the absolute rank difference between H and E), and Hi_coloc = Hi when Ei > 0 or 0 otherwise, Ei_coloc = Ei when Hi > 0 or 0 otherwise.
[0191] Image granularity features include: Granularity, which returns a measurement value for each granulometry instance within the granulometry set range.
[0192] 6. Feature selection and construction of pathomics features:
[0193] To alleviate the unit dependency of each feature, all pathomics features are normalized using a standardization method. The normalization formula is as follows:
[0194]
[0195] where x is the original intensity, f(x) is the normalized intensity, μ and δ are the mean and variance, respectively, and s is an optional scaling factor (default set to 1).
[0196] To identify the key pathological features associated with disease progression in LUAD patients after receiving EGFR-TKI treatment, a systematic two-step procedure was implemented. First, the maximum relevance minimum redundancy (mRMR) method was used to remove highly redundant features, i.e., features with a correlation coefficient greater than 0.8. This step ensured that the maximum amount of information and non-redundant variables were retained for subsequent analysis. Next, the least absolute shrinkage and selection operator (LASSO)-Cox regression model was applied to identify the features most closely associated with disease progression. LASSO is a robust regression method that integrates feature selection and regularization by shrinking the coefficients of variables with less information to zero, thereby excluding them from the model and thus minimizing overfitting. This method is particularly suitable for high-dimensional datasets, especially when there is multicollinearity among features. The LASSO model was fine-tuned using five-fold cross-validation, and the optimal lambda value lambda.min was determined from the lasso_fit output. Finally, based on the pathomics features selected by LASSO-cox regression, a Cox proportional hazards regression model was used to construct a prognostic biomarker.
[0197] 7. Prognostic performance of pathomics features:
[0198] The critical value of the pathomics score was determined in the training cohort using the X-tile software (version 3.6.1, Yale University School of Medicine). This predefined threshold was then applied to the validation cohort to stratify all patients into high and low pathomics score groups for prognostic analysis. The difference in PFS between the high and low pathomics score groups was assessed using the Kaplan-Meier method, and statistically compared by the log-rank test. The correlation between the total score based on the pathomics score and PFS was also analyzed and visualized using a scatter plot. Finally, the predictive performance of the pathomics score in the training and validation cohorts was evaluated and visualized using a confusion matrix.
[0199] 8. Development and validation of individualized recurrence prediction model:
[0200] To improve the prediction accuracy of PFS, pathomics features were combined with clinical risk factors, including laboratory parameters and clinical staging variables, for univariate and multivariate Cox regression analysis to identify independent prognostic factors. Variables with p-value < 0.05 in univariate analysis were included in the multivariate Cox regression model. Stepwise regression method guided by Akaike information criterion (AIC) combined with forward and backward selection strategies was used to identify independent predictors of disease progression after EGFR-targeted therapy. Subsequently, these predictors were integrated into a composite predictive biomarker. Subsequently, nomograms were constructed to visually demonstrate the predictive ability of the model for PFS, thus enabling individualized risk stratification and disease progression prediction. In addition, time-independent receiver operating characteristic (ROC) curves were constructed for the training and validation cohorts to assess the performance of the combined model in predicting disease progression. Calibration curves were generated to assess the agreement between the predicted and observed probabilities of progression based on the combined model. In addition, decision curve analysis was performed to assess the clinical utility of the model developed in this study.
[0201] 9. Statistical analysis:
[0202] Statistical analysis was performed using R software (version 3.6.3) and Python (version 3.7.0). Student's t-test or Mann-Whitney U-test was used for comparison of continuous variables. Chi-square test or Fisher's exact probability method was used for comparison of categorical variables. Kaplan-Meier (KM) curve analysis using log-rank test was used to identify different clinical outcomes. Cox regression analysis was used for univariate and multivariate analysis, and the hazard ratio (HR) and its 95% confidence interval were calculated. All statistical analyses were two-sided, and p-values less than 0.05 were considered statistically significant. The optimal cutoff value for continuous markers was determined using X-tile software (version 3.6.1; Yale University School of Medicine).
[0203] Example 1: Characteristics of the study cohort
[0204] This example summarizes the characteristics of the patients included.
[0205] According to Figure 1A total of 122 patients with pathologically confirmed LUAD and treated with EGFR-TKI were enrolled according to the inclusion and exclusion criteria of the study. The mean age of patients in the training cohort (n=84) and the validation cohort (n=38) was 66.62±10.93 and 67.13±11.85, respectively. The median progression-free survival (mPFS) was 13.2 months with a 95% confidence interval (CI) of 12.3-15.0 months. Among the 122 patients, 58 (47.5%) had disease progression at the time of mPFS of 13.2 months. There was no significant difference in the baseline characteristics between the training cohort and the validation cohort.
[0206] Example 2: Construction of Pathomics Features
[0207] This example relates to extracting pathomics features for the cases, and the analysis flowchart is shown in Figure 2
[0208] For each patient, a total of 90 pathomics features were extracted using CellProfiler. Among the 90 pathomics features, 25 features with a correlation coefficient greater than 0.8 were removed using the mRMR method. Then, the remaining 65 features were used as the input of LASSO-Cox regression to identify the features most closely related to disease progression after EGFR-TKI treatment. The feature selection process of LASSO-Cox in the training cohort is shown in Figure 3 Figure 3 Table 1. Five features and their coefficients
[0209] Table 1. Five features and their coefficients
[0210] feature coefficient Correlation_RWC_Eosin_Hematoxylin 0.3546915 Granularity_11_Eosin 0.2825601 Granularity_2_Hematoxylin -0.3331479 Granularity_4_Hematoxylin 0.2121504 Granularity_5_Eosin 0.4502989
[0211] Subsequently, the pathomic score of each patient was constructed using Cox proportional hazards regression. The pathomic score distributions of the training cohort 1.02 (0.63, 1.43) and the validation cohort 1.01 (0.72, 1.48) were similar (p=0.9493). As shown in Figure 3 A-B, there was a statistically significant difference in the pathomic score between patients with disease progression and patients without progression in the training cohort (p=8.9e-05) and the validation cohort (p=0.017).
[0212] Example 3: Association of Pathomics Score with Disease Progression in Patients Receiving EGFR-TKI Treatment
[0213] This example relates to stratifying patients using the pathomicsScore value and performing correlation analysis and validation of the stratification results.
[0214] The optimal cutoff value was determined to be 1.25 using the pathomicsScore in the training cohort, which corresponds to the highest standardized log-rank statistic. Subsequently, patients in the training cohort and the validation cohort were stratified into a low pathomicsScore group and a high pathomicsScore group. The median progression-free survival (mPFS) of patients in the high pathomicsScore group was 7.0 months, while the median progression-free survival (mPFS) of patients in the low pathomicsScore group was 15.4 months. Kaplan-Meier survival analysis showed that the PFS of patients in the low pathomicsScore group was significantly longer than that of patients in the high pathomicsScore group Figure 4 C-D). Given that the pathomicsScore is negatively correlated with disease progression, the low pathomicsScore is classified as low risk, and the high pathomicsScore is classified as high risk for the sake of clarity and subsequent analysis.
[0215] To evaluate the relationship between the total score predicted by the pathomicsScore and PFS, a correlation analysis was performed. As shown in the scatter plot Figure 4 E-F), the total score of LUAD patients was significantly negatively correlated with PFS. In addition, the pathomicsScore was used to predict disease progression after EGFR-TKI treatment. The prediction results are presented in the form of a confusion matrix, as shown in Figure 4 G-H.
[0216] Example 4: Validation of the results of the pathomicsScore based on progression-free survival
[0217] To further evaluate the clinical interpretability of the pathomicsScore, subgroup analysis was performed based on key clinicopathological variables. In the training cohort Figure 5 ) and the validation cohort Figure 6 ), Kaplan-Meier curves of the high-risk group and the low-risk group were plotted according to different EGFR mutation types, smoking history, CYFRA21-1 levels, and pathological stages, respectively. Figure 6 B is the KM curve result of the high-risk group and the low-risk group among smokers. In these two groups, due to the small number of samples, the results are not significant. In addition, due to the insufficient number of samples with pathological stage III in both the training cohort and the validation cohort, there is no statistical significance, and therefore the KM curve result of pathological stage III is not shown.
[0218] These analyses demonstrate the robustness of the pathomicsScore in stratifying patients across different clinical subgroups, highlighting its potential as a valuable prognostic tool in individualized treatment strategies.
[0219] To further investigate the correlation between the presence of TP53 co-mutation and prognosis, Kaplan-Meier survival analysis was performed on 34 patients who received next-generation sequencing (NGS). Among them, 12 patients had TP53 co-mutation. There was no statistically significant difference in PFS between patients with and without TP53 co-mutation (P=0.287), suggesting that TP53 co-mutation may not have a significant prognostic effect in this cohort. However, due to the limited sample size, a definitive conclusion cannot be drawn, and it is necessary to verify it in a larger independent cohort.
[0220] Example 5: Establishment and verification of individualized disease progression prediction model
[0221] This example relates to the construction of a combined prediction model using a pathomics model and a clinical data-based model, and the analysis of the prediction performance of the combined model.
[0222] First, single-factor and multi-factor Cox regression analysis was used to identify independent prognostic factors for PFS, and the results are summarized in Table 2.
[0223] Table 2. Univariate and multivariate Cox regression analysis of progression-free survival
[0224]
[0225]
[0226] Abbreviations: ECOG PS, Eastern Cooperative Oncology Group performance status; EGFR-TKIs, epidermal growth factor receptor tyrosine kinase inhibitors; VPI, visceral pleural invasion; CEA, carcinoembryonic antigen; CA19-9, carbohydrate antigen 19-9; SCC, squamous cell carcinoma antigen; CYFRA21-1, cytokeratin 19 fragment; NSE, neuron-specific enolase; ProGRP, pro-gastrin-releasing peptide.
[0227] Single-factor Cox analysis showed that smoking history, pathological stage, histological subtype, CYFRA21-1, and pathomicsScore were potential predictors of disease progression (p<0.05). Multi-factor Cox analysis showed that preoperative pathological stage and pathomicsScore were independent prognostic factors for disease progression (p<0.05). Figure 7 A displays the forest plot of single-factor and multi-factor Cox regression analysis, respectively.
[0228] Subsequently, a combined prediction model including two independent prognostic factors was constructed and presented in the form of a nomogram ( Figure 7 B). The model showed good predictive performance for disease progression. In the training and validation cohorts, the area under the ROC curve (AUC) was 0.789 and 0.728, respectively. The ROC curve results of the combined model, pathomicsScore model, and pathological staging model are shown in Figure 2. Figure 7 The combined model achieved specificities of 0.677 and 0.714 in the training and validation cohorts, sensitivities of 0.909 and 1, and accuracies of 0.738 and 0.789, respectively. Details of the pathomic model, pathological staging model, and combined model are shown in Table 3.
[0229] Figure 7 EF presented two representative cases outside the study cohort, demonstrating the application of the model in predicting disease progression.
[0230] Table 3. Predictive performance of the pathomic scoring model, pathological staging model, and combined model in the training and validation cohorts
[0231]
[0232] To evaluate the prediction performance of the model for progression-free survival (PFS), time-dependent ROC curves at 0.5, 1, and 1.5 years were plotted in the training and validation cohorts ( Figure 8 AF). The AUC for pathomicsScore in predicting 1.5-year PFS was 0.743 in the training cohort and 0.725 in the validation cohort. In contrast, the AUC for the pathological stage model in predicting 1.5-year PFS was 0.645 in the training cohort and 0.644 in the validation cohort. Compared with the models using pathological stage and pathological stage alone, the combined model integrating pathological stage and pathological stage significantly improved the AUC for predicting 1.5-year PFS to 0.789 and 0.728 in the training cohort and validation cohort, respectively.
[0233] In addition, to evaluate the consistency between the disease progression predicted by the model and the actual clinical outcomes, calibration curves of 0.5-year, 1-year, and 1.5-year PFS of the combined model were constructed ( Figure 8 GH). Calibration curves showed good agreement between model-predicted disease progression and actual clinical outcomes in both the training and validation cohorts.
[0234] Finally, a decision curve analysis was performed on the model. The results showed that the net benefit of the combined model was higher than the net benefit of intervening in all patients and not intervening in all patients, further highlighting the clinical utility of the combined model ( Figure 8I).
[0235] The prediction performance of the combined model is further visualized using a confusion matrix, as shown in Figure 8 J-K.
[0236] All documents referred to in the present application are incorporated herein by reference as if each individual document were incorporated by reference. In addition, it is to be understood that the application can be carried out by specifically different embodiments and that each disclosed embodiment can be implemented with or without the corresponding benefits disclosed herein.
Claims
1. A method for constructing a prognostic prediction model for patients with lung adenocarcinoma receiving EGFR-TKI targeted therapy, characterized in that: The method comprises the steps of: (A1) Pathological omics data from lung adenocarcinoma patients receiving EGFR-TKI targeted therapy were used as training sets; (A2) Preprocessing the pathological genomics data, including the following steps: (a2.1) selecting an area with the highest tumor cell density from the pathological data; the area with the highest tumor cell density is an area with a tumor cell density of ≥100 per unit area; (a2.2) extracting pathomic features from the area with the highest tumor cell density; (a2.3) normalizing the pathomic features; (a2.4) screening the pathological features to obtain preferred pathological features; (A3) Using the preferred pathomic features for Cox proportional hazards regression model training to obtain a pathomic feature model (pathomicScore).
2. The method according to claim 1, wherein The preferred pathological features are: (1)Granularity_5_Eosin; (2)Correlation_RWC_Eosin_Hematoxylin; (3)Granularity_2_Hematoxylin; (4) Granularity_11_Eosin; and (5)Granularity_4_Hematoxylin.
3. A system for predicting the prognosis of patients with lung adenocarcinoma, characterized in that: include: An input module, wherein the input module is configured to input data, wherein the input data is pathological omics data of the object to be tested; A prediction module, the prediction module being configured to execute a prediction model for the prognosis of lung adenocarcinoma patients receiving EGFR-TKI targeted therapy, thereby obtaining a pathology score based on the pathology data, and determining the prognosis of the subject based on the pathology score to obtain a prediction result; the prediction model is constructed using the method of claim 1; An output module is configured to output the prediction result of the prediction module.
4. A method for constructing a combined prediction model for the prognosis of patients with lung adenocarcinoma receiving EGFR-TKI targeted therapy, characterized in that: The method comprises the steps of: (B1) constructing a pathomic feature model (pathomicScore) using the method of claim 1; (B2) Constructing a clinical data scoring model, including the following steps: (B2.1) providing clinical data of the lung adenocarcinoma patient; (B2.2) screening the clinical data using a regression method to obtain optimal clinical data; (B2.3) converting the preferred clinical data into numerical values to obtain a clinical data scoring model; (B3) Integrating the models described in step (B1) and step (B2) to obtain a combined model.
5. A combined prediction system for the prognosis of patients with lung adenocarcinoma receiving EGFR-TKI targeted therapy, characterized in that: The system comprises: An input module, wherein the input module is configured to input data, wherein the input data includes: pathological genomics data and clinical data of the subject to be tested; A prediction module is configured to process the input data and convert the input data into a score; the prediction module includes: (M1) a prediction unit based on pathogenomic features, wherein the prediction unit (M1) is configured to execute a prediction model for the prognosis of lung adenocarcinoma patients based on pathogenomics, thereby obtaining a pathogenomics score for the subject to be tested; the prediction model for the prognosis of lung adenocarcinoma patients based on pathogenomics is constructed using the method of claim 1; (M2) a prediction unit based on clinical data, wherein the (M2) prediction unit is configured to convert the clinical data factors into a score; Integrating the scores obtained by the prediction units (M1) and (M2) to obtain a combined score, and judging the prognosis of the subject receiving EGFR-TKI targeted therapy according to the combined score, thereby obtaining a prediction result; An output module is configured to output the prediction result of the prediction module.
6. A method for categorizing and / or evaluating the prognosis of lung adenocarcinoma patients receiving EGFR-TKI targeted therapy, characterized in that: The method comprises the steps of: (C1) providing pathological omics data and / or clinical data of a test subject; (C2) inputting the pathomic data into a pathomicScore model to obtain a pathomic score, and / or inputting the pathomic data and clinical data into a combination model to obtain a combination score; the pathomicScore model is constructed using the method of claim 1, and the combination model is constructed using the method of claim 4; (C3) comparing the pathological omics score and / or combined score with a reference value or standard value, thereby classifying lung adenocarcinoma patients receiving EGFR-TKI targeted therapy and / or evaluating the prognosis of patients receiving EGFR-TKI targeted therapy.
7. The method according to claim 6, wherein The method comprises the steps of: (C1) providing pathological omics data of a test subject; (C2) inputting the pathomic data into a pathomicScore model to obtain a pathomic score; (C3) comparing the pathological omics score with a reference value or a standard value, thereby classifying lung adenocarcinoma patients undergoing EGFR-TKI targeted therapy and / or evaluating the prognosis of EGFR-TKI targeted therapy; Among them, if the pathological omics score is lower than the reference value or standard value, it means that the test subject is a subject with a good prognosis after receiving EGFR-TKI targeted therapy; if the pathological omics score is higher than the reference value or standard value, it means that the test subject is a subject with a poor prognosis after receiving EGFR-TKI targeted therapy.
8. The method according to claim 6, wherein The method comprises the steps of: (C1) providing pathological omics data and clinical data of a test subject; (C2) inputting the pathomic data and clinical data into a combined model to obtain a combined score; (C3) comparing the combined score with a reference value or a standard value to classify lung adenocarcinoma patients undergoing EGFR-TKI targeted therapy and / or evaluate the prognosis of EGFR-TKI targeted therapy; Among them, if the combined score is lower than the reference value or standard value, it means that the test subject is a subject with a good prognosis after receiving EGFR-TKI targeted therapy; if the combined score is higher than the reference value or standard value, it means that the test subject is a subject with a poor prognosis after receiving EGFR-TKI targeted therapy.
9. The method according to claim 5, wherein The evaluation includes predicting the progression-free survival of the subject in a specific number of years; the number of years is selected from the following group: 0.5 years, 1 year, 1.5 years, or a combination thereof.
10. A method for adjusting medication / treatment regimen, characterized in that: include: (I) providing pathological omics data and / or clinical data from a subject; (II) inputting the pathomic data into a pathomicScore model to obtain a pathomic score, and / or inputting the pathomic data and clinical data into a combination model to obtain a combination score; the pathomicScore model is constructed using the method of claim 1, and the combination model is constructed using the method of claim 4; (III) adjusting medication / treatment regimen according to the pathomic score and / or combined score.