NSCLC immunotherapy curative effect prediction method and system based on self-supervised deep learning pathomics
By constructing a generative pre-training model based on self-supervised deep learning, the problem of dependence on labeled data in the prior art is solved, and the accuracy of predicting immunotherapy efficacy in patients with NSCLC and the generalization ability of the model is improved.
Patent Information
- Application Number
- CN202510115315.9
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-01-24
- Publication Date
- 2025-05-27
AI Technical Summary
When predicting the efficacy of immunotherapy in patients with NSCLC, the prior art relies on a large amount of labeled data, resulting in the need for data labeling resources, time and expertise, limiting the generalization ability and prediction accuracy of the model.
The pathologic omics method based on self-supervised deep learning is adopted to construct generative pre-trained models, including patch-level and WSI-level pathological feature extraction models for self-supervised deep learning, as well as immunotherapy efficacy prediction models with fine-tuning of supervised weights, to reduce dependence on labeled data.
The self-supervised learning of the pathological images is extracted, the transfer learning ability of the model is improved, the dependence on labeled data is reduced, and the accuracy of predicting the efficacy of immunotherapy in patients with NSCLC is effectively improved.
Smart Images

Figure CN120048516A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the field of prediction of immunotherapy efficacy, and particularly to a method and system for predicting the efficacy of NSCLC immunotherapy based on self-supervised deep learning pathomics. Background Art
[0002] As the main treatment option for patients with non-small cell lung cancer (NSCLC) lacking targeted treatment options, PD-(L)1 has a durable response limited to a small subset of patients. Tumor mutational burden (TMB) and PD-L1 expression, as FDA-approved efficacy prediction markers, cannot fully explain the differences in immunotherapy efficacy, which may be attributed to the intrinsic complexity of the tumor microenvironment and cancer immune response.
[0003] Histopathological sections stained with hematoxylin and eosin (H&E) are commonly used in clinical diagnosis of malignant tumors. Digital slide scanners can digitize entire H&E-stained specimens into high-resolution whole-slide images (WSIs) containing cell structures and microenvironments. Through artificial intelligence, the biological information rich in WSIs can be captured and analyzed, and significant progress has been made in tasks such as cancer diagnosis, subtype classification, prediction of mutations, and prognosis prediction.
[0004] A pathomics model based on the ViT-RNN architecture was used to predict the efficacy of immunotherapy, achieving an accuracy of 81% in the validation cohort, but the accuracy dropped to 60% in an independent external validation set. The improvement of external validation performance usually relies on a large amount of labeled data for supervised learning, but data annotation requires a large amount of resources, time, and accurate annotation of professional knowledge, which has currently become a major obstacle to the widespread application of computational pathology.
[0005] Therefore, how to provide a method and system for predicting the efficacy of NSCLC immunotherapy based on self-supervised deep learning pathomics that can effectively reduce the dependence of pathomics models on labeled data, improve the generalization ability, and effectively improve the accuracy of immunotherapy efficacy prediction is an urgent problem for those skilled in the art. Summary of the Invention
[0006] In view of this, the present invention proposes a method and system for predicting the efficacy of NSCLC immunotherapy based on self-supervised deep learning pathomics.
[0007] To achieve the above object, the present invention adopts the following technical solutions:
[0008] A method for predicting the efficacy of NSCLC immunotherapy based on self-supervised deep learning pathomics, comprising:
[0009] Step 1: Obtain high-resolution panoramic section images of HE-stained histopathological sections, perform segmentation and staining preprocessing on the high-resolution panoramic section images, and divide them into a training set and a validation set according to a preset ratio;
[0010] Step 2: Construct a generative pre-training model based on self-supervised deep learning pathomics, and based on the training set, use the PFS after immunotherapy as the research outcome to train an NSCLC immunotherapy efficacy prediction model based on generative pre-training; among them, the generative pre-training model includes two parts. The first part is a basic model for patch-level and WSI-level pathological feature extraction based on self-supervised deep learning, and the second part is an immunotherapy efficacy prediction model based on supervised weight fine-tuning;
[0011] Step 3: Input the validation set into the NSCLC immunotherapy efficacy prediction model to obtain the NSCLC patient immunotherapy efficacy prediction results, and perform model performance analysis during training, internal validation, and external validation.
[0012] Optionally, in Step 1, the segmentation and staining preprocessing of the high-resolution panoramic section images is specifically as follows:
[0013] Divide the high-resolution panoramic section images into 1024×1024 pixel patch blocks, remove the blank background area, perform color normalization using the Macenko algorithm, and according to the PFS time after receiving immunotherapy, divide the patients with disease progression within 5.5 months after immunotherapy into one group, and divide the patients without disease progression within 5.5 months after immunotherapy into one group.
[0014] Optionally, in Step 2, the patch-level and WSI-level pathological feature extraction model based on self-supervised deep learning is specifically as follows:
[0015] Use the first Vision Transformer model for patch-level feature extraction, and use the Patchaggregator algorithm for patch-level feature aggregation;
[0016] Use the second Vision Transformer model to perform WSI-level feature extraction on the aggregated patch-level features;
[0017] Among them, the structures of the first Vision Transformer model and the second Vision Transformer model are both integrated by a context encoder, a target encoder, and a predictor, and the target encoder is used to dynamically allocate probabilities.
[0018] Optionally, in Step 2, the immunotherapy efficacy prediction model based on supervised weight fine-tuning is specifically as follows:
[0019] Fine-tune the weights of the model for patch-level feature extraction learning in self-supervised deep learning in the training queue, use the third Vision Transformer model obtained by fine-tuning for patch-level feature extraction, and use the Patchaggregator algorithm for patch-level feature aggregation;
[0020] Fine-tune in the training queue according to the transfer learning results of WSI-level feature learning in self-supervised deep learning, use the Transformer encoder algorithm for WSI-level feature extraction, and obtain the immunotherapy classification result through a Class tokens.
[0021] Optionally, in step 3, perform model performance analysis during training, internal validation, and external validation, including but not limited to the following: AUC, accuracy, sensitivity, specificity, survival analysis, and multivariate Cox regression analysis.
[0022] The present invention also provides a NSCLC immunotherapy efficacy prediction system based on self-supervised deep learning pathomics using a NSCLC immunotherapy efficacy prediction method based on self-supervised deep learning pathomics, including:
[0023] Training set and validation set acquisition module: used to obtain high-resolution panoramic slice images of HE-stained tissue pathological sections, perform segmentation and staining preprocessing on the high-resolution panoramic slice images, and divide them into a training set and a validation set according to a preset ratio;
[0024] Efficacy prediction model training module: used to construct a generative pre-training model based on self-supervised deep learning pathomics, and based on the training set, with the PFS after immunotherapy as the research outcome, train a NSCLC immunotherapy efficacy prediction model based on the generative pre-training; wherein, the generative pre-training model includes two parts, the first part is a basic model for patch-level and WSI-level pathological feature extraction based on self-supervised deep learning, and the second part is an immunotherapy efficacy prediction model based on supervised weight fine-tuning;
[0025] Efficacy prediction and model performance analysis module: used to input the validation set into the NSCLC immunotherapy efficacy prediction model to obtain the NSCLC patient immunotherapy efficacy prediction result, and perform model performance analysis during training, internal validation, and external validation.
[0026] As can be seen from the above technical solutions, compared with the prior art, the present invention proposes a method and system for predicting the efficacy of NSCLC immunotherapy based on self-supervised deep learning pathomics. By constructing a GPT model based on self-supervised deep learning pathomics, it can learn the self-features of pathological images and serve as a feature extractor; and it has good transfer learning ability. When performing downstream tasks on this basis, only weight fine-tuning is required to achieve good performance, realizing the effective improvement of the accuracy of immunotherapy efficacy prediction while reducing the dependence on labeled data. BRIEF DESCRIPTION OF THE DRAWINGS
[0027] In order to more clearly illustrate the technical solutions in the embodiments of the present invention or the prior art, the following will briefly introduce the drawings required for use in the description of the embodiments or the prior art. Obviously, the drawings in the following description are only the embodiments of the present invention. For those of ordinary skill in the art, other drawings can be obtained according to the provided drawings without creative efforts.
[0028] Figure 1 It is a schematic flowchart of the method of the present invention.
[0029] Figure 2 It is a schematic structural diagram of the GPT model based on self-supervised deep learning pathomics of the present invention. DETAILED DESCRIPTION OF THE EMBODIMENTS
[0030] The following will clearly and completely describe the technical solutions in the embodiments of the present invention with reference to the drawings in the embodiments of the present invention. Obviously, the described embodiments are only a part of the embodiments of the present invention, rather than all of the embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those of ordinary skill in the art without creative efforts belong to the scope of protection of the present invention.
[0031] Embodiment 1:
[0032] Embodiment 1 of the present invention discloses a method for predicting the efficacy of NSCLC immunotherapy based on self-supervised deep learning pathomics, as Figure 1 shown, including:
[0033] Step 1: Obtain high-resolution panoramic slice images of HE-stained tissue pathological sections, and perform preprocessing such as segmentation and staining on the high-resolution panoramic slice images, and divide them into a training set and a validation set according to a preset ratio.
[0034] Performing preprocessing such as segmentation and staining on the high-resolution panoramic slice images specifically includes:
[0035] Use the pyvips module to segment high-resolution panoramic slice images into 1024×1024 pixel patches, remove the blank background area, perform color normalization using the Macenko algorithm to reduce staining errors caused by color imbalance, and divide patients who experienced disease progression within 5.5 months of immunotherapy into one group and patients who did not experience disease progression within 5.5 months of immunotherapy into another group according to the PFS time after receiving immunotherapy.
[0036] The inclusion criteria for the training set and validation set are as follows:
[0037] (a) Patients with pathologically and radiologically diagnosed advanced NSCLC, without other types of malignancies, or NSCLC patients who had recurrence or metastasis within 6 months after surgery; (b) Received PD-(L)1 inhibitors as first-line or second-line treatment; (c) Corresponding H&E-stained histological specimens were available; (d) Complete clinical and pathological information and survival follow-up data were retained. The primary endpoint was the PFS (progression-free survival) of PD-(L)1 inhibitors, which was defined as the time from the start of PD-(L)1 inhibitors to disease progression or death.
[0038] Step 2:
[0039] Construct a generative pre-training (GPT) model based on self-supervised deep learning pathologyomics, as Figure 2 shown, and based on the training set, use the PFS after immunotherapy as the research outcome to train a NSCLC immunotherapy efficacy prediction model based on generative pre-training; among them, the generative pre-training model consists of two parts. The first part is a basic model for patch-level and WSI-level pathological feature extraction based on self-supervised deep learning, and the second part is an immunotherapy efficacy prediction model based on supervised weight fine-tuning.
[0040] Self-Supervised can learn the self-features of pathological images and serve as a feature extractor, with good transfer learning ability. When performing downstream tasks on this basis, only weight fine-tuning is required to achieve good performance.
[0041] The patch-level and WSI-level pathological feature extraction model based on self-supervised deep learning is specifically:
[0042] Use the first Vision Transformer model to perform patch-level feature extraction and use the Patchaggregator algorithm to perform patch-level feature aggregation;
[0043] Use the second Vision Transformer model to perform WSI-level feature extraction on the aggregated patch-level features;
[0044] Among them, the structures of the first Vision Transformer model and the second Vision Transformer model are both integrated by a context encoder, an object encoder, and a predictor, aiming to predict the representations of various object blocks from a single context block in the same image. The object encoder is used to dynamically allocate probabilities.
[0045] The immune therapy efficacy prediction model based on supervised weight fine-tuning is specifically as follows:
[0046] Fine-tune the weights of the model for patch-level feature extraction learning in self-supervised deep learning in the training cohort, use the fine-tuned third Vision Transformer model for patch-level feature extraction, and use the Patchaggregator algorithm for patch-level feature aggregation;
[0047] Fine-tune according to the transfer learning results of WSI-level feature learning in self-supervised deep learning in the training cohort, use the Transformer encoder algorithm for WSI-level feature extraction, and obtain the immune therapy classification result through a Class tokens.
[0048] Step 3: Input the validation set into the NSCLC immune therapy efficacy prediction model to obtain the NSCLC patient immune therapy efficacy prediction result, and perform model performance analysis during training, internal validation, and external validation.
[0049] Perform model performance analysis during training, internal validation, and external validation, including but not limited to the following: AUC, accuracy, sensitivity, specificity, survival analysis, and multivariate Cox regression analysis.
[0050] The model performance analysis can also be evaluated by calculating precision, recall, balanced accuracy, F1 score, and AUROC, etc.
[0051] Example 2:
[0052] Embodiment 2 of the present invention discloses a specific application of a method for predicting the efficacy of NSCLC immunotherapy based on self-supervised deep learning pathomics, as follows:
[0053] The dataset comes from Shandong Cancer Hospital and Shandong Provincial Hospital: including 10,539 WSIs of 5 tumor types, among which, the WSIs of patients with lung cancer, esophageal cancer, breast cancer, cervical cancer, and ovarian cancer are represented. All WSIs have not been manually annotated. The Shandong Cancer Hospital cohort is divided into a training set and a validation set according to a ratio of 7:3, and the Shandong Provincial Hospital cohort is an external independent validation set.
[0054] Patch-level training and WSI-level training are performed on different types of tumors respectively. The self-supervised patho-GPT model (a GPT model based on self-supervised deep learning pathomics) is initially trained on the cervical cancer dataset, followed by the lung cancer, ovarian cancer, and esophageal cancer datasets. As the sample size increases, the number of iterations required to stabilize the loss value gradually decreases, indicating an improvement in the learning ability of patho-GPT.
[0055] The training cohort for immunotherapy consists of 833,311 patches extracted from 579 HE-stained pathological sections. During 54 iterations, the loss value decreased from 0.658 to 0.342.
[0056] In the internal validation cohort, the AUC of patho-GPT is 0.774, the accuracy is 0.828, the sensitivity is 0.882, and the specificity is 0.667. The KM survival curve also suggests that grouping according to the patho-GPT model (R group: immunotherapy-resistant group; S group: immunotherapy-sensitive group) can effectively screen the population that benefits from immunotherapy (P < 0.0001). Multivariate cox regression analysis indicates that patho-GPT has independent predictive ability in the internal validation cohort, with a hazard ratio of 0.58 (95% CI 0.39 - 0.86, P = 0.0049 (whether the influence of each independent variable on the dependent variable is significant)).
[0057] In the independent external validation cohort, the AUC generated by patho-GPT is 0.752, the accuracy is 0.758, the sensitivity is 0.767, and the specificity is 0.737. In the external validation cohort, the KM survival curve also suggests that grouping according to the patho-GPT model can effectively screen the population that benefits from immunotherapy, P = 0.0024. After adjusting for the influence of confounding factors by multivariate Cox regression analysis, patho-GPT also has independent predictive ability in the external validation cohort (HR = 0.41, 95% CI 0.23 - 0.72, P = 0.0019).
[0058] Example 3:
[0059] Example 3 of the present invention discloses a NSCLC immunotherapy efficacy prediction system based on self-supervised deep learning pathomics for predicting the efficacy of NSCLC immunotherapy using a method based on self-supervised deep learning pathomics, including:
[0060] A training set and validation set acquisition module: used to obtain high-resolution panoramic section images of HE-stained tissue pathological sections, perform preprocessing, and divide them into a training set and a validation set according to a preset ratio;
[0061] Efficacy prediction model training module: used to construct a GPT model based on self-supervised deep learning pathomics, and based on the training set, train an NSCLC immunotherapy efficacy prediction model based on GPT; among them, the GPT model based on self-supervised deep learning pathomics includes two parts, the first part is a patch-level and WSI-level pathological feature extraction model based on self-supervised deep learning, and the second part is an immunotherapy outcome prediction model based on supervised fine-tuning;
[0062] Efficacy prediction and model performance analysis module: used to input the validation set into the NSCLC immunotherapy efficacy prediction model, obtain the NSCLC immunotherapy efficacy prediction result, and perform model performance analysis during the training, internal validation, and external validation processes.
[0063] The embodiments of the present invention disclose an NSCLC immunotherapy efficacy prediction method and system based on self-supervised deep learning pathomics. By constructing a GPT model based on self-supervised deep learning pathomics, it can learn the self-features of pathological images and serve as a feature extractor; and it has good transfer learning ability. When performing downstream tasks on this basis, only weight fine-tuning is required to achieve good performance, realizing the effective improvement of the immunotherapy efficacy prediction accuracy on the basis of reducing the dependence on labeled data.
[0064] In this specification, each embodiment is described in a progressive manner. The key point of each embodiment is to illustrate the differences from other embodiments. The same or similar parts among the embodiments can be referred to each other. For the devices disclosed in the embodiments, since they correspond to the methods disclosed in the embodiments, the description is relatively simple, and the relevant parts can be referred to the description of the method part.
[0065] The above description of the disclosed embodiments enables those skilled in the art to implement or use the present invention. Various modifications to these embodiments will be obvious to those skilled in the art. The general principles defined herein can be implemented in other embodiments without departing from the spirit or scope of the present invention. Therefore, the present invention will not be limited to these embodiments shown herein, but will conform to the widest scope consistent with the principles and novel features disclosed herein.
Claims
1. A method for predicting the efficacy of NSCLC immunotherapy based on self-supervised deep learning pathogenomics, characterized in that: include: Step 1: obtaining a high-resolution panoramic slice image of a tissue pathology section stained with HE, and performing segmentation and staining preprocessing on the high-resolution panoramic slice image, and dividing the image into a training set and a validation set according to a preset ratio; Step 2: construct a generative pre-training model based on self-supervised deep learning pathology omics, and based on the training set, train a NSCLC immunotherapy efficacy prediction model based on generative pre-training with PFS after immunotherapy as the research outcome; wherein the generative pre-training model includes two parts, the first part is a basic model for patch-level and WSI-level pathological feature extraction based on self-supervised deep learning, and the second part is an immunotherapy efficacy prediction model based on supervised weight fine-tuning; Step 3: Input the validation set into the NSCLC immunotherapy efficacy prediction model to obtain the immunotherapy efficacy prediction results for NSCLC patients, and perform model performance analysis during training, internal validation, and external validation.
2. The method for predicting the efficacy of NSCLC immunotherapy based on self-supervised deep learning pathogenomics according to claim 1, characterized in that: In step 1, the high-resolution panoramic slice image is segmented and stained preprocessed, specifically: The high-resolution panoramic slice images were segmented into patches of 1024×1024 pixels, and the blank background areas were removed. Color standardization was performed using the Macenko algorithm. According to the PFS time after receiving immunotherapy, patients whose disease progressed within 5.5 months of immunotherapy were divided into one group, and patients whose disease did not progress within 5.5 months of immunotherapy were divided into another group.
3. The method for predicting the efficacy of NSCLC immunotherapy based on self-supervised deep learning pathogenomics according to claim 1, characterized in that: In step 2, the patch-level and WSI-level pathological feature extraction models based on self-supervised deep learning are as follows: The first Vision Transformer model is used to extract patch-level features, and the Patch aggregator algorithm is used to aggregate patch-level features. The second Vision Transformer model is used to extract WSI-level features from the aggregated patch-level features; The structures of the first Vision Transformer model and the second Vision Transformer model are both integrated by a context encoder, a target encoder and a predictor, and the target encoder is used to dynamically allocate probabilities.
4. The method for predicting the efficacy of NSCLC immunotherapy based on self-supervised deep learning pathogenomics according to claim 1, characterized in that: In step 2, the immunotherapy efficacy prediction model based on supervised weight fine-tuning is as follows: Fine-tune the weights of the model learned by patch-level feature extraction in the self-supervised deep learning in the training queue, use the fine-tuned third Vision Transformer model to extract patch-level features, and use the Patchaggregator algorithm to aggregate patch-level features; According to the transfer learning results of WSI-level feature learning in the self-supervised deep learning, fine-tuning is performed on the training cohort, WSI-level features are extracted using the Transformer encoder algorithm, and the immunotherapy classification results are obtained through a Class token.
5. The method for predicting the efficacy of NSCLC immunotherapy based on self-supervised deep learning pathogenomics according to claim 1, characterized in that: In step 3, model performance analysis is performed during training, internal validation, and external validation, including but not limited to the following: AUC, accuracy, sensitivity, specificity, survival analysis, and multivariate Cox regression analysis.
6. A NSCLC immunotherapy efficacy prediction system based on self-supervised deep learning pathology omics using a NSCLC immunotherapy efficacy prediction method based on self-supervised deep learning pathology omics as described in any one of claims 1 to 5, characterized in that: include: The training set and validation set acquisition module is used to acquire high-resolution panoramic slice images of HE-stained tissue pathological slices, and segment and stain the high-resolution panoramic slice images, and divide them into training sets and validation sets according to a preset ratio; Therapeutic effect prediction model training module is used to construct a generative pre-training model based on self-supervised deep learning pathology omics, and based on the training set, train a NSCLC immunotherapy efficacy prediction model based on generative pre-training with PFS after immunotherapy as the research outcome; wherein, the generative pre-training model includes two parts, the first part is a basic model for patch-level and WSI-level pathological feature extraction based on self-supervised deep learning, and the second part is an immunotherapy efficacy prediction model based on supervised weight fine-tuning; The efficacy prediction and model performance analysis module is used to input the validation set into the NSCLC immunotherapy efficacy prediction model to obtain the immunotherapy efficacy prediction results for NSCLC patients, and perform model performance analysis during training, internal validation and external validation.