A method for intelligently predicting the efficacy of PD-L1 inhibitors from gastric cancer H&E staining images

By applying semi-supervised learning and multi-instance learning methods in H&E stained images of gastric cancer, combined with deep convolutional neural networks, an intelligent prediction system was built, which solved the consistency problem of evaluation of PD-L1 expression level in gastric cancer, and improved the prediction accuracy of the efficacy of PD-1/PD-L1 inhibitors.

CN115330722BActive Publication Date: 2025-06-17FUZHOU UNIV +1
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202210974249.7
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2022-08-15
Publication Date
2025-06-17
Estimated Expiration
2042-08-15

AI Technical Summary

Technical Problem

The prior art has low consistency in assessing PD-L1 expression levels in gastric cancer, resulting in difficulty in predicting the efficacy of PD-1/PD-L1 inhibitors.

Method used

A cancerous region segmentation model based on semi-supervised learning and a multi-instance learning image classification model, combined with a deep convolutional neural network, an intelligent prediction system for gastric cancer H&E stained images is constructed to generate a therapeutic effect discrimination matrix to assist clinical decision-making.

Benefits of technology

It improves the specific localization of cancerous areas in H&E staining images of gastric cancer and the prediction accuracy of PD-L1 inhibitor efficacy, reduces the subjective error of clinicians, and assists clinical decision-making.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN115330722B_ABST
    Figure CN115330722B_ABST
Patent Text Reader

Abstract

The present invention provides a method for intelligently predicting the efficacy of PD-L1 inhibitors from gastric cancer H&E staining images. A cancerous region segmentation model based on semi-supervised learning is proposed to segment the cancerous regions in H&E staining images. Secondly, image patches containing cancerous regions are taken out, and the class labels of each image patch are obtained according to the efficacy labels of PD-L1 inhibitors marked by clinicians, and an image classification model based on multi-instance learning is used for prediction at the image patch level. Finally, an efficacy discrimination matrix for the entire H&E staining image is generated based on the prediction results at the image patch level, and an image classification model based on a deep convolutional neural network is used for efficacy prediction at the slice level. By applying this technical solution, the prediction results of the efficacy of PD-L1 inhibitors in gastric cancer patients can be obtained to assist doctors in making clinical decisions.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of image recognition, and in particular to a method for intelligently predicting the efficacy of PD-L1 inhibitors from H&E staining images of gastric cancer. Background Art

[0002] According to the statistics of the International Agency for Research on Cancer (IARC) in 2020, gastric cancer is the fifth most common malignant tumor in the world, and its mortality rate ranks fourth among all malignant tumors. The main treatments for gastric cancer include surgery, endoscopic therapy, chemotherapy, gene therapy, traditional Chinese medicine treatment, and comprehensive treatment. In recent years, immune checkpoint inhibitors represented by PD-1 / PD-L1 monoclonal antibodies have made breakthrough progress in tumor immunotherapy. At present, the prediction of the efficacy of immunotherapy focuses on the expression of PD-L1 in the tumor microenvironment. The interpretation standard of PD-L1 expression level in gastric cancer uses immunohistochemistry to score CPS (comprehensive positive score). The larger the value, the higher the sensitivity of the tumor to immunotherapy. At the 2020 European Society for Medical Oncology Annual Meeting (ESMO), the detailed results of CheckMate-649, the largest and only successful global Phase III study in the field of gastric cancer to date, were released. The results show:

[0003] (1) Among all patients, drug O (nivolumab, an immunotherapy drug for gastric cancer approved in China) combined with chemotherapy was used as the first-line treatment for gastric cancer patients, which significantly prolonged the overall survival compared with chemotherapy alone.

[0004] (2) In patients with high tumor PD-L1 expression (CPS ≥ 5), O-drug combined with chemotherapy showed a greater overall survival benefit, and a significant progression-free survival benefit was also observed.

[0005] As one of the evaluation methods for PD-L1 companion diagnosis, CPS has important guiding significance for the clinical use of PD-L1 inhibitors for immunotherapy. However, due to the differences in staining intensity of PD-L1 immunohistochemical staining sections and the errors caused by subjective factors such as experience and fatigue of pathologists, the consistency of CPS scores among different pathologists is relatively low. Even the same pathologist may have different results when diagnosing the same case at different times. In response to this problem, in recent years, domestic and foreign researchers have used artificial intelligence technology to quantify PD-L1 immune scores, aiming to assist case doctors in improving the consistency and efficiency of interpretation. However, compared with the TPS scores of non-small cell lung cancer and melanoma, and the IC scores of triple-negative breast cancer, the CPS score of gastric cancer is more difficult. Therefore, most of the current research focuses on TPS scores and IC scores, and there is little research on the intelligent prediction of CPS scores for gastric cancer.

[0006] Meanwhile, the expression level of PD-L1, tumor-infiltrating lymphocytes (TIL), defective mismatch repair (dMMR) and microsatellite instability-high (MSI-H), tumor mutation burden (TMB), and gut commensal bacteria affect the efficacy of PD-1 / PD-L1 inhibitors on tumors from different aspects. These factors are interrelated, which also poses difficulties for clinicians in choosing whether to use PD-1 / PD-L1 inhibitors for immunotherapy in gastric cancer patients. Therefore, there is an urgent need for a biomarker to evaluate the efficacy of PD-1 / PD-L1 inhibitors to assist doctors in making clinical decisions. Summary of the Invention

[0007] In view of this, the purpose of the present invention is to provide a method for intelligent prediction of the efficacy of PD-L1 inhibitors from gastric cancer H&E staining images, obtain the prediction results of the efficacy of PD-L1 inhibitors in gastric cancer patients, and assist doctors in making clinical decisions.

[0008] To achieve the above object, the present invention adopts the following technical solutions: A method for intelligent prediction of the efficacy of PD-L1 inhibitors from gastric cancer H&E staining images. First, a cancerous region segmentation model based on semi-supervised learning is proposed to segment the cancerous regions in the H&E staining images. Secondly, the image patches containing the cancerous regions are taken out, and the class labels of each image patch are obtained according to the efficacy labels of PD-L1 inhibitors marked by clinicians, and an image classification model based on multi-instance learning is used for prediction at the image patch level. Finally, a efficacy discrimination matrix of the entire H&E staining image is generated based on the prediction results at the image patch level, and an image classification model based on deep convolutional neural network is used for efficacy prediction at the slice level.

[0009] In a preferred embodiment, first, the H&E stained image is cut into image patches of 512×512 pixels. According to the gastric cancer classification, some image patches are selected to label the gastric cancer region, obtaining some labeled image patches and unlabeled image patches. Secondly, DeepLabV3+ is selected as the segmentation backbone network, and the EfficientNet-B3 network pre-trained through the ImageNet dataset is used to replace the encoder structure in the DeepLabV3+ network to construct a meanteacher model. Finally, for the labeled image patches, they are input into the student network, and the cross-entropy loss function combined with FocalLoss is used to constrain between the obtained prediction results and the true labels, that is, MixedLoss. For the unlabeled image patches, they are respectively input into the teacher network and the student network, and the obtained prediction results are constrained by the consistency loss function, that is, ConsistencyLoss. Among them, the student network adjusts and improves the network optimization objective through MixedLoss combined with Consistency Loss, and the weight parameters of the teacher network are obtained from the exponential moving average of the weight parameters of the student network. At the same time, various auxiliary image transformations are added to the training data.

[0010] In a preferred embodiment, first, based on the patient's survival period, progression-free survival period, and degree of tumor shrinkage, the efficacy label of the gastric cancer patient corresponding to the H&E stained image after using the PD-L1 inhibitor is evaluated, denoted as significant efficacy 1 and no significant efficacy 0. Secondly, the gastric cancer region segmentation model selects the image patches containing the cancerous region in the H&E stained image and the 4 image patches around it, and retains the position information of all the selected image patches in the original H&E stained image. Thirdly, according to the PD-L1 inhibitor efficacy label, the class labels of the selected image patches are given, and an image classification model based on multi-instance learning is used for training. For the H&E stained images with no significant efficacy, the labels of all the selected image patches are set to 0, while for the H&E stained images with significant efficacy, the initial labels of the selected image patches are set to 1. After each iteration of the classification model, the classification model is used to predict all the image patches, and some image patches with the highest prediction probability of the classification model are respectively selected from the H&E stained images with no significant efficacy and significant efficacy for the next stage of training until the model converges. Finally, a 128×128 pixel efficacy discrimination matrix is initialized, and the efficacy discrimination matrix is assigned values according to the position information and the prediction probability of the classification model. The pixel values of the unassigned pixels are all set to 0, and this efficacy discrimination matrix is the marker for predicting the efficacy of the PD-L1 inhibitor.

[0011] In a preferred embodiment, first, for the efficacy discrimination matrix and its corresponding efficacy labels, three different deep convolutional neural networks, namely EfficientNet-B3, ResNet50, and MobileNetV3, are used for training, and the cross-entropy loss function is used to adjust and improve the optimization objective of the deep neural network; secondly, the predicted class probabilities of the three deep convolutional neural networks are averaged, and the classification threshold is selected through the median of the predicted probabilities in the training set to obtain the final prediction result of the efficacy of PD-L1 inhibitors for gastric cancer patients; finally, the LOG-RANK test is used to evaluate the differences in the survival period and progression-free survival period of patients with different prediction results on the test set, so as to evaluate the accuracy of the prediction results.

[0012] Compared with the prior art, the present invention has the following beneficial effects:

[0013] (1) By means of semi-supervised learning, the present invention fully reduces the workload of annotating cancerous regions while visually displaying the specific locations of cancerous regions in gastric cancer H&E staining images.

[0014] (2) By means of multi-instance learning, the present invention obtains image patch-level labels from the slice-level labels of the efficacy of PD-L1 inhibitors for training, and constructs a slice-level efficacy discrimination matrix by combining the cancerous regions and efficacy prediction results of the image patches as a biomarker for predicting the efficacy of PD-L1 inhibitors.

[0015] (3) The present invention extracts features of different dimensions from the biomarkers for predicting the efficacy of PD-L1 inhibitors through a deep convolutional neural network to obtain the prediction result of the efficacy of gastric cancer patients after using PD-L1 inhibitors, assisting doctors in making clinical decisions. BRIEF DESCRIPTION OF THE DRAWINGS

[0016] Figure 1 is the overall framework flowchart of the preferred embodiment of the present invention;

[0017] Figure 2 is the flowchart of the gastric cancer region segmentation algorithm based on semi-supervised learning in the preferred embodiment of the present invention;

[0018] Figure 3 is the flowchart of the algorithm for obtaining biomarkers of the efficacy of PD-L1 inhibitors by multi-instance learning in the preferred embodiment of the present invention;

[0019] Figure 4 is the algorithm for predicting the efficacy of PD-L1 inhibitors based on a deep convolutional neural network in the preferred embodiment of the present invention. DETAILED DESCRIPTION OF THE EMBODIMENTS

[0020] The present invention will be further described below with reference to the accompanying drawings and embodiments.

[0021] It should be noted that the following detailed description is exemplary and is intended to provide further explanation of the present application. Unless otherwise specified, all technical and scientific terms used herein have the same meaning as commonly understood by those of ordinary skill in the technical field to which this application belongs.

[0022] It should be noted that the terms used herein are only for describing specific embodiments and are not intended to limit the exemplary embodiments according to the present application; as used herein, unless the context clearly indicates otherwise, the singular forms are also intended to include the plural forms. In addition, it should also be understood that when the terms "comprising" and / or "including" are used in this specification, they indicate the presence of features, steps, operations, devices, components, and / or combinations thereof.

[0023] In view of the problems that current pathologists have low consistency in the CPS score of gastric cancer patients and clinicians have difficulty predicting the efficacy of PD-1 / PD-L1 inhibitors for gastric cancer patients, the present invention proposes a method for intelligently predicting the efficacy of PD-L1 inhibitors from H&E staining images of gastric cancer, aiming to assist pathologists in evaluating the PD-L1 expression level and helping clinicians make decisions at the same time.

[0024] To achieve the above object, the present invention first studies a cancerous region segmentation model based on semi-supervised learning to segment the cancerous regions in H&E staining images. Secondly, the image patches containing the cancerous regions are taken out, and the class labels of each image patch are obtained according to the efficacy labels of PD-L1 inhibitors marked by clinicians, and an image classification model based on multi-instance learning is used for prediction at the image patch level. Finally, a efficacy discrimination matrix for the entire H&E staining image is generated based on the prediction results at the image patch level, and an image classification model based on a deep convolutional neural network is used for efficacy prediction at the slice level to assist doctors in pathological diagnosis and immunotherapy evaluation. The overall framework flowchart is as Figure 1 shown.

[0025] Specifically, the algorithm framework of the gastric cancer region segmentation algorithm based on semi-supervised learning is as Figure 2As shown below. First, the H&E staining images are cut into image patches of 512×512 pixels. According to the gastric cancer classification, some image patches are selected to label the gastric cancer regions, obtaining some labeled image patches and unlabeled image patches. Secondly, the present invention selects DeepLabV3+ as the segmentation backbone network, and uses the EfficientNet-B3 network pre-trained through the ImageNet dataset to replace the encoder structure in the DeepLabV3+ network to construct a meanteacher model. Finally, for the labeled image patches, they are input into the student network, and the cross-entropy loss function combined with FocalLoss is used to constrain between the obtained prediction results and the true labels, that is, MixedLoss. For the unlabeled image patches, they are respectively input into the teacher network and the student network, and the obtained prediction results are constrained by the consistency loss function, that is, Consistency Loss. Among them, the student network adjusts and improves the network optimization objective through MixedLoss combined with ConsistencyLoss, and the weight parameters of the teacher network are obtained from the exponential moving average of the weight parameters of the student network. At the same time, various auxiliary image transformations are added to the training data to amplify the diversity of the training data and further improve the generalization performance of the model.

[0026] Specifically, the algorithm flow of an algorithm for obtaining PD-L1 inhibitor efficacy markers based on multi-instance learning is as Figure 3 shown below. First, based on clinical data such as the survival period, progression-free survival period, and tumor shrinkage degree of patients, the efficacy labels of gastric cancer patients corresponding to the H&E staining images after using PD-L1 inhibitors are evaluated, denoted as significant efficacy 1 and no significant efficacy 0. Secondly, the gastric cancer region segmentation model selects the image patches containing the cancerous regions in the H&E staining images and the 4 image patches around them, and retains the position information of all the selected image patches in the original H&E staining images. Thirdly, according to the PD-L1 inhibitor efficacy labels, the class labels of the selected image patches are given, and a multi-instance learning-based image classification model is used for training. For the H&E staining images with no significant efficacy, the labels of all the selected image patches are set to 0, while for the H&E staining images with significant efficacy, the initial labels of the selected image patches are set to 1. After each iteration of the classification model, the classification model is used to predict all the image patches, and some image patches with the highest prediction probabilities of the classification model are selected from the H&E staining images with no significant efficacy and significant efficacy respectively for the next stage of training until the model converges. Finally, a 128×128 pixel efficacy discrimination matrix is initialized, and the efficacy discrimination matrix is assigned values according to the position information and the prediction probabilities of the classification model. The pixel values of the unassigned pixels are all set to 0, and this efficacy discrimination matrix is the marker for predicting the efficacy of PD-L1 inhibitors.

[0027] Specifically, the algorithm flow of a PD-L1 inhibitor efficacy prediction algorithm based on a deep convolutional neural network is as follows Figure 4 shown. First, for the efficacy discrimination matrix and its corresponding efficacy labels, three different deep convolutional neural networks, namely EfficientNet-B3, ResNet50, and MobileNetV3, are used for training, and the cross-entropy loss function is used to adjust and improve the optimization objective of the deep neural network. Secondly, the predicted class probabilities of the three deep convolutional neural networks are averaged, and the classification threshold is selected through the median of the predicted probabilities in the training set to obtain the final prediction result of the efficacy of PD-L1 inhibitors for gastric cancer patients. Finally, the LOG-RANK test is used to evaluate the differences in the survival period and progression-free survival period of patients with different prediction results on the test set, so as to evaluate the accuracy of the prediction results and continuously adjust and improve the performance of the model.

[0028] The present invention visually shows the specific location of the cancerous area in gastric cancer H&E staining images, which can assist pathologists in improving the efficiency and accuracy of diagnosis.

[0029] The present invention proposes a method for intelligently predicting the efficacy of PD-L1 inhibitors in gastric cancer H&E staining images, which can assist pathologists in evaluating the PD-L1 expression of gastric cancer patients and assist clinicians in making decisions. At the same time, it also provides a reference for the establishment of prediction methods for other key indicators in the tumor microenvironment.

[0030] The process or manner of using the product.

[0031] (1) Retrospectively collect the cases of gastric cancer patients with complete clinical data and follow-up information. At the same time, all these cases should have been treated with PD-L1 / PD-1 inhibitors for immunotherapy. Through a professional digital slide scanner, gastric cancer digital pathology slide H&E staining images are obtained to construct a gastric cancer digital pathology slide database.

[0032] (2) H&E staining images may have problems such as slice contamination and blank slices. After excluding useless data, the images are selectively enhanced to eliminate irrelevant information in the images and enhance the detectability of relevant information.

[0033] (3) The finally obtained gastric cancer digital pathology slide H&E staining images are divided into dataset 1, dataset 2, dataset 3, and dataset 4 according to a ratio of 1:1:1:2. Based on clinical data such as the survival period, progression-free survival period, and tumor shrinkage degree of the patients, clinicians evaluate the efficacy labels of the gastric cancer patients corresponding to all H&E staining images after using PD-L1 inhibitors, denoted as significant efficacy 1 and no significant efficacy 0.

[0034] (4) Cut the H&E stained images in Dataset 1 into image patches of 512×512 pixels under a 20× magnification. Select some image patches according to the gastric cancer classification, and have professional pathologists label the cancerous regions. Combine the unlabeled image patches to train a gastric cancer region segmentation model based on semi-supervised learning.

[0035] (5) Use the gastric cancer region segmentation model to predict the H&E stained images in Dataset 2. Select the image patches containing the cancerous regions and the 4 image patches around them, and use an image classification model based on multi-instance learning for training according to the efficacy labels after PD-L1 inhibitor treatment.

[0036] (6) Combine the gastric cancer region segmentation model based on semi-supervised learning and the image classification model based on multi-instance learning to predict the H&E stained images in Dataset 3, and obtain an efficacy discrimination matrix. Correlate the PD-L1 inhibitor efficacy labels with the efficacy discrimination matrix, and use 3 image classification models based on deep convolutional neural networks for training respectively.

[0037] (7) Combine the gastric cancer region segmentation model based on semi-supervised learning and the image classification model based on multi-instance learning with the 3 image classification models based on deep convolutional neural networks to predict the H&E stained images in Dataset 4, and average the 3 PD-L1 efficacy prediction probabilities obtained to get the prediction result of the final gastric cancer PD-L1 inhibitor efficacy.

[0038] (8) Use the LOG-RANK test to evaluate the differences in the survival and progression-free survival of gastric cancer patients with different prediction results on Dataset 4, so as to evaluate the accuracy of the prediction results, continuously adjust and improve the performance of the model, and better assist doctors in making clinical decisions.

Claims

1. A method for intelligently predicting the efficacy of PD-L1 inhibitors from gastric cancer H&E staining images, characterized in that, A cancerous region segmentation model based on semi-supervised learning is proposed to segment cancerous regions in H&E stained images. Secondly, image patches containing cancerous regions are extracted, and class labels for each image patch are obtained according to the PD-L1 inhibitor efficacy labels marked by clinicians. An image classification model based on multi-instance learning is used for prediction at the image patch level. Finally, a treatment efficacy discrimination matrix for the entire H&E stained image is generated based on the prediction results at the image patch level, and an image classification model based on a deep convolutional neural network is used for treatment efficacy prediction at the slice level. Generating a treatment efficacy discrimination matrix for the entire H&E stained image based on the prediction results at the image patch level includes: First, based on the patient's survival period, progression-free survival period, and degree of tumor shrinkage, evaluate the treatment efficacy label of the gastric cancer patient corresponding to the H&E stained image after using the PD-L1 inhibitor, denoted as significant efficacy 1 and no significant efficacy 0. Secondly, the gastric cancer region segmentation model selects the image patches containing cancerous regions in the H&E stained image and the 4 image patches around them, and retains the position information of all the selected image patches in the original H&E stained image. Thirdly, according to the PD-L1 inhibitor treatment efficacy label, give the class labels of the selected image patches, and use an image classification model based on multi-instance learning for training. For H&E stained images with no significant efficacy, the labels of all selected image patches are set to 0, while for H&E stained images with significant efficacy, the initial labels of the selected image patches are set to 1. After each iteration of the classification model, use the classification model to predict all the image patches, and select the partial image patches with the highest prediction probability of the classification model from the H&E stained images with no significant efficacy and significant efficacy respectively for the next stage of training until the model converges. Finally, initialize a 128×128 pixel treatment efficacy discrimination matrix, and assign values to the treatment efficacy discrimination matrix according to the position information and the prediction probability of the classification model. The pixel values of the unassigned pixels are all set to 0, and this treatment efficacy discrimination matrix is the marker for predicting the efficacy of the PD-L1 inhibitor.

2. The method for intelligently predicting the efficacy of PD-L1 inhibitors from gastric cancer H&E staining images according to claim 1, characterized in that, First, cut the H&E stained images into image patches of 512×512 pixels, select some image patches according to the gastric cancer classification to label the gastric cancer regions, and obtain some labeled image patches and unlabeled image patches; secondly, select DeepLab V3+ as the segmentation backbone network, and use the EfficientNet-B3 network pre-trained through the ImageNet dataset to replace the encoder structure in the DeepLab V3+ network to construct a mean teacher model; finally, for the labeled image patches, input them into the student network, and use the cross-entropy loss function combined with Focal Loss to constrain between the obtained prediction results and the true labels, that is, Mixed Loss; for the unlabeled image patches, input them into the teacher network and the student network respectively, and use the consistency loss function to constrain the obtained prediction results, that is, Consistency Loss; among them, the student network adjusts and improves the network optimization objective through Mixed Loss combined with Consistency Loss, and the weight parameters of the teacher network are obtained from the exponential moving average of the weight parameters of the student network; at the same time, various auxiliary image transformations are added to the training data.

3. The method for intelligently predicting the efficacy of PD-L1 inhibitors from gastric cancer H&E staining images according to claim 1, characterized in that, First, train the efficacy discrimination matrix and the corresponding efficacy labels using 3 different deep convolutional neural networks, namely EfficientNet-B3, ResNet50, and MobileNetV3, and use the cross-entropy loss function to adjust and improve the optimization objective of the deep neural network; secondly, average the predicted class probabilities of the 3 deep convolutional neural networks, and select the classification threshold through the median of the predicted probabilities in the training set to obtain the final prediction result of the efficacy of PD-L1 inhibitors for gastric cancer patients; finally, use the LOG-RANK test to evaluate the differences in the survival period and progression-free survival period of patients with different prediction results on the test set to evaluate the accuracy of the prediction results.

Citation Information

Patent Citations

  • A breast cancer pathological tissue image segmentation method based on semi-supervised k-means algorithm

    CN109102510A

  • Colorectal cancer digital pathological image discrimination method and system based on weak supervised learning

    CN113221978A