A method and system for training and predicting a gastric cancer HER2 status prediction model

By acquiring and labeling gastric cancer tissue slice images and training a deep learning model, the problems of high cost and long time required for gastric cancer HER2 status prediction were solved, achieving efficient and accurate HER2 status assessment.

CN116110608BActive Publication Date: 2026-06-02SHUNDE HOSPITAL SOUTHERN MEDICAL UNIV (THE FIRST PEOPLES HOSPITAL OF SHUNDE FOSHAN)

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
SHUNDE HOSPITAL SOUTHERN MEDICAL UNIV (THE FIRST PEOPLES HOSPITAL OF SHUNDE FOSHAN)
Filing Date
2023-01-18
Publication Date
2026-06-02

AI Technical Summary

Technical Problem

Existing methods for predicting HER2 status in gastric cancer are costly and time-consuming, and there is a lack of effective HER2 status assessment methods in gastric cancer, especially compared to HER2 status assessment methods for breast cancer, which have not been widely used in gastric cancer.

Method used

By acquiring a sample image set, including a first stained image and a second stained image, state labeling processing is performed on the second stained image, and combined with pathological annotation processing, a gastric cancer HER2 state prediction model is trained. Deep learning and computer vision technologies are used to improve prediction accuracy.

Benefits of technology

This method enables efficient and low-cost prediction of HER2 status in gastric cancer, improving the accuracy and efficiency of prediction while reducing the time and cost associated with immunohistochemical staining.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN116110608B_ABST
    Figure CN116110608B_ABST
Patent Text Reader

Abstract

The application discloses a kind of stomach cancer HER2 state prediction model training, prediction method and system, wherein, training method includes: obtaining sample image set, the sample image set includes first dyeing image and second dyeing image;Second dyeing image in the sample image set is marked with state, and image label is obtained;The sample image set is pathologically annotated, and training sample set is obtained;The training sample set and the image label are input into untrained stomach cancer HER2 state prediction model and are trained, and the stomach cancer HER2 state prediction model of training completion is obtained.The embodiment of the application can improve the accuracy of stomach cancer HER2 state prediction model, and can be widely applied in artificial intelligence technical field.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to the field of artificial intelligence technology, and in particular to a training and prediction method and system for a gastric cancer HER2 status prediction model. Background Technology

[0002] Human epidermal growth factor receptor 2 (HER2)-positive gastric cancer is a rare and specialized type of gastric cancer, accounting for approximately 10% of all gastric cancers. In recent years, numerous clinical studies have confirmed that HER2-positive tumors possess unique biological characteristics, making them suitable for specific treatment strategies. Accurate prediction of HER2 status plays an increasingly important role in these applications. Unfortunately, while the HERONE challenge, based on H&E staining to assess HER2 status in breast cancer, is gaining momentum, this area remains unexplored in gastric cancer. Related techniques assess samples using hematoxylin and eosin (HE) staining, followed by prediction of HER2 status using techniques such as immunohistochemical staining (IHC) and fluorescence in situ hybridization (FISH). However, prediction using IHC is costly, time-consuming, and yields poor predictive results. In summary, the technical problems existing in these techniques urgently need to be addressed. Summary of the Invention

[0003] In view of this, embodiments of the present invention provide a training and prediction method and system for a gastric cancer HER2 status prediction model, so as to improve the accuracy of gastric cancer HER2 status prediction and reduce prediction time and cost.

[0004] On one hand, the present invention provides a training method for a gastric cancer HER2 status prediction model, the method comprising:

[0005] Obtain a sample image set, which includes a first stained image and a second stained image;

[0006] The second stained image in the sample image set is subjected to state labeling processing to obtain image labels;

[0007] The sample image set is subjected to pathological annotation processing to obtain the training sample set;

[0008] The training sample set and the image labels are input into the untrained gastric cancer HER2 status prediction model for training, resulting in a trained gastric cancer HER2 status prediction model.

[0009] Optionally, the acquisition of the sample image set, the sample image set including a first stained image and a second stained image, includes:

[0010] Collect a set of gastric cancer tissue section samples;

[0011] The gastric cancer tissue section sample set was stained and a sample image set was obtained by scanning the sections. The sample image set includes a first stained image and a second stained image.

[0012] Optionally, the step of performing state labeling processing on the second stained image in the sample image set to obtain image labels includes:

[0013] The second stained image in the sample image set is subjected to HER2 state interpretation processing according to a unified standard to obtain the sample state;

[0014] The sample states are used as labels for the corresponding images to obtain image labels.

[0015] Optionally, the step of performing pathological annotation processing on the sample image set to obtain a training sample set includes:

[0016] The first stained image in the sample image set is subjected to region annotation processing to obtain a region stained image set;

[0017] The region-stained image set and the second-stained image in the sample image set are cropped to obtain a cropped image set.

[0018] The cropped image set is labeled based on image features and histopathology to obtain a training sample set.

[0019] Optionally, the step of inputting the training sample set and the image labels into an untrained gastric cancer HER2 status prediction model for training, to obtain a trained gastric cancer HER2 status prediction model, includes:

[0020] The training sample set and the image labels are input into the untrained HER2 state prediction model for gastric cancer to obtain the state prediction results.

[0021] The training loss value is determined based on the state prediction result and the image label;

[0022] The parameters of the gastric cancer HER2 status prediction model are updated based on the loss value to obtain the trained gastric cancer HER2 status prediction model.

[0023] Optionally, the cropping process of the stained image of the region and the second stained image in the sample image set to obtain a cropped image set includes:

[0024] The background region is cropped from the stained region image and the second stained region image respectively to obtain the first cropped image and the second cropped image;

[0025] Based on the first cropped image, the second cropped image is subjected to shape alignment and resolution adjustment to obtain the third cropped image;

[0026] The first and third cropped images are subjected to tile cropping processing to obtain a cropped image set.

[0027] Optionally, the step of inputting the training sample set and the image labels into an untrained gastric cancer HER2 state prediction model to obtain the state prediction result includes:

[0028] The gastric cancer HER2 status prediction model includes an identification module and a prediction module;

[0029] The tumor identification process is performed on the region-stained image set in the training sample set by the identification module to obtain the identification result;

[0030] Based on the prediction module and the recognition results, the training sample set is classified, predicted, and aggregated for voting to obtain the state prediction result.

[0031] Optionally, the step of performing classification prediction and aggregation voting on the training sample set based on the prediction module and the recognition results to obtain the state prediction result includes:

[0032] The prediction module includes a residual neural network and a classification unit;

[0033] The training sample set is filtered based on the recognition results to obtain the prediction dataset;

[0034] The predicted dataset is input into a residual neural network for classification to obtain classification labels;

[0035] The proportion of the classification labels is statistically analyzed, and the statistical results are input into several state classification units to predict the state, thus obtaining a prediction result set.

[0036] The predicted result set is aggregated and voted on to obtain the state prediction result.

[0037] On the other hand, embodiments of the present invention also provide a method for predicting the HER2 status of gastric cancer, the method comprising:

[0038] Obtain tissue sections of the gastric cancer to be predicted;

[0039] The gastric cancer tissue section to be predicted is digitally scanned to obtain scanned images;

[0040] The scanned image is input into the gastric cancer HER2 state prediction model obtained by the training method of the gastric cancer HER2 state prediction model as described in any of the preceding items, and the state prediction result is obtained.

[0041] On the other hand, embodiments of the present invention also provide a training system for a gastric cancer HER2 status prediction model, the system comprising:

[0042] The first module is used to acquire a sample image set, which includes a first stained image and a second stained image.

[0043] The second module is used to perform state labeling processing on the second stained image in the sample image set to obtain image labels;

[0044] The third module is used to perform pathological annotation processing on the sample image set to obtain a training sample set;

[0045] The fourth module is used to input the training sample set and the image labels into the untrained gastric cancer HER2 state prediction model for training, so as to obtain the trained gastric cancer HER2 state prediction model.

[0046] Compared with existing technologies, the present invention, employing the above technical solution, has the following technical advantages: In this embodiment, image labels are obtained by performing state labeling processing on the second stained images in the sample image set, serving as the prediction output labels for the gastric cancer HER2 state prediction model. Then, a training sample set is obtained by performing pathological annotation processing on the sample image set to train the gastric cancer HER2 state prediction model, thereby improving the accuracy of the gastric cancer HER2 state prediction model. Furthermore, this embodiment can also perform state prediction using the gastric cancer HER2 state prediction model, thereby reducing prediction costs and time, and improving prediction efficiency. Attached Figure Description

[0047] To more clearly illustrate the technical solutions in the embodiments of this application, the accompanying drawings used in the description of the embodiments will be briefly introduced below. Obviously, the accompanying drawings described below are only some embodiments of this application. For those skilled in the art, other drawings can be obtained based on these drawings without creative effort.

[0048] Figure 1 This is a flowchart of a training method for a gastric cancer HER2 status prediction model provided in an embodiment of this application;

[0049] Figure 2 These are the first and second stained images corresponding to a gastric cancer tissue slice in a sample image set provided in this application embodiment;

[0050] Figure 3 This is a flowchart of a pathological annotation of a first stained image provided in an embodiment of this application;

[0051] Figure 4 This is a schematic diagram illustrating the classification and labeling of image blocks according to an embodiment of this application;

[0052] Figure 5 This is a flowchart of a method for pathological annotation of a sample image set, provided in an embodiment of this application;

[0053] Figure 6 This is a structural diagram of a prediction module provided in an embodiment of this application. Detailed Implementation

[0054] To make the objectives, technical solutions, and advantages of this application clearer, the following detailed description is provided in conjunction with the accompanying drawings and embodiments. It should be understood that the specific embodiments described herein are merely illustrative and not intended to limit the scope of this application.

[0055] First, let's analyze some of the terms used in this application:

[0056] H&E staining: Hematoxylin-eosin staining, or HE staining for short, is one of the commonly used staining methods in paraffin sectioning. Hematoxylin is an alkaline staining solution, which mainly tints the chromatin in the cell nucleus and nucleic acids in the cytoplasm with a purple-blue color; eosin is an acidic dye, which mainly tints the components in the cytoplasm and extracellular matrix with a red color. HE staining is the most basic and widely used technique in histology, embryology, and pathology teaching and research.

[0057] IHC staining: Immunohistochemistry (IHC) is a process that utilizes the principle of specific binding between antigens and antibodies. It uses a chemical reaction to make the chromogenic agent (fluorescein, enzyme, metal ion, isotope) labeled with antibodies appear colored to identify antigens (peptides and proteins) in tissue cells.

[0058] Targeted therapy: Targeted therapy is a treatment approach that targets known cancer-causing sites at the cellular and molecular level. Trastuzumab is the earliest and most well-supported targeted therapy drug for gastric cancer.

[0059] Human epidermal growth factor receptor 2 (HER2): Human epidermal growth factor receptor 2 (HER2) is the main target of trastuzumab. Overexpression of HER2 promotes cell proliferation and tumorigenesis by activating various signaling pathways. Blocking HER2 can improve the prognosis of patients with HER2-positive tumors through multiple mechanisms.

[0060] Machine learning (ML) is the study of how computers can simulate or implement human learning behavior to acquire new knowledge or skills and reorganize existing knowledge structures to continuously improve their performance.

[0061] Deep Learning: Deep learning (DL) is a new research direction in the field of machine learning (ML). It features more complex algorithms, higher data dimensions, and a wider range of operational scenarios.

[0062] Computer vision (CV) is a simulation of biological vision using computers and related equipment. Its main task is to process acquired images or videos to obtain three-dimensional information about the corresponding scene, just as humans and many other types of organisms do every day.

[0063] Image segmentation: Image segmentation is a research area in computer vision. It is the technique and process of dividing an image into several specific regions with unique properties and extracting targets of interest.

[0064] Semantic segmentation: Semantic segmentation is a type of image segmentation that infers relevant knowledge or semantics from images and segments the images according to the semantics. Each semantic is a category.

[0065] Human epidermal growth factor receptor 2 (HER2)-positive gastric cancer is a rare and special type of gastric cancer, accounting for approximately 10% of all gastric cancers. This type of gastric cancer has a poor prognosis, and conventional chemotherapy is not very effective. Therefore, early and accurate identification and prediction of HER2-positive gastric cancer is of great significance. The general procedure for predicting HER2 in gastric cancer involves evaluating the obtained sample using hematoxylin and eosin (HE) staining, followed by prediction of HER2 status using techniques such as immunohistochemical staining (IHC) and fluorescence in situ hybridization (FISH). Among these, IHC is currently the most widely used and important auxiliary method. In recent years, emerging technologies such as NGS-based liquid biopsy have been used for the prediction and assessment of HER2-positive gastric cancer; however, these new technologies have poor accessibility and are currently difficult to popularize.

[0066] Digital pathology refers to the process of acquiring full-field digital slices of stained tissue slides using whole-slide scanning equipment, and then objectively analyzing these digital images using artificial intelligence models trained with different algorithms. This effectively assists pathologists in their daily work. The process of using artificial intelligence to analyze tissue slices is also known as computational pathology. In recent years, numerous clinical studies have confirmed that HER2-positive tumors possess unique biological characteristics, making them suitable for specific treatment strategies. Accurate prediction of HER2 status plays an increasingly important role, and AI-based computational pathology methods hold promise for improving the efficiency and accuracy of HER2 status assessment in gastric cancer. Unfortunately, despite the HERONE challenge, which uses H&E staining to assess HER2 status in breast cancer, this field remains largely unexplored in gastric cancer.

[0067] In view of this, refer to Figure 1 This invention provides a training method for a gastric cancer HER2 status prediction model, the training method comprising:

[0068] S101. Obtain a sample image set, the sample image set including a first stained image and a second stained image;

[0069] S102. Perform state labeling processing on the second stained image in the sample image set to obtain image labels;

[0070] S103. Perform pathological annotation processing on the sample image set to obtain a training sample set;

[0071] S104. Input the training sample set and the image labels into the untrained gastric cancer HER2 status prediction model for training, and obtain the trained gastric cancer HER2 status prediction model.

[0072] In this embodiment of the invention, a sample image set containing a first stained image and a second stained image is first collected, wherein the first stained image is an HE stained image and the second stained image is an IHC stained image. Then, this embodiment of the invention interprets the HER2 status of the second stained image, using HER2 negative or positive as the label for the slice image; this is also the final label predicted by the gastric cancer HER2 status prediction model of this invention. Next, this embodiment of the invention performs pathological annotation processing on all images in the sample image set, uses the annotated images as a training sample set to train the gastric cancer HER2 status prediction model, and optimizes the gastric cancer HER2 status prediction model using the image labels obtained above, resulting in a trained gastric cancer HER2 status prediction model. This embodiment of the invention, by drawing on the deep learning research results of two mainstream directions in computer vision—image semantic segmentation and image classification—and supplementing them with data augmentation and ensemble learning ideas and technologies, combined with the significant characteristics of medical images—high resolution and small sample size—innovatively designs and implements a gastric cancer HER2 status prediction model, thereby improving the accuracy and accessibility of HER2 status assessment.

[0073] Further, as a preferred embodiment, the acquisition of the sample image set, the sample image set including a first stained image and a second stained image, includes:

[0074] Collect a set of gastric cancer tissue section samples;

[0075] The gastric cancer tissue section sample set was stained and a sample image set was obtained by scanning the sections. The sample image set includes a first stained image and a second stained image.

[0076] In this embodiment of the invention, gastric cancer tissue samples from multiple centers over the past 10 years were collected, totaling 251 positive cases and 1000 negative cases. All sections were scanned to obtain corresponding H&E staining images and IHC staining images, i.e., the first staining image and the second staining image. For example... Figure 2 As shown, the left image is an H&E staining image of a gastric cancer tissue section, and the right image is its corresponding IHC staining image.

[0077] Further, as a preferred embodiment, the step of performing state labeling processing on the second stained image in the sample image set to obtain image labels includes:

[0078] The second stained image in the sample image set is subjected to HER2 state interpretation processing according to a unified standard to obtain the sample state;

[0079] The sample states are used as labels for the corresponding images to obtain image labels.

[0080] In this embodiment of the invention, the second stained image, namely the IHC stained image, in the sample image set is used to determine the HER2 status. That is, multiple pathology experts jointly determine the HER2 status of the IHC stained image of the slide according to a unified standard, and use HER2 negative or positive as the label of the slide. The image label is used as the final label of the gastric cancer HER2 status prediction model of the present invention.

[0081] As a further preferred embodiment, the step of performing pathological annotation processing on the sample image set to obtain a training sample set includes:

[0082] The first stained image in the sample image set is subjected to region annotation processing to obtain a region stained image set;

[0083] The region-stained image set and the second-stained image in the sample image set are cropped to obtain a cropped image set.

[0084] The cropped image set is labeled based on image features and histopathology to obtain a training sample set.

[0085] In this embodiment of the invention, a deep learning image annotation tool is used to distinguish between tumor regions and non-tumor regions in the first stained image (H&E stained image) in the sample image set, resulting in a set of region stained images, such as... Figure 3 As shown, (a) is an H&E staining image of a gastric cancer tissue section sample; (b) is the tumor region marked on the H&E staining image by a pathologist; and (c) is an image converted from the marked H&E staining image into an image suitable for model training. In this embodiment of the invention, the marked region staining image set and the second staining image in the sample image set are cropped to obtain a cropped image set. Finally, by combining image features and histopathology, pathologists designed four-category labels for the cropped image set, resulting in a training sample set. The four-category labels are shown in Table 1, which contains the four-category labels for the IHC patch annotations, as shown below:

[0086]

[0087] Table 1

[0088] like Figure 4As shown, A indicates a weak label, B indicates a strong label (<10%), C indicates a strong label (10%-49%), and D indicates a strong label (≥50%). The purpose of the four-category labels is as follows: 10% is the currently defined cutoff point (i.e., stained cells accounting for ≥10% of tumor cells are considered positive), and 50% is the physical cutoff point. Performing four-category classification during the annotation process, using 10% and 50% as the critical points for area classification, avoids individual differences arising from subjective judgment of the specific proportions and retains as much annotation information as possible for later reference. Since the second stained image in the region stained image set and the sample image set has already been cropped in the above steps (i.e., the small patches of the IHC stained image and the H&E stained image are aligned), the label of each small patch of the IHC stained image can be found in its corresponding small patch of the H&E stained image.

[0089] Further, as a preferred embodiment, the step of inputting the training sample set and the image labels into an untrained gastric cancer HER2 status prediction model for training, to obtain a trained gastric cancer HER2 status prediction model, includes:

[0090] The training sample set and the image labels are input into the untrained HER2 state prediction model for gastric cancer to obtain the state prediction results.

[0091] The training loss value is determined based on the state prediction result and the image label;

[0092] The parameters of the gastric cancer HER2 status prediction model are updated based on the loss value to obtain the trained gastric cancer HER2 status prediction model.

[0093] In this embodiment of the invention, a training sample set and image labels are input into an untrained gastric cancer HER2 state prediction model for training. The training sample set and the image labels obtained above are input into the untrained gastric cancer HER2 state prediction model to obtain state prediction results. Then, the gastric cancer HER2 state prediction model is optimized based on the image labels. Specifically, the training loss value is determined through the state prediction results and image labels. The parameters of the gastric cancer HER2 state prediction model are updated based on the training loss value to obtain a trained gastric cancer HER2 state prediction model.

[0094] Further, as a preferred embodiment, the cropping process of the stained image of the region and the second stained image in the sample image set to obtain a cropped image set includes:

[0095] The background region is cropped from the stained region image and the second stained region image respectively to obtain the first cropped image and the second cropped image;

[0096] Based on the first cropped image, the second cropped image is subjected to shape alignment and resolution adjustment to obtain the third cropped image;

[0097] The first and third cropped images are subjected to tile cropping processing to obtain a cropped image set.

[0098] like Figure 5 As shown, in this embodiment of the invention, the region-annotated stained image and the second stained image are first cropped to obtain a first cropped image and a second cropped image. Background areas without gastric cancer tissue are removed, and the image size is reduced to save memory occupied by the image file. Then, the second stained image is manually rotated to roughly align the shape of the gastric cancer tissue in the second cropped image (IHC stained image) with the shape of the gastric cancer tissue in the first cropped image (H&E stained image). The resolution of the IHC stained image is then adjusted to match that of the H&E stained image, resulting in a third cropped image. Finally, the first and third cropped images are processed by tile cropping, each cropped into multiple small tiles with a resolution of 512x512, resulting in a cropped image set. Since the gastric cancer tissue in the H&E stained image and the IHC stained image has been aligned in the above steps, the small tiles in the IHC stained image and the small tiles in the H&E stained image are also aligned.

[0099] Further, as a preferred embodiment, the step of inputting the training sample set and the image labels into an untrained gastric cancer HER2 state prediction model to obtain state prediction results includes:

[0100] The gastric cancer HER2 status prediction model includes an identification module and a prediction module;

[0101] The tumor identification process is performed on the region-stained image set in the training sample set by the identification module to obtain the identification result;

[0102] Based on the prediction module and the recognition results, the training sample set is classified, predicted, and aggregated for voting to obtain the state prediction result.

[0103] Reference Figure 6The HER2 state prediction model for gastric cancer in this embodiment of the invention includes an identification module and a prediction module. The prediction module is essentially an image segmentation task in computer vision. Therefore, this embodiment uses several deep learning models commonly used in medical image segmentation tasks for training, including FCN-8s, U-Net, SegNet, DeepLab v3+, and PSPNet. Experimentally, DeepLab v3+, with the best quantification index, was selected as the final identification model. The identification module performs tumor identification processing on the region-stained image set in the training sample set to obtain the identification results. Then, based on the identification results, the prediction module performs classification prediction and aggregation voting on the training sample set to obtain the state prediction results. The prediction module includes a ResNet50 model and several state classification models, which may include decision trees, logistic regression, support vector machines, random forests, Naive Bayes, GBDT, AdaBoost, XGBoost, etc.

[0104] Further, as a preferred embodiment, the step of performing classification prediction and aggregation voting on the training sample set based on the prediction module and the recognition results to obtain the state prediction result includes:

[0105] The prediction module includes a residual neural network and a classification unit;

[0106] The training sample set is filtered based on the recognition results to obtain the prediction dataset;

[0107] The predicted dataset is input into a residual neural network for classification to obtain classification labels;

[0108] The proportion of the classification labels is statistically analyzed, and the statistical results are input into several state classification units to predict the state, thus obtaining a prediction result set.

[0109] The predicted result set is aggregated and voted on to obtain the state prediction result.

[0110] In this embodiment of the invention, the prediction module includes a residual neural network and a classification unit. The residual neural network is a ResNet50 model, and the classification unit consists of several state classification models. This embodiment uses the ResNet50 model to predict the classification of H&E stained small patches. The input of this model is the H&E stained small patch, and the output is its corresponding classification label. All small patches obtained from each H&E stained image are processed through the residual neural network to obtain classification labels. Then, the proportion of each classification label is statistically calculated. The statistical results of the proportions are input into several state classification units for state prediction. The outputs of these models are all HER2 states: negative or positive. Finally, using the voting concept in ensemble learning, the prediction results of all state classification models are aggregated and voted on. The prediction result with the highest number of votes is the final output of the gastric cancer HER2 state prediction model.

[0111] On the other hand, embodiments of the present invention also provide a method for predicting the HER2 status of gastric cancer, the method comprising:

[0112] Obtain tissue sections of the gastric cancer to be predicted;

[0113] The gastric cancer tissue section to be predicted is digitally scanned to obtain scanned images;

[0114] The scanned image is input into the gastric cancer HER2 state prediction model obtained by the training method of the gastric cancer HER2 state prediction model as described in any of the preceding items, and the state prediction result is obtained.

[0115] In this embodiment of the invention, a gastric cancer tissue slice to be predicted is obtained. The gastric cancer tissue slice to be predicted is scanned by a scanner to obtain a digital slice, namely an H&E stained image, which is then input into the gastric cancer HER2 status prediction model obtained by the training method of the gastric cancer HER2 status prediction model. The tumor region is identified by the recognition module to obtain the H&E tumor image, and then the HER2 status (negative or positive) of the H&E tumor image is predicted by the prediction module to obtain the final status prediction result.

[0116] It is understood that the content of the above-mentioned training method embodiment for the gastric cancer HER2 status prediction model is applicable to this gastric cancer HER2 status prediction method embodiment. The specific functions implemented by this gastric cancer HER2 status prediction method embodiment are the same as those of the above-mentioned training method embodiment for the gastric cancer HER2 status prediction model, and the beneficial effects achieved are also the same as those achieved by the above-mentioned training method embodiment for the gastric cancer HER2 status prediction model.

[0117] On the other hand, embodiments of the present invention also provide a training system for a gastric cancer HER2 status prediction model, the system comprising:

[0118] The first module is used to acquire a sample image set, which includes a first stained image and a second stained image.

[0119] The second module is used to perform state labeling processing on the second stained image in the sample image set to obtain image labels;

[0120] The third module is used to perform pathological annotation processing on the sample image set to obtain a training sample set;

[0121] The fourth module is used to input the training sample set and the image labels into the untrained gastric cancer HER2 state prediction model for training, so as to obtain the trained gastric cancer HER2 state prediction model.

[0122] It is understood that the content of the above-described training system embodiment for the gastric cancer HER2 status prediction model is applicable to the embodiment of the gastric cancer HER2 status prediction method. The specific functions implemented by the training system embodiment for the gastric cancer HER2 status prediction model are the same as those of the above-described training method embodiment for the gastric cancer HER2 status prediction model, and the beneficial effects achieved are also the same as those achieved by the above-described training method embodiment for the gastric cancer HER2 status prediction model.

[0123] In related technologies, the obtained samples are usually evaluated by hematoxylin and eosin (HE) staining, and then the HER2 status is predicted by techniques such as immunohistochemical staining (IHC) and fluorescence in situ hybridization (FISH). IHC staining is expensive and generally takes 3-7 working days.

[0124] In summary, the embodiments of the present invention have the following advantages: the gastric cancer HER2 status prediction model trained by the embodiments of the present invention can directly predict the HER2 status of gastric cancer based on hematoxylin and eosin (HE) stained sections, no longer relying on the immunohistochemical staining (IHC) technology currently used in clinical practice, reducing working time and saving costs, and improving the efficiency of status prediction.

[0125] In some alternative embodiments, the functions / operations mentioned in the block diagrams may not occur in the order shown in the operation diagrams. For example, depending on the functions / operations involved, two consecutively shown blocks may actually be executed substantially simultaneously, or the blocks may sometimes be executed in reverse order. Furthermore, the embodiments presented and described in the flowcharts of this invention are provided by way of example to provide a more comprehensive understanding of the technology. The disclosed methods are not limited to the operations and logic flows presented herein. Alternative embodiments are contemplated in which the order of various operations is altered and sub-operations described as part of a larger operation are executed independently.

[0126] Furthermore, although the invention has been described in the context of functional modules, it should be understood that, unless otherwise stated, one or more of the described functions and / or features may be integrated into a single physical device and / or software module, or one or more functions and / or features may be implemented in a separate physical device or software module. It is also understood that a detailed discussion of the actual implementation of each module is unnecessary for understanding the invention. Rather, given the properties, functions, and internal relationships of the various functional modules in the apparatus disclosed herein, the actual implementation of the module will be understood within the scope of conventional skill of an engineer. Therefore, those skilled in the art can implement the invention as set forth in the claims using ordinary techniques without excessive experimentation. It is also understood that the specific concepts disclosed are merely illustrative and not intended to limit the scope of the invention, which is determined by the full scope of the appended claims and their equivalents.

[0127] In the description of this specification, references to terms such as "one embodiment," "some embodiments," "example," "specific example," or "some examples," etc., indicate that a specific feature, structure, material, or characteristic described in connection with that embodiment or example is included in at least one embodiment or example of the invention. In this specification, the illustrative expressions of the above terms do not necessarily refer to the same embodiment or example. Furthermore, the specific features, structures, materials, or characteristics described may be combined in any suitable manner in one or more embodiments or examples.

[0128] Although embodiments of the invention have been shown and described, those skilled in the art will understand that various changes, modifications, substitutions and alterations can be made to these embodiments without departing from the principles and spirit of the invention, the scope of which is defined by the claims and their equivalents.

[0129] The above is a detailed description of the preferred embodiments of the present invention, but the present invention is not limited to the embodiments described. Those skilled in the art can make various equivalent modifications or substitutions without departing from the spirit of the present invention, and these equivalent modifications or substitutions are all included within the scope defined by the claims of this application.

Claims

1. A training method for a gastric cancer HER2 status prediction model, characterized in that, The training method includes: Obtain a sample image set, which includes a first stained image and a second stained image; The second stained image in the sample image set is subjected to state labeling processing to obtain image labels; The sample image set is subjected to pathological annotation processing to obtain the training sample set; The training sample set and the image labels are input into the untrained gastric cancer HER2 status prediction model for training, and the trained gastric cancer HER2 status prediction model is obtained. The step of inputting the training sample set and the image labels into an untrained gastric cancer HER2 status prediction model for training, to obtain a trained gastric cancer HER2 status prediction model, includes: The training sample set and the image labels are input into the untrained HER2 state prediction model for gastric cancer to obtain the state prediction results. The training loss value is determined based on the state prediction result and the image label; The parameters of the gastric cancer HER2 status prediction model are updated based on the loss value to obtain the trained gastric cancer HER2 status prediction model. The step of inputting the training sample set and the image labels into an untrained gastric cancer HER2 state prediction model to obtain state prediction results includes: The gastric cancer HER2 status prediction model includes an identification module and a prediction module; The tumor identification process is performed on the region-stained image set in the training sample set by the identification module to obtain the identification result; Based on the prediction module and the recognition results, the training sample set is classified, predicted, and aggregated for voting to obtain the state prediction result. The step of performing classification prediction and aggregation voting on the training sample set based on the prediction module and the recognition results to obtain the state prediction result includes: The prediction module includes a residual neural network and a classification unit; The training sample set is filtered based on the recognition results to obtain a prediction dataset; the prediction dataset consists of H&E-colored small image patches. The predicted dataset is input into a residual neural network for classification to obtain the classification labels corresponding to the H&E-colored small map patches; The percentage statistics of the classification labels are calculated, and the percentage statistics results are input into several state classification units for state prediction to obtain a prediction result set; the prediction result set is the HER2 state predicted by each state classification unit, and the HER2 state is represented as negative or positive. The predicted result set is aggregated and voted on to obtain the state prediction result.

2. The method according to claim 1, characterized in that, The acquisition of the sample image set, which includes a first stained image and a second stained image, includes: Collect a set of gastric cancer tissue section samples; The gastric cancer tissue section sample set was stained and a sample image set was obtained by scanning the sections. The sample image set includes a first stained image and a second stained image.

3. The method according to claim 1, characterized in that, The step of performing state labeling processing on the second stained image in the sample image set to obtain image labels includes: The second stained image in the sample image set is subjected to HER2 state interpretation processing according to a unified standard to obtain the sample state; The sample states are used as labels for the corresponding images to obtain image labels.

4. The method according to claim 1, characterized in that, The step of performing pathological annotation processing on the sample image set to obtain a training sample set includes: The first stained image in the sample image set is subjected to region annotation processing to obtain a region stained image set; The stained image set of the region and the second stained image in the sample image set are cropped to obtain a cropped image set; The cropped image set is labeled based on image features and histopathology to obtain a training sample set.

5. The method according to claim 4, characterized in that, The cropping process of the stained image of the region and the second stained image in the sample image set to obtain a cropped image set includes: The background region is cropped from the stained region image and the second stained region image respectively to obtain the first cropped image and the second cropped image; Based on the first cropped image, the second cropped image is subjected to shape alignment and resolution adjustment to obtain the third cropped image; The first and third cropped images are subjected to tile cropping processing to obtain a cropped image set.

6. A method for predicting HER2 status in gastric cancer, characterized in that, The prediction method includes: Obtain tissue sections of the gastric cancer to be predicted; The gastric cancer tissue section to be predicted is digitally scanned to obtain scanned images; The scanned image is input into the gastric cancer HER2 state prediction model obtained by the training method of the gastric cancer HER2 state prediction model as described in any one of claims 1-5, and the state prediction result is obtained.

7. A training system for a gastric cancer HER2 status prediction model, characterized in that, The system is applied in the training method of the gastric cancer HER2 status prediction model as described in any one of claims 1-5, and the system comprises: The first module is used to acquire a sample image set, which includes a first stained image and a second stained image. The second module is used to perform state labeling processing on the second stained image in the sample image set to obtain image labels; The third module is used to perform pathological annotation processing on the sample image set to obtain a training sample set; The fourth module is used to input the training sample set and the image labels into the untrained gastric cancer HER2 state prediction model for training, so as to obtain the trained gastric cancer HER2 state prediction model.