HER2 state identification method and system based on IHC and HE bimodal image

Through the recognition method of IHC and HE dual-modality images, the ConvNext and Prov-GigaPath models are used for deep feature extraction, combined with the XGBoost and ABMIL models for feature aggregation, which solves the problem of insufficient information utilization in single-modality image analysis and achieves high accuracy and reliability in HER2 status interpretation.

CN120612693APending Publication Date: 2025-09-09金凤实验室
View PDF 0 Cites 3 Cited by

Patent Information

Application Number
CN202510725607.4
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-05-30
Publication Date
2025-09-09

AI Technical Summary

Technical Problem

The single-modality image analysis method in the existing technology cannot fully mine information related to HER2 status, the IHC image feature extraction capability is insufficient, and the HE image features have a low correlation with HER2, resulting in insufficient accuracy and precision in interpretation.

Method used

A recognition method based on IHC and HE dual-modal images was adopted. The ConvNext and Prov-GigaPath models were used for deep feature extraction. The XGBoost and ABMIL models were combined for feature aggregation. The dual-modal features were fused through the logistic regression model to output the HER2 status judgment results.

Benefits of technology

It significantly improves the accuracy and reliability of HER2 status interpretation, reduces data annotation costs and time, and enhances the generalization ability of the model and the credibility of interpretation.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120612693A_ABST
    Figure CN120612693A_ABST
Patent Text Reader

Abstract

The invention discloses an HER2 state recognition method and system based on an IHC and HE bimodal image, and relates to the medical image processing and analysis technology, and the method comprises the steps: carrying out the preprocessing of a full-slice image WSI containing IHC and HE, and cutting the WSI into designated image blocks; constructing a multi-task learning framework to classify image blocks obtained by cutting the IHC image, and determining HER2 expression intensity in the tumor region; calculating the area ratio of the image block corresponding to each grade in the tumor region to calculate an IHC HER2 score; taking each image block of the cut HE image as the input of a Prov-GigaPath model, and aggregating the multi-dimensional features of each image block by using an ABMIL model, and taking the aggregated multi-dimensional features as an HE HER2 score; and converting the IHC HER2 score and the HE HER2 score into one-hot codes, and outputting an HER2 judgment result by using a logistic regression model. The HER2 state interpretation result can be quickly output, a pathologist is assisted in diagnosis, and the diagnosis efficiency and accuracy are improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present application relates to the field of medical image processing and analysis technology, and in particular to a method and system for identifying HER2 status based on IHC and HE dual-modality images. Background Art

[0002] Existing HER2 status interpretation technologies mostly rely on single-modality image analysis techniques. For IHC images, some techniques employ traditional image segmentation algorithms, such as threshold segmentation and region growing, to extract positive cell regions. These techniques then use manually defined rules to calculate the proportion of positive cells and classify HER2 status. Other studies have used shallow convolutional neural networks for feature extraction and classification of IHC images, but these networks have simple structures and limited ability to extract complex image features. For HE images, HER2-related feature analysis is typically performed based on manually designed feature extraction methods combined with traditional machine learning classifiers. Alternatively, weakly supervised deep learning models are constructed to predict low HER2 expression from HE images. However, these models struggle to capture subtle information related to HER2 status within the images.

[0003] In the prior art:

[0004] 1. Insufficient utilization of single-modality image information: Single-modality IHC or HE images only contain partial information related to HER2 status. Existing single-modality analysis technologies cannot fully explore the key features in the image, affecting the accuracy of interpretation.

[0005] 2. Low correlation between HE image features and HER2: Under weak supervision conditions, existing methods find it difficult to effectively extract meaningful features from HE images, which affects the final HER2 status prediction accuracy.

[0006] 3. Inadequate IHC image feature extraction capabilities: Traditional methods such as threshold segmentation and region growing rely on manually set parameters and are poorly able to handle complex situations such as uneven staining and cell overlap, which can easily lead to misidentification or omission of positive cells. While shallow convolutional neural networks can automatically extract features, their simple structure lacks the ability to deeply mine subtle features such as cell morphology and staining intensity, resulting in low accuracy in distinguishing low HER2 expression. Summary of the Invention

[0007] The embodiments of the present application provide a HER2 status identification method and system based on IHC and HE dual-modality images. The method uses the patient's IHC and HE images as input, quickly outputs HER2 status interpretation results, assists pathologists in diagnosis, and improves diagnostic efficiency and accuracy.

[0008] The present application embodiment proposes a HER2 status identification method based on IHC and HE dual-modality images, comprising the following steps:

[0009] Preprocess the whole-slice image WSI containing IHC and HE, and cut the WSI into specified image blocks;

[0010] For the cut IHC image blocks, a multi-task learning framework is constructed. The ConvNext-based patch classification model is used to classify the image blocks obtained by cutting the IHC image. In addition, the tumor tissue / non-tumor tissue is simultaneously classified into two categories to discriminate and determine the HER2 expression intensity in the tumor region; and

[0011] Based on the classification results and HER2 expression intensity, the area ratio of the image blocks corresponding to each grade within the tumor region is calculated. The HER2 grade of the entire IHC image is determined using the XGBoost model as the IHC HER2 score.

[0012] The cut HE image blocks are input into the Prov-GigaPath model, and based on the output of the Prov-GigaPath model, the ABMIL model is used to aggregate the multidimensional features of each HE image block to obtain the HER2 grade of the entire HE image as the HEHER2 score;

[0013] The IHC HER2 score and HE HER2 score were converted into one-hot encoding and used as input to the logistic regression model to output the HER2 judgment results using the logistic regression model.

[0014] An embodiment of the present application further proposes a HER2 status identification system based on IHC and HE dual-modality images, characterized in that it includes a processor and a memory, wherein the memory stores a computer program, and when the computer program is executed by the processor, it implements the steps of the aforementioned HER2 status identification method based on IHC and HE dual-modality images.

[0015] This application fully utilizes the dual-modality information of IHC and HE images, using advanced deep learning models such as ConvNeXt networks and the Prov-Gigapath model to perform in-depth feature extraction on each image. This is combined with XGBoost and ABMIL models for precise analysis and feature aggregation. Finally, methods such as logistic regression are used to effectively fuse the dual-modality features. Compared to traditional single-modality or simple fusion methods, this method can more comprehensively and accurately capture information related to HER2 status, significantly improving the accuracy and reliability of interpretation.

[0016] The above description is only an overview of the technical solution of the present application. In order to more clearly understand the technical means of the present application, it can be implemented in accordance with the contents of the specification. In order to make the above and other purposes, features and advantages of the present application more obvious and easy to understand, the specific implementation methods of the present invention are listed below. BRIEF DESCRIPTION OF THE DRAWINGS

[0017] Various other advantages and benefits will become apparent to those skilled in the art upon reading the detailed description of the preferred embodiment below. The accompanying drawings are for illustration purposes only and are not to be considered as limiting the present invention. The same reference symbols are used throughout the drawings to represent the same components. In the drawings:

[0018] Figure 1 Schematic diagram of the process of the HER2 status identification method based on IHC and HE dual-modality images of this embodiment;

[0019] Figure 2 Schematic diagram of the architecture of the dyeing normalization network based on the PatchGAN model of this embodiment;

[0020] Figure 3 Schematic diagram of the process of the HER2 expression status prediction model based on IHC images in this embodiment;

[0021] Figure 4 FIG. 4 is a flowchart of the HER2 expression status prediction model based on HE images in this embodiment. DETAILED DESCRIPTION

[0022] Exemplary embodiments of the present disclosure will be described in more detail below with reference to the accompanying drawings. Although exemplary embodiments of the present disclosure are shown in the accompanying drawings, it should be understood that the present disclosure can be implemented in various forms and should not be limited by the embodiments set forth herein. Rather, these embodiments are provided to enable a more thorough understanding of the present disclosure and to fully convey the scope of the present disclosure to those skilled in the art.

[0023] For HER2 status interpretation, the embodiment of the present application constructs a set of IHC and HE dual-modal image analysis solutions. First, the IHC and HE full-slice images are preprocessed to unify the image quality; for IHC images, the ConvNext network is used to perform patch-level Her2 four-classification, and then the WSI level score is obtained by combining the patch statistical features with the XGBoost model; for HE images, the patch features are extracted with the help of the Prov-Gigapath model, and the WSI level score is obtained after aggregation by the ABMIL model. Finally, the scores of the two images are converted into one-hot encoding, and the logistic regression method is used for fusion analysis to output the final HER2 status interpretation result. Specifically, the embodiment of the present application proposes a HER2 status recognition method based on IHC and HE dual-modal images, such as Figure 1 As shown, the following steps are included:

[0024] In step S101, a whole-slide WSI image containing IHC and HE is preprocessed and segmented into designated image blocks. In some embodiments, preprocessing the whole-slide WSI containing IHC and HE includes segmenting the tissue region of the stained whole-slide image (WSI) using RGB thresholding and Canny edge detection techniques to detect and distinguish background and blurred areas. This process extracts all tumor and healthy tissue from the WSI, helps distinguish white background and blurred areas in the WSI image, and reduces the workload of manual annotation.

[0025] Cutting a WSI into specified image blocks involves cropping the WSI at a specified magnification into image blocks of a specified pixel size. For example, a WSI image at a 40x magnification can be cropped into image blocks of 256 × 256 pixels, with a resolution of 0.16 microns per pixel. This helps split large WSI images into smaller, more manageable blocks, improving computational efficiency and reducing resource consumption.

[0026] After cutting, the color normalization network based on the PatchGAN model is used to enhance the color of the image blocks. The color normalization network architecture based on the PatchGAN model is as follows: Figure 2 As shown in the figure, a U-Net-based network is used as the image generator, and a PatchGAN model is used as the discriminator of the color normalization network. Through color normalization, the image color is enhanced and the impact of the color on the model's generalization ability is reduced.

[0027] In step S102, a multi-task learning framework is constructed, and the image blocks obtained by cutting the IHC image are classified using the patch classification model based on ConvNext, and the tumor tissue / non-tumor tissue is synchronously classified into two categories to discriminate and determine the HER2 expression intensity in the tumor area. In a specific example, a multi-task learning architecture is constructed to achieve two-category discrimination of tumor tissue / non-tumor areas and four grades (0, 1+, 2+, 3+) of HER2 expression intensity in the tumor area, specifically following the standardized interpretation of the CSCO Breast Cancer Diagnosis and Treatment Guidelines (2024 Edition). In the embodiment of the present application, the multi-task learning framework also uses the ConvNext model to synchronously perform two-category discrimination of tumor tissue / non-tumor tissue on the image block (patch) and determine the HER2 expression intensity in the tumor area. In essence, the ConvNext model performs a six-classification result (HER2 expression intensity of 4 tumor tissues and 2 non-tumor tissue categories).

[0028] In step S103, based on the classification results and HER2 expression intensity, the area ratio of the image blocks corresponding to each grade within the tumor region is calculated, and the HER2 grade of the entire IHC image is determined by the XGBoost model as the IHC HER2 score.

[0029] In step S104, each image block of the cut HE image is used as the input of the Prov-GigaPath model, and based on the output of the Prov-GigaPath model, the ABMIL model is used to aggregate the multidimensional features of each image block to obtain the HER2 grade of the entire HE image as the HE HER2 score.

[0030] In step S105 , the IHC HER2 score and the HE HER2 score are converted into one-hot encoding and used as input to the logistic regression model to output the HER2 judgment result using the logistic regression model.

[0031] Traditional HER2 status interpretation methods rely on large amounts of finely labeled image data for model training. Manual labeling is not only costly and inefficient, but also prone to labeling errors. In IHC image analysis, the method of the present application uses the ConvNext network for patch-level classification, combined with the XGBoost model to perform WSI-level scoring based on statistical features, reducing the need for precise labeling of each cell or small area; in HE image analysis, the weakly supervised ABMIL model is used, which only requires WSI-level HER2 status labels, eliminating the need for detailed labeling of each patch. This approach significantly reduces the manpower and time costs of data labeling, while reducing the impact of labeling errors on model performance, which is conducive to the rapid construction of large-scale training data sets and accelerates model development and application.

[0032] The WSI scoring model based on XGBoost is constructed. In some embodiments, according to the classification results and the HER2 expression intensity, the area ratio of the image block corresponding to each grade in the tumor area is calculated. The HER2 grade of the entire IHC image is determined by the XGBoost model, including establishing a statistical modeling framework based on probability density estimation, integrating the spatial distribution information output in the first stage, such as Figure 3 As shown, specifically including:

[0033] Calculate the area ratio of each graded image block within the tumor area;

[0034] Construct a two-dimensional feature space and calculate the aggregation and heterogeneity index of strong positive (e.g., 3+) image blocks;

[0035] Based on the calculated aggregation and heterogeneity index, a score conversion is performed according to the preset quantitative decision tree to determine the HER2 expression intensity within the tumor area. For example, when the proportion of 3+ areas is ≥10%, it is judged to be 3+ (high expression). XGBoost is highly robust to high-dimensional features, nonlinear relationships, and data noise, and is particularly suitable for integrated modeling of multi-center heterogeneous data. Through regularization constraints and early stopping strategies (Early Stopping), the model can effectively capture common laws and local specificity across centers while ensuring generalization capabilities, providing a reliable basic model for subsequent external independent verification.

[0036] In some embodiments, as Figure 4 As shown, it also includes: using the Prov-GigaPath model, using the Vision Transformer (ViT) as a feature encoder to obtain a multi-dimensional feature vector corresponding to each image block, for example, using ViT as a feature encoder to obtain a 1536-dimensional feature vector corresponding to each image block.

[0037] The Prov-GigaPath model can be pre-trained on 1.3 billion 256×256 pathology image blocks of 171,189 WSIs based on the DINOv2 self-supervised learning method. It can automatically learn multi-level features such as cell morphology and tissue structure in pathology images. It has strong feature expression and generalization capabilities and can capture subtle features related to HER2 expression in breast cancer in pathology images.

[0038] HE image feature aggregation is based on a multi-instance approach. In this multi-instance approach, each patient's pathology image contains multiple image blocks, each of which can be considered an instance, and the entire image is a package. The 1536-dimensional features of each image block are aggregated using the ABMIL model to obtain the HER2 grade of the entire HE image. Specifically, in some embodiments, the ABMIL model is used to aggregate the multi-dimensional features of each image block to obtain the HER2 grade of the entire HE image, which includes:

[0039] Layer normalization LayerNorm is used to normalize the input multiple feature vectors. For example, the input 1536-dimensional feature vector is normalized, and then the elements of each feature vector are standardized to make the network training more stable and accelerate convergence.

[0040] Then, the Gaussian error linear unit GELU activation function is used to self-adjust the activation behavior according to the distribution of input data. The GELU function can adaptively adjust the activation behavior according to the distribution of input data, introduce nonlinear factors into the network, and enhance the expressive ability of the model.

[0041] Feature aggregation is performed through a gated attention mechanism consisting of one attention branch and three attention heads. The attention branch is used to calculate the attention weight (patch attention score) of each image block, focusing on image blocks that play a key role in HER2 expression in breast cancer and mining potential effective information. The gating mechanism controls the flow of information through learning, screens out more discriminative features, further enhances the model's ability to capture important features, and achieves effective aggregation of features from multiple image blocks.

[0042] In some embodiments, it also includes: using LabelSmoothingCrossEntropy as the loss function and using optim.AdamW as the optimizer to train the network model in the HE image step. For example, in the specific training process of the model, LabelSmoothingCrossEntropy is used as the loss function, where the smoothing parameter is set to 0.1. The traditional cross entropy loss function easily makes the model overconfident about the label during the training process. LabelSmoothingCrossEntropy suppresses the overfitting phenomenon of the model to a certain extent by smoothing the label. Use optim.AdamW as the optimizer and set the learning rate lr to 5e -4 , weight decay weight_decay is 1e -4 .

[0043] In some embodiments, outputting a HER2 determination result using a logistic regression model includes:

[0044] Logistic regression is used to learn the linear relationship between IHC and HE image features and the true HER2 status to output a HER2 judgment result, which includes HER2 expression status 0, 1+, 2+, and 3+. Specifically, logistic regression learns the linear relationship between IHC and HE image features and the true HER2 status. By learning and analyzing the weights of these features, the system comprehensively considers the information provided by both modalities and outputs the final HER2 status interpretation result, determining the patient's HER2 expression status (0, 1+, 2+, 3+), providing an accurate basis for clinical diagnosis.

[0045] During the results synthesis phase, the method of this application converts the HER2 scores of IHC and HE images into one-hot encoding and then uses logistic regression to perform a fusion analysis. This can, to a certain extent, reflect the contribution of the image features of the two modalities to the final interpretation. By analyzing the weight coefficients of the logistic regression model, it can be determined whether IHC and HE images play a key role in HER2 status judgment. In addition, the attention mechanism in the ABMIL model intuitively demonstrates the model's degree of attention to different regions when aggregating HE image patch features. These provide a visual and interpretable basis for the model's decision-making process, making it easier for users to understand the model's interpretation logic and enhance their trust in the interpretation results.

[0046] Compared with traditional single-modality or simple fusion methods, the method of the present application can capture information related to HER2 status more comprehensively and accurately, significantly improving the accuracy and reliability of the interpretation.

[0047] An embodiment of the present application further proposes a HER2 status identification system based on IHC and HE dual-modality images, comprising a processor and a memory, wherein the memory stores a computer program, and when the computer program is executed by the processor, the steps of the aforementioned HER2 status identification method based on IHC and HE dual-modality images are implemented.

[0048] Furthermore, although exemplary embodiments have been described herein, the scope includes any and all embodiments based on the present disclosure with equivalent elements, modifications, omissions, combinations (e.g., solutions that intersect various embodiments), adaptations, or changes. The present invention is not limited to the examples described in this specification or during the practice of this application, which examples are to be construed as non-exclusive.

[0049] The above description is intended to be illustrative rather than restrictive. For example, the above examples (or one or more of them) can be used in combination with each other. For example, those of ordinary skill in the art may use other embodiments when reading the above description.

[0050] The above embodiments are merely exemplary embodiments of the present disclosure. Those skilled in the art may make various modifications or equivalent substitutions to the present invention within the essence and protection scope of the present disclosure, and such modifications or equivalent substitutions should also be deemed to fall within the protection scope of the present invention.

Claims

1. A method for identifying HER2 status based on IHC and HE dual-modality images, characterized in that: The steps include: Preprocess the whole-slice image WSI containing IHC and HE, and cut the WSI into specified image blocks; For the cut IHC image blocks, a multi-task learning framework is constructed. The ConvNext-based patch classification model is used to classify the image blocks obtained by cutting the IHC image. In addition, the tumor tissue / non-tumor tissue is simultaneously classified into two categories to discriminate and determine the HER2 expression intensity in the tumor region; and Based on the classification results and HER2 expression intensity, the area ratio of the image blocks corresponding to each grade within the tumor region is calculated. The HER2 grade of the entire IHC image is determined using the XGBoost model as the IHC HER2 score. Each cut HE image block is input into the Prov-GigaPath model. Based on the output of the Prov-GigaPath model, the ABMIL model is used to aggregate the multidimensional features of each HE image block to obtain the HER2 grade of the entire HE image as the HE HER2 score. The IHC HER2 score and HE HER2 score were converted into one-hot encoding and used as input to the logistic regression model to output the HER2 judgment results using the logistic regression model.

2. The HER2 status recognition method based on IHC and HE dual-modality images according to claim 1, characterized in that: Preprocessing of whole-slice images WSI containing IHC and HE includes: The tissue area of ​​the stained WSI was segmented using RGB threshold processing and Canny edge detection method to detect and distinguish the background and blurred areas.

3. The HER2 status recognition method based on IHC and HE dual-modality images according to claim 1, wherein: Cutting the WSI into a specified image block includes: cutting the WSI into an image block with a specified pixel size at a specified magnification; and, After cutting, the color normalization network based on the PatchGAN model is used to perform color enhancement on the image blocks.

4. The HER2 status recognition method based on IHC and HE dual-modality images according to claim 1, wherein: For IHC images, based on the classification results and HER2 expression intensity, the area ratio of the image blocks corresponding to each grade within the tumor area is calculated. The HER2 grade of the entire IHC image is determined using the XGBoost model, including: Calculate the area ratio of each graded image block within the tumor area; Construct a two-dimensional feature space and calculate the aggregation and heterogeneity index of strong positive image blocks; Based on the calculated aggregation and heterogeneity indices, score transformation was performed according to a pre-defined quantitative decision tree to determine the HER2 grade within the tumor region.

5. The HER2 status recognition method based on IHC and HE dual-modality images according to claim 4, characterized in that: Also includes: For HE images, the Prov-GigaPath model is used, with the Vision Transformer (ViT) as the feature encoder, to obtain the multidimensional feature vector corresponding to each image block. The Prov-GigaPath model is pre-trained on preset pathological image blocks based on the DINOv2 self-supervised learning method.

6. The HER2 status recognition method based on IHC and HE dual-modality images according to claim 5, characterized in that: The ABMIL model is used to aggregate the multidimensional features of each HE image block to obtain the HER2 grade of the entire HE image, which includes: Use layer normalization LayerNorm to normalize the input multiple feature vectors, and use the Gaussian error linear unit GELU activation function to self-adjust the activation behavior according to the distribution of input data; Feature aggregation is performed through a gated attention mechanism consisting of 1 attention branch and 3 attention heads, where the attention branch is used to calculate the attention weight of each image block.

7. The HER2 status recognition method based on IHC and HE dual-modality images according to claim 6, characterized in that: Also includes: LabelSmoothingCrossEntropy was used as the loss function and optim.AdamW was used as the optimizer to train the ABMIL model.

8. The HER2 status recognition method based on IHC and HE dual-modality images according to claim 7, characterized in that: The HER2 judgment results output by the logistic regression model include: Logistic regression is used to learn the linear relationship between IHC and HE image features and the true HER2 status to output the HER2 judgment result.

9. The HER2 status recognition method based on IHC and HE dual-modality images according to claim 8, characterized in that: The output HER2 judgment results include HER2 expression status 0, 1+, 2+, and 3+.

10. A HER2 status recognition system based on IHC and HE dual-modality images, characterized in that: The invention comprises a processor and a memory, wherein a computer program is stored in the memory, and when the computer program is executed by the processor, the steps of the HER2 status identification method based on IHC and HE dual-modality images are implemented as described in any one of claims 1 to 9.

Citation Information

Cited By

  • Tumor HER2 expression grading method based on HE dyeing image

    CN120808885A

  • Method for grading tumor her2 expression based on he-stained images

    CN120808885B

  • Lung adenocarcinoma pathological image auxiliary diagnosis system and method based on adaptive tree framework

    CN120932868A