Tumor HER2 expression grading method based on HE dyeing image

Through multi-stage registration and multimodal semantic classification models, tumor HER2 expression grading based on HE-stained images was achieved, which solved the problems of insufficient image registration accuracy and HER2 grade discrimination stability in existing technologies, improved grading accuracy and reduced dependence on IHC staining technology.

CN120808885AActive Publication Date: 2025-10-17XUZHOU CENT HOSPITAL +1

Patent Information

Application Number
CN202511296404.4
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-09-11
Publication Date
2025-10-17
Estimated Expiration
2045-09-11

AI Technical Summary

Technical Problem

Existing technologies for grading tumor HER2 expression have problems such as insufficient accuracy in slice image registration, low quality of stained image generation, and poor stability in HER2 expression level discrimination, making it difficult to meet the needs of large-scale screening and automated analysis.

Method used

Through multi-stage registration to generate training samples, an IHC staining image generation model and a multimodal semantic classification model are constructed. By calculating the similarity of visual embedding features and semantic encoding features, end-to-end prediction from HE staining images to HER2 expression levels is achieved, reducing dependence on IHC staining technology.

Benefits of technology

It improves the accuracy and stability of HER2 expression grading, reduces medical costs, enhances the clinical interpretability of generated images and the generalization ability of the model, and supports modular deployment and end-to-end reasoning.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120808885A_ABST
    Figure CN120808885A_ABST
Patent Text Reader

Abstract

The invention relates to a tumor HER2 expression grading method based on an HE dyeing image. The method comprises the following steps: respectively loading each target HE image block in a target HE image block group into an IHC image generation model so as to generate a to-be-detected IHC image block in a structure alignment state with the current target HE image block by using the IHC image model; and loading each to-be-detected IHC image block to a multi-modal semantic classification model to perform expression grading processing by using the multi-modal semantic classification model, and generating HER2 expression grade prediction information of the current to-be-detected IHC image block after grading processing. According to the method, the tumor HER2 expression grading prediction can be realized based on the HE dyeing image, the dependence on an IHC dyeing technology is reduced, and meanwhile, the HER2 expression grading prediction precision, stability and clinical availability are improved.
Need to check novelty before this filing date? Find Prior Art

Description

TECHNICAL FIELD

[0001] The application relates to a tumor HER2 expression grading method, in particular to a tumor HER2 expression grading method based on HE staining images. BACKGROUND

[0002] Molecular typing of tumors is of great significance in clinical treatment, and HER2 protein expression level (0, 1+, 2+, 3+) is an important indicator for determining whether targeted therapy is suitable for some tumor types. In traditional HER2 typing, immunohistochemistry (IHC) staining sections are relied on, and are mainly observed and interpreted by professional pathologists under a microscope. However, the cost of obtaining HER2 expression state by IHC staining technology is high, the production cycle is long, and the quality of the stained section and the operation process are required to be high, which limits its application in large-scale screening and automatic analysis.

[0003] In contrast, hematoxylin and eosin (HE) staining images, as a standard tissue staining technique, have low acquisition cost and are widely available in medical institutions at all levels, and have higher practicality and scalability. Therefore, how to indirectly predict the HER2 expression level through HE staining images has become a research focus. However, when inferring the HER2 expression level from the HE staining image, there are still several key technical challenges, specifically: First, when aligning the HE staining image and the IHC staining image in space, the existing methods mostly adopt a single-scale registration strategy, which is difficult to achieve high-precision registration at both the whole image level and the block level. Especially in the case of tissue structure deformation, section thickness difference or image occlusion between the HE staining image and the IHC staining image, only using coarse registration or single block registration often cannot meet the requirements of position consistency for subsequent image generation and classification tasks.

[0004] Second, although some existing methods attempt to generate HE staining images to IHC staining images through image generation models, these methods generally do not construct a complete task chain of HE→IHC→grading, lack practical usability and inter-module coordination mechanism. In addition, during the image generation process, due to the high resolution and rich structure details of pathological images, the existing generative adversarial network often has the problem of mode collapse, which is manifested as repeated structure, missing details or abnormal color distribution of the output image, resulting in distortion of the staining features, and seriously affecting the clinical interpretability and reliability of the generated image. In the existing literature, there are few systematic modeling of membrane staining area structure consistency, sample diversity control and high-frequency detail restoration, which restricts the application performance of the image generation model in the medical image field.

[0005] When grading HER2 expression levels, manual annotation is required. Specifically, manual annotation has the characteristics of strong subjectivity and blurred inter-class boundaries, especially in the 1+ and 2+, 2+ and 3+ between the misjudgment, leading to the generalization performance of traditional CNN or shallow classification model in HER2 expression grading is insufficient.

[0006] The multi-modal visual language model that appeared in recent years has strong semantic modeling ability, but there is currently a lack of complete mechanism to effectively introduce it into the HER2 expression grading task, especially in the context that the HER2 level semantics is not covered by the pre-training model, direct application often leads to text supervision failure. Existing research usually ignores the introduction of domain knowledge (such as HER2 interpretation standards) into the language encoding process, or does not design a prototype guiding mechanism with semantic migration ability, making it difficult to accurately model the fine-grained level differences in tumor pathology images.

[0007] The application with publication number CN114820555A discloses a breast cancer pathological image classification method based on SENet channel attention and transfer learning. The method first uses wavelet transform combined with transfer learning to extract features from breast cancer pathological images, then uses SENet channel attention mechanism to fuse the features, and finally constructs a classifier to realize effective classification of pathological images. The application enhances the feature expression ability to solve the problems of feature loss and overfitting of traditional CNN models, and improves the accuracy of breast cancer benign and malignant classification. However, the application can only classify single-stained images and does not consider the registration and conversion between different stained sections, nor does it involve the prediction task of HER2 expression levels, making it difficult to meet the needs of molecular-level auxiliary diagnosis.

[0008] The application with publication number CN117765252A discloses a breast cancer identification system and method based on Swin Transformer and contrastive learning. The steps include: first, acquiring breast cancer X-ray image data; then, cropping the original image into a 224x224 pixel three-channel image to form an input image sample; then, building a Swin Transformer neural network structure and introducing a SwinCLR framework to pre-train the model in a SimCLR unsupervised manner using unlabeled images; then, using labeled images to supervise the training of the model; finally, outputting the corresponding breast cancer probability of the image.

[0009] As can be seen from the above description, the application with publication number CN117765252A combines the advantages of contrast learning and SwinTransformer, and can improve model performance and enhance the accuracy and training efficiency of breast cancer recognition in the absence of labeled data. However, the application mainly performs overall recognition analysis on breast cancer X-ray images, and its task positioning is the presence detection of diseases (i.e., whether cancer is present), without involving HE staining images and IHC staining image analysis at the tissue slice level, nor does it include the HER2 expression grading task. The method also does not consider the staining conversion or registration problem between images, making it difficult to apply to the scene of molecular pathology assisted diagnosis based on HE image inference of HER2 expression.

[0010] The application with publication number CN119006942A discloses a recognition and classification method and system for breast cancer pathological images, including the following steps: collecting breast cancer pathological images and performing preprocessing; extracting multi-scale texture features of the images based on a convolutional neural network, and capturing local and global texture information in breast tissue by designing convolution kernels with different receptive field sizes; then, performing adaptive weighting processing on the extracted texture features using an attention mechanism, dynamically adjusting feature weights according to the importance of different regions, and extracting key texture features; subsequently, filtering the key texture features using a sequential feature selection algorithm, and obtaining an optimal feature subset through iterative search and cross-validation; finally, constructing an ensemble learning classification model based on the selected features, and performing recognition and classification on the images to be predicted.

[0011] As can be seen from the above description, the application with publication number CN119006942A has certain effect in improving the classification accuracy and explainability of breast cancer pathological images. However, the application only performs image-level classification on the texture structure in breast cancer pathological images, and does not involve image registration between different staining methods, staining image generation, HER2 expression grading, etc., making it difficult to be used in the expression prediction scene of IHC staining, and also lacking the ability of image generation and multi-modal semantic reasoning, with relatively limited application range.

[0012] In summary, in the prior art, there are obvious deficiencies in the accuracy of slice image registration, the quality of staining image generation, and the stability of HER2 expression level discrimination, and there is an urgent need for a method covering image preprocessing, generation, and classification to meet the application requirements in medical image scenarios. SUMMARY

[0013] The purpose of the present application is to overcome the deficiencies in the prior art, and to provide a tumor HER2 expression grading method based on HE staining images, which can realize tumor HER2 expression grading prediction based on HE staining images, reduce the dependence on IHC staining images, and improve the accuracy, stability, and clinical usability of HER2 expression grading prediction.

[0014] According to the technical scheme provided by the application, a tumor HER2 expression grading method based on HE staining images is provided, and the method comprises the following steps: A reference HE staining image is provided, and at least segmentation screening is performed on the reference HE staining image to generate a target HE image block group after segmentation screening, wherein each target HE image block in the target HE image block group contains tumor tissue information; Each target HE image block in the target HE image block group is loaded into an IHC staining image generation model respectively, so as to generate a to-be-inspected IHC image block in a structural alignment state with the current target HE image block by using the IHC staining image generation model; Each to-be-inspected IHC image block is loaded into a multi-modal semantic classification model, so as to perform expression grading processing by using the multi-modal semantic classification model, and HER2 expression level prediction information of the current to-be-inspected IHC image block is generated after the grading processing, wherein When performing expression grading processing, the multi-modal semantic classification model extracts visual embedding features of the to-be-inspected IHC image block, and calculates and determines target HER2 level information in an optimal approximation state with the visual embedding features; The target HER2 level information is configured as the expression level prediction information of the current to-be-inspected IHC image block.

[0015] When calculating and determining the target HER2 level information in an optimal approximation state with the visual embedding features, the following steps are included: The feature similarity corresponding to each HER2 level semantic encoding feature in the reference semantic encoding feature group is calculated; For any visual embedding feature, when the calculated feature similarity is in a maximum state, the visual embedding feature and the corresponding HER2 level semantic encoding feature are in an optimal approximation state, and at this time, the corresponding HER2 level semantic encoding feature is configured as the optimal HER2 level semantic encoding feature; The HER2 level information in the optimal HER2 level semantic encoding feature is configured as the target HER2 level information.

[0016] The multi-modal semantic classification model at least includes a visual encoder module, a language encoder module, and a text-image alignment inference module, wherein The visual embedding features of the to-be-inspected IHC image block are extracted by using the visual encoder module; The reference semantic encoding feature set is generated by using the language encoder module, wherein when the reference semantic encoding feature set is generated, the HER2 level guide source information set is loaded into the language encoder module to semantically encode each HER2 level guide source information in the HER2 level guide source information set by using the language encoder module, and a corresponding HER2 level semantic encoding feature is generated after language encoding, and the HER2 level semantic encoding features in the reference semantic encoding feature set are independent of each other; The feature similarity between the visual embedding feature and each HER2 level semantic encoding feature is calculated by using the image-text alignment reasoning module, and the target HER2 level information is calculated and determined based on the calculated feature similarity, and the expression level prediction information of the current IHC image block is output.

[0017] The multi-modal semantic classification model further comprises a category prototype matching module, wherein, When the feature similarity calculated by the image-text alignment reasoning module is in a fuzzy approximate state, the category prototype matching module is configured to perform prototype similarity calculation processing to determine the target HER2 level information through the prototype similarity calculation processing; When performing the prototype similarity calculation processing, the prototype similarity between the visual embedding feature and the category prototype feature in the category prototype matching module is calculated respectively, and the optimal prototype similarity and the target category prototype feature corresponding to the optimal prototype similarity are determined, The target HER2 level information is generated based on the target category prototype feature, The target HER2 level information is configured as the expression level prediction information of the current IHC image block, and the category prototype matching module is configured to output the expression level prediction information of the current IHC image block.

[0018] The feature similarity calculated by the image-text alignment reasoning module comprises a cosine similarity, wherein, When the feature similarity adopts the cosine similarity, then:

[0019] wherein, is the feature similarity, is the visual embedding feature, is the HER2 level semantic encoding feature of the category is the vector dot product of the visual embedding feature and the HER2 level semantic encoding feature of the category is the vector module product of the visual embedding feature and the HER2 level semantic encoding feature of the category

[0020] ​​​​​When building a multimodal semantic classification model, it includes: Select a pre-trained framework model and configure the selected pre-trained framework model as a multimodal semantic classification basic model, wherein the multimodal semantic classification basic model includes a visual encoder module, a language encoder module, a prompt enhancement module, an image-text alignment reasoning module, and a category prototype matching basic module; Constructing a classification model training dataset to perform model training on a multimodal semantic classification basic model using the classification model training dataset, wherein the classification model training dataset includes a plurality of classification training samples, each classification training sample includes a training IHC image block and training expression level text information corresponding to the current training IHC image block; A classification model training condition is configured for training the multimodal semantic classification base model until the multimodal semantic classification base model is trained to reach a classification model target state, and thereafter, a multimodal semantic classification model is generated based on the multimodal semantic classification base model trained to reach the classification model target state, wherein: During model training, the corresponding weight parameters of the visual encoder module and the language encoder module are frozen, the prompt enhancement module fine-tunes the corresponding weight parameters through the back-propagation mechanism, and the category prototype matching basic module updates the corresponding category prototype features in the training of each batch of classification training samples through the momentum update mechanism; During model training, the prompt enhancement module is used to concatenate a learnable training vector to each HER2 grade description text in the HER2 grade description text group to generate HER2 grade training source information corresponding to each HER2 grade description text. Subsequently, the language encoder module is used to semantically encode the HER2 grade training source information to generate HER2 grade training encoding features; The visual encoder module is used to extract the training visual features of the training IHC image blocks in each classification training sample. After that, the image-text alignment inference module is configured to calculate the feature similarity between the training visual features and the corresponding feature encoding features of all HER2 grade training. The category prototype matching basic module is used to calculate and update the category prototype features corresponding to each training expression grade information.

[0021] When calculating the prototype features of each training expression level information category, we have:

[0022] in, Classify the categories within the training samples for each batch The number of classification training samples, For each annotation classification training sample The training visual features of the classification training samples, For each batch of classification training samples a training expression level text information of a classification training sample, for a class a class prototype feature corresponding to the training sample; for each class updating the class prototype feature by a momentum updating mechanism then:

[0023] wherein: is an updating momentum coefficient.

[0024] The IHC staining image generation model at least includes an image generator, loads a target HE image block into the IHC staining image generator, and generates a to-be-inspected IHC image block by the IHC staining image generator, wherein, The IHC staining image generation model is generated based on an IHC staining image generation base model after model training, wherein the IHC staining image generation base model includes an image generator and an image discriminator, and the IHC staining image generation base model is model trained by using a generated model training data set; The generated model training data set includes a plurality of generated training samples, each of which includes a training HE image block and a reference IHC image block corresponding to the training HE image block, wherein the training HE image block and the reference IHC image block in the same generated training sample are formed based on the same tumor tissue section, and the training HE image block and the reference IHC image block both contain tumor tissue information, and the training HE image block and the reference IHC image block in the same training sample are at least in a tissue structure alignment state; The IHC staining image generation base model is configured to meet the generated model training conditions required for model training until the training effect of the IHC staining image generation base model reaches a generated model target state, and thereafter, the IHC staining image generation model is generated based on the IHC staining image generation base model whose training effect reaches the generated model target state.

[0025] When the generated model training data set is generated, the generated training samples in the generated model training data set are generated respectively, wherein, When the generated training sample is generated, it includes: obtaining an HE staining section source image and a paired IHC staining section source image derived from the same tumor tissue section; Multi-stage registration is performed on the HE staining section source image and the IHC staining section source image to generate an HE staining registration image and an IHC staining registration image after multi-stage registration, wherein, When performing multi-stage registration, at least comprising sequentially performed full-map level coarse registration, meso-region level non-rigid registration and tile level fine registration, wherein, After performing the full-map level coarse registration, the HE staining slice source map and the IHC staining slice source map are in a preliminary alignment state; After performing the meso-region non-rigid registration, local offset of the tissue structure is corrected; After performing the tile level fine registration, the IHC staining registration map and the HE staining registration map are respectively generated, and the IHC staining registration map and the HE staining registration map correspond in pixel scale in cell boundary, nucleus structure and membrane dyeing region; The HE staining registration map and the IHC staining registration map are subjected to segmentation screening processing, so as to form the training HE image block and the reference IHC image block corresponding to the training HE image block in the training sample after segmentation screening, and the training HE image block and the reference IHC image block are at least in a tissue structure level alignment state.

[0026] When the HE staining registration map and the IHC staining registration map are subjected to segmentation screening processing, the following steps are included: The HE training candidate region is sequentially selected on the HE staining registration map, and the tumor region recognition is performed on the HE training candidate region, so as to determine the binary mask of each pixel in the HE training candidate region after the tumor region recognition, wherein when the binary mask is 1, it represents that the current pixel belongs to the tumor region, and when the binary mask is 0, it represents that the current pixel does not belong to the tumor region; Based on the binary mask of each pixel in the HE training candidate region, the tumor region proportion is calculated, wherein when the calculated tumor region proportion matches the region selection threshold, the current HE training candidate region is configured as the training HE image block; Based on the position state of the training HE image block, the corresponding reference IHC image block is segmented on the IHC staining registration map.

[0027] Advantages of the present application: when constructing the generation training sample of the training IHC staining image generation basic model, the HE staining slice source image and the IHC staining slice source image are subjected to multi-stage registration, so that after multi-stage registration, the IHC staining registration image and the HE staining registration image achieving key histological morphology spatial consistency at the WSI (whole slide imaging) level can be generated, and thereafter, the training HE image block and the reference IHC image block in the corresponding generation training sample can be cut based on the IHC staining registration image and the HE staining registration image. After the IHC staining image generation basic model is trained by using the training sample thus made and the IHC staining image generation model is obtained, the to-be-inspected IHC image block generated by using the IHC staining generation model can be in structural alignment with the target HE image block, which provides favorable support for subsequent expression grading processing by using the multi-modal semantic classification model, and improves the accuracy and reliability of the obtained expression level prediction information.

[0028] When generating the expression level prediction information, the IHC staining image block obtained by scanning is not directly used, thereby reducing the dependence on the IHC staining technology, significantly reducing the use frequency of the IHC staining technology, and reducing the medical cost.

[0029] By introducing the double-discriminator structure of the local discriminator-global discriminator into the IHC staining image generation basic model, and further introducing the frequency domain structure preservation loss and the sample feature diversity regular term loss into the generation loss function, the mode collapse problem in medical image generation can be effectively alleviated, the restoration ability of the membrane staining area details can be enhanced, and the fidelity and the consistency of the clinical interpretability of the generated to-be-inspected IHC image block can be improved.

[0030] When the feature similarity calculated by the image-text alignment inference module is in a fuzzy approximate state, the category prototype matching module is configured to perform prototype similarity calculation processing, so as to determine the target HER2 level information through the prototype similarity calculation processing, thereby improving the discrimination ability and the generalization ability of the model in the case that the HER2 expression level has a fuzzy boundary, and the model is particularly suitable for discriminating the 2+ and 1+ / 3+ boundary samples; A full-process system from preparing the generation training sample through multi-stage registration, training the IHC staining image generation model to the multi-modal semantic classification model is constructed, the modular deployment and the end-to-end inference are supported, whether to use the generation path or directly perform grading based on the reference HE staining can be selected according to the deployment scene, and the convenience of deployment is improved. BRIEF DESCRIPTION OF DRAWINGS

[0031] Figure 1 An embodiment flowchart of the tumor HER2 expression grading method of the present application.

[0032] Figure 2An embodiment schematic diagram for not performing multi-stage registration on the HE staining slice source image and the IHC staining slice source image.

[0033] Figure 3 An embodiment schematic diagram for generating the HE staining registration image and the IHC staining registration image after multi-stage registration of the present application.

[0034] Figure 4 An embodiment schematic diagram for generating the training HE image block and the reference IHC image block after the segmentation screening processing of the present application.

[0035] Figure 5 An embodiment schematic diagram for generating the base model of the HER2 IHC staining image of the present application.

[0036] Figure 6 An embodiment structure block diagram of the global discriminator of the present application.

[0037] Figure 7 An embodiment structure block diagram of the local discriminator of the present application.

[0038] Figure 8 An embodiment structure block diagram of the multi-modal semantic classification model of the present application. DETAILED DESCRIPTION

[0039] The present application will be further described below in combination with specific drawings and embodiments.

[0040] In order to reduce the dependence on IHC staining technology, based on the HE staining image, the present application provides a tumor HER2 expression grading method based on the HE staining image, specifically, the tumor HER2 expression grading method comprises: A reference HE staining image is provided, and at least segmentation screening processing is performed on the reference HE staining image to generate a target HE image block group after segmentation screening, wherein the target HE image block in the target HE image block group contains tumor tissue information; Each target HE image block in the target HE image block group is loaded into an IHC staining image generation model respectively, so as to generate a to-be-inspected IHC image block in a structural alignment state with the current target HE image block by using the IHC staining image generation model; Each to-be-inspected IHC image block is loaded into a multi-modal semantic classification model, so as to perform expression grading processing by using the multi-modal semantic classification model, and HER2 expression level prediction information of the current to-be-inspected IHC image block is generated after grading processing, wherein When performing expression grading processing, the multi-modal semantic classification model extracts visual embedding features of the to-be-inspected IHC image block, and calculates and determines target HER2 level information in an optimal approximation state with the visual embedding features. The target HER2 level information is configured as the expression level prediction information of the current IHC image block to be tested.

[0041] It should be noted that the present invention implements tumor HER2 expression grading based on HE staining images. Specifically, the HE staining image alone can be used to predict the HER2 expression level of a tumor. This means that the prediction of the HER2 expression level of a tumor is independent of the tumor IHC staining image or the tumor IHC staining technology. Specifically, the HE staining image provided should be a HE staining image of the tumor tissue to be graded for HER2 expression. Currently, tumor types for which HER2 expression testing is clinically significant include breast cancer, cervical cancer, and gastric cancer.

[0042] Depend on Figure 1 It can be seen that when executing the tumor HER2 expression grading method of the present invention, a reference HE staining image should be provided, wherein the reference HE staining image should be the HE staining image corresponding to the tumor HER2 expression level to be predicted. For example, for a tumor tissue sample, when it is necessary to determine the tumor HER2 expression level, only the HE staining image of the tumor tissue sample can be provided, that is, the obtained HE staining image is used as the reference HE staining image, and the reference HE staining image can be used to ultimately predict the tumor HER2 expression level.

[0043] It should be understood that the reference HE staining image can be provided by commonly used technical means in the art, such as obtaining it through a full-view pathology slide scanning system and generating a digital image file with a pyramid structure. The resolution of the reference HE staining image is preferably 100,000×80,000 pixels, and the image format is .tiff or other formats that support multi-resolution access. In order to predict the tumor HER2 expression level, the Figure 1 It can be seen that the reference HE-stained image should at least be segmented and screened to generate a target HE image block group after the segmentation and screening process. The target HE image block group includes several target HE image blocks, and each target HE image block contains tumor tissue information. That is, only when the target HE image block contains tumor tissue information is it necessary to predict the tumor HER2 expression level.

[0044] Specifically, the target HE image blocks contain tumor tissue information, specifically meaning that tumor tissue is present in all target HE-stained images. The tumor tissue present is of the aforementioned tumor type, such as breast cancer, cervical cancer, or gastric cancer. In the following description, "containing tumor tissue information" has the same meaning, and reference is made to the description herein. The method for segmenting and filtering the reference HE-stained image can be referred to in the corresponding description below. Furthermore, all target HE image blocks have the same scale.

[0045] Depend onFigure 1 It can be known that after the target HE image block group is generated, each target HE image block in the target HE image block should be loaded into the IHC staining image generation model respectively, wherein each target HE image block loaded into the IHC image block generation model can generate a corresponding IHC image block to be detected through the IHC staining image generation model, and the generated IHC image block to be detected should be in structural alignment with the target HE image currently loaded into the IHC image block generation model. Therefore, when the target HE image block group includes multiple target HE image blocks, multiple corresponding IHC image blocks to be detected can be generated respectively after being loaded into the IHC image block generation model. The IHC image blocks to be detected should be consistent with the number of target HE image blocks loaded into the IHC image block generation model and one-to-one correspondence. Generally, the size of the IHC image block to be detected is consistent with the size of the target HE image block.

[0046] It should be noted that the tumor HER2 expression grading of the present application specifically refers to determining the HER2 expression level of each IHC image block to be detected. Therefore, after the IHC image block to be detected is generated, the IHC image block to be detected should be loaded into the multi-modal semantic classification model for HER2 expression grading prediction, so as to utilize the multi-modal semantic classification model to perform expression grading processing on the current IHC image block to be detected, and the expression level prediction information of the current IHC image block to be detected can be generated after the expression grading processing. The generated expression level prediction information is the tumor HER2 expression grading prediction result of the IHC image block to be detected, thereby realizing the prediction of the tumor HER2 expression level. Since the IHC image block obtained by scanning is not directly utilized in the tumor HER2 expression grading, the dependence on the IHC staining image / IHC staining technology is reduced.

[0047] In an embodiment of the present application, when performing expression grading processing, the visual embedding features of the IHC image block to be detected should be extracted, and thereafter, the target HER2 level information in the optimal approximation state with the visual embedding features is calculated and determined. When the target HER2 level information is determined, the determined target HER2 level information can be configured as the expression level prediction information of the current IHC image block to be detected.

[0048] As can be seen from the above description, for each target HE image block, the IHC staining image generation model can generate a corresponding to-be-detected IHC image block, and then the multi-modal semantic classification model is used to perform expression grading processing on a to-be-detected HER2 image block, so that the expression level prediction information of the to-be-detected IHC image block can be determined. It should be understood that after the expression level prediction information of each to-be-detected IHC image block is determined, that is, the expression level prediction information corresponding to each target HE image block is determined. When the expression level prediction information corresponding to all target HE image blocks is determined, part / whole of the expression level prediction information can be used to provide medical diagnosis reference for clinical pathologists.

[0049] In an embodiment of the present application, when the target HER2 level information in the optimal approximation state with the visual embedding feature is calculated and determined, the following steps are included: The feature similarity corresponding to each HER2 level semantic encoding feature in the reference semantic encoding feature group is calculated. For any visual embedding feature, when the calculated feature similarity is in the maximum state, the visual embedding feature is in the optimal approximation state with the corresponding HER2 level semantic encoding feature, and at this time, the corresponding HER2 level semantic encoding feature is configured as the optimal HER2 level semantic encoding feature. The HER2 level information in the optimal HER2 level semantic encoding feature is configured as the target HER2 level information.

[0050] It should be noted that when the target HER2 level information in the optimal approximation state with the visual embedding feature is determined, the reference semantic encoding feature group should also be generated in the multi-modal semantic classification model. The way to generate the reference semantic encoding feature group will be described below. Specifically, the reference semantic encoding feature group includes a plurality of HER2 level semantic encoding features, and each HER2 level encoding semantic feature is in a mutually independent state with the remaining HER2 level encoding semantic features.

[0051] After the reference semantic encoding feature group is generated, the multi-modal semantic classification model should be configured to calculate the feature similarity between the visual embedding feature and each HER2 level semantic encoding feature. Specifically, after the feature similarity between the visual embedding feature and each HER2 level semantic encoding feature is calculated, the maximum feature similarity can be determined, at which time the visual embedding feature is in the optimal approximation state with the corresponding HER2 level semantic encoding feature. Here, the corresponding HER2 level semantic encoding feature is the HER2 level semantic encoding feature for which the maximum feature similarity is calculated, and after that, the corresponding HER2 level semantic encoding feature can be configured as the optimal HER2 level semantic encoding feature.

[0052] After determining the optimal HER2 level semantic coding feature, the HER2 level information within the optimal HER2 level semantic coding feature can be extracted, and the extracted and determined HER2 level information can be configured as the target HER2 level information, thereby achieving the determination of the target HER2 level information.

[0053] In a specific implementation, the feature similarity may be cosine similarity. When the feature similarity is cosine similarity, the method for calculating the feature similarity may be:

[0054] in, is the feature similarity, is the visual embedding feature, For category HER2-level semantic encoding features, The visual embedding features and categories are HER2-level semantic encoding features The vector dot product of The visual embedding features and categories are HER2-level semantic encoding features Vector modular multiplication of . In general, , HER2 level semantic encoding features have the same dimension as visual embedding features, for categories For the specific meaning of HER2, please refer to the following description. For the semantic coding features of HER2 level, please refer to the following corresponding description.

[0055] In one embodiment of the present invention, the IHC staining image generation model includes at least an image generator, and the target HE image block is loaded into the IHC staining image generator so that the IHC staining image generator generates an IHC image block to be tested.

[0056] Specifically, the image generator can adopt an existing commonly used form. For example, the image generator can be composed of an encoder network and a decoder network, wherein the encoder network can adopt a residual network structure (such as ResNet-101). Within the image generator, the target HE image block is feature extracted through the encoder network. The decoder network should generally be structurally symmetrical with the encoder network, so that the encoded features output by the encoder network are upsampled layer by layer and reconstructed into IHC image blocks through the decoder network, and skip connections are used to retain the low-level spatial information in the encoder network. Of course, the image generator can also adopt other forms, which shall be based on the ability to generate the required HER2 image blocks to be tested.

[0057] In an embodiment of the present application, the IHC staining image generation model is generated based on the IHC staining image generation base model after model training, wherein the IHC staining image generation base model comprises an image generator and an image discriminator, and the IHC staining image generation base model is trained by using the generated generation model training data set. The generation model training data set comprises a plurality of generation training samples, each generation training sample comprising a training HE image block and a reference IHC image block corresponding to the training HE image block, wherein the training HE image block and the reference IHC image block in the same generation training sample are formed based on the same tumor tissue section, and the training HE image block and the reference IHC image block both contain tumor tissue information, and the training HE image block and the reference IHC image block in the same training sample are at least in a state of alignment at the tissue structure level. The generation model training conditions for training the IHC staining image generation base model are configured, and the IHC staining image generation base model is trained until the generation model target state is reached, and thereafter, the IHC staining image generation model is generated based on the IHC staining image generation base model trained to reach the generation model target state.

[0058] It should be understood that when constructing the above-mentioned IHC staining image generation model, the IHC staining image generation base model should be constructed first, and the corresponding IHC staining image generation model can be generated after model training of the IHC staining image generation base model, Figure 5 An embodiment of the constructed IHC staining image generation base model is shown in FIG. 1, and it can be seen from the figure that compared with the training generated IHC staining image generation model, the image discriminator should also be provided in the IHC staining image generation base model, which is used to judge the authenticity of the generated training IHC image block, that is, the image discriminator is used to participate in the model training of the IHC staining image generation base model, and the image generator in the IHC staining image generation base model should adopt the same structure form as the image generator in the IHC staining image generation model, and the case of the image generator in the IHC staining image generation base model can be referred to the above description.

[0059] In order to train the IHC staining image generation model, a generation model training data set should be prepared. Generally, the generation model training data set should include a plurality of generation training samples, and each generation training sample includes a training HE image block and a reference IHC image block. In order to ensure that the image generator generates an IHC image block that is structurally aligned with the target HE image block, in actual implementation, the training HE image block and the reference IHC image block in the same generation training sample should come from the same tumor tissue section, and the training HE image block and the reference IHC image block are at least in a tissue structure alignment state. The tumor tissue section can be obtained from the cancerous tissue mentioned above, such as breast cancer, gastric cancer, or cervical cancer. In addition, in order to ensure that the generated IHC image block contains tumor tissue information, the training HE image block and the reference IHC image block should both contain tumor tissue information.

[0060] When training the IHC staining image generation model using the generation model training data set, the training HE image block in each generation training sample is loaded into the image generator in the IHC staining image generation model to generate a corresponding pseudo IHC image block. The reference IHC image block can be used to compare the generated pseudo IHC image block, that is, the reference IHC image block is mainly used to calculate the generation model loss function.

[0061] It should be understood that when training the IHC staining image generation model, the generation model training conditions should be configured. The configured generation model training conditions generally include the generation model loss function. In addition, the configured generation model training conditions can also include: using the Adam optimizer, setting the initial learning rate to 2e-4, the batch size to 16, and the training rounds to 200. Specifically, the meanings of the training rounds and the batch size are consistent with the prior art. After 200 rounds of model training of the IHC staining image generation model using the generation model training data set, the model training of the IHC staining image generation model can be terminated. At this time, the IHC staining image generation model can be generated based on the IHC staining image generation model trained for 200 rounds. That is, when the training reaches 200 rounds, the model training of the IHC staining image generation model reaches the generation model target state. At this time, the image discriminator in the IHC staining image generation model trained for 200 rounds is stopped working, and only the image generator is configured to be in a working state, thereby generating the IHC staining image generation model.

[0062] In addition, other ways can be used to determine whether the IHC staining image generation base model is trained to reach the generation model target state, such as after the model training of the IHC staining image generation base model reaches 200 rounds, verification is performed on the verification set, and when the loss function of the generation model verified on the verification set is less than the generation training loss threshold and tends to be stable, it is considered that the corresponding IHC staining image generation base model reaches the generation model target state, and other cases are not listed here; generally, the size of the generation training loss threshold can be selected according to actual needs.

[0063] In order to solve the common mode collapse problem (Mode Collapse) in the process of generating the to-be-inspected IHC image block from the target HE image block, in an embodiment of the present application, the image discriminator includes a local discriminator and a global discriminator, wherein, The global discriminator is based on a discriminant architecture of spatial pyramid pooling and channel-spatial self-attention mechanism, so as to configure the global discriminator to discriminate the authenticity of the image organization structure; The local discriminator is based on an edge-guided lightweight residual mechanism discriminant architecture, so as to configure the local discriminator to perform fine-grained discrimination on the membrane staining area and the cell nucleus edge.

[0064] In specific implementation, the global discriminator and the local discriminator are independent of each other, and based on the discriminant structures of the global discriminator and the local discriminator, the global discriminator and the local discriminator can form a significant difference in feature extraction mechanism and receptive field processing, forming a complementary mechanism of macro-topological verification and microscopic pathological feature discrimination.

[0065] Figure 6 An embodiment of the global discriminator is shown in FIG. 6, Figure 6 The spatial pyramid pooling adopts a four-level pyramid pooling layer, and specifically includes a global first convolution layer, a global second convolution layer, a global third convolution layer, and a global average layer, wherein the global first convolution layer can adopt 1x1 convolution, the global second convolution layer can adopt 3x3 convolution, and the global third convolution layer can adopt 6x6 convolution; During model training, the pseudo IHC image block generated by the image generator is respectively subjected to convolution processing by the global first convolution layer, the global second convolution layer, and the global third convolution layer, and at the same time, is subjected to average processing by the global average difference, after which the features after the convolution processing and the average processing are subjected to feature splicing by a feature splicing layer, the spliced features are subjected to gated reweighting by the channel-spatial self-attention mechanism, and finally the global discriminant output is obtained through the full connection layer.

[0066] Specifically, the global average processing, feature splicing, channel-space self-attention mechanism and full connection layer can adopt the existing common form, the multi-scale tissue topological features of the pseudo IHC image block can be extracted in parallel through the spatial pyramid pooling (SPP), and the deep-level dependency relationship of the tissue structure can be learned through the channel-space self-attention mechanism.

[0067] Figure 7 An embodiment of the local discriminator is shown in the middle, which adopts an edge-guided lightweight residual architecture, and innovatively fuses the traditional edge detection operator with deep feature extraction, Figure 7 In the middle, the local discriminator includes a Sobel convolution layer, a Res Block1 unit and a Res Block2 unit, wherein the Sobel convolution layer extracts the edge prior information of the pseudo HER2 image block, the Res Block1 unit and the Res Block2 unit constitute a two-level residual compression unit, the Res Block1 unit is a shallow residual module, which retains the edge details through residual connection and preliminarily fuses low-order features; the Res Block2 unit is a deep residual module, which further compresses the features and enhances the semantic expression ability, and its output is connected to a global max pooling layer to focus on the key area. Specifically, the local discriminator can avoid gradient degradation through the residual structure, and improve the local discrimination accuracy by using edge guidance.

[0068] During model training, for the generated pseudo IHC image block, the high-frequency signal of the cell membrane boundary is actively enhanced through the trainable Sobel convolution layer, and the gradient response of the nuclear membrane discontinuous region is strengthened; then a two-level residual compression unit is used to reduce the parameter amount while retaining the key details of the chromatin distribution through residual connection; finally, a global max pooling layer is used to replace the full connection layer to force the network to focus on the most discriminative local pathological features (such as nuclear cracks or abnormal chromatin condensation areas), in addition, the local discrimination output can also be obtained through the global max pooling layer.

[0069] In order to further suppress the mode collapse problem, the generation model loss function of the present application can be:

[0070] wherein, is the generation model loss function, is the reconstruction loss weight, is the reconstruction loss, is the adversarial loss, is the adversarial loss weight, is the perception loss, is the perception loss weight, is the frequency domain structure preservation loss, is the frequency domain structure preservation loss weight, a sample feature diversity regular term loss, a sample feature diversity regular term loss weight, a cell edge preservation constraint loss, a cell edge preservation constraint loss weight.

[0071] In implementation, the reconstruction loss weight , the adversarial loss weight , the perceptual loss weight , the frequency domain structure preservation loss weight , the sample feature diversity regular term loss weight and the cell edge preservation constraint loss weight may be determined according to requirements, and will not be described herein.

[0072] Specifically, when the reconstruction loss is added into the generative model loss function, the basic structure can be maintained, and for the reconstruction loss , there is:

[0073] wherein, is the pseudo IHC image block output by the generator, is the benchmark IHC image block, denotes the L1 norm, and the following all denote the same meaning, and reference can be made to the description herein.

[0074] When the adversarial loss is added into the generative model loss function, the fidelity of the image generator in generating the training IHC image block can be improved, and for the adversarial loss , there is:

[0075] wherein, is the local discriminant output of the local discriminator, is the global discriminant output of the global discriminator, and the value range of the local discriminant output and the global discriminant output is usually [0, 1], denotes the mathematical expectation.

[0076] When the perceptual loss is calculated, the VGG network can be used by the present application, and specifically, the training IHC image block and the benchmark IHC image block are loaded into the VGG network, and then the perceptual loss can be calculated, and for the perceptual loss , there is:

[0077] wherein, represents the first layer output feature map of the VGG trained feature extraction network.

[0078] The frequency domain structure preserving loss is added in the generative model loss function to constrain the high frequency components of the pseudo IHC image block and the reference IHC image block in the Fourier frequency domain, to preserve the clarity and morphological consistency of the membrane staining edge, and to improve the detail clarity of the edge and the staining area. Then, we have:

[0079] wherein, is a two-dimensional Fast Fourier Transform (FFT) operation on the pseudo IHC image block, is a two-dimensional Fast Fourier Transform operation on the reference IHC image block.

[0080] The sample feature diversity regular term loss is added in the generative model loss function to prevent the generated image from becoming homogeneous in the later training stage, that is, to prevent the IHC staining image generation base model from collapsing into a mode of outputting a single image in the training; for the sample feature diversity regular term loss Then, we have:

[0081] wherein, , is the image intermediate feature generated by the encoder network in the same batch in the training stage, is the cosine similarity between the image intermediate feature and the image intermediate feature . Specifically, when extracting the image intermediate feature, it should be the intermediate feature output by the bottom encoder in the encoder network.

[0082] The cell edge preservation constraint loss is added in the generative model loss function to preserve the cell edge, to enhance the restoration ability of the cell nucleus edge and the membrane staining contour through image gradient comparison, and for the cell edge preservation constraint loss Then, we have:

[0083] wherein, represents the edge map of the image, which can be obtained by Sobel, Laplacian, etc. gradient operator.

[0084] As can be seen from the above description, the image discriminator simultaneously adopts the local discriminator and the global discriminator, and the loss function of the generation model adopts the above form, which can effectively prevent the common mode collapse.

[0085] In an embodiment of the present application, when the generation model training data set is made, the generation training samples in the generation model training data set are made respectively, wherein, When the generation training sample is made, it includes: Obtain the HE staining slice source image and the paired IHC staining slice source image from the same tumor tissue section; The HE staining slice source image and the IHC staining slice source image are multi-stage registered to generate the HE staining registration image and the IHC staining registration image after multi-stage registration, wherein, When multi-stage registration is performed, it at least includes sequentially performing full-image level coarse registration, mesoscale region level non-rigid registration and tile level fine registration, wherein, After performing the full-image level coarse registration, the HE staining slice source image and the IHC staining slice source image are in a preliminary alignment state; After performing the mesoscale region non-rigid registration, the local offset of the tissue structure is corrected; After performing the tile level fine registration, the IHC staining registration image and the HE staining registration image are generated respectively, and the IHC staining registration image and the HE staining registration image correspond in pixel scale in the cell boundary, the nucleus structure and the membrane staining region; The HE staining registration image and the IHC staining registration image are subjected to segmentation and screening processing to form the training HE image block in the generation training sample and the reference IHCN image block corresponding to the training HE image block after segmentation and screening, and the training HE image block and the reference IHC staining image block are at least in the tissue structure level alignment state.

[0086] As can be seen from the above description, the generation model training data set includes several generation training samples, wherein the generation training samples should be made in the same way. Generally, when making a generation training sample, the HE staining slice source image and the IHC staining slice source image should be obtained, wherein the HE staining slice source image and the IHC staining slice source image should come from the same tumor tissue section. The way of obtaining the HE staining slice source image and the IHC staining slice source image can refer to the above description of the reference HE staining image, which will not be repeated here. It should be noted that the type of tumor tissue corresponding to the reference HE staining image should be consistent with the type of tumor tissue corresponding to the HE staining slice original image and the IHC staining slice original image, such as one of breast tumor tissue, gastric tumor tissue or cervical tumor tissue, Figure 2 An embodiment in which the HE staining slice source image and the IHC staining slice source image come from the breast is shown in FIG. Figure 2In the figure, the left figure is the HE-stained slice source image, and the right figure is the IHC-stained slice source image. It should be understood that the breast as the source should contain tumor tissue.

[0087] In order to solve the problem of insufficient spatial registration accuracy, the present invention performs multi-stage registration on the HE-stained slice source images and the IHC-stained slice source images, so that accurate spatial correspondence at two scales, macrostructure and cell tissue area, can be achieved through multi-stage registration. Specifically, the multi-stage registration may include sequentially performed full-image coarse registration, mid-scale region-level non-rigid registration, and tile-level fine registration. The following examples illustrate the methods and processes of full-image coarse registration, mid-scale region-level non-rigid registration, and tile-level fine registration.

[0088] When performing full-image coarse registration, one possible approach is: Since the HE-stained slice source image and the IHC-stained slice source image have high resolution, they are first downsampled to a low-resolution level (such as Level 3) to reduce computational complexity and suppress high-frequency noise interference. Afterwards, key feature points are extracted from the downsampled images. For example, a feature point detection algorithm (such as SIFT, SURF, or ORB) can be preferably used to extract the key point sets in the two images, and then:

[0089] in, is the set of key points in the image after HE dyeing and sampling, is the key point in the image after HE dyeing and sampling, is the set of key points in the image after IHC staining and sampling, are the key points in the image after IHC staining and sampling, A set of key points The number of internal key points, A set of key points The number of internal key points.

[0090] Key Points Contains its corresponding image position , main direction ,scale and local descriptors (Take SIFT as an example. SIFT detects local extreme points of the image in a Gaussian pyramid, extracts stable key point positions, and constructs a gradient histogram descriptor with consistent direction for each key point (usually a 128-dimensional vector) for subsequent feature matching and geometric alignment between images. SIFT features are highly robust to image scaling, rotation, and a certain degree of brightness changes, and are suitable for stable detection of structural areas such as cell nucleus edges and gland contours in pathological images).

[0091] For any key point , by calculating its relationship with the key points The Euclidean distance of the local descriptor: ,in, is the Euclidean distance of the local descriptor.

[0092] Establish a preliminary matching point pair set ,in, Represents the matching threshold, which is used to filter key point pairs whose Euclidean distance between descriptors is less than the threshold. In order to eliminate the influence of mismatching, the RANSAC algorithm is used to match the set of point pairs. Perform internal point screening: In each round of sampling, select the matching point pair from the set Randomly select 3 to 4 pairs of matching points to estimate the affine transformation matrix , for the Key points and Key points , such that:

[0093] For the affine transformation matrix obtained above, calculate the geometric reprojection error and calculate whether the collective reprojection error is less than the threshold , when the geometric reprojection error is less than the threshold When , the matching point pair is judged to be an internal point; among them, for the geometric reprojection error, there is: .

[0094] In the specific implementation, the set of transformation matrices with the most inliers is finally selected as the coarse alignment output. After that, the affine transformation matrix can be used to affine transform the IHC stained slice source image as a whole to the HE stained slice source image, so that the HE stained slice source image and the IHC stained slice source image are in a preliminary alignment state. That is, the HE stained slice image is used as the alignment reference here, and a HE stained preliminary aligned image and an IHC stained preliminary aligned image can be formed.

[0095] When performing mesoscale region-level non-rigid registration, one possible approach is: The HE staining preliminary alignment image and the IHC staining preliminary alignment image are respectively divided into a plurality of medium-sized image window regions (such as 10000*10000 pixels), and a local non-rigid registration operation is respectively performed in each window for correcting local deviation caused by deformation, distortion or slice thickness difference of the tissue structure. In specific implementation, an elastic deformation model based on B-spline interpolation is preferably used for local deviation correction.

[0096] Specifically, in each medium-sized image window region, a regular grid of control points is defined , and a two-dimensional deformation vector of each control point is defined as follows: , wherein, respectively represent the displacement amount of the control point in the horizontal direction and the vertical direction. The overall deformation vector of an arbitrary pixel point is obtained by weighted difference value of the B-spline basis function to the control point displacement, and then has: , wherein, represents the cubic B-spline basis function of the pixel point in the x direction, represents the cubic B-spline basis function of the pixel point in the y direction, and is a pixel-level non-rigid deformation field.

[0097] It should be noted that when the grid of control points is defined, the current medium-sized image window region needs to be uniformly grid-divided, and the grid of control points is the vertex of the i-th row and the j-th column of the divided grid; the above horizontal direction is generally the length direction of the divided grid, and the vertical direction is generally the width direction of the divided grid.

[0098] In order to solve the optimal deformation parameter, the following mesoscale registration loss function is constructed to jointly minimize the pixel registration error and the smoothness of the deformation field:

[0099] , wherein, is the mesoscale registration loss function, represents the pixel value of the HE medium-sized image window region, represents the pixel value of the HER2 medium-sized image window region, represents the second derivative of the control point deformation vector, which is used to constrain the smoothness of the deformation field, is the regularization coefficient for balancing the two terms.

[0100] It should be noted that, taking the HE staining preliminary alignment image as the reference, the position state of the IHC staining preliminary alignment image is adjusted, and when the mesoscale registration loss function is the smallest, the mesoscale region non-rigid registration is completed, so as to realize the correction of the local offset of the tissue structure. At this time, the HE staining mesoscale registration image is formed based on the HE staining preliminary alignment image, and the IHC staining mesoscale registration image can be formed based on the adjusted IHC staining preliminary alignment image.

[0101] When performing the tile-level fine registration, a feasible way is: The tile fine registration region is selected on the HE staining mesoscale registration image and the IHC staining mesoscale registration image (preferably, the size of the tile fine registration region is 1024*1024 pixels), and then the high-precision registration operation is performed by using the dense optical flow estimation algorithm based on the multi-resolution pyramid; When performing the high-precision registration operation, the optical flow target is to estimate the two-dimensional displacement vector field of each pixel, and then: , wherein, is the horizontal displacement of the tile fine registration region on the IHC staining mesoscale registration image to the tile fine registration region on the HE staining mesoscale registration image, is the vertical displacement of the tile fine registration region on the IHC staining mesoscale registration image to the tile fine registration region on the HE staining mesoscale registration image; in addition, the registered position should satisfy the photometric consistency assumption, and then: , is the gray value of the pixel point (x, y) in the tile fine registration region on the HE staining mesoscale registration image, and similarly, is the corresponding gray value in the tile fine registration region on the IHC staining mesoscale registration image.

[0102] The high-precision registration loss function is constructed, and for the high-precision registration loss function, then:

[0103] wherein, is a regularization coefficient, , is the gradient of the optical flow field, which is used to suppress local discontinuity or non-physical deformation.

[0104] It should be noted that, taking the HE staining mesoscale registration image as the reference, the position state of the IHC staining mesoscale registration image is adjusted, and when the high-precision registration loss function is the smallest, the tile-level fine registration is completed, so as to ensure that the cell boundary, the nucleus structure and the membrane dyeing region are accurately corresponding in the pixel scale. In addition, after performing the tile-level fine registration, the IHC staining registration image and the HE staining registration image are corresponding in the pixel scale in the cell boundary, the nucleus structure and the membrane dyeing region. Figure 3An embodiment of the IHC staining registration map and the HE staining registration map is shown, wherein the left drawing is the HE staining registration map, and the right drawing is the IHC staining registration map.

[0105] In an embodiment of the present application, when the HE staining registration map and the IHC staining registration map are segmented and screened, the following steps are included: HE training candidate regions are sequentially selected on the HE staining registration map, and tumor region identification is performed on the HE training candidate regions, so as to determine the binary mask of each pixel in the HE training candidate region after tumor region identification, wherein when the binary mask is 1, it indicates that the current pixel belongs to the tumor region, and when the binary mask is 0, it indicates that the current pixel does not belong to the tumor region. Based on the binary mask of each pixel in the HE training candidate region, the tumor region proportion is calculated, wherein when the calculated tumor region proportion matches the region selection threshold, the current HE training candidate region is configured as a training HE image block. Based on the position state of the training HE image block, a corresponding reference IHC image block is segmented on the IHC staining registration map.

[0106] In specific implementation, a fixed step sliding window strategy can be used to sequentially select HE training candidate regions on the HE staining registration map. In order to ensure the independence of training and testing samples, the HE training candidate regions should be selected to avoid overlapping of the HE training candidate regions. Preferably, the size of the HE training candidate region is 1024x1024 pixels, and the sliding step is 1024 pixels.

[0107] In order to eliminate non-tumor tissue or background region and improve the pertinence and effectiveness of the training data, the present application adopts a block screening mechanism based on cancer region mask. Specifically, when the used pathological image data are all derived from self-collected tumor tissue section samples, and the tumor regions therein are finely outlined at pixel level by qualified pathologists to form a cancer region mask map for modeling. On this basis, a tumor region identification model based on U-Net or other semantic segmentation architecture is trained to automatically predict potential cancer regions at full image scale, wherein the tumor region identification model and the training method and process of the tumor region identification model can be consistent with the prior art, which will not be described here.

[0108] For the HE training candidate region, the HE training candidate region is loaded into the tumor region identification model to output the binary mask of each pixel in each HE training candidate region by using the tumor region identification model, wherein when the binary mask is 1, it indicates that the current pixel belongs to the tumor region, and when the binary mask is 0, it indicates that the current pixel does not belong to the tumor region, wherein the non-tumor region can be non-cancerous tissue or background region.

[0109] Based on the binary mask of each pixel in the HE training candidate region, the tumor area ratio is calculated, and then

[0110] wherein, is the tumor area ratio, is the area of the current HE training candidate region, for example, if the size of the HE training candidate region is 1024x1024 pixels, should be 1024x1024.

[0111] After calculating the tumor area ratio of the current HE training candidate region, the tumor area ratio is compared with the region selection threshold. When the tumor area ratio is greater than the region selection threshold, it is considered that the tumor area ratio matches the region selection threshold. At this time, the current HE training candidate region is configured as a training HE image block. Since the IHC staining registration map and the HE staining registration map are highly registered in space, the corresponding reference IHC image block can be cut on the IHC staining registration map based on the position state of the training HE image block. At this time, a training HE image block and a reference IHC image block in a generated training sample are formed, Figure 4 An embodiment of a training HE image block and a reference IHC image block in a generated training sample is output in the figure. The left side of the figure is a training HE image block, and the right side of the figure is a reference IHC image block.

[0112] In specific implementation, the region selection threshold can be 0.25. Of course, the region selection threshold can also be selected as other numerical values, which can be selected as needed. It should be noted that when the reference HE image is cut and screened, the above cutting and screening description can be referred to. The difference is that only the corresponding target HE image block needs to be obtained during cutting and screening.

[0113] It should be understood that based on the above multi-stage registration and cutting and screening processing, it can be obtained that the training HE image block and the reference IHC image block in the same generated training sample should come from the same tumor tissue section, and the training HE image block and the reference IHC image block are at least in an aligned state at the tissue structure level. At the same time, the training HE image block and the reference IHC image block both contain tumor tissue information. Further, after the above multi-stage registration and cutting and screening processing, the generated training sample is prepared, and when the reference HE staining image is cut and screened, the consistency requirement of the generation task and the classification task of the present application can be met.

[0114] In order to be able to perform the above expression grading processing, in an embodiment of the present application, the multi-modal semantic classification model at least includes a visual encoder module, a language encoder module, and a text-image alignment inference module, wherein, extract visual embedding features of the to-be-inspected IHC image block by using the visual encoder module; generate a reference semantic encoding feature group by using the language encoder module, wherein when the reference semantic encoding feature group is generated, the HER2 level guide source information group is loaded into the language encoder module to perform semantic encoding on each HER2 level guide source information in the HER2 level guide source information group by using the language encoder module, and after the semantic encoding, corresponding HER2 level semantic encoding features are generated, and the HER2 level semantic encoding features in the reference semantic encoding feature group are independent of each other; calculate feature similarities between the visual embedding features and each HER2 level semantic encoding feature by using the image-text alignment reasoning module, and determine the target HER2 level information based on the calculated feature similarities, and output the expression level prediction information of the current to-be-inspected IHC image block.

[0115] Figure 8 An embodiment of the multi-modal semantic classification model is shown in FIG. 1, and as shown in the figure, the multi-modal semantic classification model can at least include a visual encoder module, a language encoder module, and an image-text alignment reasoning module, wherein the visual encoder module can adopt a structure such as ViT (Vision Transformer) or ResNet-101, and specifically can extract visual embedding features of the to-be-inspected IHC image block. The language encoder module is used to generate a reference semantic encoding feature group, and for the determined multi-modal semantic classification model, the language encoder module generates the reference semantic encoding feature group unchanged in the reasoning work, that is, the HER2 level semantic encoding features in the reference semantic encoding feature group remain stable, and of course, for different multi-modal semantic classification models, the reference semantic encoding feature group generated by the language encoder module can be different, and specifically can be selected as needed. The manner and process of generating the reference semantic encoding feature group by the language encoder module can be referred to the corresponding description below.

[0116] Specifically, the feature similarities between the visual embedding features and each HER2 level semantic encoding feature can be calculated by using the image-text alignment reasoning module, and the target HER2 level information is determined based on the calculated feature similarities, and thereafter, the expression level prediction information of the current to-be-inspected IHC image block is outputted, and specifically, the manner and process of calculating the feature similarities and determining the target HER2 level information can be referred to the corresponding description above, which will not be described here again.

[0117] In an embodiment of the present application, the multi-modal semantic classification model further includes a category prototype matching module, wherein, When the feature similarity calculated by the image-text alignment reasoning module is in a fuzzy approximate state, the category prototype matching module is configured to perform prototype similarity calculation processing to determine the target HER2 level information through the prototype similarity calculation processing. When the prototype similarity calculation processing is performed, the prototype similarity between the visual embedding feature and the class prototype feature in the class prototype matching module is calculated respectively, and the optimal prototype similarity and the target class prototype feature corresponding to the optimal prototype similarity are determined, The target HER2 level information is generated based on the target class prototype feature, The target HER2 level information is configured as the expression level prediction information of the current to-be-inspected IHC image block, and the class prototype matching module is configured to output the expression level prediction information of the current to-be-inspected IHC image block.

[0118] As can be seen from the above description, after calculating the feature similarity between the visual embedding feature and each HER2 level semantic encoding feature, the graph-text alignment reasoning module needs to select and determine the maximum feature similarity, and the HER2 level semantic encoding feature corresponding to the maximum feature similarity is configured as the optimal HER2 level semantic encoding feature. It can be understood that the gap between the selected maximum feature similarity and the suboptimal feature similarity may be small, for example, when the difference between the maximum feature similarity and the suboptimal feature similarity is not greater than a switching threshold, the feature similarity calculated by the graph-text alignment reasoning module is in a fuzzy approximation state, and the suboptimal feature similarity specifically refers to a feature similarity that is less than the maximum feature similarity but greater than other feature similarities. The size of the switching threshold can be selected as needed, and will not be described here.

[0119] When in the fuzzy approximation state, if the above-mentioned method is still used to determine the target HER2 level information, the accuracy of the output expression level prediction information will be low. In order to further improve the accuracy of generating the expression level prediction information, the present application further sets a class prototype matching module in the multi-modal semantic classification model, as shown in Figure 8 As shown, when the feature similarity calculated by the graph-text alignment reasoning module is in the fuzzy approximation state, the class prototype matching module is started to perform the prototype similarity calculation processing, that is, when reasoning, the expression grading processing is preferably performed by the graph-text alignment reasoning module.

[0120] In specific implementation, a plurality of class prototype features are stored in the class prototype matching module, and when the prototype similarity calculation is performed, the prototype similarity between the visual embedding feature and each class prototype feature is calculated. After calculating the prototype similarity between the visual embedding feature and all class prototype features, the optimal prototype similarity and the target class prototype feature corresponding to the optimal prototype similarity can be determined, wherein the optimal prototype similarity is the maximum value of the calculated prototype similarity, and after the optimal prototype similarity is determined, the corresponding target class prototype feature can be determined.

[0121] It should be noted that when calculating the prototype similarity between the visual embedding feature and the category prototype feature, the form of calculating the feature similarity can be used, specifically, when the calculation method of the feature similarity is used, then

[0122] wherein, is the prototype similarity between the visual embedding feature and the category prototype feature of the category , is the category prototype feature of the category , is the vector dot product between the visual embedding feature and the category prototype feature , is the vector module multiplication between the visual embedding feature and the category prototype feature .

[0123] Specifically, after obtaining the target category prototype feature, the target HER2 level information can be generated based on the target category prototype feature, and then the target HER2 level information is configured as the expression level prediction information of the current IHC image block, and the category prototype matching module is configured to output the expression level prediction information of the current IHC image block, that is, when the target category prototype feature and the corresponding expression level prediction information are obtained through the category prototype matching module, the category prototype matching module outputs the level prediction information.

[0124] In an embodiment of the present application, when constructing a multi-modal semantic classification model, the following steps are included: select a pre-training framework model, and configure the selected pre-training framework model as a multi-modal semantic classification base model, wherein the multi-modal semantic classification base model includes a visual encoder module, a language encoder module, a prompt enhancement module, a text-image alignment inference module, and a category prototype matching basic module; construct a classification model training data set to train the multi-modal semantic classification base model using the classification model training data set, wherein the classification model training data set includes a plurality of classification training samples, and each classification training sample includes a training IHC image block and training expression level text information corresponding to the current training IHC image block; configure a classification model training condition for training the multi-modal semantic classification base model, until the multi-modal semantic classification base model is trained to reach a classification model target state, and then generate a multi-modal semantic classification model based on the multi-modal semantic classification base model trained to reach the classification model target state, wherein During the model training, the weights of the visual encoder module and the language encoder module are frozen, the prompt enhancement module is prompted to fine-tune the corresponding weight parameters through the back propagation mechanism, and the category prototype matching basic module updates the corresponding category prototype feature in the training of each batch of classification training samples through the momentum update mechanism. During the model training, each HER2 level description text in the HER2 level description text group is spliced with a learnable training vector by the prompt enhancement module to generate HER2 level training source information corresponding to each HER2 level description text, and then the language encoder module is used to perform semantic encoding on the HER2 level training source information to generate HER2 level training encoding features. The visual encoder module is used to extract training visual features of the training IHC image block in each classification training sample, and then the graphic-text alignment inference module is configured to calculate the feature similarity between the training visual features and all HER2 level training encoding features, and the category prototype matching basic module is used to calculate and update the category prototype feature corresponding to each training expression level information.

[0125] It should be understood that for the above-mentioned multi-modal semantic classification model, a multi-modal semantic classification basic model and a classification model training data set should generally be constructed first, then the multi-modal semantic classification basic model is trained through the classification model training data set, and after the multi-modal semantic classification basic model training reaches the classification model target state, the multi-modal semantic classification model is generated based on the multi-modal semantic classification basic model trained to reach the classification model target state.

[0126] In an embodiment of the present application, a pre-trained framework model can be selected, and the selected pre-trained framework model is used as the multi-modal semantic classification basic model. The pre-trained framework model specifically refers to a model after pre-training, and the pre-trained framework model can be a MUSK (Multimodal transformer with Unified maSKed modeling) model. The MUSK model realizes cross-modal alignment of pathological images and text descriptions through contrast learning on a large-scale medical image data set (50 million pathological images and 1 billion pathological related text labels of 11577 patients). Based on the above description of the MUSK model, the visual encoder module in the multi-modal semantic classification basic model can have the ability to extract visual embedding features. For example, after loading the training IHC image block in the classification training sample into the visual encoder module, the visual encoder module can extract the training visual features of the training IHC image block. The training visual features can refer to the visual embedding features described above, which will not be described here.

[0127] In addition, the language encoder module in the multi-modal semantic classification base model has semantic encoding capability, and therefore, when the multi-modal semantic classification base model is trained, the weights of the visual encoder module and the language encoder module can be frozen, that is, during the training of the multi-modal semantic classification base model, the weights of the visual encoder module and the language encoder module remain unchanged, and only the weights of the prompt enhancement module need to be fine-tuned. The class prototype matching module does not directly participate in back propagation, but in the training of each batch of classification training samples, the corresponding class prototype features in the class prototype matching base module are updated through a momentum update mechanism. As can be seen from the above description, the image-text alignment reasoning module mainly calculates the target HER2 level information that is in an optimal approximation state with the visual embedding features, and therefore, the image-text alignment reasoning module has no weight parameters that need to be trained.

[0128] Specifically, the classification model training data set should include a plurality of classification training samples, each classification training sample including a training IHC image block and training expression level text information corresponding to the current training IHC image block, wherein the training IHC image block can be a pseudo IHC image block generated by the IHC image generation model and / or a benchmark IHC image block obtained through segmentation and screening processing. Preferably, the classification model training data set can simultaneously include training IHC image blocks using pseudo IHC image blocks and training IHC image blocks using benchmark IHC image blocks.

[0129] For the training expression level text information in each classification training sample, the training expression level text information is mainly used to describe the HER2 expression level of the training IHC image block. The training expression level text information can be a preset fixed template text, which is used to guide the alignment of the training visual features and the semantic information of the training expression level text information. The text description of the training expression level text information can include: level 1), "0 level: no obvious membrane staining or only cell nucleus staining"; level 2), "1+: cell membrane shows incomplete membrane staining, weak positive, and staining is discontinuous"; level 3), "2+: membrane staining is moderate, and more than 10% of cells show incomplete or partial strong positive membrane staining"; and level 4), "3+: more than 10% of cells show complete, strong positive membrane staining, clear boundary, and continuous staining". Figure 8 The level 1 text to level 4 text shown in the above description represents four types of training expression level text information. The level 1 text to level 4 text can be the text description corresponding to 1) to 4) respectively. As can be seen, the training expression level text information mainly describes that the HER2 expression level of the training IHC image block is one of the above level 1) to level 4).

[0130] It should be noted that the training expression level text information described above can be used as the standard input of the language encoder module, and thereafter, the text and image features are matched with the IHC image block in the multi-modal semantic classification base model to realize the prediction training of the expression level. Preferably, the training expression level text information can be flexibly extended according to the clinical classification standard, but should be consistent during the model training and inference stage to ensure the alignment stability.

[0131] To enhance the perception ability of the language encoder module to the task semantics, the present application sets a prompt enhancement module in the multi-modal semantic classification base model, which can splice a set of learnable training vectors in front of each HER2 expression level of each class through the prompt enhancement module to generate HER2 level training source information corresponding to each HER2 level description text. Figure 8 As described above, when there are level 1 text to level 4 text, four learnable training vectors are randomly generated through the prompt enhancement module, and the four learnable training vectors are spliced in front of the level 1 text to the level 4 text respectively to form four corresponding HER2 level training source information. Figure 8 In the above, the four learnable training vectors can form a prompt learnable vector sequence.

[0132] It should be noted that the dimension of the learnable training vector generated by the prompt enhancement module should be consistent with the original word vector of the level description text, and the weight parameter of the prompt enhancement module is involved in the optimization during the training stage to guide the multi-modal semantic classification base model to understand the task context.

[0133] During model training, for each classification training sample, the HER2 level training source information is loaded into the language encoder module to perform semantic encoding on the current HER2 level training source information using the language encoder, so that the HER2 level training encoding features can be generated after semantic encoding. Therefore, during the model training process, when there are four HER2 level training source information, four corresponding HER2 level training encoding features can be generated after semantic encoding of all HER2 level training source information. Of course, the training IHC image block in each classification training sample should be loaded into the visual encoder module to extract the corresponding training visual features by the visual encoder module.

[0134] It can be understood that after the multi-modal semantic classification base model is trained to reach the classification model target state, the four HER2 level training coding features are trained based on the above four HER2 levels, and four corresponding HER2 level semantic coding features can be formed, that is, the inference work refers to the semantic coding feature group which includes 4 independent HER2 level semantic coding features. When the expression level text information is trained as other conditions, the HER2 level semantic coding features can be determined, and the reference semantic coding feature group includes the corresponding number of HER2 level semantic coding features, which will not be described one by one here.

[0135] During model training, the classification training samples are generally sent into the multi-modal semantic classification base model in batches. For each classification training sample in the same batch, the training visual features of the training IHC image block in each classification training sample are extracted by using the visual encoder module. Thereafter, the graph-text alignment inference module is configured to calculate the feature similarity between the training visual features and all HER2 level training coding features. The way to calculate the feature similarity between the training visual features and the corresponding features of the HER2 level training coding features can refer to the above description of the calculation of the feature similarity, which will not be described here.

[0136] For each batch of classification training samples, each training expression level information category prototype feature can be calculated as follows:

[0137] Among them, is the number of classification training samples of category in each batch of classification training samples, is the training visual feature of the th classification training sample in each batch of classification training samples, is the training expression level text information of the th classification training sample in each batch of classification training samples, is the category prototype feature of the classification training sample corresponding to category ; For each category , the category prototype feature is updated by a momentum update mechanism, and has:

[0138] Among them: is the update momentum coefficient.

[0139] In specific implementation, the size of the update momentum coefficient can be selected as needed, such as the update momentum coefficient may be 0.1. During the above update, Specifically, it refers to assigning the right calculation result to the coordinates, that is, realizing the update of the category prototype feature The category is one of the above-mentioned levels 1) to 4).

[0140] It should be noted that when the text description of the training expression level text information is the above-mentioned levels 1) to 4), the prototype similarity between the visual embedding feature and the category prototype feature is calculated by the category prototype matching module, and the target HER2 level information is determined, which can significantly improve the discrimination ability between the boundary levels such as 1+ and 2+, 2+ and 3+, and thus further improve the accuracy of the expression grading prediction information.

[0141] It should be understood that when training the multi-modal semantic classification base model, the classification model training conditions should be configured. Generally, the configured classification model training conditions at least include a classification training loss function. Of course, the classification model training conditions can also include other necessary conditions, such as: setting the initial learning rate of the AdamW optimizer to 3e-5, and using a linear warm-up and cosine annealing strategy, using a dropout of 0.1 and a weight decay of 0.01 to prevent overfitting. During the training process, the accuracy, recall rate and F1 score of the classification should be monitored on the validation set. When the indicators do not improve for 5 consecutive epochs, the early stopping mechanism is triggered, at which time it can be considered that the model training of the multi-modal semantic classification base model has reached the classification model target state.

[0142] Specifically, when training the multi-modal semantic classification base model, a graph-text alignment loss function based on contrastive learning is introduced to improve the semantic consistency between the visual feature and the HER2 level training encoding feature, that is, the classification training loss function can adopt the graph-text alignment loss function. In an embodiment of the present application, when the classification training loss function adopts the graph-text alignment loss function, the classification training loss function adopts the CLIP-style bidirectional contrast form, while maximizing the feature similarity between each pair of training IHC image blocks and their corresponding training expression level text information, and minimizing the feature similarity with other training expression level text information. For the classification training loss function, there is:

[0143] In the formula, is the number of categories, is the training visual feature, is the HER2 level training encoding feature, is the cosine similarity function between the HER2 level training encoding feature and the training visual feature , and is a temperature hyperparameter.

[0144] For the above-described embodiments, the number of categories is 4. The temperature hyperparameter τ is used to scale the distribution of image-text similarity scores. Generally, the smaller the temperature hyperparameter τ (such as 0.01), the sharper the distribution, and the model pays more attention to high-similarity samples but is sensitive to noise. The larger the temperature hyperparameter τ (such as 0.2), the flatter the distribution, and the model is more tolerant of low-similarity samples and enhances generalization. The temperature hyperparameter τ can be set to 0.07 according to experience.

[0145] As can be known from the above-described classification training loss function, the classification training loss function can cause the multi-modal semantic classification base model to learn a one-to-one correspondence between the training visual features and the HER2 grade training encoded features, thereby significantly improving the alignment capability of the HER2 grade semantic expression.

[0146] In specific implementation, when the multi-modal semantic classification base model is trained to reach the classification model target state, thereafter, a multi-modal semantic classification model can be generated based on the multi-modal semantic classification base model trained to reach the classification model target state. As can be known from the above description, when the multi-modal semantic classification model is generated, the prompt enhancement module is in a non-working state, and the reference semantic encoded feature group generated by the language encoder module remains fixed, that is, in the inference working process, the HER2 grade semantic encoded features in the reference semantic encoded feature group will not be transformed.

[0147] In addition, a category prototype matching basic module can be generated by the category prototype matching basic module, and corresponding category prototype features are stored in the category prototype matching basic module for subsequent prototype similarity calculation. These category prototype features essentially represent the central representation of each category in the feature space, and the mechanism is as follows: During the training process, the sample features of each category are continuously aggregated, the feature center points of each category are calculated and saved through the above-described dynamic update mechanism, and finally a category prototype feature set corresponding to the number of categories (such as four prototype vectors corresponding to four categories) is formed.

[0148] As can be known from the above description, the prompt enhancement module and the category prototype matching basic module are integrated in the multi-modal semantic classification base model as lightweight additional modules, have good task migration capability and semantic adaptation capability, avoid damaging the backbone model, improve the HER2 grade classification performance, and are particularly suitable for category boundary fuzzy tasks.

[0149] ​It should be noted that after the multi-modal semantic classification model is deployed, the above method can be used to first use the image-text alignment reasoning module to perform the above calculation of the visual embedding feature and the feature similarity of each HER2 level semantic encoding feature, and output the expression level prediction information of the front IHC image block by the image-text alignment reasoning module. When the feature similarity calculated by the image-text alignment reasoning module is in a fuzzy approximate state, the processor deploying the multi-modal semantic classification model is configured to perform prototype similarity calculation processing by the category prototype matching module to determine the target HER2 level information.

[0150] Specifically, the processor deploying the multi-modal semantic classification model can also obtain the feature similarity of the visual embedding feature calculated by the image-text alignment reasoning module and each HER2 level semantic encoding feature, and then determine whether the calculated feature similarity is in a fuzzy approximate state. Therefore, the working and output state of the image-text alignment reasoning module and the category prototype matching module can be controlled by the processor deploying the multi-modal semantic classification model. Of course, the image-text alignment reasoning module can also output all the calculated feature similarities at the same time, and then the processor determines the maximum feature similarity and whether the calculated feature similarity is in a fuzzy approximate state. The working and output state of the image-text alignment reasoning module and the category prototype matching module can be selected according to actual needs, which will not be described here. It should be understood that when it is determined that the feature similarity is in a fuzzy approximate state, the target HER2 level information determined by the image-text alignment reasoning module should be ignored, and the target HER2 level information determined by the category prototype matching module should be used.

Claims

1. A method for grading tumor HER2 expression based on HE staining images, characterized by: The tumor HER2 expression grading method includes: providing a reference HE-stained image, and performing at least segmentation and screening processing on the reference HE-stained image to generate a target HE image block group after segmentation and screening, wherein all target HE image blocks in the target HE image block group contain tumor tissue information; Each target HE image block in the target HE image block group is loaded into the IHC staining image generation model, so as to generate an IHC image block to be inspected that is structurally aligned with the current target HE image block using the IHC staining image generation model; Each IHC image block to be tested is loaded into a multimodal semantic classification model to perform expression classification processing using the multimodal semantic classification model, and after the classification processing, HER2 expression level prediction information of the current IHC image block to be tested is generated, wherein, When performing expression classification processing, the multimodal semantic classification model extracts visual embedding features of the IHC image block to be tested, and calculates and determines the target HER2 level information that is in the best approximation state with the visual embedding features; The target HER2 level information is configured as the expression level prediction information of the current IHC image block to be tested.

2. The method for grading tumor HER2 expression based on HE staining images according to claim 1, characterized in that: When calculating and determining the target HER2 level information that best approximates the visual embedding feature, it includes: Calculate the feature similarity between the visual embedding feature and the corresponding semantic encoding feature of each HER2 level in the reference semantic encoding feature group; For any visual embedding feature, when the calculated feature similarity is at the maximum state, the visual embedding feature and the corresponding HER2-level semantic coding feature are in the best approximation state. At this time, the corresponding HER2-level semantic coding feature is configured as the optimal HER2-level semantic coding feature; The HER2 level information within the optimal HER2 level semantic encoding feature is configured as the target HER2 level information.

3. The method for grading tumor HER2 expression based on HE staining images according to claim 2, characterized in that: The multimodal semantic classification model includes at least a visual encoder module, a language encoder module and a picture-text alignment reasoning module, wherein: The visual encoder module is used to extract the visual embedding features of the IHC image patch to be inspected; generating a reference semantic coding feature group using a language encoder module, wherein when generating the reference semantic coding feature group, the HER2 level guidance source information group is loaded into the language encoder module, so that each HER2 level guidance source information in the HER2 level guidance source information group is semantically encoded by the language encoder module, and corresponding HER2 level semantic coding features are generated after the language encoding, and the HER2 level semantic coding features in the reference semantic coding feature group are independent of each other; The image-text alignment inference module is used to calculate the feature similarity between the visual embedding feature and the semantic encoding feature of each HER2 grade. The target HER2 grade information is determined based on the calculated feature similarity, and the expression grade prediction information of the current IHC image block to be tested is output.

4. The method for grading tumor HER2 expression based on HE staining images according to claim 3, characterized in that: The multimodal semantic classification model also includes a category prototype matching module, wherein: When the feature similarity calculated by the image-text alignment reasoning module is in a fuzzy approximate state, the category prototype matching module is configured to perform prototype similarity calculation processing to determine the target HER2 level information through the prototype similarity calculation processing; When performing the prototype similarity calculation process, the prototype similarities of the visual embedding features and the category prototype features in the category prototype matching module are calculated respectively, and the optimal prototype similarity and the target category prototype features corresponding to the optimal prototype similarity are determined. Generate target HER2 level information based on target category prototype features, The target HER2 level information is configured as the expression level prediction information of the current IHC image block to be inspected, and the category prototype matching module is configured to output the expression level prediction information of the current IHC image block to be inspected.

5. The method for grading tumor HER2 expression based on HE staining images according to claim 3, characterized in that: The feature similarity calculated by the image-text alignment inference module includes cosine similarity, in, When the feature similarity adopts cosine similarity, then: in, is the feature similarity, is the visual embedding feature, For category HER2-level semantic encoding features, The visual embedding features and categories are HER2-level semantic encoding features The vector dot product of The visual embedding features and categories are HER2-level semantic encoding features Vector modular multiplication of .

6. The method for grading tumor HER2 expression based on HE staining images according to claim 4, characterized in that: When building a multimodal semantic classification model, it includes: Select a pre-trained framework model and configure the selected pre-trained framework model as a multimodal semantic classification basic model, wherein the multimodal semantic classification basic model includes a visual encoder module, a language encoder module, a prompt enhancement module, an image-text alignment reasoning module, and a category prototype matching basic module; Constructing a classification model training dataset to perform model training on a multimodal semantic classification basic model using the classification model training dataset, wherein the classification model training dataset includes a plurality of classification training samples, each classification training sample includes a training IHC image block and training expression level text information corresponding to the current training IHC image block; A classification model training condition is configured for training the multimodal semantic classification base model until the multimodal semantic classification base model is trained to reach a classification model target state, and thereafter, a multimodal semantic classification model is generated based on the multimodal semantic classification base model trained to reach the classification model target state, wherein: During model training, the corresponding weight parameters of the visual encoder module and the language encoder module are frozen, the prompt enhancement module fine-tunes the corresponding weight parameters through the back-propagation mechanism, and the category prototype matching basic module updates the corresponding category prototype features in the training of each batch of classification training samples through the momentum update mechanism; During model training, the prompt enhancement module is used to concatenate a learnable training vector to each HER2 grade description text in the HER2 grade description text group to generate HER2 grade training source information corresponding to each HER2 grade description text. Subsequently, the language encoder module is used to semantically encode the HER2 grade training source information to generate HER2 grade training encoding features; The visual encoder module is used to extract the training visual features of the training IHC image blocks in each classification training sample. After that, the image-text alignment inference module is configured to calculate the feature similarity between the training visual features and the corresponding feature encoding features of all HER2 grade training. The category prototype matching basic module is used to calculate and update the category prototype features corresponding to each training expression grade information.

7. The method for grading tumor HER2 expression based on HE staining images according to claim 6, characterized in that: When calculating the prototype features of each training expression level information category, we have: in, Classify the categories within the training samples for each batch The number of classification training samples, For each annotation classification training sample The training visual features of the classification training samples, For each batch of classification training samples The training expression level text information of the classification training samples, For category Category prototype features corresponding to training samples; For each category , update the category prototype features through the momentum update mechanism When , there are: in: is the updated momentum coefficient.

8. The method for grading tumor HER2 expression based on HE staining images according to any one of claims 1 to 7, characterized in that: The IHC staining image generation model includes at least an image generator, and the target HE image block is loaded into the IHC staining image generator so that the IHC staining image generator generates an IHC image block to be tested, wherein: The IHC staining image generation model is generated based on the IHC staining image generation basic model after model training. The IHC staining image generation basic model includes an image generator and an image discriminator. The IHC staining image generation basic model is trained using the prepared generative model training dataset. The generative model training dataset includes a number of generative training samples, each generative training sample includes a training HE image block and a reference IHC image block corresponding to the training HE image block, wherein the training HE image block and the reference IHC image block in the same generative training sample are formed based on the same tumor tissue slice, and both the training HE image block and the reference IHC image block contain tumor tissue information, and the training HE image block and the reference IHC image block in the same training sample are aligned at least at the tissue structure level; The generative model training conditions required for model training of the IHC staining image generation base model are configured until the training effect of the IHC staining image generation base model reaches a target state of the generative model. Thereafter, an IHC staining image generation model is generated based on the IHC staining image generation base model whose training effect reaches the target state of the generative model.

9. The method for grading tumor HER2 expression based on HE staining images according to claim 8, characterized in that: When making a generative model training dataset, generate training samples in the generative model training dataset respectively, where: When generating training samples, it includes: Obtain HE-stained slice source images and paired IHC-stained slice source images from the same tumor tissue slice; The HE-stained slice source image and the IHC-stained slice source image are multi-stage registered to generate the HE-stained registered image and the IHC-stained registered image after the multi-stage registration, wherein: When performing multi-stage registration, it includes at least the whole-image coarse registration, the mid-scale region-level non-rigid registration and the tile-level fine registration in sequence, wherein: After performing full-image coarse registration, the HE-stained slice source image and the IHC-stained slice source image are in a preliminary alignment state; After performing non-rigid registration of mesoscale regions to correct for local shifts in tissue structures; After performing tile-level fine registration, IHC staining registration maps and HE staining registration maps are generated respectively, and the IHC staining registration map and the HE staining registration map are made to correspond to each other at the pixel level in terms of cell boundaries, nuclear structures, and membrane staining areas; The HE staining registration map and the IHC staining registration map are segmented and screened to generate training HE image blocks and reference IHC image blocks corresponding to the training HE image blocks within the generated training samples after segmentation and screening, and the training HE image blocks and the reference IHC image blocks are aligned at least at the tissue structure level.

10. The method for grading tumor HER2 expression based on HE staining images according to claim 9, characterized in that: When segmenting and screening the HE staining registration image and the IHC staining registration image, the following steps are included: Selecting HE training candidate regions on the HE staining registration image in sequence, and performing tumor region identification on the HE training candidate regions, so as to determine a binary mask for each pixel in the HE training candidate regions after the tumor region identification, wherein when the binary mask is 1, it indicates that the current pixel belongs to the tumor region, and when the binary mask is 0, it indicates that the current pixel does not belong to the tumor region; Calculate the tumor area ratio based on the binary mask of each pixel in the HE training candidate area, wherein when the calculated tumor area ratio matches the area selection threshold, configure the current HE training candidate area as the training HE image block; Based on the position status of the training HE image block, the corresponding benchmark IHC image block is cut on the IHC staining registration map.

Citation Information

Patent Citations

  • Breast cancer pathological image classification method based on SENet channel attention and transfer learning

    CN114820555A

  • Breast cancer recognition system and method based on Swin Transform and comparative learning

    CN117765252A

  • Identification and classification method and system for breast cancer pathological image

    CN119006942A

  • Multi-level lesion detection optimized pathological section image segmentation detection method

    CN116596952A

  • Generalized zero sample image classification method based on fused visual information

    CN116797821A

Cited By

  • Semi-subjective HER2 scoring method and device based on prior intensity constraint

    CN121074052A