Breast cancer HER2 prediction method based on multiple prompt learning

Through the multiple prompt learning method, combined with the descriptive text and image features generated by the large language model, the text-guided feature aggregation module and the global feature fusion module are used to solve the problems of high labeling costs, strong data dependence and poor results interpretation in the breast cancer HER2 score, and achieve efficient, accurate and interpretable HER2 prediction effects.

CN119993468AActive Publication Date: 2025-05-13SOUTH CHINA UNIV OF TECH

Patent Information

Application Number
CN202510480597.2
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-04-17
Publication Date
2025-05-13
Estimated Expiration
2045-04-17

AI Technical Summary

Technical Problem

The prior art has problems in breast cancer HER2 scores with high labeling costs, strong data dependence and poor interpretation of results, making it difficult to achieve efficient, accurate and interpretable HER2 predictions.

Method used

Multiple prompt learning methods are used to generate instance-level class and staining descriptive text through large language models, and combine it with image features. The text-guided feature aggregation module and global feature fusion module are used to generate HER2 features of breast cancer WSI pathological images.

Benefits of technology

It improves the accuracy and interpretability of HER2 scores, reduces dependence on large-scale data sets, improves the generalization performance of the model, and implements an efficient parameter training process.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119993468A_ABST
    Figure CN119993468A_ABST
Patent Text Reader

Abstract

The invention discloses a breast cancer HER2 prediction method based on multiple prompt learning. The method comprises the following steps: firstly, acquiring image features of a breast cancer WSI pathological image; then generating a multi-prompt descriptive text by adopting a large language model, and embedding learnable prompts into the descriptive text by using a multi-modal pre-training model to generate multi-prompts; using a text encoder to obtain multiple prompt features; performing aggregation by adopting a text-guided feature aggregation module to obtain category global visual features and dyeing degree global visual features, and inputting the category global visual features and the dyeing degree global visual features into a global feature fusion module for integration to obtain HER2 features; and finally, calculating the cosine similarity between the slice-level prompt feature and the HER2 feature, obtaining an HER2 prediction result of the breast cancer WSI pathological image, and performing optimization by using cross entropy loss. According to the method, semantic information related to text description can be effectively captured, the HER2 evaluation process of a pathologist is simulated, and the HER2 prediction accuracy is improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The invention belongs to the technical field of pathological image processing and HER2 prediction, and in particular relates to a breast cancer HER2 prediction method based on multiple prompt learning. Background Art

[0002] Breast cancer is the second most common cancer worldwide, and it is also the cancer with the highest incidence and mortality rate among women, with the fourth highest mortality rate worldwide. In clinical diagnosis, hematoxylin andeosin (H&E) is usually used to stain tumor tissues and perform morphological evaluation under an optical microscope. In addition, immunohistochemistry (IHC) is also used to evaluate the expression of biomarkers, such as human epidermal growth factor receptor 2 (HER2). The amplification of the HER2 gene leads to a significant increase in the expression of HER2 protein on the cell surface, which is closely related to the accelerated proliferation and enhanced invasiveness of tumor cells. The accurate assessment of HER2 status is not only an important basis for the diagnosis and classification of breast cancer, but also plays a key role in clinical treatment decisions. Therefore, HER2 scoring is a core link in the diagnosis, classification and treatment of breast cancer. In clinical practice, pathologists determine the HER2 score (0, 1+, 2+, 3+) based on the proportion of invasive cancer cells stained to different degrees; cases with 0 or 1+ are classified as negative; cases with 3+ are classified as positive; and cases with 2+ are considered ambiguous and require further detection of the amplification status of the HER2 gene by fluorescence in situ hybridization (FISH). HER2 scoring is not only used to guide clinical treatment plans, but also provides an important basis for scientific research and drug development related to breast cancer. Through large-scale analysis of HER2 status, the clinical effects of new targeted drugs can be evaluated and personalized treatment plans can be continuously optimized. However, traditional HER2 evaluation methods mainly rely on the subjective judgment of pathologists and have significant limitations. On the one hand, differences between and within observers can affect the consistency and reliability of diagnostic results; on the other hand, facing high-resolution full-field digital slice images (Whole Slide Image, WSI) and complex tumor tissue structures, traditional methods are difficult to achieve standardized and quantitative analysis.

[0003] With the rapid development of digital pathology and artificial intelligence technology, the use of deep learning algorithms to score HER2 can overcome the problems of traditional methods to a certain extent. There are currently three mainstream technologies: the first is a method based on small slice (patch) classification; this type of method divides the WSI into smaller patches, and then uses machine learning or deep learning algorithms to extract features and classify each patch. It can handle complex cell structures and staining patterns, but requires a large amount of small patch image annotation data to train the model. The second is a method based on multiple instance learning (MIL), which aims to solve the problem of lack of fine-grained annotation for HER2 scoring. This method uses the entire WSI image annotation and combines the multiple instance learning paradigm to extract the overall representation of the pathological image. TalhaQaiser et al. (TalhaQaiser and Nasir M Rajpoot. “Learningwhere to see: a novel attention model forautomatedimmunohistochemicalscoring”. In:IEEE transactions on medical imaging38.11 (2019), pp. 2620–2631.) introduced reinforcement learning into the HER2 scoring task. By imitating the way pathologists analyze tissue sections, they can intelligently select the most diagnostically valuable areas from large-size full-slide images for analysis instead of processing all sub-image blocks to predict HER2 scores. Wenqi Lu et al. (Wenqi Lu et al. “SlideGraph+: Whole slideimage level graphs to predict HER2 status inbreast cancer”. In: MedicalImageAnalysis 80 (2022), p. 102486.) used a graph neural network to construct node features using cell nucleus features, DAB density features, and deep features to directly predict HER2 scores from H&E stained WSI. In addition, SahanYorucSelcuk et al. (SahanYorucSelcuk et al. “Automated HER2 Scoring inBreast Cancer Images UsingDeep Learning and Pyramid Sampling”. In: BMEF (BMEFrontiers) (2024).) proposed a pyramid sampling method to introduce low-resolution images, thereby focusing on large-scale features such as tissue hierarchical structure and tumor area to help more accurately predict HER2 scores.The last type is segmentation-based methods, which focus on improving the interpretability of scores by accurately identifying and segmenting regions of interest, such as cell nuclei and tumor regions; this method can analyze the characteristics of specific structures in more detail, thereby providing more accurate biomarker expression assessments.

[0004] However, the three existing methods also have different shortcomings. The first method based on patch classifiers requires annotations for each patch in order to train the classifier, and this process requires the participation of professional pathologists, but the cost of such fine annotations is extremely high, especially when qualified pathologists are scarce. The second method based on MIL can overcome the high cost of pathological image annotation to a certain extent, but because MIL is trained based on the annotations of the entire WSI as supervisory information, it needs to rely on large-scale annotated data to obtain better results, and it is usually difficult to explain the results. Therefore, in the HER2 score prediction task, the number of samples of the existing public strictly annotated HER2 score dataset is very small, and the MIL model trained on such a dataset usually faces problems such as poor generalization performance. Finally, the segmentation-based method also consumes a lot of resources to provide fine tissue area and cell nucleus annotations, and even consumes more resources than the first method, because a patch image often contains multiple cell nuclei, and providing annotations for each cell nucleus will take more time. Summary of the invention

[0005] The main purpose of the present invention is to overcome the shortcomings and deficiencies of the prior art and to provide a breast cancer HER2 prediction method based on multiple prompt learning, which aims to use the visual description of pathological priors to guide the aggregation of image features, simulate the actual situation of pathologists when performing HER2 scoring tasks, and help improve the accuracy of HER2 scoring.

[0006] In order to achieve the above object, the present invention adopts a breast cancer HER2 prediction method based on multiple prompt learning, comprising the following steps: Obtain breast cancer WSI pathology images and divide them into multiple patches, and use image encoder to extract image features of breast cancer WSI pathology images; A large language model is used to generate instance-level category descriptive text and instance-level staining degree descriptive text; slice-level descriptive text is generated by referring to and adopting the description in the standard clinical evaluation guidelines; The instance-level category descriptive text, instance-level coloring degree descriptive text, slice-level descriptive text and their corresponding category tokens and learnable prompts are input into the multimodal pre-trained model for assembly, and the instance-level category prompts, instance-level coloring degree prompts and slice-level prompts are obtained through word segmentation and word embedding processing; Use a text encoder to perform integration operations on instance-level category hints, instance-level coloring hints, and slice-level hints, respectively, to obtain instance-level category hint features, instance-level coloring hint features, and slice-level hint features; The instance-level category hint features and image features, the instance-level coloring hint features and image features are respectively input into the text-guided feature aggregation module for aggregation to obtain the category global visual features and the coloring degree global visual features; The category global visual features and the staining degree global visual features are input into the global feature fusion module for integration to obtain the HER2 features of the breast cancer WSI pathology image; The cosine similarity between the slice-level cue features and the HER2 features of breast cancer WSI pathology images was calculated to obtain the HER2 prediction results of breast cancer WSI pathology images and optimized using cross entropy loss.

[0007] As a preferred technical solution, the obtaining of instance-level category hints, instance-level staining degree hints and slice-level hints is specifically as follows: The category tokens corresponding to the instance-level category descriptive text, the instance-level coloring degree descriptive text and the slice-level descriptive text are obtained respectively and assembled to obtain the instance-level category descriptive text, the instance-level coloring degree descriptive text and the slice-level descriptive text with category information; The learnable prompts corresponding to the instance-level category descriptive text, instance-level staining degree descriptive text and slice-level descriptive text are initialized by fixed templates respectively; The instance-level category descriptive text with category information, the instance-level coloring degree descriptive text and the slice-level descriptive text are segmented, and the corresponding learnable prompts are word embedded to obtain instance-level category prompts, instance-level coloring degree prompts and slice-level prompts.

[0008] As a preferred technical solution, the execution integration operation is specifically as follows: Use a text encoder to encode instance-level category hints, instance-level coloring hints, and slice-level hints, respectively, to obtain instance-level category hint encoding, instance-level coloring hint encoding, and slice-level hint encoding; The instance-level category hint coding, instance-level coloring hint coding and slice-level hint coding are integrated respectively. k The average operation is performed on the dimensions to obtain instance-level category hint features, instance-level staining degree hint features, and slice-level hint features.

[0009] As a preferred technical solution, the text-guided feature aggregation module includes a cross-attention layer and an output layer; The obtained category global visual features and staining degree global visual features are specifically: The instance-level category hint feature and image feature, the instance-level coloring hint feature and image feature are input into the cross attention layer to obtain the first attention feature and the second attention feature respectively; The instance-level category hint feature and the first attention feature, the instance-level coloring degree hint feature and the second attention feature are respectively sent to the output layer to obtain the category global visual feature and the coloring degree global visual feature.

[0010] As a preferred technical solution, the global feature fusion module includes an input layer, a cross attention layer, a mean layer and an output layer; The HER2 features of the WSI pathological image of breast cancer are obtained as follows: Send the category global visual feature and the color intensity global visual feature into the input layer and concatenate them in the quantity dimension to obtain the global visual feature; A new learnable category query token is defined for the slice-level hint feature, and is input into the cross-attention layer with the global visual feature to obtain the third attention feature; Use the mean layer to average the global visual features in the first dimension to obtain the global average visual features; The third attention feature and the global average visual feature are input into the output layer to generate the HER2 feature of the breast cancer WSI pathological image.

[0011] As a preferred technical solution, the HER2 prediction result of the WSI pathological image of breast cancer is obtained as follows: Calculate the cosine similarity between the slice-level cue features and the HER2 features of breast cancer WSI pathology images; The HER2 classification probability of breast cancer WSI pathology images was calculated based on cosine similarity, and the HER2 prediction results of breast cancer WSI pathology images were obtained; The parameters of the learnable hints, the context-guided feature aggregation module, and the global feature fusion module are optimized using cross-entropy loss.

[0012] On the other hand, the present invention provides a breast cancer HER2 prediction system with multiple prompt learning, which is applied to the above-mentioned breast cancer HER2 prediction method, and includes an image feature extraction module, a description text generation module, a multiple prompt generation module, a multiple prompt representation module, a multiple prompt aggregation module, a multiple prompt integration module and a HER2 prediction module; The image feature extraction module is used to obtain a breast cancer WSI pathology image and divide it into multiple patches, and use an image encoder to extract image features of the breast cancer WSI pathology image; The descriptive text generation module is used to generate instance-level category descriptive text and instance-level staining degree descriptive text using a large language model; and to generate slice-level descriptive text by referring to and using the description in the standard clinical assessment guide; The multiple prompt generation module is used to assemble instance-level category descriptive text, instance-level coloring degree descriptive text, slice-level descriptive text and their corresponding category tokens and learnable prompts into a multimodal pre-training model, and obtain instance-level category prompts, instance-level coloring degree prompts and slice-level prompts through word segmentation and word embedding processing; The multiple prompt representation module is used to use the text encoder to perform integration operations on the instance-level category prompt, the instance-level coloring degree prompt and the slice-level prompt respectively, so as to obtain the instance-level category prompt feature, the instance-level coloring degree prompt feature and the slice-level prompt feature; The multiple prompt aggregation module is used to input the instance-level category prompt feature and image feature, the instance-level coloring degree prompt feature and image feature into the text-guided feature aggregation module for aggregation, so as to obtain the category global visual feature and the coloring degree global visual feature; The multiple prompt integration module is used to input the category global visual features and the staining degree global visual features into the global feature fusion module for integration to obtain the HER2 features of the breast cancer WSI pathological image; The HER2 prediction module is used to calculate the cosine similarity between the slice-level prompt features and the HER2 features of the breast cancer WSI pathology image, obtain the HER2 prediction results of the breast cancer WSI pathology image and optimize them using the cross entropy loss.

[0013] In another aspect, the present invention provides a computer-readable storage medium storing a program, which, when executed by a processor, implements the above-mentioned breast cancer HER2 prediction method.

[0014] Another aspect of the present invention provides an electronic device, comprising: at least one processor; and a memory in communication with the at least one processor; wherein, The memory stores computer program instructions that can be executed by the at least one processor, and the computer program instructions are executed by the at least one processor so that the at least one processor can perform the above-mentioned breast cancer HER2 prediction method.

[0015] Compared with the prior art, the present invention has the following advantages and beneficial effects: 1. The method of the present invention can be trained in a parameter-efficient manner, reducing the reliance on large-scale data sets and improving the generalization performance of the model. The present invention uses a pre-trained large language model to generate multiple prompt descriptive texts, and then uses a multimodal pre-trained model to embed learnable prompts into the text to generate multiple prompts. At the same time, a text-guided feature aggregation module is designed to aggregate WSI-level visual features related to the text, and finally the WSI-level visual features related to the category and staining degree are integrated through a global feature fusion module to obtain a WSI-level HER2 feature representation.

[0016] 2. The method of the present invention can effectively capture semantic information related to text descriptions, imitate the process of pathologists evaluating HER2, guide the model to focus on tissue areas of different types and staining degrees through instance-level category prompts and instance-level staining degree prompts, and use the attention matrix of the intermediate calculated text prompt features and image features to generate a heat map, thereby improving the interpretability of the model. BRIEF DESCRIPTION OF THE DRAWINGS

[0017] In order to more clearly illustrate the technical solutions in the embodiments of the present application, the drawings required for use in the description of the embodiments will be briefly introduced below. Obviously, the drawings described below are only some embodiments of the present application. For ordinary technicians in this field, other drawings can be obtained based on these drawings without creative work.

[0018] Figure 1 The figure is a flow chart of a method for predicting HER2 in breast cancer by multiple prompt learning according to an embodiment of the present invention.

[0019] Figure 2 The figure is a flow chart of a breast cancer HER2 prediction method based on multiple prompt learning according to an embodiment of the present invention.

[0020] Figure 3 The schematic diagram is a structural diagram of a breast cancer HER2 prediction system based on multiple prompt learning according to an embodiment of the present invention.

[0021] Figure 4 Schematic diagram of the structure of a computer-readable storage medium in an embodiment of the present invention. DETAILED DESCRIPTION

[0022] In order to enable those skilled in the art to better understand the present application, the technical solutions in the embodiments of the present application will be clearly and completely described below in conjunction with the drawings in the embodiments of the present application. Obviously, the described embodiments are only part of the embodiments of the present application, rather than all of the embodiments. Based on the embodiments in the present application, all other embodiments obtained by those skilled in the art without making creative work are within the scope of protection of the present application.

[0023] Reference to "embodiments" in this application means that a particular feature, structure, or characteristic described in conjunction with the embodiments may be included in at least one embodiment of the present application. The appearance of the phrase in various locations in the specification does not necessarily refer to the same embodiment, nor is it an independent or alternative embodiment that is mutually exclusive with other embodiments. It is explicitly and implicitly understood by those skilled in the art that the embodiments described in this application may be combined with other embodiments.

[0024] The method proposed in this paper aims to use the visual description of pathology priors to guide the aggregation of image features, simulating the actual situation of pathologists when performing HER2 scoring tasks; these visual descriptions are generated by a large language model (LLM) and contain morphological information related to tissue region categories and staining levels. Through these descriptions, the model can better focus on visual features related to the task. See Figure 1 , Figure 2 In one embodiment of the present application, a method for predicting breast cancer HER2 using multiple prompt learning is provided, comprising the following steps: S1. Obtain breast cancer WSI pathology images and divide them into multiple patches (Patches), using image encoder E I Extracting image features of breast cancer WSI pathological images ,in I n For the n patches, M is the number of patches segmented from the breast cancer WSI pathology image, d is the feature dimension. In this embodiment, the image encoder E I The ViT-B / 32 (Vision Transformer) structure in the pre-trained CLIP model is used for implementation. The parameters of the image encoder need to be frozen during the operation ( Figure 2 snowflake logo).

[0025] S2. Use a large language model to generate instance-level category descriptive text and instance-level staining degree descriptive text; refer to and use the description in the standard clinical evaluation guidelines to generate slice-level descriptive text.

[0026] In this embodiment, the large language model uses GPT-4 to generate instance-level category descriptive text and instance-level staining degree descriptive text; specifically, through question-answering interaction with GPT-4, for example, "Please provide 5 visual descriptions of HER2 immunohistochemical pathology images, respectively for tissue category / staining degree", multiple descriptive texts are generated to increase diversity and richness. For slice-level descriptive text, the description in the standard clinical evaluation guide (Antonio C Wolff et al. "Human epidermal growth factor receptor 2 testing in breast cancer cer: American Society of Clinical Oncology / College of American Pathologistsclinical practice guidelinefocused update". In: Archives of pathology&laboratory medicine 142.11 (2018), pp. 1364–1382.) is referenced and used for generation.

[0027] S3. Input instance-level category descriptive text, instance-level coloring degree descriptive text, slice-level descriptive text and their corresponding category tokens and learnable prompts into the multimodal pre-trained model for assembly, and obtain instance-level category prompts, instance-level coloring degree prompts and slice-level prompts through word segmentation and word embedding processing.

[0028] Furthermore, the present invention uses a multimodal pre-training model to generate multiple prompts, specifically: S3.1. Obtain the category tokens corresponding to the instance-level category descriptive text, the instance-level staining degree descriptive text, and the slice-level descriptive text respectively and assemble them to obtain the instance-level category descriptive text, the instance-level staining degree descriptive text, and the slice-level descriptive text with category information. For example, assemble according to the following template: “Animmunohistochemical pathological image of <cls>, which is <description>", where cls is the category token and description is the descriptive text. The category tokens of the instance-level category descriptive text include Normal, Tumor, Lymphocyte, and Stroma; the category tokens of the instance-level staining degree descriptive text include None, Weak, Moderate, and Strong; the category tokens of the slice-level descriptive text are HER2 scores, including 0, 1+, 2+, and 3+.

[0029] S3.2. Initialize the instance-level category descriptive text, instance-level staining degree descriptive text, and slice-level descriptive text to obtain learnable hints corresponding to them through fixed templates p In this embodiment, the fixed template adopts the above-mentioned assembly template.

[0030] S3.3. Perform word segmentation on the instance-level category descriptive text with category information, the instance-level coloring degree descriptive text, and the slice-level descriptive text, and perform word embedding on the corresponding learnable prompts to obtain instance-level category prompts. , instance-level coloring hint And slice-level hints ,in N c is the number of category tokens corresponding to the instance-level category descriptive text, N s is the number of category tokens corresponding to the instance-level coloring descriptive text, k is the number of corresponding descriptive texts, L is the text length of the corresponding prompt, d is the embedding dimension. The prompt corresponding to each descriptive text can be expressed as: T =[ p 1, p 2, ..., p m ; cls ; x 1, x 2, ..., x n ], in, p 1, p 2, ..., p m represents a learnable hint corresponding to a descriptive text, m is the number of learnable cues, cls is the category token corresponding to the descriptive text, x 1, x 2, ..., x n is the word embedding obtained by converting the corresponding descriptive text, n is the number of word embeddings, and satisfies m + n +1= L .

[0031] In this embodiment, the multimodal pre-trained model uses the CLIP model pre-trained on 400 million image-text pairs to generate multiple prompts, where k =5, L =77.

[0032] S4. Use a text encoder to perform integration operations on instance-level category hints, instance-level coloring degree hints, and slice-level hints, respectively, to obtain instance-level category hint features, instance-level coloring degree hint features, and slice-level hint features.

[0033] In order to obtain a more comprehensive and complete semantic representation, an integration operation is performed on the generated multiple prompts, specifically: S4.1. Using a text encoder E T Encode the instance-level category hint, instance-level staining hint, and slice-level hint respectively to obtain the instance-level category hint encoding , Instance-level coloring hint encoding And slice-level hint encoding In this embodiment, the text encoder E T Use the Transformer structure in the pre-trained CLIP model.

[0034] S4.2. Ensemble the instance-level category hint coding, instance-level coloring hint coding, and slice-level hint coding, respectively. k Perform an average operation on the dimension to obtain instance-level category hint features , instance-level coloring level hint features And slice-level hint features , l c Indicates the category in the slice level hint c Among them, , mean () represents the average operation, which is an integration method.

[0035] S5. Input the instance-level category hint feature and image feature, the instance-level coloring degree hint feature and image feature into the text-guided feature aggregation module for aggregation, respectively, to obtain the category global visual feature and the coloring degree global visual feature.

[0036] However, there are still significant differences between the initialized IHC pathology image features and the prompt features; the method of calculating the aggregation weights by the dot product of image and text features may introduce additional interference, thereby weakening the performance of the prompt-based pooling strategy. Therefore, the present invention designs a text-guided feature aggregation module (Context-guided Feature Aggregator, CFA), which is a variant structure based on Transformer, which uses prompt features to guide image features to aggregate into prompt-related global visual features, including a cross-attention layer and an output layer; wherein the cross-attention layer is used to calculate the cross-attention between the prompt features and the image features to generate attention features; the output layer is used to add the attention features to the prompt features, and obtain the global visual features through the MLP layer and the residual connection. Thus, the steps of obtaining the category global visual features and the staining degree global visual features are specifically as follows: S5.1. Use the cross attention layer to calculate the cross attention between the instance-level category hint feature and the image feature, and between the instance-level coloring hint feature and the image feature, and obtain the first attention feature and the second attention feature. The process is described as: , , in, is the first attention feature, is the second attention feature, is the learnable projection layer parameter matrix, initialized to the identity matrix , to prevent the image features and hint features extracted by the image encoder from being destroyed in the initial stage.

[0037] S5.2, input the instance-level category hint feature and the first attention feature, the instance-level coloring hint feature and the second attention feature into the output layer respectively, and obtain the category global visual feature and the coloring degree global visual feature. The process is expressed as: , , in, X cls is the category global visual feature, X sin is the global visual feature of the degree of staining, MLP A multi-layer perceptron.

[0038] S6. Input the category global visual features and the staining degree global visual features into the global feature fusion module for integration to obtain the HER2 features of the breast cancer WSI pathology image.

[0039] Since the HER2 score focuses on the staining expression in the invasive cancer area, the present invention designs an effective global feature fusion module (Global Feature Fusion, GFF) to integrate the global visual features of each category and each staining degree into the HER2 feature of WSI; the global feature fusion module (GFF) includes an input layer, a cross-attention layer, a mean layer and an output layer; wherein the input layer is used to splice the global visual features of each category and each staining degree in the quantity dimension to obtain the global visual feature; the cross-attention layer is used to perform cross-attention calculations on the learnable category query token and the global visual feature to obtain the attention feature; the mean layer is used to average the global visual feature in the quantity dimension to obtain the average feature; the output layer is used to add the attention feature and the average feature, and connect them with the residual through the MLP layer to obtain the HER2 feature of WSI. Therefore, the acquisition process of the HER2 feature of the breast cancer WSI pathological image is as follows: S6.1. First, the category global visual feature and the coloring degree global visual feature are sent to the input layer for concatenation in the quantity dimension to obtain the global visual feature .

[0040] S6.2, then, a new learnable category query token is defined for the slice-level hint feature. , and the global visual feature input cross attention layer into the interaction of global visual features to obtain the third attention matrix; the process is described as: , in, is the third attention matrix.

[0041] S6.3. Using the mean layer for global visual features X G In the first dimension (i.e. N c + N s dimension) to obtain the global average visual feature mean ( X G ).

[0042] S6.4. Finally, the third attention feature The global average visual feature input and output layer are added and then passed through the MLP layer to generate the HER2 feature of the breast cancer WSI pathology image. The process is described as follows: , in, HER2 features of breast cancer WSI pathological images, It is the category token in the category global visual feature.

[0043] S7. Calculate the cosine similarity between the slice-level cue features and the HER2 features of the breast cancer WSI pathology image, obtain the HER2 prediction results of the breast cancer WSI pathology image, and optimize them using the cross entropy loss.

[0044] Specifically, the process of step S7 is: S7.1. First, calculate the slice-level hint features HER2 features in WSI pathological images of breast cancer x Cosine similarity of cos ( l c , x ), cos () is the cosine similarity function.

[0045] S7.2. Next, the HER2 classification probability of the breast cancer WSI pathology image is calculated based on the cosine similarity, and the HER2 prediction result of the breast cancer WSI pathology image is obtained, which is expressed as: , in, p ( y = c | x ) indicates HER2 signature x Medium Category c The probability of τ is the temperature coefficient used to control the slice-level hint feature With HER2 characteristics x Scale of similarity.

[0046] S7.3. Use cross entropy loss to optimize the parameters of the learnable prompt, text-guided feature aggregation module and the global feature fusion module; the cross entropy loss function is: L = - logp ( y = c | x ).

[0047] The method designed in this invention was tested on the private dataset ZJH-HER2 and the public dataset HER2C (TalhaQaiser et al. "Her 2 challenge contest: a detailed assessment ofautomated her 2 scoring algorithmsin whole slide images of breast cancer tissues". In: Histopathology 72.2 (2018), pp. 227–238). Five MIL-based pathology image classification methods were selected for comparison, namely: ABMIL (Maximilian Ilse, JakubTomczak, and Max Welling. "Attention-based deep multiple instancelearning". In: International conference onmachine learning. PMLR. 2018, pp. 2127–2136) is an early and well-known attention-based MIL method that aggregates image features by utilizing the attention mechanism.

[0048] CLAM (Ming Y Lu et al. "Data-efficient and weakly supervised computational pathology on whole-slide images". In: Nature biomedicalengineering 5.6 (2021), pp. 555–570.) is extended to multi-category problems based on ABMIL and introduces instance-level clustering constraints to optimize the feature space.

[0049] TransMIL (Zhuchen Shao et al. "Transmil: Transformer based correlatedmultiple instance learning for whole slide imageclassification". In: Advancesin neural information processing systems 34 (2021), pp. 2136–2147.) applies the self-attention mechanism to the pathological image classification task, further improving the performance of the MIL-based method.

[0050] TOP (Linhao Qu et al. "The rise of ai language pathologists: Exploring two-level prompt learn ing for few-shot weakly-supervised whole slide image classification". In: Advances in Neural Information Processing Systems 36 (2024).) uses a pre-trained visual-language model to incorporate prior text information into the image feature aggregation process, providing richer semantic information, especially in small sample scenarios, which can improve performance.

[0051] ViLa-MIL (Jiangbo Shi et al. "ViLa-MIL: Dual-scale Vision-LanguageMultiple Instance Learning for Whole Slide Image Classification". In:Proceedings of the IEEE / CVF Conference on Computer Vision and Pattern Recognition. 2024, pp. 11248–11258.), further enhances the performance of vision-language models by using dual-scale visual description text cues based on a frozen large language model (LLM).

[0052] In order to adapt TOP and ViLa-MIL to the HER2 scoring task, their original prompt templates were modified in this experiment. However, these two methods still follow the methods described in their respective papers and use LLM to generate prior text descriptions. In order to be consistent with the ViLa-MIL method, image features are extracted at 5× and 10× magnifications. The final results are shown in Table 1: Table 1 Experimental results

[0053] In this experiment, ACC and F1 Score were used to evaluate the performance of each method. As shown in Table 1, compared with the other five existing methods, our method (Ours) achieved 79.2% ACC and 77.6% F1-Score on the ZJH-HER2 dataset, and 85.7% ACC and 87.1% F1-Score on the HER2C dataset, which are significantly better than the comparison methods. It is worth noting that the prompt learning-based method (using CLIP-I and CLIP-I&T as the backbone network) is overall better than the traditional attention-based MIL method (using ResNet50 as the backbone network), which is mainly due to the text-guided image feature aggregation module, which injects rich semantic context by introducing prior text information and further aligns information between different modalities. Although TOP and ViLa-MIL also use text prompts for feature aggregation, our method performs better; this improvement comes from the multi-prompt learning strategy proposed by our method, which guides the model to focus more accurately on regions of different tissue types and staining intensities by decoupling category prompts from staining degree prompts. In addition, this method also combines the global feature fusion module to combine the staining degree information with the category information, effectively simulating the decision-making process of pathologists when scoring.

[0054] It should be noted that, for the sake of convenience, the aforementioned method embodiments are all expressed as a series of action combinations, but those skilled in the art should know that the present invention is not limited to the described order of actions, because according to the present invention, certain steps can be performed in other orders or simultaneously.

[0055] Based on the same idea as the breast cancer HER2 prediction method of multiple prompt learning in the above embodiment, the present invention also provides a breast cancer HER2 prediction system of multiple prompt learning, which can be used to execute the above breast cancer HER2 prediction method of multiple prompt learning. For ease of explanation, the structural schematic diagram of the embodiment of the breast cancer HER2 prediction system of multiple prompt learning only shows the parts related to the embodiment of the present invention. Those skilled in the art can understand that the illustrated structure does not constitute a limitation on the device, and may include more or fewer components than the illustrated one, or combine certain components, or arrange the components differently.

[0056] See also Figure 3 , in another embodiment of the present application, a breast cancer HER2 prediction system with multiple prompt learning is provided, the system comprising an image feature extraction module, a description text generation module, a multiple prompt generation module, a multiple prompt representation module, a multiple prompt aggregation module, a multiple prompt integration module and a HER2 prediction module; The image feature extraction module is used to obtain the breast cancer WSI pathology image and divide it into multiple patches, and use the image encoder to extract the image features of the breast cancer WSI pathology image; The descriptive text generation module is used to generate instance-level category descriptive text and instance-level staining degree descriptive text using a large language model; and to generate slice-level descriptive text by referring to and using the description in the standard clinical evaluation guidelines; The multiple prompt generation module is used to assemble instance-level category descriptive text, instance-level coloring degree descriptive text, slice-level descriptive text and their corresponding category tokens and learnable prompts into the multimodal pre-trained model, and obtain instance-level category prompts, instance-level coloring degree prompts and slice-level prompts through word segmentation and word embedding processing; The multiple prompt representation module is used to use the text encoder to perform integration operations on the instance-level category prompt, the instance-level coloring degree prompt and the slice-level prompt respectively, so as to obtain the instance-level category prompt feature, the instance-level coloring degree prompt feature and the slice-level prompt feature; The multiple prompt aggregation module is used to input the instance-level category prompt features and image features, the instance-level coloring degree prompt features and image features into the text-guided feature aggregation module for aggregation, so as to obtain the category global visual features and the coloring degree global visual features; The multiple prompt integration module is used to input the category global visual features and the staining degree global visual features into the global feature fusion module for integration to obtain the HER2 features of the breast cancer WSI pathological image; The HER2 prediction module is used to calculate the cosine similarity between the slice-level cue features and the HER2 features of breast cancer WSI pathology images, obtain the HER2 prediction results of breast cancer WSI pathology images and optimize them using cross entropy loss.

[0057] It should be noted that a breast cancer HER2 prediction system with multiple prompt learning of the present invention corresponds one to one with a breast cancer HER2 prediction method with multiple prompt learning of the present invention. The technical features and beneficial effects described in the embodiment of the above-mentioned breast cancer HER2 prediction method with multiple prompt learning are applicable to the embodiment of breast cancer HER2 prediction with multiple prompt learning. For specific contents, please refer to the description in the embodiment of the method of the present invention, which will not be repeated here. This is hereby declared.

[0058] In addition, in the implementation of the breast cancer HER2 prediction system with multiple prompts learning in the above embodiment, the logical division of each program module is only an example. In actual application, the above functions can be assigned to different program modules as needed, for example, for the convenience of corresponding hardware configuration requirements or software implementation. That is, the internal structure of the breast cancer HER2 prediction system with multiple prompts learning is divided into different program modules to complete all or part of the functions described above.

[0059] See also Figure 4 In another embodiment of the present invention, a computer-readable storage medium is provided, wherein a program is stored. When the program is executed by a processor, the above-mentioned breast cancer HER2 prediction method based on multiple prompt learning can be implemented, specifically: Obtain breast cancer WSI pathology images and divide them into multiple patches, and use image encoder to extract image features of breast cancer WSI pathology images; A large language model is used to generate instance-level category descriptive text and instance-level staining degree descriptive text; slice-level descriptive text is generated by referring to and adopting the description in the standard clinical evaluation guidelines; The instance-level category descriptive text, instance-level coloring degree descriptive text, slice-level descriptive text and their corresponding category tokens and learnable prompts are input into the multimodal pre-trained model for assembly, and the instance-level category prompts, instance-level coloring degree prompts and slice-level prompts are obtained through word segmentation and word embedding processing; Use a text encoder to perform integration operations on instance-level category hints, instance-level coloring hints, and slice-level hints, respectively, to obtain instance-level category hint features, instance-level coloring hint features, and slice-level hint features; The instance-level category hint features and image features, the instance-level coloring hint features and image features are respectively input into the text-guided feature aggregation module for aggregation to obtain the category global visual features and the coloring degree global visual features; The category global visual features and the staining degree global visual features are input into the global feature fusion module for integration to obtain the HER2 features of the breast cancer WSI pathology image; The cosine similarity between the slice-level cue features and the HER2 features of breast cancer WSI pathology images was calculated to obtain the HER2 prediction results of breast cancer WSI pathology images and optimized using cross entropy loss.

[0060] In another embodiment, the present application provides an electronic device for implementing the above-mentioned breast cancer HER2 prediction method of multiple prompt learning, comprising at least one processor, and a memory connected to the at least one processor in communication; wherein the memory stores computer program instructions executable by the at least one processor, and when the computer program instructions are executed by the at least one processor, the at least one processor is enabled to execute the above-mentioned breast cancer HER2 prediction method of multiple prompt learning, specifically: Obtain breast cancer WSI pathology images and divide them into multiple patches, and use image encoder to extract image features of breast cancer WSI pathology images; A large language model is used to generate instance-level category descriptive text and instance-level staining degree descriptive text; slice-level descriptive text is generated by referring to and adopting the description in the standard clinical evaluation guidelines; The instance-level category descriptive text, instance-level coloring degree descriptive text, slice-level descriptive text and their corresponding category tokens and learnable prompts are input into the multimodal pre-trained model for assembly, and the instance-level category prompts, instance-level coloring degree prompts and slice-level prompts are obtained through word segmentation and word embedding processing; Use a text encoder to perform integration operations on instance-level category hints, instance-level coloring hints, and slice-level hints, respectively, to obtain instance-level category hint features, instance-level coloring hint features, and slice-level hint features; The instance-level category hint features and image features, the instance-level coloring hint features and image features are respectively input into the text-guided feature aggregation module for aggregation to obtain the category global visual features and the coloring degree global visual features; The category global visual features and the staining degree global visual features are input into the global feature fusion module for integration to obtain the HER2 features of the breast cancer WSI pathology image; The cosine similarity between the slice-level cue features and the HER2 features of breast cancer WSI pathology images was calculated to obtain the HER2 prediction results of breast cancer WSI pathology images and optimized using cross entropy loss.

[0061] Those skilled in the art can understand that all or part of the processes in the above-mentioned embodiment methods can be completed by instructing the relevant hardware through a computer program, and the program can be stored in a non-volatile computer-readable storage medium. When the program is executed, it can include the processes of the embodiments of the above-mentioned methods. Among them, any reference to memory, storage, database or other media used in the embodiments provided in this application can include non-volatile and / or volatile memory. Non-volatile memory may include read-only memory (ROM), programmable ROM (PROM), electrically programmable ROM (EPROM), electrically erasable programmable ROM (EEPROM) or flash memory. Volatile memory may include random access memory (RAM) or external cache memory. As an illustration and not limitation, RAM is available in many forms, such as static RAM (SRAM), dynamic RAM (DRAM), synchronous DRAM (SDRAM), double data rate SDRAM (DDRSDRAM), enhanced SDRAM (ESDRAM), synchronous link (Synchlink) DRAM (SLDRAM), memory bus (Rambus) direct RAM (RDRAM), direct memory bus dynamic RAM (DRDRAM), and memory bus dynamic RAM (RDRAM).

[0062] The technical features of the above embodiments may be combined arbitrarily. To make the description concise, not all possible combinations of the technical features in the above embodiments are described. However, as long as there is no contradiction in the combination of these technical features, they should be considered to be within the scope of this specification.

[0063] The above embodiments are preferred implementation modes of the present invention, but the implementation modes of the present invention are not limited to the above embodiments. Any other changes, modifications, substitutions, combinations, and simplifications that do not deviate from the spirit and principles of the present invention should be equivalent replacement methods and are included in the protection scope of the present invention.< / description> < / cls>

Claims

1. A breast cancer HER2 prediction method based on multiple prompt learning, characterized in that: The steps include: Obtain breast cancer WSI pathology images and divide them into multiple patches, and use image encoder to extract image features of breast cancer WSI pathology images; A large language model is used to generate instance-level category descriptive text and instance-level staining degree descriptive text; slice-level descriptive text is generated by referring to and adopting the description in the standard clinical evaluation guidelines; The instance-level category descriptive text, instance-level coloring degree descriptive text, slice-level descriptive text and their corresponding category tokens and learnable prompts are input into the multimodal pre-trained model for assembly, and the instance-level category prompts, instance-level coloring degree prompts and slice-level prompts are obtained through word segmentation and word embedding processing; Use a text encoder to perform integration operations on instance-level category hints, instance-level coloring hints, and slice-level hints, respectively, to obtain instance-level category hint features, instance-level coloring hint features, and slice-level hint features; The instance-level category hint features and image features, the instance-level coloring hint features and image features are respectively input into the text-guided feature aggregation module for aggregation to obtain the category global visual features and the coloring degree global visual features; The category global visual features and the staining degree global visual features are input into the global feature fusion module for integration to obtain the HER2 features of the breast cancer WSI pathology image; The cosine similarity between the slice-level cue features and the HER2 features of breast cancer WSI pathology images was calculated to obtain the HER2 prediction results of breast cancer WSI pathology images and optimized using cross entropy loss.

2. The method for predicting HER2 in breast cancer according to claim 1, characterized in that: The obtaining of instance-level category hints, instance-level staining degree hints and slice-level hints is specifically as follows: The category tokens corresponding to the instance-level category descriptive text, the instance-level coloring degree descriptive text and the slice-level descriptive text are obtained respectively and assembled to obtain the instance-level category descriptive text, the instance-level coloring degree descriptive text and the slice-level descriptive text with category information; The learnable prompts corresponding to the instance-level category descriptive text, instance-level staining degree descriptive text and slice-level descriptive text are initialized by fixed templates respectively; The instance-level category descriptive text with category information, the instance-level coloring degree descriptive text and the slice-level descriptive text are segmented, and the corresponding learnable prompts are word embedded to obtain instance-level category prompts, instance-level coloring degree prompts and slice-level prompts.

3. The method for predicting HER2 in breast cancer according to claim 1, characterized in that: The execution integration operation is specifically as follows: Use a text encoder to encode instance-level category hints, instance-level coloring hints, and slice-level hints, respectively, to obtain instance-level category hint encoding, instance-level coloring hint encoding, and slice-level hint encoding; The instance-level category hint coding, instance-level coloring hint coding and slice-level hint coding are integrated respectively. k The average operation is performed on the dimensions to obtain instance-level category hint features, instance-level staining degree hint features, and slice-level hint features.

4. The method for predicting HER2 in breast cancer according to claim 1, characterized in that: The text-guided feature aggregation module includes a cross-attention layer and an output layer; The obtained category global visual features and staining degree global visual features are specifically: The instance-level category hint feature and image feature, the instance-level coloring hint feature and image feature are input into the cross attention layer to obtain the first attention feature and the second attention feature respectively; The instance-level category hint feature and the first attention feature, the instance-level coloring degree hint feature and the second attention feature are respectively sent to the output layer to obtain the category global visual feature and the coloring degree global visual feature.

5. The method for predicting HER2 in breast cancer according to claim 1, characterized in that: The global feature fusion module includes an input layer, a cross attention layer, a mean layer and an output layer; The HER2 features of the WSI pathological image of breast cancer are obtained as follows: Send the category global visual feature and the color intensity global visual feature into the input layer and concatenate them in the quantity dimension to obtain the global visual feature; A new learnable category query token is defined for the slice-level hint feature, and is input into the cross-attention layer with the global visual feature to obtain the third attention feature; Use the mean layer to average the global visual features in the first dimension to obtain the global average visual features; The third attention feature and the global average visual feature are input into the output layer to generate the HER2 feature of the breast cancer WSI pathological image.

6. The method for predicting HER2 in breast cancer according to claim 1, characterized in that: The HER2 prediction result of the WSI pathological image of breast cancer is obtained as follows: Calculate the cosine similarity between the slice-level cue features and the HER2 features of breast cancer WSI pathology images; The HER2 classification probability of breast cancer WSI pathology images was calculated based on cosine similarity, and the HER2 prediction results of breast cancer WSI pathology images were obtained; The parameters of the learnable hints, the context-guided feature aggregation module, and the global feature fusion module are optimized using cross-entropy loss.

7. A breast cancer HER2 prediction system based on multiple prompt learning, characterized in that: The breast cancer HER2 prediction method applied to any one of claims 1-6 comprises an image feature extraction module, a description text generation module, a multiple prompt generation module, a multiple prompt representation module, a multiple prompt aggregation module, a multiple prompt integration module and a HER2 prediction module; The image feature extraction module is used to obtain a breast cancer WSI pathology image and divide it into multiple patches, and use an image encoder to extract image features of the breast cancer WSI pathology image; The descriptive text generation module is used to generate instance-level category descriptive text and instance-level staining degree descriptive text using a large language model; and to generate slice-level descriptive text by referring to and using the description in the standard clinical assessment guide; The multiple prompt generation module is used to assemble instance-level category descriptive text, instance-level coloring degree descriptive text, slice-level descriptive text and their corresponding category tokens and learnable prompts into a multimodal pre-training model, and obtain instance-level category prompts, instance-level coloring degree prompts and slice-level prompts through word segmentation and word embedding processing; The multiple prompt representation module is used to use the text encoder to perform integration operations on the instance-level category prompt, the instance-level coloring degree prompt and the slice-level prompt respectively, so as to obtain the instance-level category prompt feature, the instance-level coloring degree prompt feature and the slice-level prompt feature; The multiple prompt aggregation module is used to input the instance-level category prompt feature and image feature, the instance-level coloring degree prompt feature and image feature into the text-guided feature aggregation module for aggregation, so as to obtain the category global visual feature and the coloring degree global visual feature; The multiple prompt integration module is used to input the category global visual features and the staining degree global visual features into the global feature fusion module for integration to obtain the HER2 features of the breast cancer WSI pathological image; The HER2 prediction module is used to calculate the cosine similarity between the slice-level prompt features and the HER2 features of the breast cancer WSI pathology image, obtain the HER2 prediction results of the breast cancer WSI pathology image and optimize them using the cross entropy loss.

8. A computer-readable storage medium storing a program, characterized in that: When the program is executed by a processor, the breast cancer HER2 prediction method according to any one of claims 1 to 6 is implemented.

9. An electronic device, characterized in that: include: at least one processor; and, a memory communicatively connected to the at least one processor; wherein, The memory stores computer program instructions that can be executed by the at least one processor, and the computer program instructions are executed by the at least one processor so that the at least one processor can execute the breast cancer HER2 prediction method as described in any one of claims 1-6.

Citation Information

Patent Citations

  • Breast cancer molecular typing prediction method under guidance of text information

    CN117995386A

  • Breast cancer classification method for WSI image and electronic equipment

    CN118918386A

  • Fine-grained multi-mode prompt learning method based on visual language pre-training model

    CN119538179A

  • Prediction model construction method based on weak supervision multi-modal contrast learning and breast cancer HER2 score prediction method

    CN119579540A

  • Pathological section image analysis method and device based on large language model

    CN119648625A

Cited By

  • HER2 state prediction method based on pathological full-slice feature learning algorithm

    CN120726027A

  • Multi-modal deep learning fusion model construction method for breast cancer

    CN121439245A

  • Breast cancer pathology image classification method and system based on multi-layer attention

    CN122416447A