A postoperative pathological image analysis method based on text language prompts

By preprocessing and segmenting whole-slice scan images of pathology, and combining them with a visual-linguistic multimodal model to generate multimodal features, the problems of spatial information loss and cross-device adaptation difficulties in traditional methods are solved, and high-precision pathological analysis is achieved.

CN120782767BActive Publication Date: 2025-11-07HUNAN UNIV
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202511262812.8
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2025-09-05
Publication Date
2025-11-07
Estimated Expiration
2045-09-05

AI Technical Summary

Technical Problem

Traditional deep learning methods suffer from problems such as loss of spatial information, insufficient utilization of semantic knowledge, and difficulty in cross-device adaptation in the analysis of whole-slice scan images of pathology, resulting in insufficient ability of the model to recognize rare pathological features.

Method used

A text-based language prompting method is used to preprocess and segment whole-section scan images of pathology, extract local and global image features, generate multimodal features by combining a visual-language multimodal model, and perform lesion analysis through image and text attention heatmaps.

Benefits of technology

It achieves high-precision and robust pathological analysis in data-scarce scenarios, breaks through the limitations of a single visual modality, and provides an interpretable and interactive cross-modal fusion solution.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120782767B_ABST
    Figure CN120782767B_ABST
Patent Text Reader

Abstract

The application relates to a postoperative pathological image analysis method based on a text language prompt. The method first cuts a pathological whole-section scanning image to obtain a plurality of first image blocks and a plurality of second image blocks; extracts features of the first image blocks as local image features; extracts and weight-sums features of all the second image blocks to obtain global image features; fuses the local image features based on cosine similarity between the global image features and the local image features to obtain overall image features; inputs a pathological whole-section scanning image to be analyzed into a visual-language multimodal large model, extracts features of an output overall visual feature text description to obtain text features; then fuses the overall image features and the text features to obtain multimodal features; and inputs the multimodal features into a classifier to obtain probability distribution of each cancer type. The method provides an innovative solution of an interpretable, interactive and cross-modal fusion for postoperative pathological analysis.
Need to check novelty before this filing date? Find Prior Art

Description

TECHNICAL FIELD

[0001] The present application relates to the technical field of pathological whole slide image analysis, in particular to a postoperative pathological image analysis method based on text language prompts. BACKGROUND

[0002] Postoperative pathology recognition is a core link for evaluating surgical effect and formulating treatment plan, and is crucial for improving the quality of life of patients. Whole slide image (WSI) contains multi-scale pathological features (such as cell morphology, tissue infiltration mode, etc.), but the extraction and accurate recognition of complex tissue microenvironment features face multiple challenges: traditional deep learning methods (such as convolutional neural network) are limited by GPU memory, and need to divide the ultra-high resolution WSI into small pieces for processing, resulting in the loss of cell spatial relationship information; the lack of training data, class imbalance and high cost of expert annotation restrict the recognition and analysis ability of the model on rare pathological features; in addition, the color and contrast differences of images collected by different devices further increase the difficulty of model adaptation. SUMMARY

[0003] Therefore, in order to solve the problems of traditional methods in the extraction of ultra-high resolution WSI features, such as loss of spatial information, insufficient use of semantic knowledge and difficulty in cross-device adaptation, it is necessary to provide a postoperative pathological image analysis method based on text language prompts, comprising:

[0004] S1: different pre-processing is performed on the pathological whole slide image to be analyzed, and the first pre-processed pathological whole slide image is non-overlappingly cut to obtain a plurality of first image blocks; the second pre-processed pathological whole slide image is uniformly divided into a plurality of second image blocks;

[0005] S2: extracting the features of each first image block as the local image features of each first image block; extracting the features of all second image blocks and performing weighted summation to obtain global image features;

[0006] S3: based on the cosine similarity between the global image features and each local image feature, the local image features are weighted and summed to obtain overall image features, and an image attention heat map is generated based on each cosine similarity;

[0007] S4: inputting the pathological whole slide image to be analyzed into a visual-linguistic multimodal large model, outputting the corresponding overall visual feature text description based on the designed prompt words, and extracting the features of the overall visual feature text description to obtain text features;

[0008] S5: fusing the overall image features and the text features based on the improved Mamba hidden state space to obtain multimodal features and a text attention heat map;

[0009] S6: inputting the multi-modal features into an MLP classifier to obtain a probability distribution of each cancer type; and performing lesion analysis based on the image attention heat map and the text attention heat map.

[0010] Preferably, different pre-processing is performed on the pathological whole section scan images to be analyzed, including:

[0011] First pre-processing, performing staining normalization processing on the pathological whole section scan images to be analyzed, and performing adaptive histogram equalization on the normalized pathological whole section scan images to obtain first pre-processed pathological whole section scan images.

[0012] Second pre-processing, performing 4 times down-sampling on the pathological whole section scan images to be analyzed to obtain a thumbnail, and the thumbnail is the second pre-processed pathological whole section scan image.

[0013] Preferably, in S2, the process of obtaining the local image features includes:

[0014] Each first image block is input into a pre-trained pathological image feature extraction model to obtain a corresponding initial local feature vector;

[0015] Each first image block is input into a pre-trained cell detection model to obtain a corresponding cell density;

[0016] Each cell density is normalized,

[0017] The normalized cell density on each first image block is weighted to obtain the local image features of each first image block.

[0018] Preferably, in S2, the process of obtaining the global image features includes:

[0019] Each second image block is input into a pre-trained pathological image feature extraction model to obtain a corresponding tissue image feature vector;

[0020] Each second image block is input into a pre-trained cell detection model to obtain a corresponding second cell density;

[0021] Each second cell density is normalized,

[0022] The normalized cell density on each second image block is weighted to obtain a final tissue image feature vector of each second image block;

[0023] The final tissue image feature vectors of each second image block are summed to obtain the global image features.

[0024] Preferably, in S3, generating the image attention heat map based on the cosine similarities comprises:

[0025] Mapping the top-left corner coordinates of each first image block in the pathological whole slide scan image to be analyzed to thumbnail coordinates according to the thumbnail sampling factor;

[0026] Linearly normalizing each cosine similarity, and mapping the normalized cosine similarity to a corresponding first color in the HSV or RGB color space;

[0027] Creating a first blank canvas with the same size as the thumbnail, and filling the corresponding first color on the first blank canvas according to the mapping coordinates to obtain a first preliminary heat map;

[0028] Superimposing the first preliminary heat map and the thumbnail with transparency to obtain the image attention heat map.

[0029] Preferably, S4 comprises:

[0030] S4.1: downsampling the pathological whole slide scan image to be analyzed to a resolution conforming to the visual-linguistic multimodal large model by the bicubic interpolation method, and converting the downsampled pathological whole slide scan image to an RGB three-channel format;

[0031] S4.2: inputting the pathological whole slide scan image processed by step S4.1 into the pre-trained visual-linguistic multimodal large model, and using “what overall visual features does this pathological whole slide scan image show?” as a prompt word, outputting a text description of the overall visual features corresponding to the pathological whole slide scan image;

[0032] S4.3: encoding the overall visual feature text description using a text encoder matched with the pre-trained visual-linguistic multimodal large model to obtain a text feature.

[0033] Preferably, in S5, the process of obtaining the multimodal feature comprises:

[0034] Step 1: concatenating all local image features and text features to obtain a mixed feature;

[0035] Step 2: dividing any one of the local image features and the text feature in the mixed feature into channels to obtain corresponding image divided features and text divided features, respectively;

[0036] Step 3: performing a channel exchange operation on the image divided features and the text divided features to obtain a first cross-modal feature and a second cross-modal feature;

[0037] Step 4: inputting the first cross-modal feature and the second cross-modal feature into a visual state space module to extract a first interaction feature and a second interaction feature, respectively;

[0038] Step 5: Map the first interaction feature and the second interaction feature to a hidden state space respectively to obtain the first hidden feature and the second hidden feature respectively; pass the first interaction feature and the second interaction feature through a gating mechanism respectively to obtain the first gating vector and the second gating vector respectively;

[0039] Step 6: Calculate a text-image attention weight based on the first hidden feature and the second hidden feature, and the calculation formula is:

[0040] ;

[0041] ;

[0042] wherein, represents a text-image attention weight between the i-th local image feature and the text feature; represents a softmax function; represents a similarity between the i-th local image feature and the text feature; represents the second hidden feature corresponding to the text feature; represents the first hidden feature corresponding to the i-th local image feature; represents a state dimension of the hidden state space; Step 7: Weighted sum all the first hidden features based on the corresponding text-image attention weights to obtain a global hidden feature; Step 8: Perform bidirectional fusion on the first hidden feature and the second hidden feature based on the first gating vector, the second gating vector and the global hidden feature to obtain the first hidden fusion feature and the second hidden fusion feature; and the fusion formula is:

[0043]

[0044]

[0045] ;

[0046] ;

[0047] wherein, represents the first hidden fusion feature of the i-th local image feature and the text feature; represents the first hidden feature corresponding to the i-th local image feature; represents the first gating vector corresponding to the i-th local image feature; represents the global hidden feature; represents the second hidden feature corresponding to the text feature; ​​​​​​It represents the Hadamardi (or Hadama) stack; This represents the second hidden fusion feature; This represents the second gating vector corresponding to the text features;

[0048] Step 9: Map the first hidden fusion feature and the second hidden fusion feature back to the original space to obtain the first fusion feature and the second fusion feature respectively;

[0049] Step 10: Repeat steps 2-9 until all local image features have been traversed, and obtain the first fused feature of the interaction between each local image feature and the text feature;

[0050] Step 11: Based on the corresponding text-image attention weights, aggregate all local image features and text features interacting to form a first fusion feature, and then fuse it with a second fusion feature to obtain multimodal features; the calculation formula is:

[0051] ;

[0052] in, Represents multimodal features; Indicates the second fusion feature; Indicates the first The first fusion feature is a combination of local image features and text features; N represents the number of the first image patches.

[0053] Preferably, in S5, the process of generating a text attention heatmap includes:

[0054] The coordinates of the top left corner of each first image block in the whole pathological slide scan image to be analyzed are mapped to thumbnail coordinates according to the thumbnail sampling factor;

[0055] The text-image attention weights are linearly normalized, and the normalized text-image attention weights are mapped to their corresponding second colors using HSV or RGB color spaces.

[0056] Create a second blank canvas of the same size as the thumbnail, and fill the second blank canvas with the corresponding second color according to the mapped coordinates to obtain the second preliminary heatmap;

[0057] The text attention heatmap is obtained by overlaying the second preliminary heatmap with the thumbnail using transparency.

[0058] Preferably, the pre-trained pathological image feature extraction model is a uni model or a conch model; the pre-trained cell detection model is a cell nucleus detection network based on YOLOv8 or Detectron2.

[0059] Preferably, the pre-trained visual-linguistic multimodal large model is a CLIP model, a BLIP-2 model or a LLaVA-Med model; and the text encoder matched with the pre-trained visual-linguistic multimodal large model is a CLIP text encoder or a medical text encoder based on BERT.

[0060] Beneficial effects: the method firstly cuts the pathological whole slice scanning image to obtain a plurality of first image blocks and a plurality of second image blocks; extracts the features of each first image block as local image features; extracts and weightedly sums the features of all second image blocks to obtain global image features; fuses the local image features based on the cosine similarity between the global image features and each local image feature to obtain overall image features, and generates an image attention heat map based on each cosine similarity; inputs the pathological whole slice scanning image to be analyzed into a visual-linguistic multimodal large model, and extracts features from the output overall visual feature text description to obtain text features; then fuses the overall image features and the text features to obtain multimodal features and a text attention heat map; inputs the multimodal features into a classifier to obtain the probability distribution of each cancer type; and analyzes the lesions based on the image attention heat map and the text attention heat map. The method breaks through the limitations of traditional methods that rely only on a single visual modality, are insufficient in representing rare pathological features, and have weak cross-device generalization capabilities, and still maintains high accuracy and robustness in data-scarce scenarios, providing an innovative solution for postoperative pathological analysis that is interpretable, interactive and cross-modal fusion. BRIEF DESCRIPTION OF DRAWINGS

[0061] In order to more clearly illustrate the technical solutions in the embodiments of the present application or the prior art, the drawings needed to be used in the embodiments or prior art description will be briefly introduced. Obviously, the drawings in the following description are only some embodiments of the present application, and other drawings can be obtained by those skilled in the art without creative labor.

[0062] Figure 1 The flowchart of the postoperative pathological image analysis method based on text language prompts in the embodiments of the present application. DETAILED DESCRIPTION

[0063] In order to make the above-mentioned purposes, features and advantages of the present application more obvious and easy to understand, the specific embodiments of the present application will be described in detail below with reference to the drawings. In the following description, a large number of specific details are set forth in order to provide a sufficient understanding of the present application. However, the present application can be implemented in many other ways different from those described herein, and those skilled in the art can make similar improvements without departing from the spirit of the present application, therefore the present application is not limited to the specific embodiments disclosed below.

[0064] In addition, the terms "first", "second", etc. are used only for descriptive purposes and should not be construed as implying or suggesting relative importance or an indicated number of the technical features. Thus, the features defined with "first", "second", etc. can explicitly or implicitly include at least one of the features. In the description of the present application, the meaning of "a plurality of" is at least two, such as two, three, etc., unless otherwise explicitly and specifically limited.

[0065] As shown in Figure 1 The embodiment provides a postoperative pathological image analysis method based on a text language prompt, and the method comprises the following steps:

[0066] S1: different pre-processing is performed on a pathological whole section scanning image to be analyzed, and non-overlapping cutting is performed on a first pre-processed pathological whole section scanning image to obtain N first image blocks; and N / 4 second image blocks are uniformly divided from a second pre-processed pathological whole section scanning image.

[0067] Specifically, the different pre-processing performed on the pathological whole section scanning image to be analyzed comprises the following steps:

[0068] First pre-processing: performing staining normalization processing on the pathological whole section scanning image to be analyzed, and performing adaptive histogram equalization (enhancing the contrast between cells and background) on the normalized pathological whole section scanning image to obtain the first pre-processed pathological whole section scanning image.

[0069] Second pre-processing: performing 4 times down-sampling on the pathological whole section scanning image to be analyzed to obtain a thumbnail, and the thumbnail is the second pre-processed pathological whole section scanning image.

[0070] Further, the staining normalization processing: a staining normalization algorithm (such as the Macenko method or the Reinhard method) is used to perform color standardization processing on the pathological whole section scanning image to be analyzed, so as to eliminate the color difference caused by different scanning devices or staining batches and ensure the consistency of the image color distribution.

[0071] Adaptive histogram equalization: adaptive histogram equalization (CLAHE, Contrast Limited Adaptive Histogram Equalization) is performed on the normalized image to enhance the contrast between the cell nucleus and the cytoplasm, highlight the cell structure details, and improve the accuracy of subsequent feature extraction.

[0072] Non-overlapping cutting: the pre-processed pathological whole section scanning image is non-overlappingly cut according to a preset image block size (for example, 256*256 pixels), to obtain a plurality of image blocks (patches). Each image block contains complete local tissue information and does not overlap, to ensure the independence and accuracy of subsequent feature extraction.

[0073] S2: extracting features of each first image block as local image features of each first image block; extracting features of all second image blocks and performing weighted summation to obtain global image features.

[0074] Specifically, the process of obtaining local image features includes:

[0075] Each first image block is input into a pre-trained pathological image feature extraction model to obtain a corresponding initial local feature vector;

[0076] Each first image block is input into a pre-trained cell detection model to obtain a corresponding cell density;

[0077] Each cell density is normalized,

[0078] The normalized cell density on each first image block is weighted to obtain the local image features of each first image block (the more dense the cell region, the higher the weight given).

[0079] The process of obtaining global image features includes:

[0080] Each second image block is input into a pre-trained pathological image feature extraction model to obtain a corresponding tissue image feature vector;

[0081] Each second image block is input into a pre-trained cell detection model to obtain a corresponding second cell density;

[0082] Each second cell density is normalized,

[0083] The normalized cell density on each second image block is weighted to obtain the final tissue image feature vector of each second image block (the more dense the cell region, the higher the weight given);

[0084] Summing up the final tissue image feature vectors of all second image blocks to obtain global image features.

[0085] In this embodiment, the pre-trained pathological image feature extraction model is a uni model or a conch model; the pre-trained cell detection model is a cell nucleus detection network based on YOLOv8 or Detectron2.

[0086] S3: Based on the cosine similarity between global image features and local image features, the local image features are weighted and summed to obtain the overall image features, and an image attention heatmap is generated based on each cosine similarity.

[0087] Specifically, the formula for calculating cosine similarity is:

[0088] ;

[0089] in, Representing global image features and the first Cosine similarity between local image features; Indicates the first Local image features, ; Represents global image features. ; Indicates the feature dimension; Indicates the modulus.

[0090] The formula for calculating the overall image features is:

[0091] ;

[0092] in, Represents overall image features, N represents the number of the first image blocks.

[0093] Image attention heatmaps generated based on cosine similarity include:

[0094] The coordinates of the top left corner of each first image patch in the whole pathological slide scan image to be analyzed. Mapped to thumbnail coordinates according to thumbnail sampling factor s. ;

[0095] The cosine similarity scores are linearly normalized. The formula for cosine similarity normalization is:

[0096] ;

[0097] in, Representing global image features and the first Normalized cosine similarity between local image features; Representing global image features and the first Cosine similarity between local image features; This indicates traversing all local image features; To represent a minimum value (e.g., 1e-6), preventing division by zero;

[0098] The normalized cosine similarity is mapped to a corresponding first color in HSV or RGB color space; the mapping rules include:

[0099] 0→blue (RGB 0,0,255);

[0100] 0.5→yellow (RGB 255,255,0);

[0101] 1→red (RGB 255,0,0);

[0102] A first blank canvas with the same size as the thumbnail is created, and the corresponding first color is filled in the first blank canvas according to the mapping coordinates to obtain a first preliminary heat map;

[0103] The first preliminary heat map is superimposed on the thumbnail in transparency to obtain an image attention heat map, and the calculation formula is:

[0104]

[0105] wherein, represents the image attention heat map; represents a transparency parameter, 0.6 in this embodiment; represents the thumbnail; represents the first preliminary heat map.

[0106] S4: inputting the pathological whole section scanning image to be analyzed into a visual-linguistic multimodal large model, outputting a corresponding overall visual feature text description based on a designed prompt word, and extracting features from the overall visual feature text description to obtain text features.

[0107] Specifically, the step includes:

[0108] S4.1: down-sampling the pathological whole section scanning image to be analyzed to a resolution (such as 224x224 pixels) conforming to the visual-linguistic multimodal large model by a bicubic interpolation method, and converting the down-sampled pathological whole section scanning image into an RGB three-channel format;

[0109] S4.2: inputting the pathological whole section scanning image processed in step S4.1 into a pre-trained visual-linguistic multimodal large model, and using "what overall visual features does this pathological whole section scanning image show?" as a prompt word to output an overall visual feature text description corresponding to the pathological whole section scanning image; the overall visual feature text description includes macroscopic histological features (such as cell density, structure morphology, staining depth, etc.) in the pathological whole section scanning image described in a natural language form;

[0110] ​S4.3: encode the overall visual feature text description using a text encoder matched with the pre-trained visual-linguistic multimodal large model to obtain text features The text features will serve as the key semantic representation for subsequent text-guided fusion and classification.

[0111] In this embodiment, the pre-trained visual-linguistic multimodal large model is a CLIP model, a BLIP-2 model, or an LLaVA-Med model; and the text encoder matched with the pre-trained visual-linguistic multimodal large model is a CLIP text encoder or a BERT-based medical text encoder.

[0112] S5: fuse the overall image features and text features based on the improved Mamba hidden state space to obtain multimodal features and text attention heat maps.

[0113] Specifically, the process of obtaining the multimodal features includes:

[0114] Step 1: concatenate all local image features and text features to obtain mixed features, denoted as:

[0115] ;

[0116] wherein, denotes the mixed features, ; denotes the text features; denotes the first local image feature; denotes the Nth local image feature; N represents the number of first image blocks.

[0117] Step 2: perform channel division (divided into 4 segments) on any one of the local image features and the text features in the mixed features to obtain corresponding image divided features and text divided features, respectively;

[0118] The th image divided feature is denoted as: wherein, , , , denote the th local image feature divided into 4 segments;

[0119] The text divided feature is denoted as: wherein, , , , denote the text features divided into 4 segments.

[0120] Step 3: Perform a channel exchange operation on the image division features and the text division features to obtain first cross-modal features and second cross-modal features; the formula is represented as:

[0121] ;

[0122] ;

[0123] wherein, represents the i-th first cross-modal feature; represents the j-th first cross-modal feature; represents the second cross-modal feature; represents the channel exchange operation.

[0124] Step 4: Input the first cross-modal features and the second cross-modal features into the visual state space module respectively to extract first interaction features and second interaction features respectively; the formula is represented as:

[0125] ; ;

[0126] wherein, represents the i-th first interaction feature; represents the j-th first interaction feature; represents the second interaction feature; represents the visual state space module.

[0127] Step 5: Map the first interaction features and the second interaction features to the hidden state space respectively to obtain first hidden features and second hidden features respectively; pass the first interaction features and the second interaction features through the gating mechanism respectively to obtain first gating vectors and second gating vectors respectively; the mapping formula is:

[0128] ;

[0129] ;

[0130] wherein, represents the i-th first hidden feature corresponding to the i-th local image feature, ; represents the second hidden feature corresponding to the text feature, ; represents the state dimension of the hidden state space; represents the Linear function; represents the Norm function; The gating mechanism is represented as:

[0131]

[0132] ;

[0133] ​ ;

[0134] in, Indicates the first The first gating vector corresponding to each local image feature ; This represents the second gating vector corresponding to the text features. ; This represents the activation function.

[0135] Step 6: Calculate the text-image attention weights based on the first and second hidden features. The calculation formula is as follows:

[0136] ;

[0137] ;

[0138] in, Indicates the first Text-image attention weights between local image features and text features; This represents the softmax function; Indicates the first The similarity between local image features and text features; This represents the second hidden feature corresponding to the text feature; Indicates the first The first hidden feature corresponding to each local image feature; This represents the state dimension of the hidden state space.

[0139] Step 7: Based on the corresponding text-image attention weights, sum all the first hidden features to obtain the global hidden features; the aggregation formula is:

[0140] ;

[0141] in, Represents globally hidden features. .

[0142] Step 8: Based on the first gating vector, the second gating vector, and the global hidden features, perform bidirectional fusion of the first hidden features and the second hidden features to obtain the first hidden fused features and the second hidden fused features; the fusion formula is:

[0143] ;

[0144] ;

[0145] in, Indicates the first the first hidden fusion feature of the i-th local image feature and the text feature; represents the first hidden feature corresponding to the i-th local image feature; represents the first hidden feature corresponding to the i-th local image feature; represents the first gating vector corresponding to the i-th local image feature; represents the first gating vector corresponding to the i-th local image feature; represents the global hidden feature; represents the second hidden feature corresponding to the text feature; represents the Hadamard product; represents the second hidden fusion feature; represents the second gating vector corresponding to the text feature.

[0146] Step 9: map the first hidden fusion feature and the second hidden fusion feature back to the original space respectively to obtain the first fusion feature and the second fusion feature respectively; the mapping formula is:

[0147] ;

[0148] ;

[0149] wherein, represents the second fusion feature; represents the i-th local image feature and the text feature; represents the i-th local image feature and the text feature; represents the inverse mapping function of the Linear function; represents the i-th image partition feature; represents the i-th image partition feature; represents the text partition feature.

[0150] Step 10: repeat steps 2-9 until all local image features are traversed to obtain the first fusion feature of each local image feature and the text feature interaction;

[0151] Step 11: based on the corresponding text-image attention weight, the first fusion features of all local image features and the text feature interaction are weighted and aggregated, and the second fusion feature is fused to obtain a multi-modal feature; the calculation formula is:

[0152] ;

[0153] wherein, represents the multi-modal feature; represents the second fusion feature; represents the i-th local image feature and the text feature; represents the i-th local image feature and the text feature; N represents the number of first image blocks.

[0154] Further, in S5, the process of generating the text attention heat map includes:

[0155] mapping the top-left coordinates of each first image block in the pathological whole-section scanned image to be analyzed to thumbnail coordinates according to the thumbnail sampling factor;

[0156] linearly normalizing each text-image attention weight, and mapping the normalized each text-image attention weight to a corresponding second color (mapping rule is the same as above) in the HSV or RGB color space;

[0157] creating a second blank canvas with the same size as the thumbnail, and filling the corresponding second color on the second blank canvas according to the mapping coordinates to obtain a second preliminary heat map;

[0158] superimposing the second preliminary heat map and the thumbnail in transparency to obtain a text attention heat map.

[0159] S6: inputting the multi-modal feature into an MLP classifier to obtain a probability distribution of each cancer type; and performing lesion analysis based on the image attention heat map and the text attention heat map.

[0160] Specifically, the probability distribution of each cancer type is calculated according to the following formula:

[0161]

[0162] wherein, P represents the probability distribution of each cancer type, K represents the type of cancer; softmax represents a softmax function; MLP represents a multi-layer perceptron; multi-modal feature represents a multi-modal feature.

[0163] Further, the lesion analysis based on the image attention heat map and the text attention heat map includes:

[0164] The red area in the two attention heat maps reflects the potential lesion position, and the pathologist can intuitively compare the two attention heat maps, combine with the knowledge of cell morphology, accurately locate the lesion, evaluate the infiltration range and the surgical margin, and thus improve the accuracy and efficiency of postoperative pathological diagnosis.

[0165] The postoperative pathological image analysis method based on text language prompts provided by the embodiment has the following beneficial effects:

[0166] ​The method is based on pathological whole-section scanning image after staining normalization and adaptive histogram equalization processing, first, the high cell density area in the first image block is weighted by the joint pathological image feature extraction model and the cell detection model, and a plurality of local image features are obtained, which can significantly suppress background noise and strengthen local key features; then the high cell density area in the first image block is weighted by the joint pathological image feature extraction model and the cell detection model, and the global image feature is obtained. After obtaining the "local weighting + global" dual visual representation, the image attention heat map which can directly locate the lesion is further generated based on the cosine similarity. At the same time, the system calls the visual-language multimodal large model to automatically generate the text features of the overall visual description of the pathological section by using the fixed prompt word; then the Mamba hidden state space model is used to complete the deep bidirectional fusion of the text features and all image block features with linear complexity, and the text-guided multimodal features and text attention heat map are obtained. Finally, the multimodal features are sent to the MLP classifier, and the probability distribution of each cancer type is output at one time, realizing the efficient diagnosis of end-to-end. Through the "image-text double heat map" linkage analysis, the pathologist can quickly compare the local cell morphological abnormalities and the global semantic area emphasized by the language model in the same interface, which can significantly improve the lesion positioning accuracy and diagnosis confidence. The scheme breaks through the limitations of traditional methods, such as relying on only a single visual mode, insufficient representation of rare pathological features, and weak cross-device generalization ability, and still maintains high accuracy and robustness in data-scarce scenarios, providing an interpretable, interactive and cross-modal fusion innovative solution for postoperative pathological evaluation.

[0167] The technical features of the above-described embodiments can be combined in any manner. To make the description concise, not all possible combinations of the technical features in the above-described embodiments are described, but as long as the combinations of the technical features do not contradict, they should be considered within the scope of the present disclosure.

[0168] The above-described embodiments only express several implementation manners of the present application, and the description is more specific and detailed, but it should not be understood as a limitation on the patent scope of the application. It should be noted that for ordinary skilled persons in the art, without departing from the concept of the present application, a number of modifications and improvements can be made, which are within the scope of the present application. Therefore, the patent protection scope of the present application should be subject to the appended claims.

Claims

1. A method for post-operative pathology image analysis based on textual language cues, the method comprising: receiving a plurality of images of a specimen; receiving a textual language cue; and analyzing the plurality of images of the specimen based on the textual language cue. Comprise: S1: different preprocessing is carried out on the pathological whole section scanning image to be analyzed respectively, and the first preprocessing pathological whole section scanning image is non-overlapping cutting, obtaining a plurality of first image blocks; The second preprocessing pathological whole section scanning image is uniformly divided into a plurality of second image blocks; S2: extracting the features of each first image block as the local image features of each first image block; extracting the features of all second image blocks and performing weighted summation to obtain global image features; S3: based on the cosine similarity between the global image features and each local image feature, the local image features are weighted and summed to obtain the overall image features, and the image attention heat map is generated based on each cosine similarity; S4: input the pathological whole section scanning image to be analyzed into the visual-linguistic multi-modal large model, output the corresponding overall visual feature text description based on the designed prompt word, and extract the features of the overall visual feature text description to obtain the text features; S5: based on the improved Mamba hidden state space, the overall image features and text features are fused to obtain multi-modal features and text attention heat map; S6: input the multi-modal features into the MLP classifier to obtain the probability distribution of each cancer type; and based on the image attention heat map and the text attention heat map, the lesion analysis is carried out.

2. The text language prompt based post-operative pathology image analysis method of claim 1, wherein, The different preprocessing of the pathological whole section scanning image to be analyzed comprises: First preprocessing, the pathological whole section scanning image to be analyzed is subjected to dye normalization processing, and the normalized pathological whole section scanning image is subjected to adaptive histogram equalization to obtain the first preprocessing pathological whole section scanning image; Second preprocessing, the pathological whole section scanning image to be analyzed is subjected to 4 times downsampling to obtain a thumbnail, and the thumbnail is the second preprocessing pathological whole section scanning image.

3. The text language prompt based post-operative pathology image analysis method of claim 1, wherein, In S2, the process of obtaining the local image features comprises: Each first image block is input into a pre-trained pathological image feature extraction model to obtain a corresponding initial local feature vector; Each first image block is input into a pre-trained cell detection model to obtain a corresponding cell density; The cell density is normalized, Based on the normalized cell density on each first image block, the corresponding initial local feature vector is weighted to obtain the local image features of each first image block.

4. The text language prompt based post-operative pathology image analysis method of claim 1, wherein, In S2, the process of obtaining the global image features comprises: Each second image block is input into a pre-trained pathological image feature extraction model to obtain a corresponding tissue image feature vector; Each second image block is input into a pre-trained cell detection model to obtain a corresponding second cell density; The second cell density is normalized, Based on the normalized cell density on each second image block, the corresponding tissue image feature vector is weighted to obtain the final tissue image feature vector of each second image block; The final tissue image feature vectors of each second image block are summed to obtain the global image features.

5. The text language prompt based post-operative pathology image analysis method of claim 2, wherein, In S3, generating the image attention heat map based on each cosine similarity comprises: The top-left corner coordinates of each first image block in the pathological whole section scanning image to be analyzed are mapped to the thumbnail coordinates according to the thumbnail sampling factor; Linearly normalizing each cosine similarity, and mapping each normalized cosine similarity to a corresponding first color in the HSV or RGB color space; Creating a first blank canvas with the same size as the thumbnail, and filling the corresponding first color in the first blank canvas according to the mapping coordinates to obtain a first preliminary heat map; Superimposing the first preliminary heat map and the thumbnail with transparency to obtain the image attention heat map.

6. The text language prompt based post-operative pathology image analysis method of claim 1, wherein S4 Comprise: S4.1: down-sampling the pathological whole section scanning image to be analyzed to the resolution conforming to the visual-linguistic multi-modal large model by bicubic interpolation method, and converting the down-sampled pathological whole section scanning image into an RGB three-channel format; S4.2: inputting the pathological whole section scanning image processed in step S4.1 into the pre-trained visual-linguistic multi-modal large model, and using "what overall visual features does this pathological whole section scanning image show?" as a prompt word to output the overall visual feature text description corresponding to the pathological whole section scanning image; S4.3: encoding the overall visual feature text description by using a text encoder matched with the pre-trained visual-linguistic multi-modal large model to obtain a text feature.

7. The text language prompt based post-operative pathology image analysis method of claim 2, wherein, In S5, the process of obtaining the multi-modal feature comprises: Step 1: splicing all local image features and text features to obtain a mixed feature; Step 2: dividing any one of the local image features and the text feature in the mixed feature into channels respectively to obtain corresponding image divided features and text divided features respectively; Step 3: performing a channel exchange operation on the image divided features and the text divided features to obtain a first cross-modal feature and a second cross-modal feature; Step 4: inputting the first cross-modal feature and the second cross-modal feature into a visual state space module respectively to extract a first interaction feature and a second interaction feature respectively; Step 5: mapping the first interaction feature and the second interaction feature to a hidden state space respectively to obtain a first hidden feature and a second hidden feature respectively; and passing the first interaction feature and the second interaction feature through a gating mechanism respectively to obtain a first gating vector and a second gating vector respectively; Step 6: calculating a text-image attention weight based on the first hidden feature and the second hidden feature, and the calculation formula is: ; ; wherein, represents a text-image attention weight between the th local image feature and the text feature; represents a softmax function; represents a similarity between the th local image feature and the text feature; represents a second hidden feature corresponding to the text feature; represents a first hidden feature corresponding to the th local image feature; represents a state dimension of a hidden state space; Step 7: weighting and summing all first hidden features based on the corresponding text-image attention weight to obtain a global hidden feature; Step 8: bidirectionally fusing the first hidden feature and the second hidden feature based on the first gating vector, the second gating vector and the global hidden feature to obtain a first hidden fusion feature and a second hidden fusion feature; and the fusion formula is: ; ; wherein, denotes a first hidden fusion feature of the th local image feature and the text feature; denotes a first hidden feature corresponding to the th local image feature; denotes a first gating vector corresponding to the th local image feature; denotes a global hidden feature; denotes a second hidden feature corresponding to the text feature; denotes a Hadamard product; denotes a second hidden fusion feature; denotes a second gating vector corresponding to the text feature; Step 9: mapping the first hidden fusion feature and the second hidden fusion feature back to the original space respectively to obtain a first fusion feature and a second fusion feature respectively; Step 10: repeatedly executing steps 2-9 until all local image features are traversed to obtain the first fusion feature of the interaction between each local image feature and the text feature; Step 11: weighting and aggregating all first fusion features of the interaction between the local image features and the text features based on the corresponding text-image attention weight, and fusing the second fusion feature to obtain a multi-modal feature; and the calculation formula is: ; wherein, denotes a multi-modal feature; denotes a second fused feature; denotes a first fused feature of the local image features and text features; N denotes the number of first image patches.

8. The text language prompt-based post-operative pathology image analysis method of claim 7, wherein, In S5, the process of generating the text attention heat map includes: mapping the top-left corner coordinates of each first image block in the pathological whole-section scanning image to be analyzed to thumbnail coordinates according to the thumbnail sampling factor; linearly normalizing each text-image attention weight, and mapping the normalized text-image attention weight to a corresponding second color in the HSV or RGB color space; creating a second blank canvas with the same size as the thumbnail, and filling the corresponding second color on the second blank canvas according to the mapping coordinates to obtain a second preliminary heat map; superimposing the second preliminary heat map and the thumbnail with transparency to obtain the text attention heat map.

9. The text language prompt-based post-operative pathology image analysis method according to claim 3 or 4, characterized in that, The pre-trained pathological image feature extraction model is a uni model or a conch model; the pre-trained cell detection model is a cell nucleus detection network based on YOLOv8 or Detectron2.

10. The text language prompt based post-operative pathology image analysis method of claim 6, wherein, The pre-trained visual-linguistic multimodal large model is a CLIP model, a BLIP-2 model or an LLaVA-Med model; and the text encoder matched with the pre-trained visual-linguistic multimodal large model is a CLIP text encoder or a medical text encoder based on BERT.

Citation Information

Patent Citations

  • Space channel adaptive accident prediction method and system based on eye movement attention guidance

    CN118470484A

  • Expression recognition method and system based on three-stage multi-mode visual language prompt

    CN119763171A