Pathological image cell classification method based on visual language large model and prompt learning

By using a large visual language model and cue learning method, a cell classification network for pathological images was constructed, which solved the problem of differences in annotation systems between different datasets, achieved more efficient and accurate cell classification, and improved the model's generalization ability.

CN119478937BActive Publication Date: 2025-11-18UNIV OF SCI & TECH OF CHINA
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202411631977.3
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2024-11-15
Publication Date
2025-11-18
Estimated Expiration
2044-11-15

AI Technical Summary

Technical Problem

Existing pathological image cell classification models suffer from incomplete prompts and semantic ambiguity due to differences in annotation systems across different datasets, making it difficult to achieve generalization and accurate classification across datasets.

Method used

We employ a visual language large model and cue learning approach, constructing a pathological image cell classification network through an instance-aware cue learning module, a cue conditionalization module, and a local visual text matching module. We leverage shared knowledge from multiple datasets and capture rich contextual information to enhance feature interaction and semantic association, and construct a loss function to optimize the model.

Benefits of technology

The model's cell classification performance on multiple datasets has been improved, the problems of incomplete prompts and semantic ambiguity have been resolved, and more efficient cell classification accuracy and generalization ability have been achieved.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119478937B_ABST
    Figure CN119478937B_ABST
Patent Text Reader

Abstract

The application discloses a pathological image cell classification method based on a visual language large model and prompt learning, and comprises the following steps: a cell classification framework is constructed by using an instance perception prompt learning module, a prompt conditioning module, a local visual text matching module and an image coding and decoding network, specifically as follows: in the instance perception prompt learning module, the visual language large model is used to extract the features of an input image and text respectively, so that instance-specific information is fully mined, and instance perception text features are obtained; the text features are input into the prompt conditioning module to modulate the image features, and the text features are also input into the local visual text matching module to strengthen the semantic association with the image features; a network overall loss function is constructed; and an optimizer is used to iteratively train the classification network. The application solves the technical problems of incomplete prompts and ambiguous cell feature semantics caused by large instance differences in the prior art, so that the classification performance of the model on multiple data sets can be improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention belongs to the field of pathological image cell classification technology, specifically a pathological image cell classification method based on a large visual language model and cue learning. Background Technology

[0002] As a crucial component of pathological image analysis, automated cell classification assists pathologists in disease diagnosis, thereby improving the efficiency and accuracy of pathological diagnosis. It holds significant research value and broad application prospects. In recent years, with the continuous accumulation of data resources in the field of cell classification, pathological image cell classification methods have achieved certain successes. However, the inherent morphological and staining heterogeneity of cells in pathological images poses a severe challenge to traditional methods trained using single datasets, resulting in limited performance of such cell classification models, especially in their generalization ability. Therefore, the research community is currently keen to explore the construction of generalizable models that can be applied across datasets, aiming to enhance the universality and robustness of models by integrating diverse data sources. Unfortunately, this strategy has always been hampered by the problem of differences in annotation systems between datasets caused by different sources of pathological tissues. This problem directly leads to semantic inconsistencies during model training and generalization, thus becoming a key obstacle to the development of generalizable models.

[0003] To address the challenge of varying datasets, a range of solutions have been proposed by the research community, such as multi-task learning strategies (Graham, S.; Vu, QD; Jahanifar, M.; Raza, SEA; Minhas, F.; Snead, D.; and Rajpoot, N. One model is all you need: Multi-task learning enables simultaneous histology image segmentation and classification. MedicalImage Analysis, 2023, 83: 102685.) and label ensemble strategies (Zhang, W.; Zhang, J.; Wang, X.; Yang, S.; Huang, J.; Yang, W.; Wang, W.; and Han, X. Merging nucleus datasets by correlation-based cross-training. Medical Image Analysis, 2023, 84: 102705.), which aim to extract more generalized image features and fully utilize the category semantic information from different annotation systems. However, these existing strategies often face limitations due to the surge in computational resources or fail to achieve an end-to-end cell classification framework for pathological images. Recently, a prompt-based ensemble universal model—UniCell (Huang, J.; Li, H.; Wan, X.; and Li, G. UniCell: Universal Cell Nucleus Classification via Prompt Learning. In Proceedings of the AAAI Conference on Artificial Intelligence, 2024, 38(3): 2348-2356.)—was proposed. The core design concept of this framework is to discover and utilize the shared cell knowledge among datasets, while avoiding the redundant training problem common in multi-source data ensemble processing. Specifically, this universal model first adopts advanced learnable cell category text prompting technology to achieve a unified annotation system for different datasets. Then, based on the characteristics of different dataset annotation systems, it selectively combines category text prompts to generate dataset prompts, thereby improving the generalization performance of the cell classification model.

[0004] However, despite the remarkable progress made by the aforementioned methods in representing cellular semantic information, text-based cell classification still faces two major challenges: instance variability and semantic ambiguity. Specifically, pathological image instances often exhibit significant intra- and inter-dataset differences in cell category composition. Therefore, dataset-based prompting methods that rely on fixed patterns cannot flexibly utilize contextual information from rich image instances, resulting in incomplete or even inappropriate prompts. Simultaneously, the semantic relationships between numerous cells within a single image instance and their true categories are extremely complex. Consequently, the aforementioned prompt-based general cell classification models are prone to generating semantically ambiguous cell features and struggle to learn generalizable cell representations to accurately identify different cell categories from multiple datasets. Summary of the Invention

[0005] This invention aims to address the shortcomings of existing technologies by proposing a pathological image cell classification method based on a large visual language model and cue learning. This method seeks to solve the technical problems of incomplete cues caused by differences in image instances and semantic ambiguity of cell features in existing technologies, thereby improving the cell classification performance of the model on multiple datasets.

[0006] To achieve the above-mentioned objectives, the present invention adopts the following technical solution:

[0007] The present invention provides a method for classifying cell types in pathological images based on a large visual language model and cue learning, characterized by the following steps:

[0008] Step 1: Obtain the labeled data After processing and standardizing a dataset of pathological images, the preprocessed images are obtained. A dataset of pathological images; Representing the The total number of preprocessed image instances corresponding to the dataset; the preprocessed... The first pathological image dataset Zhang's pathological images are recorded as , The corresponding actual annotation is recorded as ;and ,in, Indicates the first The true coordinates of the cell center, Indicates the first The true labeling of each cell category, and ∈{1,2,…,K}; K represents the total number of cell categories; P represents the total number of cells in all categories. for The true labeling of cell categories is The total number of cells;

[0009] Get Corresponding example hints ;

[0010] Get Category hints for all cell types in a dataset of pathological images ;

[0011] Step 2: Construct a large-scale visual language model and cue learning cell classification network, including: an instance-aware cue learning module, a cue conditionalization module, a local visual-text matching module, and the third... An image encoding / decoding network ;

[0012] Step 2.1, the instance-aware prompting learning module according to... , as well as ,get Instance-aware category text features ;in, This represents the instance-aware category text feature corresponding to the k-th cell category;

[0013] Step 2.2, the An image encoding / decoding network encoder pairs in Processing is performed to obtain Image features ;

[0014] Step 2.3, the conditional prompting module according to as well as ,get Conditional image features ;

[0015] Step 2.4, the An image encoding / decoding network decoder pairs in Processing is performed to obtain local cell image features. ;

[0016] Step 2.5, the An image encoding / decoding network Classification head pairs in Process and generate Prediction results , including: the Predicted center coordinates of each cell and the Predicted category of each cell ; , Indicates the predicted total number of cells;

[0017] Step 2.6, according to the... An image encoding / decoding network right Prediction results Using equation (1), the predicted center coordinates and category of the predicted cells that match the real cells are obtained:

[0018] (1)

[0019] In equation (1), Represents a linear summation distributive function; and Representing the first An image encoding / decoding network The obtained number The image contains the first The predicted center coordinates and category of each predicted cell that matches the real cell;

[0020] Step 3: Construct the loss function:

[0021] Step 3.1: Construct using equation (2) Classification loss :

[0022] (2)

[0023] In equation (2), This represents the balanced cross-entropy loss function; This represents the mean absolute error loss function;

[0024] Step 3.2: Construct using equation (3) Hungarian matching loss :

[0025] (3)

[0026] In equation (3), Indicates the first An image encoding / decoding network The obtained number In the prediction results of the first image The true cell category corresponding to each unmatched predicted cell; , and This represents three loss weighting coefficients;

[0027] Step 3.3: The local visual text matching module uses equation (4) to construct the first... The local visual text matching loss function corresponding to each dataset :

[0028] (4)

[0029] In equation (4), Represents the feature similarity measurement function. It is a temperature parameter;

[0030] Step 3.4: Construct the first... using equation (5) The overall loss corresponding to each dataset :

[0031] (5)

[0032] In equation (5), , and They are respectively , and Weighting coefficients;

[0033] Step 4: Train a pathological image cell classification network based on a large visual language model and cue learning using the Adam optimizer, and minimize... With the goal of improving the classification network for pathological images based on a large visual language model and cue learning, we iterate the network and select the optimal classification network after several rounds of training to identify cells in the images.

[0034] The characteristic of the pathological image cell classification network based on visual language large model and cue learning described in this invention is that step 2.1 includes:

[0035] Step 2.1.1, Definition The learnable token is and The learnable token is ;

[0036] Step 2.1.2: Generate using equation (6) Refined category text tokens :

[0037] (6)

[0038] In equation (6), This represents a set of learnable parameters for a control mapping to learn token scale; Represents a mapping function; This indicates element-wise multiplication.

[0039] Step 2.1.3: Use a text encoder to... and Encode it, and get the corresponding result Instance text embedding and text embedding matrix of K cell categories ;in, The category text embedding represents the cell category as actually labeled k;

[0040] Step 2.1.4: Use a visual encoder to... Encode to obtain Image embedding Thus As a query for the Transformer decoder As the key and value of the Transformer decoder, and by the Transformer decoder... Perform calibration and generate Visual domain cue embedding ;

[0041] Step 2.1.5: Generate using equation (7) Enhanced instance text embedding :

[0042] (7)

[0043] In equation (7), Represents a set of control visual field cues embedded The parameters to be learned for the scale;

[0044] Step 2.1.6: Using equation (8) to obtain Specific instance text features :

[0045] (8)

[0046] In equation (8), Indicates a specificity mask; This represents the cross-attention operation of the mask; Indicates the layer normalization function; This represents a feedforward neural network;

[0047] Step 2.1.7, and By splicing them together, we can obtain Instance-aware category text features .

[0048] Furthermore, step 2.3 includes:

[0049] Step 2.3.1: Using equation (9) to obtain Two condition vectors and :

[0050] (9)

[0051] In equation (9), and This represents two different multilayer perceptrons;

[0052] Step 2.3.2, using equation (10) to... Conditional adjustments are performed to obtain cue-conditional image features. :

[0053] (10).

[0054] The present invention provides an electronic device, including a memory and a processor, wherein the memory is used to store a program that supports the processor in executing the pathological image cell classification method, and the processor is configured to execute the program stored in the memory.

[0055] The present invention discloses a computer-readable storage medium on which a computer program is stored, wherein the computer program is executed by a processor to perform the steps of the pathological image cell classification method.

[0056] Compared with the prior art, the beneficial effects of the present invention are as follows:

[0057] 1. This invention proposes an instance-aware prompting learning module; this module utilizes a large visual language model and prompting learning technology to not only make full use of shared knowledge from multiple datasets, but also flexibly capture rich contextual information from different pathological image instances, avoiding incomplete or inappropriate prompts due to differences in image instances.

[0058] 2. This invention proposes a cue conditionalization module; this module uses feature-level linear modulation technology to constrain cell image features with semantic information corresponding to cell categories, thereby enabling full interaction and fusion of features from visual and linguistic modalities, making the subsequent language supervision optimization process more effective.

[0059] 3. This invention proposes a local visual text matching loss; this module explicitly utilizes language-guided supervision information to enhance the semantic association between local cell images and corresponding category text prompts, avoiding the problem of non-generalization of cell features due to semantic ambiguity, thereby improving the cell classification performance of the model on multiple datasets. Attached Figure Description

[0060] Figure 1 This is a flowchart of the method of the present invention. Detailed Implementation

[0061] The present invention will now be described in detail with reference to the accompanying drawings and specific embodiments.

[0062] In this embodiment, a pathological image cell classification method based on a large visual language model and cue learning is described, such as... Figure 1 As shown, it mainly includes the following steps:

[0063] Step 1: Obtain the labeled data After processing and standardizing a dataset of pathological images, the preprocessed images are obtained. A dataset of pathological images;

[0064] Step 1.1: In this example, we use the cell classification datasets CoNSeP, MoNuSAC, Lizard, and OCELOT. For any of these datasets, we randomly divide it into two subsets. Each time, we use one subset as the test set, and then randomly select 20% of the data from the other subset as the validation set. The remainder is used as the training set, and training is performed using only the image data from the training set.

[0065] Step 1.2: Augment the training set of the cell classification dataset. Augmentation methods include rotation by 0°, 90°, 180°, and 270°, as well as vertical and horizontal flipping. The rotation and flipping operations are combined, increasing the amount of training data by 8 times. Then, each pathological image is standardized and normalized using equations (1) and (2). The specific calculation method is as follows:

[0066] (1)

[0067] (2)

[0068] In equations (1) and (2), Pathological images The standardization results and and represent the mean and variance of all pathological image pixels in the dataset, respectively. This is the result after image normalization.

[0069] Step 1.3: Use the standardized and normalized data as the preprocessed dataset; use the preprocessed data as the... The first pathological image dataset Zhang's pathological images are recorded as , The corresponding actual annotation is recorded as ;and , of which The true annotation of the coordinates of the cell center is recorded as follows: , No. The true label of each cell category is recorded as follows: ∈{1,2,…,K}; K represents the total number of cell categories; P represents the total number of cells in all categories. for The cell category is actually labeled as The total number of cells;

[0070] Get Corresponding example hints ; obtain Category hints for the overall cell category in a dataset of pathological images ;

[0071] Step 2: Construct a large-scale visual language model and cue learning cell classification network, including: an instance-aware cue learning module, a cue conditionalization module, a local visual-text matching module, and the third... An image encoding / decoding network ;

[0072] Step 2.1, the instance-aware prompt learning module according to , as well as ,get Instance-aware category text features ,in, This represents the instance-aware category text feature corresponding to the k-th cell category;

[0073] Step 2.1.1, Obtain The example text description is "A H&E image of [Cancer_type]", where Cancer_type represents a disease type, and the text description for the k-th cell category is [Category_k], where Category_k represents the cell category name. Learnable tokens are introduced as prefixes to both the example text description and the category text description. and This creates example text prompts. and category text hints ;

[0074] Step 2.1.2, using formula (3) to promote arrive Information sharing, updating category text tokens to generate refined category text tokens :

[0075] (3)

[0076] In equation (3), This represents a set of learnable parameters for a control mapping to learn token scale; Represents a mapping function; This represents element-wise multiplication; Concatenated with the category text description ;Will Concatenated with the instance text description ;

[0077] Step 2.1.3: Use a text encoder to... and Encode it, and get the corresponding result Instance text embedding and text embedding matrix of K cell categories ;in, The category text embedding represents the cell category as actually labeled k;

[0078] Step 2.1.4: Use a visual encoder to... Encode to obtain Image embedding Thus As a query for the Transformer decoder As the key and value of the Transformer decoder, and using Equation (4) to... Perform calibration and generate Visual domain cue embedding :

[0079] (4)

[0080] Step 2.1.5: Embed the visual domain cue using equation (5). and instance text embedding Perform feature fusion to generate Enhanced instance text embedding :

[0081] (5)

[0082] In equation (5), Represents a set of control visual field cues embedded The parameters to be learned for the scale;

[0083] Step 2.1.6: Embed the enhanced instance text using equation (6). and K category text embedding matrix Perform dataset-specific feature fusion to obtain dataset-specific instance text features. :

[0084] (6)

[0085] In equation (6), This represents a specific mask; This represents the cross-attention operation of the mask; Indicates the layer normalization function; This represents a feedforward neural network;

[0086] Step 2.1.7, and By splicing them together, we can obtain Instance-aware category text features .

[0087] Step 2.2, the An image encoding / decoding network encoder pairs in Processing is performed to obtain Image features ;

[0088] In this embodiment, the encoder consists of a Swin-Transformer feature extractor and three Transformer encoders;

[0089] Step 2.3, the conditional prompt module based on as well as ,get Conditional image features ;

[0090] Step 2.3.1, using equation (7) to... Perform mapping processing to obtain two condition vectors. and :

[0091] (7)

[0092] In equation (7), and This represents two different multilayer perceptrons;

[0093] Step 2.3.2: Use equation (8) to analyze image features. Conditional adjustments are performed to obtain cue-conditional image features. :

[0094] (8)

[0095] Step 2.4, the An image encoding / decoding network decoder pairs in Processing is performed to obtain local cell image features. In this embodiment, the decoder consists of multiple Transformer decoders.

[0096] Step 2.5, the An image encoding / decoding network Classification head pairs in Process and generate Prediction results ,Include: The predicted cell, the first Predicted center coordinates of each cell and the Predicted category of each cell ;

[0097] The classification head consists of two feedforward neural networks, and their specific network structure information is shown in Table 1; where FFN1 is the regression branch and FFN2 is the classification branch; Indicates the first The number of cell categories contained in each dataset;

[0098] Table 1: Classification Header Network Structure Information

[0099]

[0100] Step 2.6, according to the... An image encoding / decoding network right Prediction results Using equation (1), the predicted center coordinates and category of the predicted cells that match the real cells are obtained:

[0101] (1)

[0102] In equation (1), Represents a linear summation distributive function; and Representing the first An image encoding / decoding network The obtained number The image contains the first The predicted center coordinates and category of each predicted cell that matches the real cell;

[0103] Step 3: Construct the loss function:

[0104] Step 3.1: Construct using equation (10) Classification loss :

[0105] (10)

[0106] In equation (10), express Accurate labeling of cells in the middle; Represents image encoding / decoding network right The prediction results; This represents the balanced cross-entropy loss function; This represents the mean absolute error loss function;

[0107] Step 3.2: Construct using equation (11) Hungarian matching loss :

[0108] (3)

[0109] In equation (11), Indicates the first An image encoding / decoding network The obtained number In the prediction results of the first image The true cell category corresponding to each unmatched predicted cell, i.e., the background class; , and This represents three loss weighting coefficients. In this embodiment, , and Set them to 5, 2, and 2 respectively;

[0110] Step 3.3: Construct the first... using equation (12) The local visual text matching loss function corresponding to each dataset :

[0111] (12)

[0112] In equation (12), Represents the feature similarity measurement function. It is a temperature parameter;

[0113] Step 3.4: Construct the first equation using equation (13). The overall loss corresponding to each dataset :

[0114] (13)

[0115] In equation (13), , and They are respectively , and The weighting coefficients; in this embodiment, , and Set them to 1, 0.1 and 1 respectively.

[0116] Step 4: Train a pathological image cell classification network based on a large visual language model and cue learning using the Adam optimizer, and minimize... With the goal of improving the classification network for pathological images based on a large visual language model and cue learning, we iterate the network and select the optimal classification network after several rounds of training to identify cells in the images.

[0117] Step 4.1: Set the learning rate of the Adam optimizer to 10⁻⁴ and iterate the pathological image cell classification network based on visual language large model and cue learning.

[0118] Step 4.2: After each iteration, validate the cell classification network on the validation set;

[0119] Step 4.2.1: In this embodiment, the threshold method is first used to evaluate the prediction results. middle The cell class probability at each predicted location is compared with a preset threshold of 0.5. If the predicted class probability is greater than 0.5, the predicted result at that location is determined to be a cell; otherwise, the predicted result at that location is determined to be background.

[0120] Step 4.2.2: For the prediction results that are determined to be cells, in this embodiment, the Hungarian matching algorithm is used to assign a prediction result to each real cell based on the coordinate distance, so as to minimize the total distance between all real cells and the corresponding matched prediction cells.

[0121] Step 4.2.3, Use Score( Using ) as the evaluation index, the cell classification results are evaluated using equations (14)-(16):

[0122] (14)

[0123] (15)

[0124] (16)

[0125] In equations (14)-(16), It is a test Score, It is the first Classification of categories Score, It is all in the dataset Average classification of cell categories Score. This indicates the number of true positives detected. This indicates the number of false positives detected. This indicates the number of false negatives detected. For the ... Cell categories, It is further divided into: the number of true cases in the classification. The number of categorical false positives and the number of classified false negatives .

[0126] In this embodiment, a corresponding true label region is established based on the center coordinates of each cell, which is a circular region with a radius of 6 pixels centered on the true position coordinates of each cell. If there are more than one cell position prediction result within the true label region, the cell prediction result closest to the true position coordinates is considered a true example. If a cell position prediction result is not within any true label region, the prediction result is considered a false positive; the remaining prediction results are false negatives.

[0127] In this embodiment, an electronic device includes a memory and a processor. The memory stores a program that supports the processor in executing the methods described above, and the processor is configured to execute the program stored in the memory.

[0128] In this embodiment, a computer-readable storage medium stores a computer program, which is executed by a processor to perform the steps of the above method.

[0129] Furthermore, to quantitatively evaluate the performance of the proposed method, this embodiment demonstrates its comparison with HoverNet (Graham, S.; Vu, QD; Raza, SEA; Azam, A.; Tsang, YW; Kwak, JT; and Rajpoot, N. Hover-Net: Simultaneous segmentation and classification of nuclei in multi-tissue histology images. Medical Image Analysis, 2019, 58: 101563.), MCSpatNet (Abousamra, S.; Belinsky, D.; VanArnam, J.; Allard, F.; Yee, E.; Gupta, R.; Kurc, T.; Samaras, D.; Saltz, J.; and Chen, C. Multi-class cell detection using spatial context representation. In Proceedings of the IEEE / CVF International Conference on Computer Vision, 2021, 4005-4014.), Cellpose (Stringer, C.; Wang, T.; Michaelos, M.; andPachitariu, M. Cellpose: a generalist algorithm for cellular segmentation.Nature Methods, 2021, 18(1): 100–106.), Omnipose (Cutler, KJ; Stringer, C.;Lo, TW; Rappez, L.; Stroustrup, N.; Brook Peterson, S.; Wiggins, PA; and Mougous, JD Omnipose: a high-precision morphology-independent solution for bacterial cell segmentation. Nature Methods, 2022, 19(11): 1438–1448.), UperNet ConvNeXt (Liu, Z.; Mao, H.; Wu, C.-Y.; Feichtenhofer, C.; Darrell, T.; and Xie, S. A convnet for the 2020s. In Proceedings of the IEEE / CVFConference on Computer Vision and Pattern Recognition, 2022, 1197611986.), PGT (Huang, J.; Li, H.; Sun, W.; Wan, X.; and Li, G. Prompt-based groupingtransformer for nucleus detection and classification. In InternationalConference on Medical Image Computing and Computer-Assisted Intervention, 2023, 569-579.) and UniCell (Huang, J.; Li, H.; Wan, Learning. In Proceedings of the AAAI The performance comparison results of single-dataset and multi-dataset cell classification methods on the cell classification datasets CoNSeP, MoNuSAC, Lizard, and OCELOT, presented at the Conference on Artificial Intelligence, 2024, 38(3): 2348-2356, are shown in Table 2.

[0130] Table 2: Performance comparison of the proposed method with other single-dataset and multi-dataset cell classification methods on the cell classification datasets CoNSeP, MoNuSAC, Lizard, and OCELOT:

[0131]

[0132] As can be seen from the above results, the method proposed in this invention can perform cell classification more accurately, efficiently and automatically.

Claims

1. A method for classifying cells in pathological images based on a large visual language model and cue learning, characterized in that, Includes the following steps: Step 1: Obtain the labeled data After processing and standardizing a dataset of pathological images, the preprocessed images are obtained. A dataset of pathological images; Representing the The total number of preprocessed image instances corresponding to the dataset; the preprocessed... The first pathological image dataset Zhang's pathological images are recorded as , The corresponding actual annotation is recorded as ;and ,in, Indicates the first The true coordinates of the cell center, Indicates the first The true labeling of each cell category, and ∈{1,2,…,K}; K represents the total number of cell types; P This represents the total number of cells across all categories. for The true labeling of cell categories is The total number of cells; Get Corresponding example hints ; Get Category hints for all cell types in a dataset of pathological images ; Step 2: Construct a large-scale visual language model and cue learning cell classification network, including: an instance-aware cue learning module, a cue conditionalization module, a local visual-text matching module, and the third... An image encoding / decoding network ; Step 2.1, the instance-aware prompting learning module according to... , as well as ,get Instance-aware category text features ;in, Indicates the first k The instance-aware category text features corresponding to each cell category; Step 2.2, the An image encoding / decoding network encoder pairs in Processing is performed to obtain Image features ; Step 2.3, the conditional prompting module according to as well as ,get Conditional image features ; Step 2.4, the An image encoding / decoding network decoder pairs in Processing is performed to obtain local cell image features. ; Step 2.5, the An image encoding / decoding network Classification head pairs in Process and generate Prediction results , including: the Predicted center coordinates of each cell and the Predicted category of each cell ; , Indicates the predicted total number of cells; Step 2.6, according to the... An image encoding / decoding network right Prediction results Using equation (1), the predicted center coordinates and category of the predicted cells that match the real cells are obtained: (1) In equation (1), Represents a linear summation distributive function; and Representing the first An image encoding / decoding network The obtained number The image contains the first The predicted center coordinates and category of each predicted cell that matches the real cell; Step 3: Construct the loss function: Step 3.1: Construct using equation (2) Classification loss : (2) In equation (2), This represents the balanced cross-entropy loss function; This represents the mean absolute error loss function; Step 3.2: Construct using equation (3) Hungarian matching loss : (3) In equation (3), Indicates the first An image encoding / decoding network The obtained number In the prediction results of the first image The true cell category corresponding to each unmatched predicted cell; , and This represents three loss weighting coefficients; Step 3.3: The local visual text matching module uses equation (4) to construct the first... The local visual text matching loss function corresponding to each dataset : (4) In equation (4), Represents the feature similarity measurement function. It is a temperature parameter; Step 3.4: Construct the first... using equation (5) The overall loss corresponding to each dataset : (5) In equation (5), , and They are respectively , and Weighting coefficients; Step 4: Train a pathological image cell classification network based on a large visual language model and cue learning using the Adam optimizer, and minimize... With the goal of improving the classification network for pathological images based on a large visual language model and cue learning, we iterate the network and select the optimal classification network after several rounds of training to identify cells in the images.

2. The pathological image cell classification method based on a large visual language model and cue learning according to claim 1, characterized in that, Step 2.1 includes: Step 2.1.1, Definition The learnable token is and The learnable token is ; Step 2.1.2: Generate using equation (6) Refined category text tokens : (6) In equation (6), This represents a set of learnable parameters for a control mapping to learn token scale; Represents a mapping function; This indicates element-wise multiplication. Step 2.1.3: Use a text encoder to... and Encode it, and get the corresponding result Instance text embedding and text embedding matrix of K cell categories ;in, This indicates that the cell category is actually labeled as k Category text embedding; Step 2.1.4: Use a visual encoder to... Encode to obtain Image embedding Thus As a query for the Transformer decoder As the key and value of the Transformer decoder, and by the Transformer decoder... Perform calibration and generate Visual domain cue embedding ; Step 2.1.5: Generate using equation (7) Enhanced instance text embedding : (7) In equation (7), Represents a set of control visual field cues embedded The parameters to be learned for the scale; Step 2.1.6: Using equation (8) to obtain Specific instance text features : (8) In equation (8), Indicates a specificity mask; This represents the cross-attention operation of the mask; Indicates the layer normalization function; This represents a feedforward neural network; Step 2.1.7, and By splicing them together, we can obtain Instance-aware category text features .

3. The pathological image cell classification method based on a large visual language model and cue learning according to claim 2, characterized in that, Step 2.3 includes: Step 2.3.1: Using equation (9) to obtain Two condition vectors and : (9) In equation (9), and This represents two different multilayer perceptrons; Step 2.3.2, using equation (10) to... Conditional adjustments are performed to obtain cue-conditional image features. : (10)。 4. An electronic device, comprising a memory and a processor, characterized in that, The memory is used to store a program that supports the processor in executing the pathological image cell classification method according to any one of claims 1-3, and the processor is configured to execute the program stored in the memory.

5. A computer-readable storage medium storing a computer program thereon, characterized in that, The computer program, when run by the processor, executes the steps of the pathological image cell classification method according to any one of claims 1-3.

Citation Information

Patent Citations

  • Face attribute recognition method and system based on visual language model

    CN116778556A

  • Automatic driving decision-making method and device and medium

    CN117734729A