Multi-agent-based red blood cell image identification and classification method
Through the iterative voting method of multi-agent system combining segmentation model, visual language model and large language model, efficient identification and classification of red blood cell images is achieved, time-consuming and inefficient problems in the existing technology are solved, and it has good interpretability and low resource consumption characteristics.
Patent Information
- Application Number
- CN202411861572.9
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2024-12-17
- Publication Date
- 2025-05-23
AI Technical Summary
The existing cell classification methods require a large number of training and testing samples, which are time-consuming and not efficient enough, and the visual language model is insensitive in image feature recognition in professional fields and cannot meet the needs of medical staff.
A multi-agent-based red blood cell image recognition and classification method is used to extract cell morphological characteristics through segmentation models, combine visual language models and large language models, and generate inference trees using iterative voting to identify and classify cell images.
This method requires only a very small sample to achieve high classification capabilities, has good interpretability, significantly reduces the workload of professionals on data set annotation and meets the needs of doctors in the process of disease diagnosis.
Smart Images

Figure CN120032156A_ABST
Abstract
Description
Technical Field
[0001] The present invention belongs to the technical field of cell image classification, and specifically designs a multi-agent-based red blood cell image recognition and classification method. Background Art
[0002] The complete blood count test is one of the most valuable medical tests that provides important information about the cellular blood composition, is used to evaluate the patient's overall health and is used to detect various diseases. In the complete blood count, the correct classification of abnormal red blood cells (RBC) is crucial because cellular abnormalities are closely related to many diseases caused by red blood cell disorders (such as anemia, thalassemia, sickle cell disease, etc.). Blood smear analysis is traditionally performed through manual inspection, which is very time-consuming, requires highly trained professionals, and is easily affected by subjective factors.
[0003] In recent years, deep learning has developed rapidly, and there are many methods for cell classification. Previous cell classification methods were mainly implemented using deep neural networks. In order to achieve good classification results, this method requires a large number of training and test samples provided by professionals, and the cycle usually takes several months. The Vision Language Model (VLM) provides a possibility for solving this challenge, but since most vision language models are trained on natural images, they are not sensitive to image feature recognition in this professional field, and the recognition effect is poor, so they cannot be used by medical personnel. Fine-tuning the model requires a large amount of training data, which cannot really solve the problem.
[0004] Recently, autonomous agents based on Large Language Models (LLMs) have attracted great interest from both industry and academia. Many studies have improved the problem-solving capabilities of artificial intelligence based on LLMs by integrating discussions among multiple agents. These systems support complex interactions and processes by simulating the collaborative intelligence of human teams. Each agent in a multi-agent system is an autonomous entity that can perform specific tasks, make decisions, and collaborate to achieve common goals. Summary of the invention
[0005] In this context, the purpose of the present invention is to provide a multi-agent-based red blood cell image recognition and classification method, which first uses a segmentation model and a visual language model to extract morphological features from cell images, then uses the powerful reasoning ability of a large language model to generate an inference tree by iterative voting, and finally the inference tree infers the correct category from the cell image and its feature description. The largest public cell dataset Pathlogic was selected for experiments, and the experimental results showed that this method only requires very few samples to achieve high classification capabilities, and has good interpretability because it provides doctors with rich descriptive information during the classification process.
[0006] To achieve the above object, the technical solution adopted by the present invention is as follows:
[0007] The present invention is a multi-agent based red blood cell image recognition and classification method, comprising the following steps:
[0008] 1) Using the segmentation model to obtain the segmentation mask of the cell from the cell image, and calculating the minimum circumscribed rectangular image of the cell from the segmentation mask to obtain a single cell image;
[0009] 2) extracting numerical features from the segmentation mask of the single cell image and converting the numerical features into text descriptions;
[0010] 3) Classifying the single cell images into different types according to the numerical features, selecting a set number of image samples for each type, inputting these cell images and attribute extraction prompts into the visual question answering model, and obtaining the attributes extracted by the model and the description of the attributes;
[0011] 4) Building an inference tree for red blood cell classification based on attribute and category pairs;
[0012] 5) Use the red blood cell attribute description and the red blood cell classification inference tree to predict the category of the cell image.
[0013] Furthermore, the step 1) specifically includes: using a segmentation model to obtain a segmentation mask of the cell from the cell image, and calculating a minimum circumscribed rectangular image of the cell;
[0014] The segmentation model consists of an image encoder, a hint encoder and a mask decoder. The image encoder is used to obtain an image embedding vector, the hint encoder is used to obtain a hint embedding vector, and the mask decoder is used to obtain a segmentation mask based on the image embedding vector and the hint embedding vector.
[0015] Furthermore, the minimum bounding rectangle image refers to a directed bounding box of all parts of the mask output by the mask decoder that are equal to 1.
[0016] Furthermore, the step 2) specifically includes:
[0017] For the minimum bounding rectangle image OBB, the aspect ratio r aspect and area A cell Respectively expressed as:
[0018]
[0019] A cell =∑ i,j M(i,j);
[0020] Among them, M(i,j) is the value of the pixel point (i,j) in the cell segmentation mask, and the sum is the number of pixels occupied by the cell;
[0021] The aspect ratio r aspect and area A cell Through the conversion function f text Converted into text description T:
[0022] T=f text (r aspect ,A cell );
[0023] where f text is the conversion function.
[0024] Furthermore, the step 4) specifically includes:
[0025] S41 initially has all attributes constituting an attribute set A and all categories C, and uses a large language model LLM to select an attribute from them, and sets the attribute to distinguish one or several categories from all other categories;
[0026] S42 executes step S41 several times, counts the number of votes for each attribute and category, selects the attribute a1 and category c1 with the most votes as the <attribute, category> pair selected in this round; removes a1 and c1 from the attribute set A and all categories C respectively; and performs the next round of voting until the attribute set A is the same as the attribute set C.
[0027] S43 sorts the discriminations based on the attributes, and builds a tree from the root node downward in the sorted order, where each leaf node is a category and each non-leaf node represents a discriminant attribute, thereby obtaining the inference tree.
[0028] Furthermore, the step 5) specifically includes: first inputting the cell image into the visual question answering model, extracting the attribute description of the cell, and then inputting the attribute description into the inference tree to obtain the predicted category.
[0029] Furthermore, the mask decoder maps the image embedding vector, the cue embedding vector and the output word unit to the mask; the mask decoder updates all embeddings using cue self-attention and cross-attention in two directions; the image embedding is upsampled, the MLP maps the output word unit to a dynamic linear classifier, and then calculates the mask foreground probability of each image position.
[0030] Furthermore, before the prompt embedding vector enters the mask decoder, a set of learnable output word-grams are spliced on the prompt embedding vector, and the output word-grams are composed of two parts: intersection-over-union word-grams and mask word-grams; finally, the prompt embedding vector, the intersection-over-union word-grams and the mask word-grams are spliced together and input into the mask decoder.
[0031] Furthermore, the GPT 3.5-turbo model is used as the large language model LLM.
[0032] Beneficial effects of the present invention:
[0033] The present invention can save the traditional deep learning model training process and no longer consume a lot of time and resources; the present invention only requires very few samples to achieve high classification capabilities and can significantly reduce the workload of professionals for data set annotation; the present invention has good interpretability because it provides doctors with rich descriptive information during the classification process, thus meeting the needs of doctors in the process of disease diagnosis. BRIEF DESCRIPTION OF THE DRAWINGS
[0034] Figure 1 This is the overall process framework diagram of the multi-agent-based red blood cell image recognition and classification method of the present application;
[0035] Figure 2 is a flowchart of image segmentation and numerical feature extraction in an embodiment of the present application;
[0036] Figure 3 is a voting flow chart of attribute category pairs in an embodiment of the present application;
[0037] Figure 4 is a box plot of the area of the mask obtained by the segmentation model in the embodiment of the present application;
[0038] Figure 5 It is a box plot of the aspect ratio of the mask obtained by the segmentation model in the embodiment of the present application. DETAILED DESCRIPTION
[0039] In order to facilitate the understanding of those skilled in the art, the present invention is further described below in conjunction with embodiments and drawings. The contents mentioned in the implementation modes are not intended to limit the present invention.
[0040] Reference Figure 1The present invention is a multi-agent based red blood cell image recognition and classification method, comprising the following steps:
[0041] 1) Use the segmentation model to obtain the segmentation mask of the cell from the cell image, and calculate the minimum bounding rectangle image of the cell from the segmentation mask.
[0042] 2) Calculate the basic numerical features of the red blood cell image from the minimum bounding rectangle image obtained in step 1) and convert them into text descriptions.
[0043] 3) Select a certain number of image samples from each category of the dataset, input these cell images and attribute extraction prompts into the visual question answering model, and obtain the attributes extracted by the model and the description of the attributes.
[0044] 4) Design prompt sentences that are easy for the model to understand for extracting red blood cell image attributes and attribute classification, input the attribute classification prompts into the large language model, and after multiple rounds of iterative voting, finally build a tree with the attribute and category pairs selected by the vote.
[0045] 5) First, input the cell images in the dataset into the visual question answering model, extract the attribute description of the cell, and then input the attribute description into the inference tree to obtain the predicted category.
[0046] Furthermore, the segmentation model in the step 1) is composed of an image encoder, a prompt encoder and a mask decoder. The image is input into the image encoder to obtain an image embedding vector, and the prompt is input into the prompt encoder to obtain a prompt embedding vector. Finally, both are input into the mask decoder to output a segmentation mask. The segmentation model of the present invention is implemented using a segmentation anything model (SAM).
[0047] The image encoder uses a visual encoder (ViT) trained with a masked autoencoder (MAE), which outputs an image embedding of shape C×H×W after 16x downsampling of the input image.
[0048] Hints in the hint encoder are points or axis-aligned boxes, which are represented by positional encodings in this invention. Point hints are the sum of the positional encoding of the point and one of two learned embeddings, which represent whether the point is in the foreground or background. Axis-aligned boxes are represented by an embedding pair: (1) the positional encoding of its top-left corner is added to the learned embedding representing "top-left corner"; (2) the same structure, but using the learned embedding representing "bottom-right corner".
[0049] The mask decoder maps the image embedding vector, the cue embedding vector, and the output token to the mask. It is modified from the Transformer decoder, and the modified decoder block uses cue self-attention and cross-attention in two directions (cue-image embedding and vice versa) to update all embeddings. After running both blocks, the image embedding is upsampled, and the MLP maps the output token to a dynamic linear classifier, which then calculates the mask foreground probability for each image position.
[0050] Before the hint embedding enters the mask decoder, a set of learnable output tokens are concatenated on it. This output token consists of two parts: one is the intersection over union (iou) token, which will be separated out later to predict the reliability of the iou. It is supervised by the MSE loss between the iou calculated by the model and the iou between the mask calculated by the model and the real mask; the other is the mask token, which will also be separated out later to participate in the prediction of the final mask. The mask token is supervised by a weighted combination of focal loss and dice loss 20:1. The final hint embedding and the above two tokens are concatenated together and input into the mask decoder.
[0051] Furthermore, the minimum bounding rectangle in step 1) refers to the oriented bounding box (OBB) of all parts of the mask output by the mask decoder that are equal to 1. The minimum bounding rectangle can more accurately reflect the shape characteristics of the object. The basic principle is to use the least squares method to fit the point set on the convex hull, and then find the intersection of the fitted straight line and the edge on the convex hull to obtain the four vertices of the minimum bounding rectangle.
[0052] Specifically, its input is a point set consisting of all pixels in the mask that are equal to 1, and finally returns a bounding rectangle of the cell outline, which contains information such as the center, width, height, rotation angle, etc. of the minimum bounding rectangle.
[0053] It can be formally described as: given an input cell image I, through the segmentation model f seg Get the cell segmentation mask M:
[0054] M=f seg (I);
[0055] Next, based on the segmentation mask M, the minimum bounding rectangle is calculated, denoted as OBB(M), and its output is the length L, width W, and area A of the minimum bounding rectangle OBB And the direction angle θ:
[0056] OBB(N)=(L,M,θ)
[0057] A OBB=L×W
[0058] Furthermore, the basic numerical features in step 2) refer to the aspect ratio and area of the red blood cells. The calculation process is as follows: for the minimum bounding rectangle image OBB(M), the aspect ratio r aspect and area A cell Respectively expressed as:
[0059]
[0060] Among them, M(i,j) is the value of the pixel point (i,j) in the cell segmentation mask, and the sum is the number of pixels occupied by the cell. Finally, the aspect ratio r aspect and area A cell Through the conversion function f text Converted into text description T:
[0061] T=f text (r aspect ,A cell )
[0062] The conversion function f text It is based on the threshold value obtained from the statistical data of a batch of randomly selected image samples to convert to the corresponding text description, and the data statistics of aspect ratio and area are analyzed using box plots. Figure 4 Box plot for area statistics. Figure 5 Box plot of aspect ratio statistics.
[0063] Furthermore, the attribute description extraction in step 3) uses the visual question answering model to extract features from the cell image. This step is divided into two sub-steps, namely attribute extraction and attribute description. The attribute extraction step is as follows: 1) Partial sampling is performed from each category of the data set. 2) These samples and the attribute extraction prompts are input into the visual question answering model, and the visual question answering model will output a list of candidate attributes. This process is repeated several times, and the results of each experiment are merged and then deduplicated. In this way, a set of attributes for cell classification can be obtained.
[0064] Furthermore, the attribute extraction in step 3) uses a visual question answering model, which is used to extract visual attributes and attribute descriptions in the present invention. The visual question answering model combines a large language model and a visual language model in a framework, wherein the large language model is based on an autoregressive structure and is trained on a large-scale Internet text corpus, and its model weights encode rich knowledge. Through this knowledge, it can guide specific tasks through input and output examples related to the current task without fine-tuning the weights. This ability is called contextual learning. When the prompt design is reasonable, the large language model can show sufficient reasoning ability when answering questions. The visual language model consists of image and text encoders, which are jointly trained to predict the correct pairing of noisy images and text pairs; the visual language model maps images to a category in a given vocabulary according to the maximum cosine similarity. Since the image-text representation space is well aligned, as long as the class name of the test set is known a priori, they can have excellent zero-sample performance on unseen data sets. Therefore, the visual question answering model not only learns to align visual and text patterns, but also learns to generate titles for image-text pairs. At inference time, a visual question answering model receives an image and a text question as input, and outputs a text answer.
[0065] Furthermore, the iterative voting in step 4) utilizes the powerful reasoning ability of the large language model LLM to build these attributes into an inference tree using a voting-based method. Initially, all attributes constitute the attribute set A and all categories C. Then let the large language model LLM select an attribute from them, which can distinguish one or several categories from all other categories. This step is performed n times, and then the number of votes for each attribute and category is counted, and the attribute a1 and category c1 with the most votes are selected as the <attribute, category> pair selected in this round. Then a1 and c1 are removed from A and C respectively. The next round of voting is carried out until the attribute set is The algorithm flow is shown in Table 1.
[0066] Table 1
[0067]
[0068] Next, we sorted the attributes based on their discrimination, and observed that SAM was better than the visual language model VLM in distinguishing attributes, and the discrimination of a single cell (single) was usually higher than that of multiple cells (multiple). Therefore, we prioritized the combination of SAM>VLM and single>multiple when sorting, and built a tree from the root node downwards in the sorted order. Each leaf node was a category, and each non-leaf node (including the root node) represented a discriminant attribute.
[0069] Furthermore, the prediction in step 5) refers to predicting the category of the cell image. First, the cell image is segmented using the process in step 1 to obtain a segmentation mask, and two attributes, area and aspect ratio, are obtained from the segmentation mask. Next, the attributes obtained from step 3 and the image are input into the visual question answering model to obtain a description of these attributes. Finally, the attribute description of the cell image is input into the inference tree to obtain the final predicted category.
[0070] The present invention selected Pathlogic, the largest public blood cell dataset currently, for experiments, and sampled an average of 100 samples for each category for experiments. Each cell image is 80*80 pixels in size. The segmentation model uses SAM. In the SAM stage, since the cells are pre-placed in the middle, the center point of the image is selected as the point prompt of SAM. After SAM performs segmentation, the segmentation mask is obtained. The segmentation result example is shown in the figure below. Next, the two attributes of the mask, the area and the aspect ratio, are calculated, and the distribution of each category on the two attributes is analyzed, as shown in the figure. The results in the figure show that 99.7% of the fragmented cells have an area less than 450, and 98.6% of the round cells have an aspect ratio less than 1.1. Since these two types of cells have significantly different distributions in the two attributes, the two attributes will be merged in the subsequent reasoning tree in the form of text prompts.
[0071] We select llava1.6-34B as the VLM to perform VQA queries on attributes. We first select some samples from each category and ask the VLM to describe their features in detail. Then we select representative attribute features from the VLM's answers. These attribute features and the two attributes in the SAM stage together form the attribute set, which is used for the subsequent construction of the inference tree and inference execution.
[0072] In the inference tree building stage, the GPT 3.5-turbo model is used as the selected LLM. The number of samples for each step is n=5. In the prompt, tell the LLM each cell category and attribute set and let it select the category and attribute that can be distinguished from all other categories. Finally, calculate the number of votes for each attribute and category from the n answers, then save the category and attribute with the highest votes and delete it from the attribute set and category set, so that a round of voting is completed, and repeat this process until the attribute set is empty. The experimental results are shown in the table. After obtaining the set of attribute-category pairs, sort them according to the sorting rules described above, and then build them into an inference tree.
[0073] Table 2 Attribute-category pair voting process
[0074]
[0075] After completing the above steps, 100 samples were selected from each category to conduct experiments on this inference tree. The experimental results are as follows:
[0076] Table 3 Accuracy of methods in various categories of PathLogic
[0077]
[0078]
[0079] In addition, different classification methods were selected to experiment on the task, and the experimental results are shown in the table. In the supervised method, EfficientNetB0, which is commonly used in previous related work, and a combination of powerful feature extractors dinov2 and mlp were selected. In the zero-sample inference method, clip was selected for the experiment. By calculating the similarity between the image and each category name, the prediction result is the category name with the highest similarity. In the method of the present invention, the experiment was conducted on the FineR method which is similar to the method. The difference of this method is mainly that the inference tree step is missing. It allows LLM to directly predict the category name in all attribute descriptions. In order to facilitate comparison, the method is modified from an open set to a closed set and the category name is directly output as the prediction result in the last step.
[0080] Table 4 Experimental results of the method of the present invention and other methods
[0081]
[0082] Experimental results show that the accuracy of the method of the present invention in experiments on large red blood cell datasets is 79.50%, which exceeds zero-sample reasoning and previous classification methods that do not require further fine-tuning and additional training data. It not only has low resource consumption without further fine-tuning and additional training data, but also does not require a large number of annotation labels, effectively reducing the workload of professionals and has a good balance between efficiency and quality.
[0083] The present invention combines the powerful recognition and reasoning capabilities of multiple large models to build a multi-agent framework for red blood cell recognition and classification. First, it abandons the traditional deep learning paradigm of training models-model reasoning, and instead makes full use of the large language model's large knowledge reserves and the visual question-answering model's good recognition capabilities for images and texts, organically combining them to extract attributes and attribute descriptions that are beneficial to classification from cell images. The present invention also designs an iterative voting algorithm for these attributes and their corresponding categories. Through multiple rounds of voting, the attributes are ranked and the attributes are built into an inference tree, which is used to predict the category of the cell image based on the attributes.
[0084] There are many practical application ways of the present invention. The above is only the preferred implementation mode of the present invention. It should be pointed out that for ordinary technicians in this technical field, several improvements can be made without departing from the principle of the present invention. These improvements should also be regarded as the scope of protection of the present invention.
Claims
1. A multi-agent based red blood cell image recognition and classification method, characterized in that: The steps include: 1) Using the segmentation model to obtain the segmentation mask of the cell from the cell image, and calculating the minimum circumscribed rectangular image of the cell from the segmentation mask to obtain a single cell image; 2) extracting numerical features from the segmentation mask of the single cell image and converting the numerical features into text descriptions; 3) Classifying the single cell images into different types according to the numerical features, selecting a set number of image samples for each type, inputting these cell images and attribute extraction prompts into the visual question answering model, and obtaining the attributes extracted by the model and the description of the attributes; 4) Building an inference tree for red blood cell classification based on attribute and category pairs; 5) Use the red blood cell attribute description and the red blood cell classification inference tree to predict the category of the cell image.
2. The multi-agent based red blood cell image recognition and classification method according to claim 1, characterized in that: The step 1) specifically includes: using a segmentation model to obtain a segmentation mask of a cell from a cell image, and calculating a minimum circumscribed rectangular image of the cell; The segmentation model consists of an image encoder, a hint encoder and a mask decoder. The image encoder is used to obtain an image embedding vector, the hint encoder is used to obtain a hint embedding vector, and the mask decoder is used to obtain a segmentation mask based on the image embedding vector and the hint embedding vector.
3. The multi-agent based red blood cell image recognition and classification method according to claim 2, characterized in that: The minimum bounding rectangle image refers to a directed bounding box of all parts equal to 1 in the mask output by the mask decoder.
4. The multi-agent based red blood cell image recognition and classification method according to claim 1, characterized in that: The step 2) specifically includes: For the minimum bounding rectangle image OBB, the aspect ratio r aspect and area A cell Respectively expressed as: A cell =∑ i,j M(i,j); Among them, M(i,j) is the value of the pixel point (i,j) in the cell segmentation mask, and the sum is the number of pixels occupied by the cell; The aspect ratio r aspect and area A cell Through the conversion function f text Convert to text description T: T = f text (r aspect ,A cell ); where f text is the conversion function.
5. The multi-agent based red blood cell image recognition and classification method according to claim 1, characterized in that: The step 4) specifically includes: S41 initially has all attributes constituting an attribute set A and all categories C, and uses a large language model LLM to select an attribute from them, and sets the attribute to distinguish one or several categories from all other categories; S42 executes step S41 several times, counts the number of votes for each attribute and category, selects the attribute a1 and category c1 with the most votes as the <attribute, category> pair selected in this round; removes a1 and c1 from the attribute set A and all categories C respectively; and performs the next round of voting until the attribute set A is the same as the attribute set C. S43 sorts the discriminations based on the attributes, and builds a tree from the root node downward in the sorted order, where each leaf node is a category and each non-leaf node represents a discriminant attribute, thereby obtaining the inference tree.
6. The multi-agent based red blood cell image recognition and classification method according to claim 1, characterized in that: The step 5) specifically includes: first inputting the cell image into the visual question answering model, extracting the attribute description of the cell, and then inputting the attribute description into the inference tree to obtain the predicted category.
7. The multi-agent based red blood cell image recognition and classification method according to claim 3, characterized in that: The mask decoder maps the image embedding vector, the cue embedding vector and the output token to a mask; the mask decoder updates all embeddings using cue self-attention and cross-attention in two directions; The image embedding is upsampled, and the MLP maps the output tokens to a dynamic linear classifier, which then computes the mask foreground probability for each image location.
8. The multi-agent based red blood cell image recognition and classification method according to claim 1, characterized in that: Before the hint embedding vector enters the mask decoder, a set of learnable output word-grams are spliced on the hint embedding vector, and the output word-grams are composed of two parts: intersection-over-union word-grams and mask word-grams; finally, the hint embedding vector, the intersection-over-union word-grams and the mask word-grams are spliced together and input into the mask decoder.
9. The multi-agent based red blood cell image recognition and classification method according to claim 1, characterized in that: Use the GPT 3.5-turbo model as the large language model LLM.