Image classification method and system based on multiple agents

Through a multi-agent system based on a large language model, combining visual image features with language information, we can achieve contextual semantic understanding of images and text, solve the adaptability and accuracy problems of existing image classification systems in multimodal tasks and complex scenarios, and improve the robustness and adaptability of image classification.

CN120673109APending Publication Date: 2025-09-19BEIJING SHENZHI HENGJI TECH CO LTD
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510515942.1
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-04-23
Publication Date
2025-09-19

AI Technical Summary

Technical Problem

When faced with multimodal task requirements and complex scenarios, existing image classification systems have problems such as strong data dependence, insufficient semantic understanding, weak generalization ability for small sample tasks, strong dependence on OCR regular classification, and high dependence on the quality of OCR results.

Method used

A multi-agent system based on a large language model is adopted, combining visual image features and language information. Through the redundant mechanism of multiple agents, contextual semantic understanding of images and texts is achieved, and dynamic collaboration is performed for image classification, which reduces the impact of OCR recognition errors and enhances the adaptability to different tasks.

Benefits of technology

It improves the accuracy and robustness of image classification, can flexibly respond to new rules and new scenarios, reduces dependence on specific data sets, and enhances classification capabilities in small sample tasks.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120673109A_ABST
    Figure CN120673109A_ABST
Patent Text Reader

Abstract

The invention discloses an image classification method and system based on multiple agents. The method comprises the steps of obtaining an image and a classification prompt word uploaded by a user based on a user interaction interface; the received image is preprocessed, and format conversion is carried out; converting the classification cue words into corresponding model thinking frameworks based on a large language model; respectively calling a visual model and a search tool based on the model thinking framework, and determining a classification result according to a visual identification result and a search result; and judging whether the classification result meets a model thinking framework or not through a large language model, and returning the classification result to the user interaction interface when the classification result meets the model thinking framework. Through the technical scheme of the invention, the dependence of a single model on a specific data set is reduced, the adaptability of different tasks is enhanced, new rules and new scenes can be flexibly coped with, and the accuracy and robustness of image classification are improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the field of image recognition technology, and in particular to a multi-agent based image classification method and a multi-agent based image classification system. Background Art

[0002] Image classification is a core task in computer vision, aiming to classify input images into predefined categories based on their content. Traditional image classification systems typically rely on deep learning models such as convolutional neural networks (CNNs), which are capable of learning powerful feature representations from large amounts of annotated data. However, with the increasing demand for multimodal tasks and the increasing complexity of classification scenarios, the limitations of single visual models in semantic understanding, task collaboration, and multimodal data processing are becoming increasingly apparent.

[0003] Currently, there are two existing image classification systems:

[0004] Solution 1: Image classification based on deep learning model

[0005] This approach directly classifies images by training a deep learning model (typically a convolutional neural network, or CNN). The model takes the original image as input and generates classification results through multiple layers of feature extraction and classifiers (such as fully connected layers or output layers). Common network architectures include ResNet, VGG, EfficientNet, and models based on Visual Transformers (ViT).

[0006] Implementation process:

[0007] Data preprocessing: standardize the image (such as resizing and data augmentation).

[0008] Model training: Use large-scale labeled datasets (such as ImageNet and enterprise-customized datasets) to train classification models.

[0009] Inference phase: The model inputs new images and outputs the predicted category and confidence level.

[0010] Solution 2: Regularized classification based on image OCR results

[0011] This solution uses OCR (Optical Character Recognition) technology to extract text from images and classify them using a rules engine. Text recognition primarily relies on OCR detection and recognition models. Common detection models include Faster RCNN and PSENet, while common recognition models include CRNN and TrOCR. Classification rules are determined by business needs.

[0012] The disadvantages of option 1 include:

[0013] Current disadvantage 1: Deep learning models are highly dependent on data and the training process is cumbersome;

[0014] Current disadvantage 2: It only focuses on image pixel information and lacks understanding of image semantics;

[0015] Current disadvantage 3: Weak generalization ability for small sample tasks.

[0016] The disadvantages of option 2 include:

[0017] Current disadvantage 1: OCR regularization classification is highly dependent on rules and difficult to adapt to new scenarios;

[0018] Current disadvantage 2: It is highly dependent on the quality of OCR results. Text recognition errors will lead to classification failure. Summary of the Invention

[0019] In response to the above problems, the present invention provides a multi-agent based image classification method and system. Through a multi-agent system based on a large language model, visual image features and language information are combined to achieve contextual semantic understanding of the combination of images and text, thereby dynamically collaborating to achieve image classification for different tasks under small sample tasks. At the same time, the redundant mechanism of multi-agents can reduce the impact of OCR recognition errors, reduce the dependence of a single model on a specific data set, enhance the adaptability of different tasks, and can flexibly respond to new rules and new scenarios, thereby improving the accuracy and robustness of image classification.

[0020] To achieve the above objectives, the present invention provides a multi-agent based image classification method, comprising:

[0021] Obtain images and classification prompt words uploaded by users based on the user interaction interface;

[0022] Preprocess the received image and convert its format;

[0023] Converting the classification prompt words into corresponding model thinking frameworks based on a large language model;

[0024] Based on the model thinking framework, the visual model and the search tool are respectively called, and the classification result is determined according to the visual recognition result and the search result;

[0025] The large language model is used to determine whether the classification result satisfies the model thinking framework, and if so, the classification result is returned to the user interaction interface.

[0026] In the above technical solution, preferably, the specific process of obtaining the image and classification prompt words uploaded by the user based on the user interaction interface includes:

[0027] A visual user interaction interface built based on the streamlit framework obtains images and classification prompt words uploaded by users, or receives the images and classification prompt words according to a service request initiated by the user.

[0028] In the above technical solution, preferably, the received image is preprocessed and format converted, and the specific process includes:

[0029] Perform image preprocessing on the received image and convert the image in base64 or binary stream format into a format that can be processed by the visual model.

[0030] In the above technical solution, preferably, the process of converting the classification prompt words into corresponding model thinking framework based on the large language model includes:

[0031] Determine the semantic information of the classification prompt words based on the large language model, and convert it into a model thinking framework including the question-thinking-action-observation-answer logical process according to the classification logic corresponding to the semantic information;

[0032] In the model thinking framework, the question process is used to confirm the questions that the user requires to be answered, the thinking process is used to confirm how to answer the user's questions, the action process is used to confirm the way to achieve the goal, the observation process is used to confirm the results returned by the action process, and the answer process is used to return the final answer.

[0033] In the above technical solution, preferably, the visual model and the search tool are respectively called based on the model thinking framework, and the classification result is determined according to the visual recognition result and the search result. The specific process includes:

[0034] According to the logical process of the model thinking framework, calling the visual model to perform OCR recognition on the image to obtain the OCR recognition result, and calling the search tool to search and obtain the classification rules that match the model thinking framework;

[0035] The OCR recognition results are classified and matched based on the classification rules to obtain a classification result of the image.

[0036] The present invention further proposes a multi-agent based image classification system, which applies the multi-agent based image classification method disclosed in any one of the above technical solutions, including:

[0037] The data acquisition module is used to obtain images and classification prompt words uploaded by users based on the user interaction interface;

[0038] Image conversion module, used to pre-process the received image and perform format conversion;

[0039] A framework conversion module, configured to convert the classification prompt words into corresponding model thinking frameworks based on a large language model;

[0040] A tool calling module, configured to call the visual model and the search tool respectively based on the model thinking framework, and determine the classification result according to the visual recognition result and the search result;

[0041] A result returning module is used to determine whether the classification result satisfies the model thinking framework through the large language model, and return the classification result to the user interaction interface if it satisfies the model thinking framework.

[0042] In the above technical solution, preferably, the data acquisition module is specifically used to:

[0043] A visual user interaction interface built based on the streamlit framework obtains images and classification prompt words uploaded by users, or receives the images and classification prompt words according to a service request initiated by the user.

[0044] In the above technical solution, preferably, the image conversion module is specifically used to:

[0045] Perform image preprocessing on the received image and convert the image in base64 or binary stream format into a format that can be processed by the visual model.

[0046] In the above technical solution, preferably, the frame conversion module is specifically used to:

[0047] Determine the semantic information of the classification prompt words based on the large language model, and convert it into a model thinking framework including the question-thinking-action-observation-answer logical process according to the classification logic corresponding to the semantic information;

[0048] In the model thinking framework, the question process is used to confirm the questions that the user requires to be answered, the thinking process is used to confirm how to answer the user's questions, the action process is used to confirm the way to achieve the goal, the observation process is used to confirm the results returned by the action process, and the answer process is used to return the final answer.

[0049] In the above technical solution, preferably, the tool calling module is specifically used to:

[0050] According to the logical process of the model thinking framework, calling the visual model to perform OCR recognition on the image to obtain the OCR recognition result, and calling the search tool to search and obtain the classification rules that match the model thinking framework;

[0051] The OCR recognition results are classified and matched based on the classification rules to obtain a classification result of the image.

[0052] Compared with the existing technology, the beneficial effects of the present invention are: through a multi-agent system based on a large language model, visual image features and language information are combined to achieve contextual semantic understanding of the combination of images and text, so as to dynamically collaborate to achieve image classification for different tasks under small sample tasks. At the same time, the redundant mechanism of multiple agents can reduce the impact of OCR recognition errors, reduce the dependence of a single model on a specific data set, enhance the adaptability of different tasks, and can flexibly respond to new rules and new scenarios, thereby improving the accuracy and robustness of image classification. BRIEF DESCRIPTION OF THE DRAWINGS

[0053] Figure 1 A schematic diagram of a system framework of a multi-agent based image classification method disclosed in one embodiment of the present invention;

[0054] Figure 2 A schematic diagram of the multi-agent module flow of a multi-agent based image classification method disclosed in one embodiment of the present invention. DETAILED DESCRIPTION

[0055] To make the objectives, technical solutions, and advantages of the embodiments of the present invention more clear, the technical solutions in the embodiments of the present invention will be clearly and completely described below in conjunction with the accompanying drawings in the embodiments of the present invention. Obviously, the described embodiments are part of the embodiments of the present invention, not all of the embodiments. Based on the embodiments of the present invention, all other embodiments obtained by ordinary technicians in this field without making creative efforts shall fall within the scope of protection of the present invention.

[0056] The present invention will be described in further detail below with reference to the accompanying drawings:

[0057] like Figure 1 and Figure 2 As shown, a multi-agent based image classification method provided by the present invention includes:

[0058] Obtain images and classification prompt words uploaded by users based on the user interaction interface;

[0059] Preprocess the received image and convert its format;

[0060] Convert classification prompt words into corresponding model thinking framework based on the large language model;

[0061] Based on the model thinking framework, the visual model and search tool are called separately, and the classification results are determined based on the visual recognition results and search results;

[0062] The large language model is used to determine whether the classification results meet the model thinking framework, and if so, the classification results are returned to the user interaction interface.

[0063] In this implementation, a multi-agent system based on a large language model is used to combine visual image features with language information to achieve contextual semantic understanding of images and text, thereby dynamically collaborating to achieve image classification for different tasks under small sample tasks. At the same time, the redundant mechanism of the multi-agent can reduce the impact of OCR recognition errors, reduce the dependence of a single model on a specific data set, enhance the adaptability of different tasks, and be able to flexibly respond to new rules and new scenarios, thereby improving the accuracy and robustness of image classification.

[0064] During implementation, the system primarily consists of a front-end interface and a back-end service framework. The front-end interface can be a visual interface built on the Streamlit framework, allowing users to upload images and prompts or initiate service requests to the server. The back-end provides services via FastAPI. The back-end business logic is implemented by the LangGraph agent service framework, responsible for analyzing image features and returning classification results.

[0065] The front-end visualization page is based on Streamlit, a Python-based visualization framework with advantages such as rapid development, high interactivity, and easy deployment. If the user does not need the visualization page, they can also directly initiate a request to the back-end service by sending a request.

[0066] The backend provides services through the FastAPI architecture. FastAPI is a high-performance Python-based web framework. Compared to the traditional Flask framework, FastAPI has faster running speed, higher coding efficiency, and more robust code logic.

[0067] Specifically, the backend service mainly includes two modules: data processing module and LangGraph agent module.

[0068] The data processing module is primarily responsible for preprocessing received images. The LangGraph agent module is the core module of this invention and primarily comprises a large language model (gpt-3.5-turbo), a visual model (QwenVL2), a rule base, and a search tool. Furthermore, the large language model can be replaced with models such as Claude, Illama, or QWen. The visual model can be replaced with models such as InternVL and CogVLM2.

[0069] Among them, the multi-agent system based on the large language model has strong contextual understanding capabilities and can utilize a large amount of pre-trained knowledge, eliminating the need to train from scratch for each task. Multiple agents can collaborate dynamically, reducing the dependence of a single model on a specific data set and enhancing its adaptability to different tasks. Through the introduction of the large language model, the system can combine visual (image features) and language (label description) information to achieve semantic understanding of the combination of images and text. The large language model can parse and supplement the feature information extracted from the visual model, thereby improving the accuracy of classification. The large language model has strong natural language understanding capabilities and can achieve dynamic classification through contextual reasoning and knowledge supplementation when rules are missing or changed.

[0070] Among them, multi-agents can reason through language models and combine known knowledge to complete small-sample tasks, and have stronger transfer learning capabilities. Collaboration between agents greatly improves the efficiency of small-sample learning, such as completing classification through knowledge completion or knowledge transfer of similar tasks. Multi-agent systems can flexibly respond to new rules and scenario changes through the division of labor among different agents (such as detection, interpretation, and reasoning). Moreover, multi-agent systems can use redundancy mechanisms (such as multimodal fusion) to reduce the impact of OCR errors. Image classification agents can cooperate with text analysis agents to improve robustness through mutual verification of images and texts.

[0071] By integrating image features, OCR results, and specific classification rules through a multi-agent framework, we can accurately classify even when key fields are damaged. For newly added classification categories, simply add the corresponding rules to the rule library, without the need for model training or code changes.

[0072] In the above embodiment, preferably, the specific process of obtaining the image and classification prompt words uploaded by the user based on the user interaction interface includes:

[0073] A visual user interaction interface built based on the streamlit framework, which obtains images and classification prompts uploaded by users, or receives images and classification prompts based on service requests initiated by users.

[0074] In the above embodiment, preferably, the received image is pre-processed and format converted, and the specific process includes:

[0075] Perform image preprocessing on the received image and convert the image in base64 or binary stream format into a format that can be processed by the visual model.

[0076] In the above embodiment, preferably, the classification prompt words are converted into corresponding model thinking frameworks based on the large language model. The specific process includes:

[0077] Determine the semantic information of the classification prompt words based on the large language model, and convert it into a model thinking framework including the question-thinking-action-observation-answer logical process according to the classification logic corresponding to the semantic information;

[0078] In the model thinking framework, the question process is used to confirm the questions that the user requires to be answered, the thinking process is used to confirm how to answer the user's questions, the action process is used to confirm the way to achieve the goal, the observation process is used to confirm the results returned by the action process, and the answer process is used to return the final answer.

[0079] During the implementation process, the following are the prompts for the thinking framework of the large language model:

[0080] Please answer the questions using the following format:

[0081] Question: The question the user asks you to answer;

[0082] Think: Think about how to answer users’ questions;

[0083] Action: What should you do to achieve this goal?

[0084] Observe: Observe the results returned by the action (the above process may be repeated multiple times);

[0085] Final Answer: Returns the final answer.

[0086] In the above embodiment, preferably, the visual model and the search tool are respectively called based on the model thinking framework, and the classification result is determined according to the visual recognition result and the search result. The specific process includes:

[0087] According to the logical process of the model thinking framework, the visual model is called to perform OCR recognition on the image to obtain the OCR recognition result, and the search tool is called to search for the classification rules that match the model thinking framework;

[0088] The OCR recognition results are classified and matched based on the classification rules to obtain the classification results of the image.

[0089] According to the multi-agent based image classification method disclosed in the above embodiment, during implementation, the specific process of the method is explained through the following examples.

[0090] Example:

[0091] The user uploads a picture and asks to classify whether the picture is an outpatient invoice, outpatient list or other picture.

[0092] 1. The user sends a request to the backend through the frontend page or direct request;

[0093] 2. After receiving the service request, the backend service framework first calls the data processing module to perform pre-processing operations such as format conversion on the image uploaded by the user;

[0094] 3. The large language model receives the user's question and begins iterative processing according to the corresponding model thinking framework;

[0095] Question: The user requires to classify whether the image is an outpatient invoice, outpatient list or other.

[0096] Consider: To determine the image category, we need the image’s visual features, image OCR results, and specific classification rules.

[0097] Action: To obtain the visual features and OCR results of an image, a visual language model needs to be used. To obtain classification rules, a search tool is required.

[0098] Action: Invoke the visual model, invoke the search tool;

[0099] Observation: Obtain the image's visual features, OCR results, and classification rules to determine whether the current image belongs to an outpatient invoice;

[0100] Final answer: Return outpatient invoice.

[0101] The present invention further proposes a multi-agent based image classification system, which applies the multi-agent based image classification method disclosed in any of the above embodiments, including:

[0102] The data acquisition module is used to obtain images and classification prompt words uploaded by users based on the user interaction interface;

[0103] Image conversion module, used to pre-process the received image and perform format conversion;

[0104] The framework conversion module is used to convert the classification prompt words into the corresponding model thinking framework based on the large language model;

[0105] The tool calling module is used to call the visual model and search tool based on the model thinking framework, and determine the classification results based on the visual recognition results and search results;

[0106] The result return module is used to determine whether the classification result meets the model thinking framework through the large language model, and return the classification result to the user interaction interface if it meets the requirements.

[0107] In the above embodiment, preferably, the data acquisition module is specifically used to:

[0108] A visual user interaction interface built based on the streamlit framework, which obtains images and classification prompts uploaded by users, or receives images and classification prompts based on service requests initiated by users.

[0109] In the above embodiment, preferably, the image conversion module is specifically used to:

[0110] Perform image preprocessing on the received image and convert the image in base64 or binary stream format into a format that can be processed by the visual model.

[0111] In the above embodiment, preferably, the framework conversion module is specifically used to:

[0112] Determine the semantic information of the classification prompt words based on the large language model, and convert it into a model thinking framework including the question-thinking-action-observation-answer logical process according to the classification logic corresponding to the semantic information;

[0113] In the model thinking framework, the question process is used to confirm the questions that the user requires to be answered, the thinking process is used to confirm how to answer the user's questions, the action process is used to confirm the way to achieve the goal, the observation process is used to confirm the results returned by the action process, and the answer process is used to return the final answer.

[0114] In the above embodiment, preferably, the tool calling module is specifically used to:

[0115] According to the logical process of the model thinking framework, the visual model is called to perform OCR recognition on the image to obtain the OCR recognition result, and the search tool is called to search for the classification rules that match the model thinking framework;

[0116] The OCR recognition results are classified and matched based on the classification rules to obtain the classification results of the image.

[0117] According to the multi-agent-based image classification system disclosed in the above-mentioned embodiment, the functions to be implemented by each module thereof correspond to the respective steps of the multi-agent-based image classification method disclosed in the above-mentioned embodiment. During the implementation process, operations are performed with reference to the above-mentioned embodiment, which will not be repeated here.

[0118] The above are merely preferred embodiments of the present invention and are not intended to limit the present invention. Those skilled in the art will readily appreciate that various modifications and variations of the present invention are possible. Any modifications, equivalent substitutions, or improvements made within the spirit and principles of the present invention shall be included within the scope of protection of the present invention.

Claims

1. A multi-agent based image classification method, characterized in that: include: Obtain images and classification prompt words uploaded by users based on the user interaction interface; Preprocess the received image and convert its format; Converting the classification prompt words into corresponding model thinking frameworks based on a large language model; Based on the model thinking framework, the visual model and the search tool are respectively called, and the classification result is determined according to the visual recognition result and the search result; The large language model is used to determine whether the classification result satisfies the model thinking framework, and if so, the classification result is returned to the user interaction interface.

2. The multi-agent based image classification method according to claim 1, characterized in that: The specific process of obtaining the image and classification prompt words uploaded by the user based on the user interaction interface includes: A visual user interaction interface built based on the streamlit framework obtains images and classification prompt words uploaded by users, or receives the images and classification prompt words according to a service request initiated by the user.

3. The multi-agent based image classification method according to claim 2, characterized in that: The received image is pre-processed and format converted, and the specific process includes: Perform image preprocessing on the received image and convert the image in base64 or binary stream format into a format that can be processed by the visual model.

4. The multi-agent based image classification method according to claim 3, characterized in that: The process of converting the classification prompt words into corresponding model thinking framework based on the large language model includes: Determine the semantic information of the classification prompt words based on the large language model, and convert it into a model thinking framework including the question-thinking-action-observation-answer logical process according to the classification logic corresponding to the semantic information; In the model thinking framework, the question process is used to confirm the questions that the user requires to be answered, the thinking process is used to confirm how to answer the user's questions, the action process is used to confirm the way to achieve the goal, the observation process is used to confirm the results returned by the action process, and the answer process is used to return the final answer.

5. The multi-agent based image classification method according to claim 4, characterized in that: Based on the model thinking framework, the visual model and search tool are called respectively, and the classification results are determined based on the visual recognition results and search results. The specific process includes: According to the logical process of the model thinking framework, calling the visual model to perform OCR recognition on the image to obtain the OCR recognition result, and calling the search tool to search and obtain the classification rules that match the model thinking framework; The OCR recognition results are classified and matched based on the classification rules to obtain a classification result of the image.

6. A multi-agent based image classification system, characterized in that: Applying the multi-agent based image classification method according to any one of claims 1 to 5, comprising: The data acquisition module is used to obtain images and classification prompt words uploaded by users based on the user interaction interface; Image conversion module, used to pre-process the received image and perform format conversion; A framework conversion module, configured to convert the classification prompt words into corresponding model thinking frameworks based on a large language model; A tool calling module, configured to call the visual model and the search tool respectively based on the model thinking framework, and determine the classification result according to the visual recognition result and the search result; A result returning module is used to determine whether the classification result satisfies the model thinking framework through the large language model, and return the classification result to the user interaction interface if it satisfies the model thinking framework.

7. The multi-agent based image classification system according to claim 6, characterized in that: The data acquisition module is specifically used for: A visual user interaction interface built based on the streamlit framework obtains images and classification prompt words uploaded by users, or receives the images and classification prompt words according to a service request initiated by the user.

8. The multi-agent based image classification system according to claim 7, characterized in that: The image conversion module is specifically used for: Perform image preprocessing on the received image and convert the image in base64 or binary stream format into a format that can be processed by the visual model.

9. The multi-agent based image classification system according to claim 8, characterized in that: The framework conversion module is specifically used for: Determine the semantic information of the classification prompt words based on the large language model, and convert it into a model thinking framework including the question-thinking-action-observation-answer logical process according to the classification logic corresponding to the semantic information; In the model thinking framework, the question process is used to confirm the questions that the user requires to be answered, the thinking process is used to confirm how to answer the user's questions, the action process is used to confirm the way to achieve the goal, the observation process is used to confirm the results returned by the action process, and the answer process is used to return the final answer.

10. The multi-agent based image classification system according to claim 9, characterized in that: The tool calling module is specifically used for: According to the logical process of the model thinking framework, calling the visual model to perform OCR recognition on the image to obtain the OCR recognition result, and calling the search tool to search and obtain the classification rules that match the model thinking framework; The OCR recognition results are classified and matched based on the classification rules to obtain a classification result of the image.