Image classification data construction method based on multi-modal Agent
Through the image classification data construction method based on multimodal Agent, the interaction between multimodal large language model and Agent is used to realize the automated construction, expansion and creation of image classification data sets, solving the problem of not being able to achieve full automation and batch processing in the existing technology, and significantly improving operation efficiency and accuracy.
Patent Information
- Application Number
- CN202510082819.5
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-01-20
- Publication Date
- 2025-05-13
AI Technical Summary
The prior art cannot fully automate, batch scale and create image classification data sets, and cannot effectively process and integrate real-world images.
The image classification data construction method based on multimodal Agent is adopted, and the image classification data set is automatically constructed, expanded and created through the interaction between multimodal large language model and Agent. The method includes steps such as image evaluation, data set information extraction, image classification and quality improvement.
Fully automation and batch processing of image classification data sets are realized, which significantly reduces manpower consumption, reduces data set expansion and creation costs, improves operational efficiency, and maintains a high degree of accuracy.
Smart Images

Figure CN119992195A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of image data processing, and in particular to a method for constructing image classification data based on a multimodal agent. Background Art
[0002] The methods of expanding and building image datasets in traditional image classification dataset expansion and creation are usually manual, and there are semi-automatic methods that can use machine assistance, but they all require human participation almost throughout the process and are not fully automated. Or automatic image generation or synthesis, but the generated or synthesized images cannot fully consider the diverse perspectives, lighting, and environmental conditions, resulting in the quality of the dataset being inferior to that of the dataset built using real-world photos.
[0003] Currently, there is no process that can fully automate the construction of image datasets, nor is there a process that can process and seamlessly integrate captured or recorded real-world images into these datasets; the purpose of the present invention is to address the defects of Agent application in the expansion and creation of image classification datasets, fill the gap in the Agent field in this regard, and allow users to provide images at will when expanding and creating image classification datasets, including randomly taken photos, images and videos obtained from the Internet, etc. These images are filtered, parameters are calibrated, and processed according to the image dataset specifications, thereby achieving full automation, batch expansion, and creation of image classification datasets. Summary of the invention
[0004] The purpose of this section is to summarize some aspects of embodiments of the present invention and briefly introduce some preferred embodiments. Some simplifications or omissions may be made in this section and the specification abstract and the invention title of this application to avoid blurring the purpose of this section, the specification abstract and the invention title, and such simplifications or omissions cannot be used to limit the scope of the present invention.
[0005] In view of the problems existing in the above-mentioned existing method for constructing image classification data based on multimodal agent, the present invention is proposed.
[0006] Therefore, the purpose of the present invention is to provide a method for constructing image classification data based on multimodal agent, which realizes full automation and batch processing of image classification data set expansion and creation, significantly reduces manpower consumption, reduces the cost of data set expansion and creation, and improves operation efficiency compared with traditional methods while maintaining a high degree of accuracy.
[0007] In order to solve the above technical problems, the present invention provides the following technical solutions: a method for constructing image classification data based on multimodal agent, comprising the following steps:
[0008] Input image and dataset information: including:
[0009] 1) Image data, photo and video formats collected from various sources, along with details of the intended image classification dataset;
[0010] 2) Enter the dataset name directly;
[0011] The agent will perform online search and integrate pre-trained model data to retrieve relevant dataset information;
[0012] After the user inputs their needs according to the prompts, the Agent begins to interact with the large model, enabling the model to use the provided operations to develop a plan that meets the user's needs, and will automatically call operations according to the developed plan to complete the specified task.
[0013] As a preferred solution of the method for constructing image classification data based on multimodal agent described in the present invention, if the image format detected during the image input process is a video, it needs to be frame extracted and decomposed to convert the video into many images for the expansion and establishment of the data set.
[0014] As a preferred solution of the method for constructing image classification data based on multimodal agent of the present invention, the input process includes the following formula:
[0015] T=JSON_Transfer(instruction=Φ(I+P))
[0016] Where I represents the initial data provided by the user; P represents the prompt for supplementary context information;
[0017] The formula describes the task generation process. The input data I and prompt P are combined into instructions through the mapping function Φ. The instructions express the specific requirements and background information of the task. The instructions are converted into standardized JSON format tasks T through the function JSON_Transfer.
[0018] As a preferred solution of the method for constructing image classification data based on multimodal agent of the present invention, wherein: constructing the image classification data set includes: ImageAction, LabelAction and ClassifyAction;
[0019] Among them, ImageAction is specifically:
[0020] After the input image is evaluated, it is transmitted to a multimodal large language model that can read the image and make the following judgments on the image:
[0021] a) Whether the image contains a single or multiple objects, the output is single or multiple;
[0022] b) the content category in the image;
[0023] Each interaction involves thinking and analysis. After the interaction, the model will think whether the result is reasonable and conforms to the prescribed paradigm. If it conforms, the judgment result will be output and the above two image information will be output and saved. If it does not conform, the image will be read again for analysis until the model believes that the output content is reasonable after thinking, and then the above two image information will be output as text information in JSON format.
[0024] Among them, LabelAction is specifically:
[0025] Responsible for extracting and compressing data set information, conveying data set related information to a large language model that can analyze text, and using the designed prompt words to determine the following information contained in the data set:
[0026] a) The type of each class in the image classification dataset;
[0027] b) The resolution of the images in the dataset required by the target image classification dataset;
[0028] c) The size of the images in the dataset required by the target image classification dataset;
[0029] The judgment process is to conduct cyclic thinking and analysis like ImageAction, and then output the results recognized by the big model as text information in JSON format;
[0030] Among them, ClassifyAction is specifically:
[0031] Using the information about the image and image dataset fed back by the big model, we can identify the intersection of the image content and all classes of the image classification dataset through interaction with the big model, and make the following judgments on the intersection:
[0032] a) If the intersection is judged to be non-empty, classify the image according to the content of the non-empty intersection;
[0033] b) If the intersection is judged to be empty, the image is considered as not conforming to the target dataset and filtered out;
[0034] Finally, the images that match the labels are placed into the corresponding classes of the target dataset.
[0035] As a preferred solution of the method for constructing image classification data based on multimodal agent of the present invention, the quality improvement of the image during the construction process includes ProcessAction; KDEAction, KSAction; DistributionAction, WashAction; AdjustAction;
[0036] Among them, ProcessAction calls the object classification and segmentation model to crop the multi-target image into a single-target image, and makes the following optimizations for the images that meet the target dataset:
[0037] a) adjusting the image to meet the requirements of the target data set in combination with the image resolution information analyzed or input by the user;
[0038] b) If the image is a single-target image, this step does not need to be performed;
[0039] c) If the image is a multi-target image, use the target detection and segmentation model, specifically GroundingDino and SAM, to perform target detection and segmentation on the multi-target image, and perform a second classification for each segmented single target image to complete the extraction and optimization of the target in the image;
[0040] Among them, KDEAction and KSAction are used to detect the kernel density estimation distribution and perform KS test between the data set and the original data set. Assume that {x1, x2....xn} is a set of independent and identically distributed samples drawn from an unknown univariate distribution, and its density function f is at any given point x. The kernel density estimator is given by the following formula:
[0041]
[0042] Where: K is a kernel function, a non-negative function, h>0 is called the smoothing parameter of the bandwidth; the kernel function with the subscript h is called a scaling kernel, which is used to estimate the shape of the function f.
[0043] As a preferred solution of the method for constructing image classification data based on multimodal agent described in the present invention, wherein: the DistributionAction receives the processed image, interacts with the multimodal large model, analyzes the image matrix vector according to the set scene, and calculates the type of probability distribution of the image vector, and the information is saved in text format according to the given restrictions. WashAction is responsible for selecting images based on distribution comparison. The image data distribution type calculated by DistributionAction will be combined with the results of KDE detection and KS detection, and it takes the intersection of the collected image and the original data set distribution, and images that do not match the distribution will be filtered out.
[0044] As a preferred solution of the method for constructing image classification data based on multimodal agent described in the present invention, wherein: the AdjustAction is a post-processing operation, which processes the generated images according to the input resolution and adjusts them to the target resolution.
[0045] As a preferred solution of the method for constructing image classification data based on multimodal agent of the present invention, the formula used by the large language model LLM includes:
[0046]
[0047]
[0048]
[0049] in, is the task status, T k is the input task, R k is the generated response, and the memory is the collection of all generated responses. Through iteration, the understanding and solution of the task are continuously improved. The goal is to find the response that meets the expected matching conditions, that is, If a satisfying response is found in at most N iterations, the iteration stops.
[0050] Beneficial effects of the invention: The invention realizes full automation and batch processing of image classification data set expansion and creation, significantly reduces manpower consumption, and reduces the cost of data set expansion and creation. Compared with traditional methods, it improves operational efficiency while maintaining a high degree of accuracy. BRIEF DESCRIPTION OF THE DRAWINGS
[0051] In order to more clearly illustrate the technical solutions of the embodiments of the present invention, the following briefly introduces the drawings required for describing the embodiments. Obviously, the drawings described below are only some embodiments of the present invention. For ordinary technicians in this field, other drawings can be obtained based on these drawings without creative labor. Among them:
[0052] Figure 1 It is a flowchart of the method for constructing image classification data based on multimodal agent of the present invention.
[0053] Figure 2 Schematic diagram of data sources of the image classification data construction method based on multimodal agent of the present invention.
[0054] Figure 3 This is a sample diagram of collected images of the method for constructing image classification data based on multimodal agent of the present invention;
[0055] Figure 4 This is a sample diagram of a data set created by the method for constructing image classification data based on a multimodal agent in the present invention.
[0056] Figure 5 The figure is a schematic diagram of the architecture of the image classification data construction method based on multimodal agent of the present invention. DETAILED DESCRIPTION
[0057] In order to make the above-mentioned objects, features and advantages of the present invention more obvious and easy to understand, the specific implementation methods of the present invention are described in detail below in conjunction with the accompanying drawings.
[0058] In the following description, many specific details are set forth to facilitate a full understanding of the present invention, but the present invention may also be implemented in other ways different from those described herein, and those skilled in the art may make similar generalizations without violating the connotation of the present invention. Therefore, the present invention is not limited to the specific embodiments disclosed below.
[0059] Secondly, the term "one embodiment" or "embodiment" as used herein refers to a specific feature, structure, or characteristic that may be included in at least one implementation of the present invention. The term "in one embodiment" that appears in different places in this specification does not necessarily refer to the same embodiment, nor does it refer to a separate or selective embodiment that is mutually exclusive with other embodiments.
[0060] Secondly, the present invention is described in detail with reference to the schematic diagram. When describing the embodiments of the present invention in detail, for the sake of convenience, the cross-sectional diagrams showing the device structure will not be partially enlarged according to the general scale, and the schematic diagrams are only examples, which should not limit the scope of protection of the present invention. In addition, in actual production, the three-dimensional dimensions of length, width and depth should be included.
[0061] Reference Figure 1 - Figure 5 , provides a method for constructing image classification data based on multimodal agent, including the following steps:
[0062] Input image and dataset information: including:
[0063] 1) Image data, photo and video formats collected from various sources, along with details of the intended image classification dataset;
[0064] 2) Enter the dataset name directly;
[0065] The agent will perform online search and integrate pre-trained model data to retrieve relevant dataset information;
[0066] After the user inputs their needs according to the prompts, the Agent begins to interact with the large model, enabling the model to use the provided operations to develop a plan that meets the user's needs, and will automatically call operations according to the developed plan to complete the specified task.
[0067] Among them, if the image format detected during the input image process is a video, it needs to be frame extracted and decomposed to convert the video into many images for the expansion and establishment of the data set.
[0068] Specifically, the input process includes the following formula:
[0069] T=JSON_Transfer(instruction=Φ(I+P))
[0070] Where I represents the initial data provided by the user; P represents the prompt for supplementary context information;
[0071] The formula describes the task generation process. The input data I and prompt P are combined into instructions through the mapping function Φ. The instructions express the specific requirements and background information of the task. The instructions are converted into standardized JSON format tasks T through the function JSON_Transfer.
[0072] Among them,: Building an image classification dataset includes: ImageAction, LabelAction and ClassifyAction;
[0073] Among them, ImageAction is specifically:
[0074] After the input image is evaluated, it is transmitted to a multimodal large language model that can read the image and make the following judgments on the image:
[0075] a) Whether the image contains a single or multiple objects, the output is single or multiple;
[0076] b) the content category in the image;
[0077] Each interaction involves thinking and analysis. After the interaction, the model will think whether the result is reasonable and conforms to the prescribed paradigm. If it conforms, the judgment result will be output and the above two image information will be output and saved. If it does not conform, the image will be read again for analysis until the model believes that the output content is reasonable after thinking, and then the above two image information will be output as text information in JSON format.
[0078] Among them, LabelAction is specifically:
[0079] Responsible for extracting and compressing data set information, conveying data set related information to a large language model that can analyze text, and using the designed prompt words to determine the following information contained in the data set:
[0080] a) The type of each class in the image classification dataset;
[0081] b) The resolution of the images in the dataset required by the target image classification dataset;
[0082] c) The size of the images in the dataset required by the target image classification dataset;
[0083] The judgment process is to conduct cyclic thinking and analysis like ImageAction, and then output the results recognized by the big model as text information in JSON format;
[0084] Among them, ClassifyAction is specifically:
[0085] Using the information about the image and image dataset fed back by the big model, we can identify the intersection of the image content and all classes of the image classification dataset through interaction with the big model, and make the following judgments on the intersection:
[0086] a) If the intersection is judged to be non-empty, classify the image according to the content of the non-empty intersection;
[0087] b) If the intersection is judged to be empty, the image is considered as not conforming to the target dataset and filtered out;
[0088] Finally, the images that match the labels are placed into the corresponding classes of the target dataset.
[0089] Furthermore, the image quality improvement during the build process includes ProcessAction; KDEAction, KSAction; DistributionAction, WashAction; AdjustAction;
[0090] Among them, ProcessAction calls the object classification and segmentation model to crop the multi-target image into a single-target image, and makes the following optimizations for the images that meet the target dataset:
[0091] a) adjusting the image to meet the requirements of the target data set in combination with the image resolution information analyzed or input by the user;
[0092] b) If the image is a single-target image, this step does not need to be performed;
[0093] c) If the image is a multi-target image, use the target detection and segmentation model, specifically GroundingDino and SAM, to perform target detection and segmentation on the multi-target image, and perform a second classification for each segmented single target image to complete the extraction and optimization of the target in the image;
[0094] Among them, KDEAction and KSAction are used to detect the kernel density estimation distribution and perform KS test between the data set and the original data set. Assume that {x1, x2....xn} is a set of independent and identically distributed samples drawn from an unknown univariate distribution, and its density function f is at any given point x. The kernel density estimator is given by the following formula:
[0095]
[0096] Where: K is a kernel function, a non-negative function, h>0 is called the smoothing parameter of the bandwidth; the kernel function with the subscript h is called a scaling kernel, which is used to estimate the shape of the function f.
[0097] Among them, DistributionAction receives the processed image and interacts with the multimodal large model, analyzes the image matrix vector according to the set scene, and calculates the type of probability distribution of the image vector. The information is saved in text format according to the given constraints. WashAction is responsible for selecting images based on distribution comparison. The image data distribution type calculated by DistributionAction will be combined with the results of KDE detection and KS detection, and it takes the intersection of the collected image and the original data set distribution. Images that do not match the distribution will be filtered out.
[0098] Furthermore, AdjustAction is a post-processing operation that processes the generated images according to the input resolution and adjusts them to the target resolution.
[0099] The formulas used by the large language model LLM include:
[0100]
[0101]
[0102]
[0103] in, is the task status, T k is the input task, R k is the generated response, and the memory is the collection of all generated responses. Through iteration, the understanding and solution of the task are continuously improved. The goal is to find the response that meets the expected matching conditions, that is, If a satisfying response is found in at most N iterations, the iteration stops.
[0104] The agent uses multiple multimodal large language models and combines them with a variety of uniquely designed construction prompts. Users only need to provide the collected images, relevant dataset information, and requirements, and interact with the large model based on prompt words and input. The large model automatically thinks about and formulates a plan for building an image dataset, and stores the plan. It compares the response with some cases and paradigms previously provided to the agent. The agent thinks about whether the plan formulated by the large model is reasonable. If it is unreasonable, it enters the next cycle. If it is reasonable, it will call the corresponding operation based on the response given by the large model. Through the above reflection process, the specified task is completed. In this process, a series of well-developed operations are organized into execution plans to automatically expand or create a target image dataset for classification.
[0105] The present invention realizes the full automation and batch processing of image classification dataset expansion and creation, significantly reduces manpower consumption and reduces the cost of dataset expansion and creation. Compared with traditional methods, it improves the operation efficiency while maintaining a high degree of accuracy. By using 4050 images of apples, strawberries, oranges and others crawled and photographed on the Internet, an image classification dataset Fruit-3 containing three categories of apples, strawberries and oranges is constructed. Figure 4 , of which 2800 images are in the Fruit-3 category, including 1000 images of apples, 1000 images of strawberries, and 800 images of oranges. There are 1250 images that do not meet the requirements. After using Agent to create Fruit-3, we get a data set of 2918 images. Figure 5 , successfully screened out 1,250 unqualified images and the accuracy of the expanded dataset reached 99.95%. It also achieved the flexibility of the image data source, allowing users to provide images at will, including randomly taken photos, images and videos obtained from the Internet, etc. Experimental verification found that the Agent's expanded dataset has a certain improvement in the classification effect of the deep learning model after training in image classification tasks.
[0106] It should be noted that the above embodiments are only used to illustrate the technical solutions of the present invention rather than to limit it. Although the present invention has been described in detail with reference to the preferred embodiments, those skilled in the art should understand that the technical solutions of the present invention may be modified or replaced by equivalents without departing from the spirit and scope of the technical solutions of the present invention, which should all be included in the scope of the claims of the present invention.
Claims
1. A method for constructing image classification data based on multimodal agent, characterized in that: The following steps are involved: Input image and dataset information: including: 1) Image data, photo and video formats collected from various sources, along with details of the intended image classification dataset; 2) Enter the dataset name directly; The agent will perform online search and integrate pre-trained model data to retrieve relevant dataset information; After the user inputs their needs according to the prompts, the Agent begins to interact with the large model, enabling the model to use the provided operations to develop a plan that meets the user's needs, and will automatically call operations according to the developed plan to complete the specified task.
2. The method for constructing image classification data based on multimodal agent according to claim 1, characterized in that: If the image format detected during the input image process is a video, it needs to be frame extracted and decomposed to convert the video into many images for the expansion and establishment of the data set.
3. The method for constructing image classification data based on multimodal agent according to claim 1, characterized in that: The input process includes the following formula: T=JSON_Transfer(instruction=Φ(I+P)) Where I represents the initial data provided by the user; P represents the prompt for supplementary context information; The formula describes the task generation process. The input data I and prompt P are combined into instructions through the mapping function Φ. The instructions express the specific requirements and background information of the task. The instructions are converted into standardized JSON format tasks T through the function JSON_Transfer.
4. The method for constructing image classification data based on multimodal agent according to claim 3, characterized in that: Constructing image classification datasets includes: ImageAction, LabelAction, and ClassifyAction; Among them, ImageAction is specifically: After the input image is evaluated, it is transmitted to a multimodal large language model that can read the image and make the following judgments on the image: a) Whether the image contains a single or multiple objects, the output is single or multiple; b) the content category in the image; Each interaction involves thinking and analysis. After the interaction, the model will think whether the result is reasonable and conforms to the prescribed paradigm. If it conforms, the judgment result will be output and the above two image information will be output and saved. If it does not conform, the image will be read again for analysis until the model believes that the output content is reasonable after thinking, and then the above two image information will be output as text information in JSON format. Among them, LabelAction is specifically: Responsible for extracting and compressing data set information, conveying data set related information to a large language model that can analyze text, and using the designed prompt words to determine the following information contained in the data set: a) The type of each class in the image classification dataset; b) The resolution of the images in the dataset required by the target image classification dataset; c) The size of the images in the dataset required by the target image classification dataset; The judgment process is to conduct cyclic thinking and analysis like ImageAction, and then output the results recognized by the big model as text information in JSON format; Among them, ClassifyAction is specifically: Using the information about the image and image dataset fed back by the big model, we can identify the intersection of the image content and all classes of the image classification dataset through interaction with the big model, and make the following judgments on the intersection: a) If the intersection is judged to be non-empty, classify the image according to the content of the non-empty intersection; b) If the intersection is judged to be empty, the image is considered as not conforming to the target dataset and filtered out; Finally, the images that match the labels are placed into the corresponding classes of the target dataset.
5. The method for constructing image classification data based on multimodal agent according to claim 1, characterized in that: The image quality improvement during the build process includes ProcessAction; KDEAction, KSAction; DistributionAction, WashAction; AdjustAction; Among them, ProcessAction calls the object classification and segmentation model to crop the multi-target image into a single-target image, and makes the following optimizations for the images that meet the target dataset: a) adjusting the image to meet the requirements of the target data set in combination with the image resolution information analyzed or input by the user; b) If the image is a single-target image, this step does not need to be performed; c) If the image is a multi-target image, use the target detection and segmentation model, specifically GroundingDino and SAM, to perform target detection and segmentation on the multi-target image, and perform a second classification for each segmented single target image to complete the extraction and optimization of the target in the image; Among them, KDEAction and KSAction are used to detect the kernel density estimation distribution and perform KS test between the data set and the original data set. Assume that {x1, x2....xn} is a set of independent and identically distributed samples drawn from an unknown univariate distribution, and its density function f is at any given point x. The kernel density estimator is given by the following formula: Where: K is a kernel function, a non-negative function, h>0 is called the smoothing parameter of the bandwidth; the kernel function with the subscript h is called a scaling kernel, which is used to estimate the shape of the function f.
6. The method for constructing image classification data based on multimodal agent according to claim 5, characterized in that: The DistributionAction receives the processed image and interacts with the multimodal large model, analyzes the image matrix vector according to the set scene, and calculates the type of probability distribution of the image vector. The information is saved in text format according to the given constraints. WashAction is responsible for selecting images based on distribution comparison. The image data distribution type calculated by DistributionAction will be combined with the results of KDE detection and KS detection, and it takes the intersection of the collected image and the original data set distribution. Images that do not match the distribution will be filtered out.
7. The method for constructing image classification data based on multimodal agent according to claim 5, characterized in that: The AdjustAction is a post-processing operation that processes the generated images according to the input resolution and adjusts them to the target resolution.
8. The method for constructing image classification data based on multimodal agent according to claim 1, characterized in that: The formulas used by the large language model LLM include: in, is the task status, T k is the input task, R k is the generated response, and the memory is the collection of all generated responses. Through iteration, the understanding and solution of the task are continuously improved. The goal is to find the response that meets the expected matching conditions, that is, If a satisfying response is found in at most N iterations, the iteration stops.