Agent-based automatic image training data generation method

Through a multi-task and multi-model automated process and the use of agents working together, the high cost of manual annotation and data quality issues in the generation of high-definition image datasets are solved, and efficient and low-cost generation of high-definition image training data is achieved.

CN120635623APending Publication Date: 2025-09-12LINKER
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510500318.4
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-04-21
Publication Date
2025-09-12

AI Technical Summary

Technical Problem

The cost of manual annotation in the generation of existing high-definition image datasets is too high, and the existing VLLM-based automatic generation method has image size limitations and data quality issues, which affect the model training effect.

Method used

It adopts a multi-task and multi-model automated process, including image preprocessing, filtering, description generation, question-answer pair generation, and target detection. Through agent collaboration, it reduces manual intervention and generates high-definition image training data.

Benefits of technology

It significantly reduces the cost of data generation, improves data quality and automation, adapts to the characteristics of high-definition images, and the generated data is more in line with model training requirements.

✦ Generated by Eureka AI based on patent content.
Patent Text Reader

Abstract

The invention discloses an Agent-based automatic image training data generation method, which comprises the following steps of: preprocessing high-definition image data serving as a data set, cutting the high-definition image data into a size suitable for VLLM input, and recording a position; transmitting the processed image into a small machine learning algorithm to filter out a part of image with low information amount; the processed image is transmitted into a VLLM model to generate detailed description; generating question-answer pairs according to the description; according to the basic fact content generated by the question and answer pair, identifying a position by using a target detection model and outputting a target box; and integrating the question and answer pairs, the basic fact content and the target box information into training data for a deep learning model to use. According to the method, high-quality image training data can be automatically generated, and the data processing efficiency and the deep learning model training effect are improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the fields of image processing, artificial intelligence, deep learning, and more specifically to an agent-based automatic image training data generation method. Background Art

[0002] The cost of manual labeling in the generation of high-definition image datasets is too high, and the existing automatic generation methods based on VLLM (Visual Language Models) have shortcomings such as image size limitations and data quality issues, which will have a significant impact on subsequent model training. Therefore, manual intervention is required before using the high-definition image dataset to improve the data quality of the generated high-definition image dataset. Summary of the Invention

[0003] To address the shortcomings of existing technologies, the present invention aims to provide an agent-based automated image training data generation method. Through a multi-task, multi-model automated process, this method reduces manual intervention, improves data generation efficiency, and ensures the quality of generated data, particularly for generating training data for high-definition images.

[0004] To achieve the above-mentioned object, the present invention provides the following technical solution: an agent-based automatic image training data generation method, comprising the following steps:

[0005] Step 1: Preprocess the high-definition image data as the dataset, crop it to a size suitable for VLLM input, and record the position of each small image in the original image;

[0006] Step 2: The processed image is passed to a small machine learning algorithm for filtering to remove parts of the image with low information content;

[0007] Step 3: The image processed in step 2 is passed into the VLLM model to generate a detailed description of the image;

[0008] Step 4: Generate corresponding question-answer pairs based on the image description in step 3;

[0009] Step 5: Based on the basic facts generated by the question-answer pair generated in step 4, the object detection model is used to identify the location of the basic facts in the image and output the target box;

[0010] Step 6: Integrate the generated question-answer pairs, basic fact content, and their corresponding target box information into complete training data for use by the deep learning model.

[0011] As a further improvement of the present invention, the specific method of generating corresponding question-answer pairs according to the image description in step 3 in step 4 is to call the LLM in the agent, generate corresponding questions through prompt words, and generate question-answer pairs based on the provided prompt words. These questions and answers are aimed at the detailed information of the image.

[0012] As a further improvement of the present invention, the specific method of using the target detection model in step five to identify the position of the basic fact in the image and output the target box is as follows: first, use LLM to extract the named entities in the question, then input the obtained question of interest into the OMDET model in the Agent process, use the prompt word to obtain the coordinate information corresponding to the object, and complete the target detection and recognition.

[0013] As a further improvement of the present invention, the specific method of generating a detailed description of the image in step three is: first, identify the scene in the image and generate a layout description of the scene; then identify the crowd in the image and generate a relevant description of the number of people; then identify other functional buildings in the scene in the image and generate a relevant description of the functional buildings; identify the architectural style of the scene in the image and generate a description of the architectural style; and finally identify the atmosphere of the scene in the image and generate a description of the atmosphere.

[0014] As a further improvement of the present invention, the prompt words in step 4 are: the scene in the picture, description of the people or crowd in the image, lighting and illumination conditions in the scene, description of the building structure in the image, description of the signs, logos, signboard patterns, or any objects with text in the image, indicating their shape, color, number, position and displayed content or text, description of the people or crowd in the image, the people's appearance, clothing, expressions, and what they are currently doing; the direction and flow of the crowd, and whether there is a queuing area or a place where people gather.

[0015] The beneficial effects of the present invention are as follows: compared with the high cost of manual annotation in the process of generating high-definition image datasets in the background art, the agent-based automatic image training data generation method of the present invention:

[0016] Reduce costs: By automating the image preprocessing, description generation, question-answer pair generation, and object detection processes, the need for manual annotation is significantly reduced, lowering the cost of data generation.

[0017] Improve data quality: Using a multi-model cascade approach not only ensures the richness and accuracy of data, but also avoids the defects of a single model and reduces the generation of hallucinations and data noise.

[0018] Adapting to HD images: By performing specific processing on HD images, the generated data is made more consistent with the characteristics of HD images, providing more effective training data for the model when processing HD images.

[0019] Improved automation: Through the collaborative work of agents based on the OmAgent framework, the orchestration of multi-task processes is simplified and the automation level of the entire data generation process is improved. DETAILED DESCRIPTION

[0020] The present invention will be further described in detail with reference to the following examples.

[0021] This embodiment provides an agent-based automated image training data generation method. This method uses the OmAgent open source framework and combines the collaborative work of multiple models to automatically generate high-definition image training data. The specific steps are as follows:

[0022] Image preprocessing: The first task is to process high-definition images and crop them to a size suitable for VLLM input. For example, for an 8K resolution (7680×4320) image, it is cropped into multiple images suitable for model input (such as 512×512), and the position of each small image in the original image is recorded.

[0023] Image filtering: The processed image is passed to a small machine learning algorithm for filtering, removing parts of the image with low information content, such as background images such as the sky, walls, and ground. Only the parts of the image that contain useful information are retained.

[0024] Image description generation: A qualified image is fed into the VLLM model, which generates a detailed description of the image, including information such as the image content and the number of objects. These descriptions provide the basis for subsequent training data generation.

[0025] Generate question-answer pairs: Generate corresponding question-answer pairs based on the image description. For example, for an object in an image, the system can generate a question-answer pair such as "What does the sign in the upper left corner of the image show?"

[0026] The generation of question-answer pairs is based on the obtained image description, which contains important elements in the image, their positional relationships, colors, shapes, etc. Based on this image description, the LLM in the agent is called to generate corresponding questions using specific prompt words.

[0027] For example, the image description says: "This picture shows a busy train station hall, specifically St. Pancras train station in London." The following is a detailed description of the picture: 1. **Overall layout**: - The train station hall is spacious and bright, with high ceilings and modern structure, using a lot of glass and metal materials. - There is an information display screen in the center of the hall, which shows the departure and arrival information of the train for passengers to inquire. 2. **Crowd**: - The hall is crowded with passengers, some of whom are walking, while others are sitting on suitcases or on chairs waiting. - The passengers are carrying a lot of luggage, including suitcases, backpacks and handbags, indicating that this is a busy transportation hub. 3. **Facilities and services**: - There are many shops and service facilities in the hall, including a restaurant called "Searcys" on the second floor. Some customers can be seen dining through the glass. - There are also multiple Signs and information boards help passengers navigate and obtain information. 4. **Architectural style**: - The railway station's architectural style is modern, with extensive glass and metal structures, making the entire space appear bright and open. - A large glass skylight at the top of the hall allows ample natural light to filter in, increasing the sense of transparency in the space. 5. **Atmosphere**: - Despite the large number of people, the overall atmosphere appears orderly, with passengers carrying out their activities in an orderly manner. - The hall is brightly lit, creating a warm and comfortable atmosphere. This image vividly depicts the busy scene of St. Pancras Station, highlighting its importance and efficiency as a transportation hub. "Using the above description method, a progressive analysis framework of "scene → crowd → functional building → style → atmosphere" is adopted to ensure comprehensive coverage from macro layout to detailed features; this will make the subsequent design of question and answer pairs more logical and reduce the computing power required in the subsequent design of question and answer pairs.

[0028] Through the corresponding prompt words: "Based on this image description, design at least 10 question-answer pairs. These question-answer pairs need to target the detailed information of the image, which may be related to the following details: 1. The scene in the picture 2. The lighting and illumination conditions in the scene, 3. Describe the architectural structures in the image, such as columns, stairs, escalators, etc., and mention the colors and relative positions of these architectural structures. 4. Describe the signs, logos, signboard patterns, or any objects with text in the image, and point out their shape, color, quantity, position, and displayed content or text. 5. Describe the people or groups in the image, their appearance, clothing, expressions, and what they are currently doing; the direction and flow of the crowd's movement, and whether there are queuing areas or places where people gather. 6. Detailed information about elements in other images, such as furniture, equipment, decorations, etc., the appearance, function, current status, and location of these elements. Example questions: 1. Question: Based on the sign on the left side of the image, where will you arrive if you go out from this exit? Answer: Going out from the exit indicated by the sign on the left will arrive at "Midland Road". 2. Q: If I want to eat something, where should I go? A: If you want to eat something, you can go to the restaurant "SEARCYS" on the second floor. 3. Q: I want to take the train, which way should I go? A: If you want to take the train, you should go straight."

[0029] Generate a sample question:

[0030] 1. **Q: What does this image primarily depict? **A: This image shows the concourse of St Pancras International Station, showing multiple levels, people, and signage.

[0031] 2. **Q: What are the lighting conditions like in the scene? **A: The scene is well-lit, primarily from the large glass windows at the top. Natural light streams in through the windows, while some areas are artificially lit.

[0032] 3. **Q: What are the prominent architectural structures in the image? **A: The prominent architectural structures include the blue-gray metal arched frame at the top, the stairs, escalator, and the elevated area supported by white columns.

[0033] 4. **Q: What signs are shown in the picture? **A: There are several blue signs in the picture, mainly in the middle and back, showing train information, "All trains," "Départs Departures," "Way out," and other information.

[0034] 5. **Q: What does the electronic screen in the middle show? **A: The electronic screen in the middle shows the "St Pancras" logo and train departure information, including time and train destination.

[0035] 6. **Q: What are some iconic restaurants or shops? **A: There is a restaurant called "SEARCYS" on the second floor. Its sign is very eye-catching with white letters on a black background.

[0036] 7. **Q: What pillars are there in the image, and what color are they? **A: The pillars in the image are white, located slightly left of center, and support the upper structure.

[0037] 8. **Q: Where are the stairs and escalator? **A: The stairs and escalator are located behind the sign and connect the upper and lower floors.

[0038] 9. **Q: What is the color and material of the second-floor guardrail? **A: The second-floor guardrail is a combination of metal and transparent glass, with an overall silver-gray color.

[0039] 10. **Q: Where are the crowded areas in the image? **A: The crowds are mainly gathered near the electronic screens and the train platform, with some people sitting on lounge chairs.

[0040] Target detection: Based on the basic fact content generated by the question-answer pair (such as "the sign in the upper left corner of the picture"), use the target detection model (such as OMDET) to identify the location of the basic fact in the image and output the target box

[0041] First, use LLM to extract named entities in the question. For example, use the prompt "Based on the above questions or answers, give a description of the area or object of interest in each question." For the question-answer pair "**Q: What signs are there in the picture? **A: There are multiple blue signs in the picture, mainly located in the middle and back, showing train information, "Alltrains", "Départs Departures", "Way out", etc.", the output will be: "**Object of interest**: blue sign." Then input the obtained question of interest into the OMDET model in the Agent process, using the prompt "Find the blue sign", and the corresponding bbox of the object will be obtained: (50, 100, 150, 200), (300, 100, 400, 200),

[0042] (0,250,500,400)

[0043] Training data integration: Finally, the generated question-answer pairs, basic facts and their corresponding target box information are integrated into complete training data for use by the deep learning model.

[0044] Through this method, high-quality training data for high-definition images can be automatically generated without the need for a large amount of manual labeling, effectively reducing the cost of data generation while improving the efficiency and quality of data processing.

[0045] This embodiment provides the following two examples:

[0046] Example 1:

[0047] Assuming the input image is a high-definition image with 8K resolution (7680×4320), the first step is to crop the image into 135 15×9 partial images, and record the location of each image. A machine learning algorithm is then used to filter out background components such as the sky and ground, retaining image segments containing useful information. Next, the VLLM model generates a detailed description of the image, including information about the objects, scene, and location in the image. Based on this description, corresponding question-answer pairs are generated, object detection is performed, and target bounding boxes are output. Finally, the generated question-answer pairs and target bounding box information are integrated and used as training data for the deep learning model.

[0048] Example 2:

[0049] The input image is a 4K resolution (3840×2160) image. After preprocessing, it is cut into 8×6 small images, filtered, and then fed into the VLLM model. The generated image description includes the number of people and objects in the image and their attributes. Based on the description, corresponding question-answer pairs are generated, such as "What is the color of the person's clothing in the picture?" The object detection model determines the location of the person in the image and marks the object box, finally completing the integration of training data.

[0050] The above description is merely a preferred embodiment of the present invention. The scope of protection of the present invention is not limited to the above embodiment. All technical solutions based on the concept of the present invention are within the scope of protection of the present invention. It should be noted that for those skilled in the art, various improvements and modifications that do not depart from the principles of the present invention should also be considered within the scope of protection of the present invention.

Claims

1. An agent-based automated image training data generation method, characterized by: The steps include: Step 1: Preprocess the high-definition image data as the dataset, crop it to a size suitable for VLLM input, and record the position of each small image in the original image; Step 2: The processed image is passed to a small machine learning algorithm for filtering to remove parts of the image with low information content; Step 3: The image processed in step 2 is passed into the VLLM model to generate a detailed description of the image; Step 4: Generate corresponding question-answer pairs based on the image description in step 3; Step 5: Based on the basic facts generated by the question-answer pair generated in step 4, the object detection model is used to identify the location of the basic facts in the image and output the target box; Step 6: Integrate the generated question-answer pairs, basic fact content, and their corresponding target box information into complete training data for use by the deep learning model.

2. The agent-based automated image training data generation method according to claim 1, characterized in that: The specific method of generating corresponding question-answer pairs based on the image description in step 3 in step 4 is to call the LLM in the agent, generate corresponding questions through prompt words, and generate question-answer pairs based on the provided prompt words. These questions and answers are aimed at the detailed information of the image.

3. The agent-based automated image training data generation method according to claim 1 or 2, characterized in that: In step 5, the target detection model is used to identify the position of the basic fact in the image and output the target frame. The specific method is as follows: first, the named entities in the question are extracted using LLM, and then the obtained question of interest is input into the OMDET model in the Agent process. The prompt word is used to obtain the coordinate information corresponding to the object to complete the target detection and recognition.

4. The agent-based automated image training data generation method according to claim 2, characterized in that: The specific method of generating a detailed description of the image in step three is as follows: first, identify the scene in the image and generate a layout description of the scene; then, identify the people in the image and generate a description of the number of people; then, identify other functional buildings in the scene in the image and generate a description of the functional buildings; then, identify the architectural style of the scene in the image and generate a description of the architectural style; and finally, identify the atmosphere of the scene in the image and generate a description of the atmosphere.

5. The agent-based automated image training data generation method according to claim 4, characterized in that: The prompt words in step 4 are: the scene in the picture, description of the people or crowd in the image, lighting and illumination conditions in the scene, description of the building structure in the image, description of the signs, logos, signboard patterns, or any objects with text in the image, indicating their shape, color, number, position and displayed content or text, description of the people or crowd in the image, the person's appearance, clothing, expression, and what they are currently doing; the direction and flow of the crowd, and whether there is a queuing area or a place where people gather.