Multi-modal large model training method, object detection method, device and electronic device
Through the multimodal large model training method, the categories and attributes of image objects are extracted using the large language model, and the thinking chain question-and-answer sample pair is constructed, solving the complexity problem of descriptive object detection and achieving efficient end-to-end object detection.
Patent Information
- Application Number
- CN202510399045.9
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2025-03-31
- Publication Date
- 2025-07-01
- Estimated Expiration
- 2045-03-31
AI Technical Summary
In the prior art, descriptive object detection is highly complex and has low detection efficiency, making it difficult to effectively locate specific objects in the image that conform to natural language description.
The multimodal large model training method is adopted to obtain the sample image and the description text of the object's label box, and use the large language model to extract the category names and attributes of the object, construct a question-and-answer sample pair in the form of a thinking chain, adjust the model parameters until the convergence conditions are met, and end-to-end descriptive object detection is achieved.
It reduces the complexity of descriptive object detection, improves detection efficiency, and can accurately identify the area occupied by objects in the image that conforms to natural language descriptions, achieving flexible object detection capabilities.
Smart Images

Figure CN119903348B_ABST
Abstract
Description
Technical Field
[0001] This application relates to the field of artificial intelligence technology, and in particular, to a multi-modal large model training method, an object detection method, a device, and an electronic device. Background Art
[0002] In some application scenarios (such as image retrieval or visual question answering, etc.), it is often necessary to locate a specific object mentioned in the natural language description text provided by the user in an image (which can be called descriptive object detection). For example, descriptive object detection can be achieved through the integration of multiple models. For instance, when it is necessary to detect "an athlete located on the green grass and wearing non-white stockings", an object detection model for detecting "person" can be used to locate the people in the image, then screened through a classification model that can determine "whether wearing stockings", and finally, a scene recognition model is used to confirm whether the target is on the green grass. However, this way of integrating multiple models increases the complexity of implementing descriptive object detection and the detection efficiency is relatively low. Summary of the Invention
[0003] The purpose of the embodiments of this application is to provide a multi-modal large model training method, an object detection method, a device, and an electronic device to reduce the complexity of implementing descriptive object detection and improve the detection efficiency. The specific technical solutions are as follows:
[0004] In the first aspect implemented by this application, a multi-modal large model training method is provided. The method includes:
[0005] Obtain a plurality of sample images and first sample description texts of object annotation boxes of a specified category in each sample image;
[0006] For each sample image, use a first large language model and a first text prompt to extract the category name of the object described by the first sample description text corresponding to the sample image and multiple attributes of the object, and combine the obtained category name with at least one of the obtained attributes to obtain multiple second sample description texts corresponding to the sample image; wherein, the first text prompt is used to indicate: extracting the category name of the object described by the input description text and the attributes of the object;
[0007] For each object annotation box in the sample image, respectively determine whether the object annotation box matches each attribute used to combine to obtain multiple second sample description texts corresponding to the sample image;
[0008] Construct a sample question that includes each second sample description text, and use the matching results between each attribute included in the second sample description text and each object bounding box to construct a sample answer in the form of a chain of thought corresponding to the sample question, obtaining a question-answer sample pair; wherein, the sample question is used to indicate the position of the image region in the input image that conforms to the second sample description text; the sample answer includes a first inference process text for describing the inference process of the multimodal large model; the first inference process text includes: each inference step and the first execution order between each inference step; according to the first execution order, each inference step is respectively: extracting the category name of the object described by the description text included in the input question and the attributes of the object, detecting the position of the image region occupied by all objects belonging to the category represented by the extracted category name in the input image, respectively determining the matching results between each image region and each extracted attribute, and determining the position of the image region that matches each extracted attribute;
[0009] Input the sample image and the sample questions in each question-answer sample pair into the multimodal large model with the initial structure to obtain a predicted answer;
[0010] Based on the difference between the obtained predicted answer and the sample answer in the corresponding question-answer sample pair, adjust the parameters of the multimodal large model with the initial structure until the preset convergence condition is reached, obtaining a trained multimodal large model.
[0011] Optionally, before the step of, for each sample image, using the first large language model and the first text prompt to extract the category name of the object described by the first sample description text corresponding to the sample image and multiple attributes of the object, and combining the obtained category name with at least one of the obtained attributes to obtain multiple second sample description texts corresponding to the sample image, the method further includes:
[0012] For each sample image, obtain the category name of the object in the object bounding box of the specified category in the sample image as the preset category name;
[0013] The step of, for each sample image, using the first large language model and the first text prompt to extract the category name of the object described by the first sample description text corresponding to the sample image and multiple attributes of the object, and combining the obtained category name with at least one of the obtained attributes to obtain multiple second sample description texts corresponding to the sample image, includes:
[0014] For each sample image, input the first text prompt, the first sample description text corresponding to the sample image, and the preset category name into the first large language model, extract the category name of the object described by the first sample description text corresponding to the sample image and multiple attributes of the object, and combine the obtained category name with at least one of the obtained attributes to obtain multiple second sample description texts corresponding to the sample image; wherein, the obtained category name is consistent with the category represented by the preset category name.
[0015] Optionally, the obtaining multiple sample images and the first sample description text of the object annotation box of the specified category in each sample image includes:
[0016] Obtain multiple sample images and the positions of the object annotation boxes of the specified category in each sample image.
[0017] For each sample image, input the sample image and the second text prompt into the second large language model to obtain the first sample description text of the object annotation box of the specified category in the sample image; wherein, the second text prompt is used to indicate: generate the description text of the object annotation box in the input image.
[0018] Optionally, the step of, for each sample image, using the first large language model and the first text prompt to extract the category name of the object described by the first sample description text corresponding to the sample image and multiple attributes of the object, and combining the obtained category name with at least one of the obtained attributes to obtain multiple second sample description texts corresponding to the sample image includes:
[0019] For each sample image, input the first sample description text corresponding to the sample image and the first text prompt into the first large language model, and extract the category name of the object described by the first sample description text corresponding to the sample image and multiple attributes of the object.
[0020] Input the extracted category name, the extracted attributes, and the third text prompt into the first large language model to obtain multiple second sample description texts corresponding to the sample image; wherein, the third text prompt is used to indicate: combine the input category name with at least one of the input attributes to obtain multiple description texts.
[0021] Optionally, the step of, for each object annotation box in the sample image, respectively determining whether the object annotation box matches each attribute used to combine to obtain multiple second sample description texts corresponding to the sample image includes:
[0022] Using an image-text matching model and a fourth text prompt, respectively detect whether the object bounding boxes of each specified category in the sample image match each text description to be matched; wherein, each text description to be matched is obtained by respectively combining each attribute for combining to obtain multiple second sample description texts corresponding to the sample image with the extracted category name; the fourth text prompt is used to indicate: determining whether each object bounding box in the input image and the input text description match.
[0023] In a second aspect of the implementation of this application, a target detection method is further provided, and the method includes:
[0024] Obtain an image to be detected and a text prompt to be utilized including a text description to be detected; wherein, the text prompt to be utilized is used to indicate: determining the position of the image area occupied by the objects in the image to be detected that conform to the text description to be detected.
[0025] Input the image to be detected and the text prompt to be utilized into a pre-trained multi-modal large model to obtain a detection result in the form of a chain of thought; wherein, the multi-modal large model is trained based on any one of the above multi-modal large model training methods; the detection result includes: a second inference process text for describing the inference process of the multi-modal large model; the second inference process text includes: each inference step and the second execution order between each inference step; according to the second execution order, each inference step is respectively: extracting the category name of the object described by the input text prompt and the attributes of the object, detecting the position of the image area occupied by all the objects in the input image that belong to the category represented by the extracted category name, respectively determining the matching result between each image area and each extracted attribute, and determining the position of the image area that matches each extracted attribute.
[0026] In a third aspect of the implementation of this application, a multi-modal large model training device is further provided, and the device includes:
[0027] A sample acquisition module, configured to acquire a plurality of sample images and a first sample description text of the object bounding boxes of the specified category in each sample image.
[0028] A description text generation module, configured to, for each sample image, use a first large language model and a first text prompt to extract the category name of the object described by the first sample description text corresponding to the sample image and multiple attributes of the object, and combine the obtained category name with at least one of the obtained attributes to obtain multiple second sample description texts corresponding to the sample image; wherein, the first text prompt is used to indicate: extracting the category name of the object described by the input description text and the attributes of the object.
[0029] A matching result determination module, configured to respectively determine, for each object bounding box in the sample image, whether the object bounding box matches each attribute used to combine multiple second sample description texts corresponding to the sample image;
[0030] A question-and-answer sample pair construction module, configured to construct a sample question including each second sample description text, and use the matching results between each attribute included in the second sample description text and each object bounding box to construct a sample answer in the form of a thought chain corresponding to the sample question, obtaining a question-and-answer sample pair; wherein, the sample question is used to indicate the position of the image region in the input image that conforms to the second sample description text; the sample answer includes a first inference process text for describing the inference process of the multimodal large model; the first inference process text includes: each inference step and the first execution order between each inference step; according to the first execution order, each inference step is respectively: extracting the category name of the object described by the description text included in the input question and the attributes of the object, detecting the position of the image region occupied by all objects belonging to the category represented by the extracted category name in the input image, respectively determining the matching results between each image region and each extracted attribute, and determining the position of the image region that matches each extracted attribute;
[0031] A predicted answer determination module, configured to input the sample image and the sample question in each question-and-answer sample pair into the multimodal large model with an initial structure to obtain a predicted answer;
[0032] A model parameter adjustment module, configured to adjust the parameters of the multimodal large model with an initial structure based on the difference between the obtained predicted answer and the sample answer in the corresponding question-and-answer sample pair until a preset convergence condition is reached, obtaining a trained multimodal large model.
[0033] Optionally, the device further includes:
[0034] A preset category name acquisition module, configured to, before, for each sample image, using a first large language model and a first text prompt to extract the category name of the object described by the first sample description text corresponding to the sample image and multiple attributes of the object, and combining the obtained category name with at least one of the obtained attributes to obtain multiple second sample description texts corresponding to the sample image, for each sample image, acquire the category name of the object in the object bounding box of the specified category in the sample image as the preset category name;
[0035] The described description text generation module is specifically configured to, for each sample image, input the first text prompt, the first sample description text corresponding to the sample image, and the preset category name into a first large language model, extract the category name of the object described by the first sample description text corresponding to the sample image and multiple attributes of the object, and combine the obtained category name with at least one of the obtained attributes to obtain multiple second sample description texts corresponding to the sample image; wherein, the obtained category name is the same as the category represented by the preset category name.
[0036] Optionally, the sample acquisition module is specifically configured to acquire multiple sample images and the positions of the object bounding boxes of the specified category in each sample image; for each sample image, input the sample image and a second text prompt into a second large language model to obtain a first sample description text of the object bounding box of the specified category in the sample image; wherein, the second text prompt is used to indicate: generating a description text of the object bounding box in the input image.
[0037] Optionally, the description text generation module is specifically configured to, for each sample image, input the first sample description text corresponding to the sample image and the first text prompt into a first large language model, extract the category name of the object described by the first sample description text corresponding to the sample image and multiple attributes of the object; input the extracted category name, the extracted attributes, and a third text prompt into the first large language model to obtain multiple second sample description texts corresponding to the sample image; wherein, the third text prompt is used to indicate: combining the input category name with at least one of the input attributes to obtain multiple description texts.
[0038] Optionally, the matching result determination module is specifically configured to use an image-text matching model and a fourth text prompt to respectively detect whether each object bounding box of the specified category in the sample image matches each text to be matched; wherein, each text to be matched is obtained by respectively combining each attribute used to combine to obtain multiple second sample description texts corresponding to the sample image with the extracted category name; the fourth text prompt is used to indicate: determining whether each object bounding box in the input image and the input description text match.
[0039] In the fourth aspect of the implementation of this application, a target detection device is further provided, and the device includes:
[0040] A data acquisition module, configured to acquire an image to be detected and a text prompt to be utilized including a description text to be detected; wherein, the text prompt to be utilized is used to indicate: determining the position of the image area occupied by the object in the image to be detected that conforms to the description text to be detected.
[0041] A detection result determination module, configured to input the image to be detected and the text prompt to be utilized into a pre-trained multi-modal large model, so as to obtain a detection result in the form of a chain of thought; wherein, the multi-modal large model is trained by using any one of the above-mentioned multi-modal large model training methods; the detection result includes: a second inference process text for describing the inference process of the multi-modal large model; the second inference process text includes: each inference step and a second execution order between each inference step; according to the second execution order, each inference step is respectively: extracting the category name of the object described by the input text prompt and the attributes of the object, detecting the positions of the regions of the image occupied by all the objects belonging to the category represented by the extracted category name in the input image, respectively determining the matching results between each image region and each extracted attribute, and determining the positions of the image regions that match each extracted attribute.
[0042] An embodiment of the present application further provides an electronic device, including:
[0043] A memory, configured to store a computer program;
[0044] A processor, configured to implement any one of the above-mentioned multi-modal large model training methods or object detection methods when executing the program stored in the memory.
[0045] An embodiment of the present application further provides a computer-readable storage medium, in which a computer program is stored, and when the computer program is executed by a processor, any one of the above-mentioned multi-modal large model training methods or object detection methods is implemented.
[0046] An embodiment of the present application further provides a computer program product containing instructions, which when running on a computer, causes the computer to execute any one of the above-mentioned multi-modal large model training methods or object detection methods.
[0047] Advantageous effects of the embodiments of the present application:
[0048] A multi-modal large model training method provided by an embodiment of this application can obtain multiple sample images and first sample description texts of object annotation frames of a specified category in each sample image; for each sample image, use a first large language model and a first text prompt to extract the category name of the object described by the first sample description text corresponding to the sample image and multiple attributes of the object, and combine the obtained category name with at least one of the obtained attributes to obtain multiple second sample description texts corresponding to the sample image; wherein, the first text prompt is used to indicate: extract the category name of the object described by the input description text and the attributes of the object; for each object annotation frame in the sample image, respectively determine whether the object annotation frame matches each of the extracted attributes; construct a sample question containing each second sample description text, and use the matching result of each attribute included in the second sample description text and each object annotation frame to construct a sample answer in the form of a chain of thought corresponding to the sample question, obtaining a question-answer sample pair; wherein, the sample question is used to indicate determining the position of the image area in the input image that conforms to the second sample description text; the sample answer contains a first inference process text for describing the inference process of the multi-modal large model; the first inference process text includes: each inference step and the first execution order between each inference step; according to the first execution order, each inference step is respectively: extract the category name of the object described by the description text included in the input question and the attributes of the object, detect the position of the image area occupied by all objects belonging to the category represented by the extracted category name in the input image, respectively determine the matching result of each image area and each of the extracted attributes, and determine the position of the image area that matches each of the extracted attributes; input the sample image and the sample question in each question-answer sample pair into the multi-modal large model with an initial structure to obtain a predicted answer; based on the difference between the obtained predicted answer and the sample answer in the corresponding question-answer sample pair, adjust the parameters of the multi-modal large model with the initial structure until the preset convergence condition is reached, obtaining a trained multi-modal large model.
[0049] Based on the above processing, according to the obtained sample images and the descriptive text (i.e., the first sample descriptive text) for describing the object annotation boxes of the specified categories in the sample images, the first large language model can be used to generate the second sample descriptive text, so as to increase the number of the obtained second sample descriptive texts and enrich the diversity of the second sample descriptive texts. Furthermore, sample questions containing each second sample descriptive text can be constructed to instruct the multi-modal large model to determine the positions of the image regions in the input images that conform to the second sample descriptive text. And the matching results between each object annotation box in each sample image and each attribute used to combine and obtain multiple second sample descriptive texts corresponding to the sample image can be combined to construct a sample answer in the form of a chain of thought, and thus a question-answer sample pair in the form of a chain of thought can be obtained. The sample answer in the form of a chain of thought can include: the specific steps of reasoning to obtain the positions of the image regions in the sample image that match each attribute in the second sample descriptive text.
[0050] Correspondingly, training the multi-modal large model based on the question-answer sample pair in the form of a chain of thought can enable the multi-modal large model to learn the ability to reason in the form of a chain of thought, determine the positions of the image regions occupied by the objects in the image that conform to the input descriptive text, and generate an answer in the form of a chain of thought. Subsequently, using the trained multi-modal large model, end-to-end descriptive object detection can be achieved without integrating multiple models, which can reduce the complexity of implementing descriptive object detection and improve the detection efficiency.
[0051] Of course, it is not necessary for any product or method implementing the present application to achieve all the above-mentioned advantages simultaneously. BRIEF DESCRIPTION OF THE DRAWINGS
[0052] In order to more clearly illustrate the technical solutions in the embodiments of the present application or the prior art, the following will briefly introduce the drawings required for use in the description of the embodiments or the prior art. Obviously, the drawings in the following description are only some embodiments of the present application, and those of ordinary skill in the art can also obtain other embodiments according to these drawings.
[0053] Figure 1 The first flowchart of the multi-modal large model training method provided by the embodiment of the present application;
[0054] Figure 2 The schematic diagram of a sample image provided by the embodiment of the present application;
[0055] Figure 3 The second flowchart of the multi-modal large model training method provided by the embodiment of the present application;
[0056] Figure 4 The flowchart of combining to obtain the second sample descriptive text provided by the embodiment of the present application;
[0057] Figure 5 A schematic flowchart of a process for constructing question - answer sample pairs provided by an embodiment of the present application;
[0058] Figure 6 A schematic flowchart of a target detection method provided by an embodiment of the present application;
[0059] Figure 7 A schematic structural diagram of a multi - modal large model training device provided by an embodiment of the present application;
[0060] Figure 8 A schematic structural diagram of a target detection device provided by an embodiment of the present application;
[0061] Figure 9 A schematic structural diagram of an electronic device provided by an embodiment of the present application. Detailed implementation manners
[0062] Next, the technical solutions in the embodiments of the present application will be clearly and completely described in conjunction with the accompanying drawings in the embodiments of the present application. Obviously, the described embodiments are only a part of the embodiments of the present application, rather than all the embodiments. Based on the embodiments in the present application, all other embodiments obtained by those of ordinary skill in the art based on the present application belong to the scope of protection of the present application.
[0063] To reduce the complexity of implementing descriptive target detection and improve the detection efficiency, an embodiment of the present application provides a multi - modal large model training method. Refer to Figure 1 , Figure 1 The first schematic flowchart of the multi - modal large model training method provided by an embodiment of the present application. The multi - modal large model training method may include:
[0064] Step S101: Obtain a plurality of sample images and the first sample description text of the object annotation boxes of the specified category in each sample image.
[0065] Step S102: For each sample image, use the first large - language model and the first text prompt to extract the category name of the object described by the first sample description text corresponding to the sample image and multiple attributes of the object, and combine the obtained category name with at least one of the obtained attributes to obtain a plurality of second sample description texts corresponding to the sample image.
[0066] Among them, the first text prompt is used to indicate: extracting the category name of the object described by the input description text and the attributes of the object.
[0067] Step S103: For each object bounding box in the sample image, determine whether the object bounding box matches each attribute used to combine multiple second sample description texts corresponding to the sample image.
[0068] Step S104: Construct a sample question containing each second sample description text, and use the matching results between each attribute included in the second sample description text and each object bounding box to construct a sample answer in the form of a thought chain corresponding to the sample question, obtaining a question-answer sample pair.
[0069] Among them, the sample question is used to indicate the position of the image region in the input image that conforms to the second sample description text; the sample answer contains a first inference process text for describing the inference process of the multimodal large model; the first inference process text includes: each inference step and the first execution order between each inference step; according to the first execution order, each inference step is respectively: extracting the category name of the object described by the description text included in the input question and the attributes of the object, detecting the position of the image region occupied by all objects belonging to the category represented by the extracted category name in the input image, respectively determining the matching results between each image region and each extracted attribute, and determining the position of the image region that matches each extracted attribute.
[0070] Step S105: Input the sample image and the sample questions in each question-answer sample pair into the multimodal large model with an initial structure to obtain a predicted answer;
[0071] Step S106: Based on the difference between the obtained predicted answer and the sample answer in the corresponding question-answer sample pair, adjust the parameters of the multimodal large model with an initial structure until a preset convergence condition is reached, obtaining a trained multimodal large model.
[0072] Based on the above processing, according to the obtained sample image and the description text (i.e., the first sample description text) for describing the object bounding box of the specified category in the sample image, the first large language model can be used to generate second sample description texts to increase the number of obtained second sample description texts and enrich the diversity of the second sample description texts. Furthermore, sample questions containing each second sample description text can be constructed to instruct the multimodal large model to determine the position of the image region in the input image that conforms to the second sample description text. And by combining the matching results between each object bounding box in each sample image and each attribute used to combine multiple second sample description texts corresponding to the sample image, a sample answer in the form of a thought chain can be constructed, and thus a question-answer sample pair in the form of a thought chain can be obtained. The sample answer in the form of a thought chain can include: the specific steps for inferring the position of the image region in the sample image that matches each attribute in the second sample description text.
[0073] Correspondingly, training a multimodal large model based on a chain-of-thought form of question-answer sample pairs enables the multimodal large model to learn to reason in the chain-of-thought form, determine the position of the image area occupied by the object that conforms to the input description text in the image, and generate an answer in the chain-of-thought form. Subsequently, using the trained multimodal large model, end-to-end descriptive object detection can be achieved without integrating multiple models, which can reduce the complexity of implementing descriptive object detection and improve the detection efficiency.
[0074] For step S101, multiple sample images and sample description texts (i.e., first sample description texts) of the object annotation boxes of a specified category in each sample image can be obtained. The specified category can be set according to the category of the target object to be detected currently. For example, if the target object to be detected currently is an athlete, the specified category can be an athlete; or, if the target object to be detected currently is a vehicle, the specified category can be a vehicle. The object annotation box of the specified category in an image can represent the bounding box of the image area occupied by the object belonging to the specified category in the image. The first sample description text of an object annotation box is used to describe the image content of the image area indicated by the object annotation box. The number of object annotation boxes of the specified category in a sample image is the same as the number of objects belonging to the specified category in the sample image, and can be one or more.
[0075] Each sample image and the object annotation boxes of the specified category in each sample image can come from a public dataset (e.g., a public object detection dataset). Or, sample images can also be collected using an image acquisition device, and the objects belonging to the specified category in each of the collected sample images can be manually annotated to obtain the object annotation boxes of the specified category in each sample image. The first sample description texts of the object annotation boxes of the specified category in each sample image can also come from a public dataset, or the description texts of the object annotation boxes of the specified category in each sample image can be determined manually to obtain the first sample description texts. For example, the sample image can be a photo of a football field with multiple athletes playing a football game. The specified category can be an athlete. Correspondingly, the object annotation boxes of the specified category in this sample image can be the bounding boxes of the image areas occupied by each athlete. As Figure 2 shown, Figure 2 is a schematic diagram of a sample image provided by an embodiment of the present application. The sample image contains three football players, namely A, B, and C. The object annotation boxes of the specified category in this sample image include: the object annotation box 21 of football player A, the object annotation box 22 of football player B, and the object annotation box 23 of football player C. Football players B and C are wearing green uniforms ( Figure 2 represented by the hatched rectangular boxes). Football player A is wearing a yellow uniform (Figure 2 The rectangular box filled with vertical lines (represented by the rectangular box filled with vertical lines in the figure) and black stockings ( Figure 2 represented by the rectangular box filled with black in the figure), and is running, then the description text of the object annotation box 21 can be "In the framed area, there is a football player wearing a yellow uniform and black stockings, and he is running".
[0076] In one implementation, for the object annotation boxes of a specified category in each sample image, a model with the ability to generate region descriptions can also be used to generate the description text of the object annotation boxes of the specified category in the sample image, obtaining the first sample description text. Step S101 includes:
[0077] Step 1: Obtain a plurality of sample images and the positions of the object annotation boxes of a specified category in each sample image.
[0078] Step 2: For each sample image, input the sample image and the second text prompt into the second large language model to obtain the first sample description text of the object annotation boxes of the specified category in the sample image.
[0079] Among them, the second text prompt is used to indicate: generate the description text of the object annotation box in the input image.
[0080] In this implementation, the second text prompt can instruct the second large language model to generate the description text of the object annotation box in the input image. In one implementation, the second text prompt can include information about the position of the object annotation box in the input image (such as the coordinates of the object annotation box). For example, the second text prompt can be "Please generate the description text of the image area in the rectangular box with the upper left vertex coordinates (a, b) and the lower right vertex coordinates (c, d) in the input image". In another implementation, an object annotation box can be added to the input image, and the image with the added object annotation box is input into the second large language model for processing, and the second text prompt can instruct the second large language model to generate the description text of the image area in the added object annotation box in the input image. For example, the second text prompt can be "Please generate the description text of the image area in the added object annotation box in the input image".
[0081] For each sample image, the position of the sample image and the object bounding box of the specified category in the sample image can be obtained. Furthermore, the sample image and the second text prompt can be input into the second large language model to obtain the first sample description text of the object bounding box of the specified category in the sample image. The second large language model can be a pre-trained multi-modal large language model with the ability to describe regions, and it can be any multi-modal large language model that can generate the description text of any image region in the image, without specific limitation. For example, the second large language model can be Qwen2.5-VL or ChatGPT.
[0082] Based on the above processing, for each sample image, the second large language model can be used to generate the description text of the object bounding box in the sample image (i.e., the first sample description text). In this way, it can be ensured that the first sample description text can be obtained. Subsequently, based on the sample image and the first sample image, it is possible to construct a question-answer sample pair in the form of a chain of thought, and then, to train the multi-modal large model. It is ensured that the trained multi-modal large model can be used subsequently to achieve end-to-end descriptive object detection, and there is no need to integrate multiple models, which can reduce the complexity of implementing descriptive object detection and improve the detection efficiency.
[0083] For step S102, for each sample image, the first large language model can be used to process the first sample description text corresponding to the sample image in combination with the first text prompt, and extract the category name of the object described by the first sample description text corresponding to the sample image and multiple attributes of the object. To ensure the accuracy of the obtained category name, the category name extracted by the first large language model is consistent with the category name of the object in the object bounding box of the specified category in the sample image. For specific details, reference can be made to the relevant descriptions of steps S107 and S1021 in the subsequent embodiments. The first large language model can be any large language model capable of implementing step S102, without specific limitation. For example, the first large language model can be Qwen2.5-72B or ChatGPT.
[0084] The category name of the object described by a descriptive text can represent the category to which the object (also referred to as the subject) described by the descriptive text belongs. For example, if the object described by a descriptive text is an athlete on a football field, the category name of the object described by the descriptive text can be "football player"; or, if the object described by a descriptive text is a little girl in a park, the category name of the object described by the descriptive text can be "girl". The object described by a descriptive text can be an objective entity. For example, it can be a person, an animal, a car, or a plant. The attributes of the object described by a descriptive text can be attribute information that can be visually perceived. For example, it can include appearance, material, interaction relationship, or action state. For example, a first sample descriptive text is "In the framed area, there is a football player wearing a yellow uniform and black stockings, and he is running". For this first sample descriptive text, the extracted category name can be "football player", and the extracted attributes can be: wearing a yellow uniform, wearing black stockings, and running.
[0085] And the obtained category name can be combined with at least one of the obtained attributes to obtain multiple second sample descriptive texts corresponding to the sample image. For example, at least one attribute can be selected from the obtained attributes to obtain multiple attribute sets (also referred to as non-empty attribute subsets). A preset number of attribute sets can be selected from the multiple obtained attribute sets, and the selected attribute sets can be combined with the obtained category name to obtain multiple second sample descriptive texts corresponding to the sample image. The preset number is greater than 1 and not greater than the total number of attribute sets. If there are N obtained attributes, the total number of attribute sets is . For example, the extracted category name is "football player", and the extracted attributes are: wearing a yellow uniform, wearing black stockings, and running. Any one, any two, or all three of the three extracted attributes can be selected and combined with the extracted category name. For example, if "wearing a yellow uniform" and "wearing black stockings" are selected, the combined second sample descriptive text can be "a football player wearing a yellow uniform and black stockings"; if "running" is selected, the combined second sample descriptive text can be "a football player who is running".
[0086] In this application, the following two methods can be used to instruct the first large language model to extract the category name and attributes, and to combine the obtained category name with the obtained attributes:
[0087] In the first method, the first text prompt and the first large language model can be used to extract the category name and attributes. And the third text prompt and the first large language model can be used to combine the obtained category name with the obtained attributes to obtain multiple second sample descriptive texts. Step S102 includes:
[0088] Step 1: For each sample image, input the first sample description text and the first text prompt corresponding to the sample image into the first large language model, and extract the category name of the object described by the first sample description text corresponding to the sample image and multiple attributes of the object.
[0089] Step 2: Input the extracted category name, the extracted attributes, and the third text prompt into the first large language model to obtain multiple second sample description texts corresponding to the sample image.
[0090] Among them, the third text prompt is used to indicate: combine the input category name with at least one of the input attributes to obtain multiple description texts.
[0091] In the embodiment of the present application, for each sample image, the first large language model can, according to the indication of the first text prompt, extract the category name of the object described by the first sample description text corresponding to the sample image and multiple attributes of the object, and can, according to the indication of the third text prompt, combine the extracted category name with at least one of the extracted attributes to obtain multiple second sample description texts corresponding to the sample image. In this way, the decomposition of the first sample description text can be realized through the first large language model, and multiple new description texts (i.e., second sample description texts) can be combined. The processing ability of the first large language model for text can be utilized to conveniently and efficiently realize the expansion of the description text. Furthermore, it is ensured that a question-and-answer sample pair in the form of a chain of thought can be constructed subsequently to realize the training of the multi-modal large model. Correspondingly, by using the trained multi-modal large model, end-to-end descriptive object detection can be realized, without integrating multiple models, which can reduce the complexity of realizing descriptive object detection and improve the detection efficiency.
[0092] The first text prompt is used to indicate: extract the category name of the object described by the input description text and the attributes of the object. The first text prompt may include: the format of outputting the extracted category name and the extracted attributes, the extraction rules that need to be satisfied for extracting the category name of the object described by the input description text and the attributes of the object, and examples of the category name and attributes extracted according to the input description text. In this way, it can be ensured that the category name and attributes that meet the requirements can be extracted. The following is an example of the first text prompt:
[0093] "You are a helpful assistant. I will provide you with a description based on a certain target box and its corresponding original object detection category name. Please help me extract visual information in the format [subject, first attribute, second attribute,..., Nth attribute], where "subject" represents the target entity, and "first attribute", "second attribute",..., "Nth attribute" represent the attribute information related to the subject. The extraction must be as strict, diverse, and concise as possible. Here are some rules: 1. The extracted subject should be the same as the given original object detection category name, belonging to synonyms or near-synonyms. If not, the output should be abandoned; 2. The subject refers to entities such as people, animals, cars, objects, plants, etc., rather than non-entities such as the environment and atmosphere; 3. Extract objective descriptions, rather than subjective ones such as inferences, guesses, or emotional descriptions; 4. If the "subject" or "attribute" cannot be accurately extracted, the output should be abandoned; 5. Attributes refer to appearance, material, action relationship, state, etc., and they cannot be functional words, inferences, guesses, or emotions."
[0094] For example:
[0095] Original detection category name: Boy
[0096] Region description: This little boy is wearing a blue helmet with white and black stripe patterns on it. The helmet seems to be firmly fastened under his chin. He shows a smile, baring his teeth. The boy has short, light brown hair and is wearing a dark blue T-shirt. The boy is riding a bicycle. From his smile, it can be seen that he enjoys this activity. He is located among a group of children and bicycles, indicating that a group activity or a family outing is taking place."
[0097] Output result:
[0098] [Little boy, wearing a blue helmet, showing a smile, having short, light brown hair, wearing a dark blue T-shirt, riding a bicycle]."
[0099] Please extract visual information for the following region description and category name:
[0100] Original detection category name: {Detection category name of the box to be processed}
[0101] Region description: {Region description of the box to be processed}".
[0102] In the above example, the extracted subject can represent the extracted category name, the original detection category name can represent the category name of the object in the object annotation box of the specified category in the sample image to be processed, and the region description is the description text. For each sample image, the detection category name of the box to be processed is the category name of the object in the object annotation box of the specified category in this sample image, and the region description of the box to be processed is the first sample description text corresponding to this sample image."
[0103] The third text prompt is used to indicate that the input category name is combined with at least one of the input attributes to obtain multiple description texts. The third text prompt may include examples of combining the input category name with at least one of the input attributes to obtain multiple description texts. In this way, it can be ensured that the combined description texts meet the requirements. The following is an example of a third text prompt:
[0104] "You are a helpful assistant. I will give you a set of visual information in the format [subject, attribute1, attribute2,..., attributeN], where "subject" represents the target entity, and "attribute1, attribute2,..., attributeN" represent N attributes of the target. Please help me complete the following tasks: Combine the N attributes into any non-empty subset and combine them with the subject to be merged into the target description respectively. Ensure that the final expression should be fluent and accurate. For example:
[0105] Visual relationship:
[0106] [‘Boy, wearing a blue helmet, showing a smile, and riding a bicycle’]
[0107] Result:
[0108] [‘Boy, wearing a blue helmet’]: The boy wearing a blue helmet
[0109] [‘Boy, showing a smile’]: The boy showing a smile
[0110] [‘Boy, riding a bicycle’]: The boy riding a bicycle
[0111] [‘Boy, wearing a blue helmet, showing a smile’]: The boy wearing a blue helmet and showing a smile
[0112] [‘Boy, wearing a blue helmet, riding a bicycle’]: The boy wearing a blue helmet and riding a bicycle
[0113] [‘Boy, showing a smile, riding a bicycle’]: The boy riding a bicycle and showing a smile
[0114] [‘Boy, wearing a blue helmet, showing a smile, riding a bicycle’]: The boy wearing a blue helmet, riding a bicycle and showing a smile.
[0115] Please synthesize the target description for the following visual relationship:
[0116] Visual relationship: {Visual relationship to be processed}".
[0117] In the above example, the visual relationship to be processed is the extracted category name and the extracted attributes, and the target description is the second sample description text.
[0118] In Method 2, the first text prompt is used to indicate: extract the category name of the object described in the input description text and the attributes of the object, and combine at least one of the extracted category name and the extracted attributes to obtain the description text. For example, the first text prompt can be "You are a helpful assistant. I will provide you with a description based on a certain target box and its corresponding original object detection category name. Please help me extract visual information according to it in the format of [subject, first attribute, second attribute,..., Nth attribute], where "subject" represents the target subject, and "first attribute", "second attribute",..., "Nth attribute" represent the attribute information related to the subject. The visual information must be extracted as strictly, diversely, and concisely as possible. The following are some rules: 1. The extracted subject should be the same as the given original object detection category name, belonging to synonyms or near-synonyms. If they are not the same, the output should be abandoned; 2. The subject refers to entities such as people, animals, cars, objects, plants, etc., rather than non-entities such as environment and atmosphere; 3. Extract objective descriptions, rather than subjective ones such as inferences, guesses, or emotional descriptions; 4. If the "subject" or "attribute" cannot be accurately extracted, the output should be abandoned; 5. Attributes refer to appearance, material, action relationship, state, etc., and they cannot be function words, inferences, guesses, or emotions. Combine the N extracted attributes into any non-empty subset and combine them with the extracted subject to form the target description respectively. Ensure that the final expression is fluent and accurate. Please process based on the following region description and category name:
[0119] Original detection category name: {Detection category name of the box to be processed}
[0120] Region description: {Region description of the box to be processed}
[0121] For step S103, for each object annotation box in the sample image, determine whether the object annotation box matches each attribute used to combine to obtain the multiple second sample description texts corresponding to the sample image. An object annotation box matching an attribute means that the image content in the object annotation box conforms to the attribute.
[0122] In one implementation, it can be determined manually whether the object annotation box matches each attribute used to combine to obtain the multiple second sample description texts corresponding to the sample image.
[0123] In another implementation, step S103 includes: using an image-text matching model and a fourth text prompt to detect whether each object annotation box of a specified category in the sample image matches each description text to be matched.
[0124] Among them, each description text to be matched is obtained by combining each attribute used to combine multiple second sample description texts corresponding to the sample image with the extracted category name respectively; the fourth text prompt is used to indicate: determining whether each object bounding box in the input image matches the input description text.
[0125] In the embodiments of the present application, for each object bounding box in the sample image, it can be determined whether the object bounding box matches each attribute used to combine multiple second sample description texts corresponding to the sample image through a model (which can be called an image-text matching model) that can judge whether the image content in the object bounding box in the image matches the description text. That is, it can be determined whether the object bounding box matches each attribute used to combine multiple second sample description texts corresponding to the sample image through an image-text matching model with the ability to determine the matching result of the referential expression. The image-text matching model can be any multi-modal large language model capable of having the ability to determine the matching result of the referential expression, without specific limitation. For example, the image-text matching model can be Qwen2.5-VL. That an object bounding box matches a description text can mean that the image content in the object bounding box conforms to the description text.
[0126] Each attribute used to combine multiple second sample description texts corresponding to the sample image and the extracted category name can be combined in advance to obtain each description text to be matched. For example, the extracted category name can be a football player, and the attributes used to combine multiple second sample description texts corresponding to the sample image can include: wearing a yellow uniform, wearing black stockings, and running. Accordingly, each description text to be matched can be: a football player wearing a yellow uniform, a football player wearing black stockings, a football player running.
[0127] Using the image-text matching model, according to the indication of the fourth text prompt, it can be detected whether each object bounding box of a specified category in the sample image matches each description text to be matched. For example, the description text to be matched, the sample image, and the fourth text prompt can be input into the image-text matching model to obtain a matching result output by the image-text matching model, which characterizes whether each object bounding box of a specified category in the sample image matches. All description texts to be matched can be input into the image-text matching model at once. Or, all description texts to be matched can also be input into the image-text matching model in multiple times until all description texts to be matched are detected. The number of description texts to be matched input into the image-text matching model each time can be set as needed, without specific limitation.
[0128] In one implementation, the fourth text prompt may include information about the positions of the bounding boxes of each object in the input image (e.g., the coordinates of the bounding boxes). For example, the fourth text prompt may be "Please respectively determine whether the rectangular box with the top-left vertex coordinates (a1, b1) and the bottom-right vertex coordinates (c1, d1) in the input image, and the rectangular box with the top-left vertex coordinates (a2, b2) and the bottom-right vertex coordinates (c2, d2) match each of the input text descriptions to be matched." In another implementation, object bounding boxes may be added to the input image, and the image with the added object bounding boxes is input into the image-text matching model for processing. The fourth text prompt may instruct the image-text matching model to determine whether the added object bounding boxes in the input image match each of the input text descriptions to be matched. For example, the fourth text prompt may be "Please respectively determine whether each of the added object bounding boxes in the input image matches each of the input text descriptions to be matched."
[0129] For example, the obtained matching results may be presented in the form of an attribute table, as shown in Table 1. The object bounding boxes in the sample image may include: the first box, the second box, and the third box. Each text description to be matched may be: a football player in a yellow uniform, a football player wearing black stockings, a football player running. "√" indicates a match, and "×" indicates a non-match. Taking the matching result between the first box and the football player in a yellow uniform as an example, the first box matches the football player in a yellow uniform. Correspondingly, it also means that the first box matches the attribute "in a yellow uniform".
[0130] Table 1
[0131]
[0132] Based on the above processing, the matching results of each object bounding box in each sample image with each attribute used to combine to obtain the corresponding multiple second sample description texts of the sample image can be obtained, that is, the matching results of each attribute included in each second sample description text with each object bounding box in each sample image can be obtained. In this way, it is ensured that subsequent construction of question-answer sample pairs can be achieved based on the obtained matching results to realize the training of the multimodal large model. It is ensured that subsequent use of the trained multimodal large model can achieve end-to-end descriptive object detection, and thus there is no need to integrate multiple models, which can reduce the complexity of implementing descriptive object detection and improve the detection efficiency.
[0133] For step S104, for each second sample description text, a sample question containing the second sample description text can be constructed. The constructed sample question can be used to indicate: determining the position of the image region in the input image that conforms to the second sample description text. For example, a second sample description text can be "an athlete wearing non-white stockings on the green grass", and the sample question containing this second sample description text can be "Please provide the coordinates of all target boxes corresponding to this description in the figure: 'an athlete wearing non-white stockings on the green grass'. Please reason in the form of a chain of thought to obtain the final answer".
[0134] And the matching results of each attribute included in the second sample description text with each object annotation box can be used to construct a sample answer in the form of a chain of thought corresponding to the sample question. The sample answer contains the first reasoning process text for describing the reasoning process of the multimodal large model. The first reasoning process text can include: each reasoning step and the first execution order between each reasoning step. According to the first execution order, each reasoning step is respectively: extracting the category name of the object described by the description text contained in the input question and the attributes of the object (which can be called description decomposition), detecting the position of the image region occupied by all objects belonging to the category represented by the extracted category name in the input image (which can be called open-set object detection), respectively determining the matching results of each image region with each of the extracted attributes (which can be called per-frame per-attribute matching), and determining the position of the image region that matches each of the extracted attributes (which can be called summarization).
[0135] During the process of description decomposition, the category names extracted are the category names used when combining to obtain the second sample description text in step S102, and the attributes extracted are the attributes used when combining to obtain the second sample description text in step S102. During the process of open-set object detection, the position of the detected image region is the position of the object annotation box of the specified category in the sample image corresponding to the second sample description text, and the position of the object annotation box can be represented by the coordinates of the upper left vertex and the lower right vertex. During the process of frame-by-frame and attribute-by-attribute matching, the matching result of each image region and each extracted attribute is the matching result determined in step S103 between each object annotation box in the sample image corresponding to the second sample description text and each attribute included in the second sample description text. During the process of summarization, the position of the image region that matches each extracted attribute is the position of the object annotation box that matches each attribute included in the second sample description text. It is possible to search for the object annotation box that matches each attribute included in the second sample description text in the matching results determined in step S103, and the position of the found object annotation box can be used as the position determined in the summarization. For example, it is possible to search for the object annotation box that matches each attribute included in the second sample description text in the attribute table. For example, if the second sample description text is "a football player wearing a yellow uniform and black stockings", it is possible to search in Table 1 for the object annotation box that matches the attributes "wearing a yellow uniform" and "wearing black stockings", that is, the first box. Correspondingly, the position of the first box can be used as the position determined in the summarization.
[0136] For example, for the second sample description text "an athlete wearing non-white stockings on the green field", the constructed sample answer is as follows:
[0137] "Step 1: (Description Decomposition)
[0138] First, split the main body and sub-attributes of the description, and they can be extracted as:
[0139] Main body: 'athlete', attributes: 'athlete wearing non-white stockings; athlete moving on the green field'.
[0140] Step 2: (Open-Set Object Detection)
[0141] All bounding boxes of the main body athlete in the image are: [008,080,242,494]; [260,081,460,511]; [521,085,772,476].
[0142] Step 3: (Frame-by-Frame and Attribute-by-Attribute Matching)
[0143] For all bounding boxes and all attributes in the description, the matching results are as follows:
[0144] [008,080,242,494]: Athletes wearing non - white thigh - high socks: No; Athletes playing on a green field: Yes;
[0145] [260,081,460,511]: Athletes wearing non - white thigh - high socks: No; Athletes playing on a green field: Yes;
[0146] [521,085,772,476]: Athletes wearing non - white thigh - high socks: Yes; Athletes playing on a green field: Yes;
[0147] Step 4: (Summary)
[0148] In summary, the coordinates of the region bounding box described by this sentence are: [521,085,772,476].
[0149] In this way, the descriptive object detection task can be split into a four - step reasoning paradigm of description decomposition, open - set object detection, per - box per - attribute matching, and summary. Subsequently, by using the constructed question - answering sample pairs in the form of a chain of thought to train the multi - modal large model, the multi - modal large model can learn the ability to reason step by step according to the above four reasoning steps, determine the position of the object region in the image that conforms to the input description text, and generate an answer in the form of a chain of thought.
[0150] For steps S105 and S106, input the sample image and the sample questions in each question - answering sample pair into the multi - modal large model with the initial structure to obtain a predicted answer. The question - answering sample pair corresponds to the obtained predicted answer. Furthermore, the difference between the calculated predicted answer and the sample answer in the corresponding question - answering sample pair can be calculated. For example, the similarity between the calculated predicted answer and the sample answer in the corresponding question - answering sample pair can be calculated to represent the difference between the two. According to the difference between the obtained predicted answer and the sample answer in the corresponding question - answering sample pair, the parameters of the multi - modal large model with the initial structure can be adjusted until the preset convergence condition is reached, and a trained multi - modal large model is obtained. For example, the preset convergence condition can be that the number of times of adjusting the parameters of the multi - modal large model reaches a preset number, or it can also be that the difference between the obtained predicted answer and the sample answer in the corresponding question - answering sample pair is less than a preset difference.
[0151] In this way, by using the trained multi - modal large model, end - to - end descriptive object detection can be achieved. Without integrating multiple models, it can reduce the complexity of implementing descriptive object detection and improve the detection efficiency. Moreover, the multi - modal large model can reason step by step according to the reasoning steps of description decomposition, open - set object detection, per - box per - attribute matching, and summary, and can accurately and logically achieve complex descriptive object detection.
[0152] It can be seen that in the solution provided by the present application, a descriptive object detection paradigm based on the chain of thought is proposed, and a corresponding descriptive object detection dataset in the form of a chain of thought is constructed, so as to make full use of the generative characteristics and logical reasoning ability of the multimodal large model, and can make full use of the complementarity of image and text information to accurately understand the objects and their attributes in complex scenes, and can improve the performance of the descriptive object detection task. The multimodal large model trained by the multimodal large model training method provided by the present application can realize the detection of open-set objects with natural language description constraints, providing a more flexible and powerful object detection ability for the intelligent vision system.
[0153] In one embodiment, referring to Figure 3 , Figure 3 is the second process schematic diagram of the multimodal large model training method provided by the embodiment of the present application. Before step S102, the multimodal large model training method further includes:
[0154] Step S107: For each sample image, obtain the class name of the object in the object annotation box of the specified class in the sample image as the preset class name.
[0155] Step S102 includes:
[0156] Step S1021: For each sample image, input the first text prompt, the first sample description text corresponding to the sample image, and the preset class name into the first large language model, extract the class name of the object described by the first sample description text corresponding to the sample image and multiple attributes of the object, and combine the obtained class name with at least one of the obtained attributes to obtain multiple second sample description texts corresponding to the sample image.
[0157] Among them, the obtained class name is the same as the class represented by the preset class name.
[0158] In the embodiments of the present application, affected by the accuracy of the first large language model, the category name of the object described in the input description text recognized by the first large language model may not be accurate. To ensure the accuracy of the category name output by the first large language model, for each sample image, the category name of the object in the object annotation box of the specified category in the sample image can be obtained as the preset category name. The preset category name can thus be used as the accurate category name. The first text prompt, the first sample description text corresponding to the sample image, and the preset category name can be input into the first large language model. The first large language model can extract a category name that is consistent with the category represented by the preset category name. If, during the process of processing the first sample description text corresponding to a sample image, the category name extracted by the first large language model is inconsistent with the preset category name, the output is discarded. For example, the first large language model can determine whether the extracted category name belongs to a synonym or a near-synonym of the preset category name. If so, it is determined that the extracted category name is consistent with the preset category name and can be output; otherwise, it is determined that the extracted category name is inconsistent with the preset category name and the output is discarded.
[0159] Based on the above processing, the category name of the object in the object annotation box of the specified category in the sample image (i.e., the preset category name) can be used as a benchmark to ensure the accuracy of the category name extracted by the first large language model. Furthermore, the accuracy of the combined second sample description text can be ensured. Correspondingly, the accuracy of the constructed question-answer sample pair can also be ensured. Furthermore, it can be ensured that the training of the multi-modal large model can be achieved. It is ensured that the end-to-end descriptive object detection can be realized by using the trained multi-modal large model later, without the need to integrate multiple models, which can reduce the complexity of realizing descriptive object detection and improve the detection efficiency.
[0160] In one embodiment, refer to Figure 4 , Figure 4 which is a schematic flowchart of a process for combining to obtain a second sample description text provided by the embodiments of the present application. For each sample image, the first text prompt, the box detection category name (i.e., the preset category name in the above embodiments), and the box region description (i.e., the first sample description text corresponding to the sample image in the above embodiments) can be input into the large language model (i.e., the first large language model in the above embodiments) for visual information extraction. That is, the category name (which can also be referred to as the subject) of the object described in the first sample description text corresponding to the sample image and multiple attributes of the object (such as Figure 4 the first attribute, the second attribute,..., the Nth attribute) are extracted. And the large language model can be used to combine the extracted category name with at least one of the obtained attributes to obtain a description text (i.e., multiple second sample description texts corresponding to the sample image) that includes all attributes or any subset of attributes, a total of Article. The sub-attribute set is other attribute sets in the attribute set in the above embodiment except the attribute set containing all attributes. In this way, the decomposition of the subject and N sub-attributes can be realized, and any sub-attribute set and the subject can be combined, so as to generate the second sample description text more efficiently and with finer granularity.
[0161] In one embodiment, see Figure 5 , Figure 5 , which is a schematic flowchart of a process for constructing a question-and-answer sample pair provided by an embodiment of the present application. It includes the following steps:
[0162] Step S501: Obtain a traditional object detection data set. For a sample image in the object detection data set, for a certain category (i.e., the specified category in the above embodiment), there can be M boxes (i.e., the object annotation boxes of the specified category in the above embodiment) in the sample image.
[0163] Step S502: For any one or several boxes in the sample image, use a multi-modal model with region description ability to generate a detailed region description. That is, steps 1 and 2 in the above embodiment.
[0164] Step S503: For each region description, use a large language model to extract the subject and N attributes, and arbitrarily combine them into non-empty attribute subsets, and synthesize them with the subject. That is, step S102 in the above embodiment.
[0165] Step S504: Through the referential expression matching model, perform consistency verification on the M boxes and N attributes to generate an attribute table. The referential expression matching model is the graphic-text matching model in the above embodiment. That is, use the graphic-text matching model and the fourth text prompt to respectively detect whether each object annotation box of the specified category in the sample image matches each text to be matched.
[0166] Step S505: Format conversion to generate descriptive object detection training data in the form of a thought chain. That is, step S104 in the above embodiment.
[0167] Based on the same inventive concept, an embodiment of the present application also provides an object detection method. See Figure 6 , Figure 6 , which is a schematic flowchart of an object detection method provided by an embodiment of the present application. The object detection method includes:
[0168] Step S601: Obtain an image to be detected and a text prompt to be used containing the text description to be detected.
[0169] Among them, the text prompt to be used is used to indicate: determining the position of the image area occupied by the object that meets the text description to be detected in the image to be detected.
[0170] Step S602: Input the image to be detected and the text prompt to be utilized into a pre-trained multi-modal large model to obtain a detection result in the form of a chain of thought.
[0171] Among them, the multi-modal large model is trained based on any of the above multi-modal large model training methods; the detection result includes: a second inference process text for describing the inference process of the multi-modal large model; the second inference process text includes: each inference step and the second execution order between each inference step; according to the second execution order, each inference step is respectively: extracting the category name of the object described by the input text prompt and the attributes of the object, detecting the positions of the image regions occupied by all the objects belonging to the category represented by the extracted category name in the input image, respectively determining the matching results between each image region and each of the extracted attributes, and determining the positions of the image regions that match each of the extracted attributes.
[0172] In the embodiments of the present application, an image to be detected for target detection currently (i.e., the image to be detected) and a text prompt to be utilized including a description text (i.e., the text prompt to be detected) that the target to be detected currently conforms to can be obtained. The text prompt to be utilized can be used to indicate: determining the positions of the image regions occupied by the objects in the image to be detected that conform to the text prompt to be detected. For example, a text prompt to be detected can be "athletes wearing black stockings on the green grassland", and the text prompt to be utilized including this text prompt to be detected can be "Please provide the coordinates of all target boxes corresponding to this description in the figure: 'athletes wearing black stockings on the green grassland'. Please reason in the form of a chain of thought to obtain the final answer".
[0173] Furthermore, input the image to be detected and the text prompt to be utilized into a pre-trained multi-modal large model. The multi-modal large model can extract the category name of the object described by the text prompt to be detected in the text prompt to be utilized and the attributes of the object (i.e., perform description decomposition), and detect the positions of the image regions occupied by all the objects belonging to the category represented by the extracted category name in the image to be detected (i.e., perform open-set object detection), respectively determine the matching results between each image region and each of the extracted attributes (i.e., perform per-frame and per-attribute matching), and determine the positions of the image regions that match each of the extracted attributes (i.e., perform summarization). That is, the multi-modal large model can reason step by step according to the inference process of description decomposition, open-set object detection, per-frame and per-attribute matching, and summarization to determine the positions of the image regions occupied by the objects in the image to be detected that conform to the text prompt to be detected, and output an answer in the form of a chain of thought.
[0174] Based on the above processing, during the training process, the multi-modal large model can learn to reason in the form of a chain of thought, determine the position of the object region in the image that conforms to the input descriptive text, and generate an answer in the form of a chain of thought. By using the trained multi-modal large model, end-to-end descriptive object detection can be achieved without integrating multiple models, which can reduce the complexity of implementing descriptive object detection and improve the detection efficiency.
[0175] The embodiment of the present application also provides a multi-modal large model training device. Refer to Figure 7 , Figure 7 which is a schematic structural diagram of a multi-modal large model training device provided by the embodiment of the present application. The device includes:
[0176] A sample acquisition module 701, configured to acquire a plurality of sample images and a first sample description text of the object annotation box of a specified category in each sample image;
[0177] A description text generation module 702, configured to, for each sample image, use a first large language model and a first text prompt to extract the category name of the object described by the first sample description text corresponding to the sample image and multiple attributes of the object, and combine the obtained category name with at least one of the obtained attributes to obtain a plurality of second sample description texts corresponding to the sample image; wherein, the first text prompt is used to indicate: extracting the category name of the object described by the input description text and the attributes of the object;
[0178] A matching result determination module 703, configured to respectively determine whether each object annotation box in the sample image matches each attribute used to combine to obtain a plurality of second sample description texts corresponding to the sample image;
[0179] The Q&A sample pair construction module 704 is used to construct a sample question containing each second sample description text, and construct a sample answer in the form of a chain of thought corresponding to the sample question by using the matching result of each attribute contained in the second sample description text and each object annotation box, so as to obtain a Q&A sample pair; wherein, the sample question is used to indicate the position of the image area in the input image that conforms to the second sample description text; the sample answer contains the first reasoning process text for describing the reasoning process of the multi-modal large model; the first reasoning process text includes: each reasoning step and the first execution order between each reasoning step; according to the first execution order, each reasoning step is respectively: extracting the category name of the object described by the description text contained in the input question and the attributes of the object, detecting the position of the image area occupied by all objects belonging to the category represented by the extracted category name in the input image, respectively determining the matching result of each image area and each extracted attribute, and determining the position of the image area that matches each extracted attribute.
[0180] The predicted answer determination module 705 is used to input the sample image and the sample question in each Q&A sample pair into the multi-modal large model with the initial structure to obtain the predicted answer.
[0181] The model parameter adjustment module 706 is used to adjust the parameters of the multi-modal large model with the initial structure based on the difference between the obtained predicted answer and the sample answer in the corresponding Q&A sample pair until the preset convergence condition is reached, and obtain the trained multi-modal large model.
[0182] Based on the multi-modal large model training device provided by the embodiments of the present application, the second sample description text can be generated by using the first large language model according to the obtained sample image and the description text (i.e., the first sample description text) for describing the object annotation box of the specified category in the sample image, so as to increase the number of the obtained second sample description texts and enrich the diversity of the second sample description texts. Furthermore, a sample question containing each second sample description text can be constructed to instruct the multi-modal large model to determine the position of the image area in the input image that conforms to the second sample description text. And the matching result of each object annotation box in each sample image and each attribute used to combine to obtain the multiple second sample description texts corresponding to the sample image can be combined to construct a sample answer in the form of a chain of thought, and thus a Q&A sample pair in the form of a chain of thought can be obtained. The sample answer in the form of a chain of thought can include: the specific steps of reasoning to obtain the position of the image area in the sample image that matches each attribute in the second sample description text.
[0183] Correspondingly, training a multimodal large model based on a chain-of-thought form of Q&A sample pairs enables the multimodal large model to learn to reason in the chain-of-thought form, determine the position of the object region in the image that conforms to the input descriptive text, and generate an answer in the chain-of-thought form. Subsequently, using the trained multimodal large model, end-to-end descriptive object detection can be achieved without integrating multiple models, which can reduce the complexity of implementing descriptive object detection and improve the detection efficiency.
[0184] In one embodiment, the apparatus further includes: a preset category name acquisition module, configured to, for each sample image, before using the first large language model and the first text prompt to extract the category name of the object described by the first sample description text corresponding to the sample image and multiple attributes of the object, and combining the obtained category name with at least one of the obtained attributes to obtain multiple second sample description texts corresponding to the sample image, obtain the category name of the object in the object annotation box of the specified category in the sample image as the preset category name;
[0185] The description text generation module 702 is specifically configured to, for each sample image, input the first text prompt, the first sample description text corresponding to the sample image, and the preset category name into the first large language model, extract the category name of the object described by the first sample description text corresponding to the sample image and multiple attributes of the object, and combine the obtained category name with at least one of the obtained attributes to obtain multiple second sample description texts corresponding to the sample image; wherein, the obtained category name is the same as the category represented by the preset category name.
[0186] In one embodiment, the sample acquisition module 701 is specifically configured to acquire multiple sample images and the positions of the object annotation boxes of the specified category in each sample image; for each sample image, input the sample image and the second text prompt into the second large language model to obtain the first sample description text of the object annotation box of the specified category in the sample image; wherein, the second text prompt is used to indicate: generate the description text of the object annotation box in the input image.
[0187] In one embodiment, the description text generation module 702 is specifically configured to, for each sample image, input the first sample description text and the first text prompt corresponding to the sample image into the first large language model, and extract the class name of the object described by the first sample description text corresponding to the sample image and multiple attributes of the object; input the extracted class name, the extracted attributes, and the third text prompt into the first large language model to obtain multiple second sample description texts corresponding to the sample image; wherein, the third text prompt is used to indicate: combining the input class name with at least one of the input attributes to obtain multiple description texts.
[0188] In one embodiment, the matching result determination module 703 is specifically configured to use the image-text matching model and the fourth text prompt to respectively detect whether the object bounding boxes of each specified class in the sample image match each text to be matched; wherein, each text to be matched is obtained by respectively combining each attribute used to combine to obtain multiple second sample description texts corresponding to the sample image with the extracted class name; the fourth text prompt is used to indicate: determining whether each object bounding box in the input image and the input description text match.
[0189] An embodiment of the present application further provides an object detection device. Refer to Figure 8 , Figure 8 which is a schematic structural diagram of an object detection device provided by an embodiment of the present application. The device includes:
[0190] A data acquisition module 801, configured to acquire an image to be detected and a text prompt to be utilized including a description text to be detected; wherein, the text prompt to be utilized is used to indicate: determining the position of the image area occupied by the object in the image to be detected that conforms to the description text to be detected.
[0191] A detection result determination module 802, configured to input the image to be detected and the text prompt to be utilized into a pre-trained multimodal large model to obtain a detection result in the form of a chain of thought; wherein, the multimodal large model is trained by using the multimodal large model training method described in any one of the above; the detection result includes: a second inference process text for describing the inference process of the multimodal large model; the second inference process text includes: each inference step and the second execution order between each inference step; according to the second execution order, each inference step is respectively: extracting the class name of the object described by the input text prompt and the attributes of the object, detecting the position of the image area occupied by all objects belonging to the class represented by the extracted class name in the input image, respectively determining the matching result between each image area and each extracted attribute, and determining the position of the image area that matches each extracted attribute.
[0192] Based on the target detection device provided by the embodiments of the present application, during the training process, the multi-modal large model can learn the ability to reason in the form of a chain of thought, determine the position of the image area occupied by the object that conforms to the input description text in the image, and generate an answer in the form of a chain of thought. By using the trained multi-modal large model, end-to-end descriptive target detection can be achieved without integrating multiple models, which can reduce the complexity of implementing descriptive target detection and improve the detection efficiency.
[0193] The embodiments of the present application also provide an electronic device, as Figure 9 shown, including:
[0194] A memory 901 for storing a computer program;
[0195] A processor 902 for implementing any of the above multi-modal large model training methods or target detection methods when executing the program stored on the memory 901.
[0196] And the above electronic device may further include a communication bus and / or a communication interface, and the processor 902, the communication interface, and the memory 901 complete communication with each other through the communication bus.
[0197] The communication bus mentioned in the above electronic device may be a Peripheral Component Interconnect (PCI) bus or an Extended Industry Standard Architecture (EISA) bus, etc. This communication bus can be divided into an address bus, a data bus, a control bus, etc. For the sake of representation, only a thick line is used in the figure, but it does not mean that there is only one bus or one type of bus.
[0198] The communication interface is used for communication between the above electronic device and other devices.
[0199] The memory may include a Random Access Memory (RAM), or may also include a Non-Volatile Memory (NVM), such as at least one disk memory. Optionally, the memory may also be at least one storage device located far from the aforementioned processor.
[0200] The above-mentioned processor may be a general-purpose processor, including a Central Processing Unit (CPU), a Network Processor (NP), etc.; it may also be a Digital Signal Processor (DSP), an Application Specific Integrated Circuit (ASIC), a Field-Programmable Gate Array (FPGA), or other programmable logic devices, discrete gate or transistor logic devices, discrete hardware components.
[0201] In another embodiment provided by this application, a computer-readable storage medium is further provided. A computer program is stored in the computer-readable storage medium, and when the computer program is executed by a processor, the steps of any of the above-mentioned multi-modal large model training methods or object detection methods are implemented.
[0202] In another embodiment provided by this application, a computer program product containing instructions is further provided. When it runs on a computer, it causes the computer to execute any of the multi-modal large model training methods or object detection methods in the above embodiments.
[0203] In the above embodiments, it can be implemented in whole or in part by software, hardware, firmware, or any combination thereof. When implemented using software, it can be implemented in whole or in part in the form of a computer program product. The computer program product includes one or more computer instructions. When the computer program instructions are loaded and executed on a computer, the processes or functions described in the embodiments of this application are generated in whole or in part. The computer may be a general-purpose computer, a special-purpose computer, a computer network, or other programmable devices. The computer instructions may be stored in a computer-readable storage medium, or transmitted from one computer-readable storage medium to another computer-readable storage medium. For example, the computer instructions may be transmitted from one website, computer, server, or data center to another website, computer, server, or data center by wire (such as coaxial cable, optical fiber, Digital Subscriber Line (DSL)) or wirelessly (such as infrared, wireless, microwave, etc.). The computer-readable storage medium may be any available medium that can be accessed by a computer, or a data storage device such as a server or data center that includes one or more integrated available media. The available medium may be a magnetic medium (for example, a floppy disk, a hard disk, a magnetic tape), an optical medium (for example, a DVD), or a Solid State Disk (SSD), etc.
[0204] It should be noted that in this text, relational terms such as first and second are only used to distinguish one entity or operation from another entity or operation, and do not necessarily require or imply any actual relationship or order between these entities or operations. Moreover, the term "comprising", "including" or any other variant thereof is intended to cover non-exclusive inclusion, so that a process, method, article or device comprising a series of elements not only includes those elements, but also includes other elements not expressly listed, or also includes elements inherent to such process, method, article or device. Without further limitation, an element defined by the statement "comprising an..." does not exclude the presence of additional identical elements in the process, method, article or device comprising the said element.
[0205] Each embodiment in this specification is described in a related manner. For the same or similar parts among the embodiments, reference can be made to each other. Each embodiment focuses on the differences from other embodiments. In particular, for the embodiments of the object detection method, device, electronic device, computer-readable storage medium and computer program product, since they are basically similar to the embodiments of the multi-modal large model training method, the description is relatively simple, and the relevant parts can refer to the partial description of the method embodiments.
[0206] The above are only the preferred embodiments of the present application and are not intended to limit the protection scope of the present application. Any modifications, equivalent replacements, improvements, etc. made within the spirit and principle of the present application are included in the protection scope of the present application.
Claims
1. A multimodal large model training method, characterized in that: The method comprises: Acquire a plurality of sample images and a first sample description text of an object annotation box of a specified category in each sample image; For each sample image, the first large language model and the first text prompt are used to extract the category name of the object described in the first sample description text corresponding to the sample image and multiple attributes of the object, and the obtained category name is combined with at least one of the obtained attributes to obtain multiple second sample description texts corresponding to the sample image; wherein the first text prompt is used to indicate: extract the category name of the object described in the input description text and the attributes of the object; For each object annotation box in the sample image, determining whether the object annotation box matches each attribute of a plurality of second sample description texts corresponding to the sample image that are combined; Construct a sample question containing each second sample description text, and use the matching result of each attribute contained in the second sample description text and each object annotation box to construct a sample answer in the form of a thinking chain corresponding to the sample question, so as to obtain a question-answer sample pair; wherein, the sample question is used to indicate the position of the image area in the input image that matches the second sample description text; the sample answer contains a first reasoning process text for describing the multimodal large model reasoning process; the first reasoning process text includes: each reasoning step and a first execution order between each reasoning step; according to the first execution order, each reasoning step is respectively: extracting the category name and the attribute of the object described by the description text contained in the input question, detecting the position of the image area occupied by all objects belonging to the category represented by the extracted category name in the input image, respectively determining the matching result of each image area and each extracted attribute, and determining the position of the image area that matches each extracted attribute; Input the sample image and the sample question in each question-answer sample pair into the multimodal large model of the initial structure to obtain the predicted answer; Based on the difference between the predicted answer and the sample answer in the corresponding question-answer sample pair, the parameters of the multimodal large model of the initial structure are adjusted until the preset convergence condition is reached, thereby obtaining a trained multimodal large model; The step of obtaining a plurality of sample images and a first sample description text of an object annotation box of a specified category in each sample image comprises: Obtain multiple sample images and the position of the object annotation box of the specified category in each sample image; For each sample image, the sample image and the second text prompt are input into the second largest language model to obtain a first sample description text of an object annotation box of a specified category in the sample image; wherein the second text prompt is used to indicate: generate a description text of the object annotation box in the input image.
2. The method according to claim 1, characterized in that Before extracting, for each sample image, the category name of the object described in the first sample description text corresponding to the sample image and a plurality of attributes of the object using the first large language model and the first text prompt, and combining the obtained category name with at least one of the obtained attributes to obtain a plurality of second sample description texts corresponding to the sample image, the method further includes: For each sample image, obtaining a category name of an object in an object annotation box of a specified category in the sample image as a preset category name; The method of extracting the category name of the object described in the first sample description text corresponding to the sample image and multiple attributes of the object by using the first large language model and the first text prompt for each sample image, and combining the obtained category name with at least one of the obtained attributes to obtain multiple second sample description texts corresponding to the sample image includes: For each sample image, the first text prompt, the first sample description text corresponding to the sample image and the preset category name are input into the first large language model to extract the category name of the object described by the first sample description text corresponding to the sample image and multiple attributes of the object, and the obtained category name is combined with at least one of the obtained attributes to obtain multiple second sample description texts corresponding to the sample image; wherein the obtained category name is consistent with the category represented by the preset category name.
3. The method according to claim 1, characterized in that The method of extracting the category name of the object described in the first sample description text corresponding to the sample image and multiple attributes of the object by using the first large language model and the first text prompt for each sample image, and combining the obtained category name with at least one of the obtained attributes to obtain multiple second sample description texts corresponding to the sample image includes: For each sample image, input the first sample description text and the first text prompt corresponding to the sample image into the first large language model, and extract the category name of the object described by the first sample description text corresponding to the sample image and multiple attributes of the object; The extracted category name, the extracted attributes, and the third text prompt are input into the first language model to obtain multiple second sample description texts corresponding to the sample image; wherein the third text prompt is used to indicate: combine the input category name with at least one of the input attributes to obtain multiple description texts.
4. The method according to claim 1, characterized in that: The step of determining, for each object annotation frame in the sample image, whether the object annotation frame matches each attribute of a plurality of second sample description texts used for combining to obtain the corresponding sample image, includes: Using the image-text matching model and the fourth text prompt, it is respectively detected whether the object annotation box of each specified category in the sample image matches each description text to be matched; wherein, each description text to be matched is: each attribute of multiple second sample description texts corresponding to the sample image is combined with the extracted category name; the fourth text prompt is used to indicate: determine whether each object annotation box in the input image matches the input description text.
5. A target detection method, characterized in that: The method comprises: Acquire an image to be detected and a text prompt to be used containing a description text to be detected; wherein the text prompt to be used is used to indicate: determine the position of the image area occupied by the object that meets the description text to be detected in the image to be detected; The image to be detected and the multimodal large model to be pre-trained using the text prompt are input to obtain a detection result in the form of a thought chain; wherein the multimodal large model is trained based on the method described in any one of claims 1-4; the detection result includes: a second reasoning process text for describing the reasoning process of the multimodal large model; the second reasoning process text includes: each reasoning step and a second execution order between each reasoning step; according to the second execution order, each reasoning step is respectively: extracting the category name of the object described by the input text prompt and the attributes of the object, detecting the position of the image area occupied by all objects belonging to the category represented by the extracted category name in the input image, respectively determining the matching result of each image area with each extracted attribute, and determining the position of the image area that matches each extracted attribute.
6. A multimodal large model training device, characterized in that: The device comprises: A sample acquisition module, used to acquire a plurality of sample images and a first sample description text of an object annotation box of a specified category in each sample image; The description text generation module is used to extract the category name of the object described by the first sample description text corresponding to the sample image and multiple attributes of the object by using the first large language model and the first text prompt for each sample image, and combine the obtained category name with at least one of the obtained attributes to obtain multiple second sample description texts corresponding to the sample image; wherein the first text prompt is used to indicate: extract the category name of the object described by the input description text and the attributes of the object; a matching result determination module, for determining, for each object annotation box in the sample image, whether the object annotation box matches each attribute of a plurality of second sample description texts corresponding to the sample image; A question-answer sample pair construction module is used to construct a sample question containing each second sample description text, and to use the matching result between each attribute contained in the second sample description text and each object annotation box to construct a sample answer in the form of a thinking chain corresponding to the sample question, so as to obtain a question-answer sample pair; wherein, the sample question is used to indicate the position of the image area in the input image that matches the second sample description text; the sample answer contains a first reasoning process text for describing the multimodal large model reasoning process; the first reasoning process text includes: each reasoning step and a first execution order between each reasoning step; according to the first execution order, each reasoning step is respectively: extracting the category name and the attribute of the object described by the description text contained in the input question, detecting the position of the image area occupied by all objects belonging to the category represented by the extracted category name in the input image, respectively determining the matching result of each image area with each extracted attribute, and determining the position of the image area that matches each extracted attribute; A predicted answer determination module is used to input the sample image and the sample question in each question-answer sample pair into the multimodal macro model of the initial structure to obtain a predicted answer; A model parameter adjustment module is used to adjust the parameters of the multimodal large model of the initial structure based on the difference between the obtained predicted answer and the sample answer in the corresponding question and answer sample pair until the preset convergence condition is reached to obtain the trained multimodal large model; The sample acquisition module is specifically used to obtain multiple sample images and the position of the object annotation box of the specified category in each sample image; for each sample image, the sample image and the second text prompt are input into the second largest language model to obtain the first sample description text of the object annotation box of the specified category in the sample image; wherein the second text prompt is used to indicate: generate a description text of the object annotation box in the input image.
7. The device according to claim 6, characterized in that The device also includes: a preset category name acquisition module, for extracting the category name of the object described in the first sample description text corresponding to the sample image and a plurality of attributes of the object by using the first large language model and the first text prompt for each sample image, and combining the obtained category name with at least one of the obtained attributes to obtain a plurality of second sample description texts corresponding to the sample image, and obtaining, for each sample image, the category name of the object in the object annotation box of the specified category in the sample image as the preset category name; The description text generation module is specifically used to input the first text prompt, the first sample description text corresponding to the sample image and the preset category name into the first large language model for each sample image, extract the category name of the object described by the first sample description text corresponding to the sample image and multiple attributes of the object, and combine the obtained category name with at least one of the obtained attributes to obtain multiple second sample description texts corresponding to the sample image; wherein the obtained category name is consistent with the category represented by the preset category name; and / or, The description text generation module is specifically used to input, for each sample image, a first sample description text and a first text prompt corresponding to the sample image into a first large language model, extract the category name of the object described by the first sample description text corresponding to the sample image and multiple attributes of the object; input the extracted category name, the extracted attributes, and a third text prompt into the first large language model to obtain multiple second sample description texts corresponding to the sample image; wherein the third text prompt is used to indicate: combining the input category name with at least one of the input attributes to obtain multiple description texts; and / or, The matching result determination module is specifically used to use the image-text matching model and the fourth text prompt to respectively detect whether the object annotation box of each specified category in the sample image matches each description text to be matched; wherein the description text to be matched is: each attribute of the multiple second sample description texts corresponding to the sample image is combined with the extracted category name; the fourth text prompt is used to indicate: determine whether each object annotation box in the input image matches the input description text.
8. A target detection device, characterized in that: The device comprises: A data acquisition module is used to acquire an image to be detected and a text prompt to be used containing a description text to be detected; wherein the text prompt to be used is used to indicate: determining the position of an image area occupied by an object in the image to be detected that matches the description text to be detected; A detection result determination module is used to input the image to be detected and the text prompt to be used into a pre-trained multimodal large model to obtain a detection result in the form of a thought chain; wherein the multimodal large model is trained based on the method described in any one of claims 1-4; the detection result includes: a second reasoning process text for describing the reasoning process of the multimodal large model; the second reasoning process text includes: each reasoning step and a second execution order between each reasoning step; according to the second execution order, each reasoning step is respectively: extracting the category name of the object described by the input text prompt and the attributes of the object, detecting the position of the image area occupied by all objects belonging to the category represented by the extracted category name in the input image, respectively determining the matching result of each image area with each extracted attribute, and determining the position of the image area that matches each extracted attribute.
9. An electronic device, characterized in that: include: Memory, used to store computer programs; A processor, for implementing the method described in any one of claims 1 to 4 or the method described in claim 5 when executing a program stored in a memory.
Citation Information
Patent Citations
Visual question and answer method, system and device based on thinking chain and storage medium
CN117891965A
Image detailed description method based on large model fusion refined scene graph thinking chain
CN118865388A