Registration understanding data generation method and device, storage medium and electronic equipment
Generate text descriptions through multimodal large language model and combine them with the computed and interchangeable model. This solves the problem of time-consuming and cost-consuming construction of data sets, and realizes efficient automated generation and data quality improvement, which promotes the application of professional fields.
Patent Information
- Application Number
- CN202510900118.8
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-07-01
- Publication Date
- 2025-08-01
- Estimated Expiration
- Not applicable · inactive patent
AI Technical Summary
In the prior art, it is time-consuming and costly to construct data sets, and it is difficult to widely use in professional fields. The traditional image enhancement method is not suitable for joint graphics and text tasks, and the data generated by a single pre-trained model is poor.
The target text description is generated through a multimodal large language model, and combined with the reference expression understanding model to calculate the interchange ratio, select the target annotation box with the highest interchange ratio as the bounding box, and generate the reference expression understanding data.
It realizes efficient and automated generation of data, reduces labeling costs, improves data quality and construction efficiency, and promotes the application of vertical industries.
Smart Images

Figure CN120411976A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of data processing, and particularly to a method, apparatus, storage medium, and electronic device for generating referring expression understanding data. Background Art
[0002] Referring Expression Comprehension (REC), also known as visual grounding, is a technology that combines image content with natural language descriptions to locate specific targets in an image through text information. Compared with traditional object detection that only relies on image features, referring expression comprehension integrates rich text semantic information and can more accurately identify the described targets. The data required for the referring expression comprehension task usually includes images, target positions (bounding boxes or segmentation masks), and corresponding natural language descriptions. The application of public datasets has promoted the development of referring expression comprehension technology.
[0003] However, the construction of datasets for professional scenarios still faces huge challenges. Data collection and annotation are complex and time-consuming, which limits the wide application of referring expression comprehension technology in professional fields.
[0004] Therefore, how to efficiently generate referring expression comprehension data has become a technical problem that needs to be solved urgently by those skilled in the art. Summary of the Invention
[0005] In view of the above problems, the present invention provides a method, apparatus, storage medium, and electronic device for generating referring expression understanding data that overcome the above problems or at least partially solve the above problems. The technical solutions are as follows:
[0006] A method for generating referring expression understanding data includes:
[0007] Obtain a target detection image, wherein at least one target annotation box is marked in the target detection image;
[0008] Input the target detection image into a multi-modal large language model, and trigger the multi-modal large language model to generate a text description of a specified category target in the target detection image through a preset query statement, and obtain the target text description output by the multi-modal large language model;
[0009] Input the target text description and the target detection image into a referring expression understanding model, so that the referring expression understanding model generates an initial bounding box for the specified category target in the target detection image according to the target text description;
[0010] Calculate the intersection over union of the initial bounding box and each target annotation box in the target detection image;
[0011] Select the target annotation box with the highest intersection over union (IoU) as the target bounding box of the specified category object in the target detection image;
[0012] Combine the target bounding box and the target text description to generate the anaphora understanding data related to the target detection image.
[0013] Optionally, before selecting the target annotation box with the highest intersection over union (IoU) as the target bounding box of the specified category object in the target detection image, the method further includes:
[0014] Verify whether the highest intersection over union (IoU) is greater than or equal to a preset intersection over union (IoU) threshold. If not, discard the initial bounding box. If so, perform the step of selecting the target annotation box with the highest intersection over union (IoU) as the target bounding box of the specified category object in the target detection image.
[0015] Optionally, the step of selecting the target annotation box with the highest intersection over union (IoU) as the target bounding box of the specified category object in the target detection image includes:
[0016] In the case where there are two or more target annotation boxes with the same highest intersection over union (IoU), calculate the center point distances between the initial bounding box and each of the target annotation boxes with the same highest intersection over union (IoU), and select the target annotation box with the smallest center point distance as the target bounding box of the specified category object in the target detection image.
[0017] Optionally, the step of selecting the target annotation box with the highest intersection over union (IoU) as the target bounding box of the specified category object in the target detection image includes:
[0018] In the case where there are two or more target annotation boxes with the same highest intersection over union (IoU), calculate the aspect ratio differences between the initial bounding box and each of the target annotation boxes with the same highest intersection over union (IoU), and select the target annotation box with the smallest aspect ratio difference as the target bounding box of the specified category object in the target detection image.
[0019] Optionally, before inputting the target text description and the target detection image into the anaphora understanding model, the method further includes:
[0020] Obtain an open-source anaphora understanding dataset, where the open-source anaphora understanding dataset includes multiple pieces of anaphora understanding data;
[0021] Based on a preset loss function and optimization strategy, use each piece of the anaphora understanding data in the open-source anaphora understanding dataset to train a pre-constructed anaphora understanding model to obtain the trained anaphora understanding model.
[0022] Optionally, after combining the target bounding box and the target text description to generate the referential expression understanding data related to the target detection image, the method further includes:
[0023] Combining the open-source referential expression understanding dataset and the referential expression understanding data related to the target detection image to retrain the referential expression understanding model, and obtaining the trained referential expression understanding model.
[0024] Optionally, the preset query statement is a prompt word used to guide the multi-modal large language model to generate a text description of the appearance features and / or behavior features of the specified category of targets in the target detection image.
[0025] A referential expression understanding data generation device includes: a target detection image acquisition unit, a target text description acquisition unit, an initial bounding box generation unit, an intersection over union calculation unit, a target bounding box selection unit, and a referential expression understanding data generation unit.
[0026] The target detection image acquisition unit is configured to acquire a target detection image, where at least one target annotation box is marked in the target detection image.
[0027] The target text description acquisition unit is configured to input the target detection image into a multi-modal large language model, and trigger the multi-modal large language model to generate a text description of the specified category of targets in the target detection image through a preset query statement, and obtain the target text description output by the multi-modal large language model.
[0028] The initial bounding box generation unit is configured to input the target text description and the target detection image into a referential expression understanding model, so that the referential expression understanding model generates an initial bounding box for the specified category of targets in the target detection image according to the target text description.
[0029] The intersection over union calculation unit is configured to calculate the intersection over union of the initial bounding box and each target annotation box in the target detection image.
[0030] The target bounding box selection unit is configured to select the target annotation box with the highest intersection over union as the target bounding box of the specified category of targets in the target detection image.
[0031] The referential expression understanding data generation unit is configured to combine the target bounding box and the target text description to generate the referential expression understanding data related to the target detection image.
[0032] A computer-readable storage medium stores a program, and when the program is executed by a processor, the referential expression understanding data generation method is implemented.
[0033] An electronic device, the electronic device includes at least one processor, at least one memory connected to the processor, and a bus; wherein, the processor and the memory complete communication with each other through the bus; the processor is used to call program instructions in the memory to execute the above-mentioned referential expression understanding data generation method.
[0034] By means of the above technical solution, the referential expression understanding data generation method, device, storage medium and electronic device provided by the present invention obtain a target detection image, wherein at least one target annotation box is marked in the target detection image; input the target detection image into a multi-modal large language model, and trigger the multi-modal large language model to generate a text description of the specified category target in the target detection image through a preset query statement, and obtain the target text description output by the multi-modal large language model; input the target text description and the target detection image into the referential expression understanding model, so that the referential expression understanding model generates an initial bounding box for the specified category target in the target detection image according to the target text description; calculate the intersection over union of the initial bounding box and each target annotation box in the target detection image; select the target annotation box with the highest intersection over union as the target bounding box of the specified category target in the target detection image; combine the target bounding box and the target text description to generate referential expression understanding data related to the target detection image. The present invention automatically generates a target text description through a multi-modal large language model, and combines the target detection image with the referential expression understanding model to accurately match the bounding box, realizing the efficient and automatic generation of referential expression understanding data.
[0035] The above description is only an overview of the technical solution of the present invention. In order to be able to understand the technical means of the present invention more clearly, it can be implemented according to the content of the specification. And in order to make the above and other purposes, features and advantages of the present invention more obvious and understandable, the specific embodiments of the present invention are specifically given below. Brief Description of the Drawings
[0036] By reading the detailed description of the preferred embodiments below, various other advantages and benefits will become clear to those of ordinary skill in the art. The drawings are only for the purpose of showing the preferred embodiments and are not considered to be a limitation of the present invention. And throughout the drawings, the same reference numerals are used to represent the same components. In the drawings:
[0037] Figure 1 Shows a schematic flow chart of an implementation manner of the referential expression understanding data generation method provided by an embodiment of the present invention;
[0038] Figure 2 Shows a schematic structural diagram of the referential expression understanding data generation device provided by an embodiment of the present invention;
[0039] Figure 3 The figure shows a schematic structural diagram of an electronic device provided by an embodiment of the present invention. Detailed implementation manners
[0040] Hereinafter, exemplary embodiments of the present invention will be described in more detail with reference to the accompanying drawings. Although the exemplary embodiments of the present invention are shown in the drawings, it should be understood that the present invention can be implemented in various forms and should not be limited by the embodiments set forth herein. On the contrary, these embodiments are provided so that the present invention can be more thoroughly understood and the scope of the present invention can be fully conveyed to those skilled in the art.
[0041] Referring Expression Comprehension (REC), also known as Visual Grounding, is a technology that combines images and natural language descriptions to locate specific objects or regions in images through text information. Different from traditional object detection that only relies on image features, visual grounding integrates text information and can handle more diverse and fine-grained tasks, such as accurately locating a target by describing "a person wearing red clothes". The training and testing of the referring expression comprehension task rely on annotated data with target positions and corresponding natural language descriptions. Currently, commonly used public datasets include RefCOCO (Referring Expressions for COCO, object localization in reference images, a dataset for locating objects in images through language descriptions), RefCOCO+ (an improved dataset for object localization in reference images, containing more descriptions and more complex language expressions, aiming to improve the model's understanding of objects), RefCOCOg (Referring Expressions for COCO with Gender, object localization in reference images, an extended version of RefCOCO and RefCOCO+, focusing on the expression of gender information and providing richer language descriptions), ReferItGame (a game-based reference object localization dataset where participants locate objects in images through language descriptions, aiming to study how humans perform visual grounding through language), and Flickr30K Entities (Flickr30K entity dataset, containing 30,000 images from Flickr, providing annotations for different objects in images, suitable for tasks such as image understanding and visual question answering), etc. However, in the vertical industry field, the construction of dataset annotation for specific application scenarios is time-consuming and costly, becoming an important bottleneck restricting the application of referring expression comprehension technology.
[0042] In recent years, with the rapid development of large language models (LLMs) and multimodal large language models, the task of referring expression understanding has received extensive attention. However, the lack of a large amount of high-quality training data severely restricts the improvement of model performance. There are mainly three challenges: First, the cost of data annotation is high. Not only target bounding box annotation is required, but also detailed text description is needed, which greatly increases the workload. Second, traditional image enhancement methods are difficult to apply to the task of referring expression understanding that combines text and images. Enhancement operations such as random flipping will cause the description to not match the image target. Third, the data quality generated by a single pre-trained model is difficult to guarantee. The target description and bounding box are not precise enough, affecting the model training effect.
[0043] Based on this, in the embodiments of the present invention, a method for generating referring expression understanding data is provided. First, by obtaining a target detection image annotated with a target annotation box, inputting it into a multimodal large language model, and automatically generating a text description of a target of a specified category using a preset query statement. Subsequently, inputting the generated text description and the target detection image into a referring expression understanding model, generating an initial bounding box according to the text description, and calculating the intersection over union (IoU) between it and each target annotation box in the image, and selecting the annotation box with the highest IoU as the target bounding box. Finally, combining the target bounding box and the text description, generating referring expression understanding data corresponding to the target detection image. It can be seen that the present invention can effectively integrate the multimodal large language model and the referring expression understanding model, realizing the efficient and automated generation of referring expression understanding data, significantly improving the efficiency and accuracy of data construction, thereby reducing the annotation cost and improving the data quality, and promoting the development of referring expression understanding technology in vertical industry fields.
[0044] As Figure 1 shown, a schematic flowchart of an implementation manner of the method for generating referring expression understanding data provided by the embodiments of the present invention, the method may include:
[0045] S100. Obtain a target detection image, where at least one target annotation box is annotated in the target detection image.
[0046] Wherein, the target detection image refers to the input image for the target detection task. The target detection image may include one or more target objects to be detected. The target detection image includes at least one target annotation box for identifying the position and range of the target object in the image.
[0047] A target annotation box is a rectangular box that spatially locates a specific target object in a target detection image. It is typically represented by the coordinates of its upper left corner and width and height (or upper left and lower right corner coordinates). Each target annotation box corresponds to a specific target object in the image, indicating its location and size. A target detection image can contain multiple target annotation boxes to label multiple different target objects.
[0048] S110. Input the target detection image into the multimodal large language model, trigger the multimodal large language model through a preset query statement to generate a text description of the specified category target in the target detection image, and obtain the target text description output by the multimodal large language model.
[0049] A large multimodal language model refers to a large language model capable of understanding and processing data from multiple modalities (such as text, images, and audio). This model not only understands and generates natural language but also integrates non-textual information, such as visual information, to achieve image-text fusion reasoning and generation tasks. Optionally, the large multimodal language model provided in this embodiment of the present invention may be CogVLM2 (a second-generation large multimodal model designed to enhance visual and language comprehension capabilities).
[0050] A preset query statement refers to a pre-designed text input within a multimodal large language model that is used to direct the model to focus on specific content or perform a specific task. By providing a clear query statement to the multimodal large language model, the embodiments of the present invention trigger the model to analyze and generate textual descriptions of objects of specified categories within an input image.
[0051] Optionally, the preset query statement is a prompt word used to guide the multimodal large language model to generate a text description of the appearance and / or behavioral features of a specified category of objects in the object detection image. For example, the preset query statement may be "Describe a cat_name in the picture in a short sentence," where cat_name is the category name of an object in the image. For example, when the category name is person, the prompt word generated is "Describe a person in the picture in a short sentence."
[0052] A designated category refers to a specific category of target explicitly identified in the target detection image, such as "person," "car," or "animal." The multimodal large language model focuses on this category using a preset query statement, identifying relevant targets in the image and generating corresponding text descriptions.
[0053] Among them, text description generation refers to the process in which a multi-modal large language model, based on the input object detection image and text query, fuses visual information and language knowledge to generate a semantically rich and natural language-compliant text description of the specified object in the object detection image.
[0054] Among them, the target text description is the natural language text output by the multi-modal large language model for the specified category of objects, usually a sentence that concisely and accurately describes the features, behaviors, or states of the objects. For example: describing information such as the clothing and actions of a person.
[0055] Specifically, in the embodiment of the present invention, the object detection image can be input into the multi-modal large language model, and the preset query statement is used to guide the multi-modal large language model to focus on the specified category of objects in the object detection image, triggering the multi-modal large language model to call its multi-modal understanding ability to analyze and semantically extract the specified category of objects. Subsequently, the multi-modal large language model generates a text description for the specified category of objects based on the image visual information and language knowledge, and outputs the target text description, thereby realizing the effective conversion from visual signals to semantic texts, which helps to automatically generate high-quality text description data and is beneficial to improving the accuracy and reliability of subsequent generated referential expression understanding data.
[0056] For the sake of easy understanding, an example is given here: Suppose the selected object detection image contains multiple people, and the preset query statement "Describe a person in the picture with a short sentence" is input into the multi-modal large language model CogVLM2 (the second-generation multi-modal large model, aiming to improve the ability of visual and language understanding). CogVLM2 (the second-generation multi-modal large model, aiming to improve the ability of visual and language understanding) identifies and focuses on a target, and automatically generates the target text description: "A young student wearing a white and blue school uniform and headphones is using a computer", and this target text description reflects the visual features and scene states of the specified category of objects in the object detection image.
[0057] S120. Input the target text description and the object detection image into the referential expression understanding model, so that the referential expression understanding model generates an initial bounding box for the specified category of objects in the object detection image according to the target text description.
[0058] Among them, the referential expression understanding model refers to a pre-trained model that can understand and parse referential expressions (i.e., descriptive phrases or sentences in natural language used to uniquely specify a target or region in an image). The referential expression understanding model combines language information and image content to accurately locate the image region referred to by the text description, usually outputting in the form of generating corresponding bounding boxes. Optionally, the referential expression understanding model provided by the embodiments of the present invention can be CLIP-VG (Contrastive Language-Image Pre-training for Visual Generation, a model that combines image processing and natural language processing, aiming to effectively combine visual content with text information through contrastive learning).
[0059] Among them, the initial bounding box refers to the target localization box generated in the target detection image based on the referential expression understanding model according to the target text description. The initial bounding box is used to roughly identify the image region where the target referred to by the text description is located, providing a basis for subsequent fine-grained target localization.
[0060] Specifically, the embodiments of the present invention can input the target text description generated by the multi-modal large language model and the corresponding target detection image into the referential expression understanding model together, so that the referential expression understanding model can utilize its understanding ability of language referential expressions, combine image visual information, analyze the target features in the target text description and locate the relevant regions, complete the mapping from natural language description to image space localization, and then generate the initial bounding box describing the target, thereby providing a key localization basis for subsequent target detection tasks.
[0061] For ease of understanding, an example is given here: Suppose the target text description generated by the multi-modal large language model is "A young student wearing a white and blue school uniform and wearing headphones is using a computer", and this description and the corresponding target detection image are input into the referential expression understanding model together. The referential expression understanding model analyzes the features described in the text, combines the image information, and identifies the target region that meets the description in the target detection image, generating an initial bounding box covering the student.
[0062] S130. Calculate the intersection over union (IOU) between the initial bounding box and each target annotation box in the target detection image.
[0063] Among them, the intersection over union (IOU) is a metric for evaluating the overlapping degree between the initial bounding box and each target annotation box in the target detection image respectively. The intersection over union is the ratio of the intersection area of two boxes to their union area, and the value range is from 0 to 1. The closer the intersection over union is to 1, the more coincident the two boxes are, and the closer the intersection over union is to 0, the lower the coincidence degree of the two boxes.
[0064] Specifically, in the embodiments of the present invention, the intersection over union (IoU) between the initial bounding box and each target annotation box in the target detection image is calculated by comparing the spatial overlap degree between the initial bounding box generated by the referring expression understanding model and each target annotation box to evaluate their matching relationship. By statistically calculating the intersection area and union area of the two boxes, an IoU value is obtained, which is used to quantify the similarity between the initial bounding box and each target annotation box, providing a basis for subsequent matching and evaluation.
[0065] S140. Select the target annotation box with the highest IoU as the target bounding box of the specified category target in the target detection image.
[0066] Among them, the target bounding box refers to a rectangular box drawn around the image area of the specified category target in the target detection task. The target bounding box is used to accurately locate the position and range of the target in the image, facilitating subsequent operations such as recognition, classification, or tracking.
[0067] Specifically, in the embodiments of the present invention, the IoU between the initial bounding box and all target annotation boxes can be calculated to compare the matching degrees of different target annotation boxes with the initial bounding box, and the target annotation box with the highest IoU is selected as the final localization box of the specified category target, that is, the target bounding box, so as to ensure that the finally selected target bounding box is the closest to the initial bounding box generated by the referring expression understanding model in space, thereby accurately reflecting the position of the specified target in the image.
[0068] S150. Combine the target bounding box and the target text description to generate referring expression understanding data related to the target detection image.
[0069] Among them, the referring expression understanding data refers to a data set containing an image, a text description (referring expression) corresponding to a specific target in the image, and an accurate position annotation (target bounding box) of the target.
[0070] Specifically, in the embodiments of the present invention, the bounding box of the located target can be paired with the text description of the target to form a complete referring expression understanding data. This referring expression understanding data contains both visual information (image and target area) and language information (target description), thereby providing a basis for multi-modal training or inference for the referring expression understanding model, enabling the model to accurately understand and locate the target described in language. For example: Suppose in a campus scene image, the bounding box of the target "student wearing school uniform" has been selected. Combining the target text description "A young student wearing a headset and a white and blue school uniform is using a computer", pair the bounding box with the text description to generate a referring expression understanding data. This referring expression understanding data contains the complete image, text description, and the corresponding target bounding box, which can provide input for model training or inference.
[0071] The method for generating reference expression understanding data provided by the present invention includes: obtaining a target detection image, wherein at least one target annotation box is marked in the target detection image; inputting the target detection image into a multi-modal large language model, triggering the multi-modal large language model to generate a text description of a specified category target in the target detection image through a preset query statement, and obtaining the target text description output by the multi-modal large language model; inputting the target text description and the target detection image into a reference expression understanding model, so that the reference expression understanding model generates an initial bounding box for the specified category target in the target detection image according to the target text description; calculating the intersection over union (IoU) between the initial bounding box and each target annotation box in the target detection image; selecting the target annotation box with the highest IoU as the target bounding box of the specified category target in the target detection image; and generating reference expression understanding data related to the target detection image by combining the target bounding box and the target text description. The present invention automatically generates a target text description through a multi-modal large language model, and accurately matches the bounding box by combining the target detection image and the reference expression understanding model, realizing the efficient and automatic generation of reference expression understanding data.
[0072] Optionally, based on one or more corresponding embodiments above, in another optional embodiment provided by the embodiments of the present invention, before selecting the target annotation box with the highest IoU as the target bounding box of the specified category target in the target detection image, the method may further include: Figure 1
[0073] Checking whether the highest IoU is greater than or equal to a preset IoU threshold. If not, the initial bounding box is discarded. If so, the step of selecting the target annotation box with the highest IoU as the target bounding box of the specified category target in the target detection image is executed.
[0074] Wherein, the preset IoU threshold is a critical value for judging the matching degree between the initial bounding box and the target annotation box. The preset IoU threshold defines the minimum requirement for the IoU between the initial bounding box and the target annotation box. Only when the IoU of the initial bounding box reaches or exceeds this threshold is it considered a valid match, otherwise the initial bounding box will be discarded.
[0075] In practical applications, the embodiments of the present invention first calculate the maximum IoU between the initial bounding box and the target annotation box, and compare this maximum IoU with the preset IoU threshold. If the maximum IoU is lower than the preset IoU threshold, it means that the overlap degree between the initial bounding box and any target annotation box is insufficient and the positioning is inaccurate, so the initial bounding box is discarded. On the contrary, the target annotation box corresponding to this maximum IoU is selected as the final target bounding box. It can be understood that the level of the preset IoU threshold can be adjusted according to the requirements of the specific task for the positioning accuracy. If the positioning accuracy requirement is high, the threshold is set larger; otherwise, it is smaller.
[0076] In the embodiment of the present invention, by determining whether the highest intersection over union (IoU) reaches a preset IoU threshold before selecting the target bounding box, initial bounding boxes with a low overlap degree with the target annotation box can be effectively filtered out, avoiding the misuse of low-quality boxes. This not only improves the positioning accuracy and reliability of the target bounding box, but also enhances the stability and robustness of the entire multi-modal target detection and referential expression understanding system, ensuring that the generated referential expression understanding data has higher quality and practical value.
[0077] Optionally, in the embodiment of the present invention, when there are two or more target annotation boxes with the same highest IoU, the center point distance between the initial bounding box and each of the target annotation boxes with the same highest IoU can be calculated, and the target annotation box with the smallest center point distance is selected as the target bounding box of the specified category target in the target detection image.
[0078] Among them, the center point distance refers to the Euclidean distance between the center point of the initial bounding box and the center point of the target annotation box, which is used to measure the spatial proximity of these two boxes in the image. In the embodiment of the present invention, the center coordinates of the initial bounding box and each of the two target annotation boxes with the same highest IoU can be obtained respectively (for example: the average value of the upper left corner and lower right corner coordinates of the box), and then the straight-line distance between these two center points can be calculated. The smaller the center point distance, the closer the positions of the two bounding boxes are.
[0079] Specifically, in the embodiment of the present invention, when the IoU between two or more target annotation boxes and the initial bounding box is the same, the optimal match cannot be distinguished only by the IoU value. At this time, by calculating the distance between the center points of the initial bounding box and these target annotation boxes, the target annotation box with the smallest center point distance is selected as the final target bounding box to ensure that the selected target bounding box is closer to the initial bounding box in terms of spatial position, thereby improving the accuracy and rationality of target positioning.
[0080] Optionally, in the embodiment of the present invention, when there are two or more target annotation boxes with the same highest IoU, the aspect ratio difference between the initial bounding box and each of the target annotation boxes with the same highest IoU can be calculated, and the target annotation box with the smallest aspect ratio difference is selected as the target bounding box of the specified category target in the target detection image.
[0081] Among them, the aspect ratio difference refers to the difference between the aspect ratio of the initial bounding box and the aspect ratio of the target annotation box, which is used to measure the shape similarity between the initial bounding box and the target annotation box. The aspect ratio is the ratio of the width to the height of the bounding box, and the aspect ratio difference can be represented by calculating the absolute difference between the aspect ratios of the initial bounding box and the target annotation box. The smaller the difference, the more similar the shapes of the initial bounding box and the target annotation box are.
[0082] Specifically, when there are two or more target annotation boxes with the same intersection over union (IoU) with the initial bounding box in the embodiments of the present invention, the optimal match cannot be distinguished solely by the IoU value. At this time, by calculating the aspect ratio difference between the initial bounding box and these target annotation boxes, the target annotation box with the smallest aspect ratio difference is selected as the final target bounding box. This ensures that the selected target bounding box is closer to the initial bounding box in shape, thereby improving the accuracy and rationality of target positioning.
[0083] Optionally, based on the above Figure 1 corresponding one or more embodiments, in another optional embodiment provided by the embodiments of the present invention, before step S120, the method may further include:
[0084] Obtain an open-source reference expression understanding dataset, where the open-source reference expression understanding dataset includes multiple reference expression understanding data. Based on a preset loss function and an optimization strategy, use each reference expression understanding data in the open-source reference expression understanding dataset to train a pre-constructed reference expression understanding model to obtain a trained reference expression understanding model.
[0085] Among them, the open-source reference expression understanding dataset refers to a publicly released dataset that can be freely obtained and used and contains multiple reference expression understanding data. Among them, each reference expression understanding data includes an image, a corresponding target region (bounding box), and a natural language expression (reference expression) describing the target.
[0086] Among them, the preset loss function refers to a function predefined during model training for measuring the error between the model's predicted output and the true annotation. The loss function guides the update of model parameters by quantifying the difference between the prediction result and the target, thereby achieving the improvement of model performance. The preset loss function may include cross-entropy loss or mean squared error.
[0087] Among them, the optimization strategy refers to an algorithm used to adjust model parameters to minimize the loss function during model training. The optimization strategy provided by the embodiments of the present invention may include selecting a specific optimization algorithm (such as SGD (Stochastic Gradient Descent), Adam (Adaptive Moment Estimation), etc.), setting the learning rate, batch size, and other hyperparameters.
[0088] Embodiments of the present invention can obtain an open-source dataset containing multiple pieces of referential expression understanding data. Using this referential expression understanding data as training samples, combined with a pre-set loss function, by measuring the error between the model prediction and the true annotation, the referential expression understanding model is guided to continuously optimize. With a reasonable optimization strategy, the parameters of the referential expression understanding model are gradually adjusted, and finally a trained referential expression understanding model is obtained, enabling it to accurately understand and locate the specified target in the referential expression, and improving the performance and practical value of the multi-modal expression understanding task.
[0089] Optionally, embodiments of the present invention can also, after generating referential expression understanding data related to the object detection image by combining the target bounding box and the target text description, retrain the referential expression understanding model by combining the open-source referential expression understanding dataset and the referential expression understanding data related to the object detection image, to obtain a trained referential expression understanding model.
[0090] Embodiments of the present invention retrain the referential expression understanding model by combining the referential expression understanding data generated from the object detection image and the open-source referential expression understanding dataset, which can effectively expand the diversity and coverage of the training data, and enhance the adaptability of the model to different scenarios and expression methods. By fusing the true annotation and the automatically generated data, the generalization performance and understanding accuracy of the model are improved, thereby further enhancing the ability to locate and understand the text of the specified category target, and achieving a more accurate and robust referential expression understanding effect.
[0091] Although the operations are depicted in a particular order, this should not be construed as requiring that the operations be performed in the particular order shown or in sequential order. In certain circumstances, multitasking and parallel processing may be advantageous.
[0092] It should be understood that the various steps recited in the method embodiments of the present invention can be executed in different orders and / or in parallel. In addition, the method embodiments may include additional steps and / or omit the steps shown. The scope of the present invention is not limited in this regard.
[0093] Corresponding to the above method embodiments, embodiments of the present invention also provide a device for generating referential expression understanding data, the structure of which is as Figure 2 shown, and may include: an object detection image acquisition unit 10, a target text description acquisition unit 20, an initial bounding box generation unit 30, an intersection over union calculation unit 40, a target bounding box selection unit 50, and a referential expression understanding data generation unit 60.
[0094] The object detection image acquisition unit 10 is configured to acquire an object detection image, wherein at least one target annotation box is annotated in the object detection image.
[0095] The target text description obtaining unit 20 is configured to input the target detection image into the multi-modal large language model, trigger the multi-modal large language model to generate the text description of the specified category target in the target detection image through the preset query statement, and obtain the target text description output by the multi-modal large language model.
[0096] The initial bounding box generating unit 30 is configured to input the target text description and the target detection image into the referential expression understanding model, so that the referential expression understanding model generates an initial bounding box for the specified category target in the target detection image according to the target text description.
[0097] The intersection over union calculating unit 40 is configured to calculate the intersection over union of the initial bounding box and each target annotation box in the target detection image.
[0098] The target bounding box selecting unit 50 is configured to select the target annotation box with the highest intersection over union as the target bounding box of the specified category target in the target detection image.
[0099] The referential expression understanding data generating unit 60 is configured to generate referential expression understanding data related to the target detection image by combining the target bounding box and the target text description.
[0100] Optionally, the referential expression understanding data generating device may further include: an initial bounding box verification unit.
[0101] The initial bounding box verification unit is configured to verify whether the highest intersection over union is greater than or equal to the preset intersection over union threshold before the target bounding box selecting unit 50 selects the target annotation box with the highest intersection over union as the target bounding box of the specified category target in the target detection image. If not, the initial bounding box is discarded. If so, the target bounding box selecting unit 50 is triggered.
[0102] Optionally, the target bounding box selecting unit 50 may be specifically configured to calculate the center point distance between the initial bounding box and each target annotation box with the same highest intersection over union in the case where there are two or more target annotation boxes with the same highest intersection over union, and select the target annotation box with the smallest center point distance as the target bounding box of the specified category target in the target detection image.
[0103] Optionally, the target bounding box selecting unit 50 may be specifically configured to calculate the aspect ratio difference between the initial bounding box and each target annotation box with the same highest intersection over union in the case where there are two or more target annotation boxes with the same highest intersection over union, and select the target annotation box with the smallest aspect ratio difference as the target bounding box of the specified category target in the target detection image.
[0104] Optionally, the referential expression understanding data generating device may further include: a referential expression understanding model training unit.
[0105] The referring expression understanding model training unit is used to obtain an open-source referring expression understanding dataset before the initial bounding box generation unit 30 inputs the target text description and the target detection image into the referring expression understanding model. The open-source referring expression understanding dataset includes multiple referring expression understanding data. Based on a preset loss function and optimization strategy, each piece of referring expression understanding data in the open-source referring expression understanding dataset is used to train a pre-constructed referring expression understanding model to obtain a trained referring expression understanding model.
[0106] Optionally, the referring expression understanding data generation device may further include: a referring expression understanding model retraining unit.
[0107] The referring expression understanding model retraining unit is used to, after the referring expression understanding data generation unit 60 combines the target bounding box and the target text description to generate referring expression understanding data related to the target detection image, combine the open-source referring expression understanding dataset and the referring expression understanding data related to the target detection image to retrain the referring expression understanding model to obtain a trained referring expression understanding model.
[0108] Optionally, the preset query statement is a prompt word used to guide the multi-modal large language model to generate a text description of the appearance features and / or behavior features of the specified category target in the target detection image.
[0109] The referring expression understanding data generation device provided by the present invention is used to: obtain a target detection image, where at least one target annotation box is marked in the target detection image; input the target detection image into a multi-modal large language model, trigger the multi-modal large language model to generate a text description of the specified category target in the target detection image through a preset query statement, and obtain the target text description output by the multi-modal large language model; input the target text description and the target detection image into the referring expression understanding model, so that the referring expression understanding model generates an initial bounding box for the specified category target in the target detection image according to the target text description; calculate the intersection over union of the initial bounding box and each target annotation box in the target detection image; select the target annotation box with the highest intersection over union as the target bounding box of the specified category target in the target detection image; combine the target bounding box and the target text description to generate referring expression understanding data related to the target detection image. The present invention automatically generates a target text description through a multi-modal large language model and accurately matches the bounding box by combining the target detection image and the referring expression understanding model, realizing the efficient and automatic generation of referring expression understanding data.
[0110] Regarding the device in the above embodiments, the specific manner in which each unit performs operations has been described in detail in the embodiments related to the method, and will not be elaborated here.
[0111] The above-mentioned referential expression understanding data generation device includes a processor and a memory. The above-mentioned target detection image acquisition unit 10, target text description acquisition unit 20, initial bounding box generation unit 30, intersection over union calculation unit 40, target bounding box selection unit 50, and referential expression understanding data generation unit 60, etc. are all stored in the memory as program units, and the processor executes the above program units stored in the memory to implement corresponding functions.
[0112] The processor contains a kernel, and the kernel retrieves the corresponding program units from the memory. One or more kernels can be set. By adjusting the kernel parameters, the multimodal large language model and the referential expression understanding model are integrated, realizing the efficient and automated generation of referential expression understanding data, and significantly improving the efficiency and accuracy of data construction.
[0113] An embodiment of the present invention provides a computer-readable storage medium, on which a program is stored, and when the program is executed by a processor, it implements the above-mentioned referential expression understanding data generation method.
[0114] An embodiment of the present invention provides a processor, and the processor is used to run a program, wherein when the program runs, it executes the above-mentioned referential expression understanding data generation method.
[0115] As Figure 3 shown, an embodiment of the present invention provides an electronic device 1000. The electronic device 1000 includes at least one processor 1001, at least one memory 1002 connected to the processor 1001, and a bus 1003; wherein, the processor 1001 and the memory 1002 complete mutual communication through the bus 1003; the processor 1001 is used to call program instructions in the memory 1002 to execute the above-mentioned referential expression understanding data generation method. The electronic device herein can be a server, a PC, a PAD, a mobile phone, etc.
[0116] The present invention also provides a computer program product, which is suitable for executing a program initialized with the steps of the referential expression understanding data generation method when executed on an electronic device.
[0117] The present invention is described with reference to the flowcharts and / or block diagrams of methods, devices, electronic devices (systems), and computer program products according to embodiments of the present invention. It should be understood that each process and / or block in the flowchart and / or block diagram, as well as the combination of processes and / or blocks in the flowchart and / or block diagram, can be implemented by computer program instructions. These computer program instructions can be provided to the processor of a general-purpose computer, a special-purpose computer, an embedded processor, or other programmable devices to generate a machine, so that the instructions executed by the processor of the computer or other programmable devices generate for realizing in the process Figure 1 a process or multiple processes and / or blocks Figure 1means for the functions specified in one or more boxes.
[0118] In a typical configuration, an electronic device includes one or more processors (CPUs), memory, and a bus. The electronic device may also include an input / output interface, a network interface, etc.
[0119] The memory may include non-permanent memory in the form of computer-readable media, random access memory (RAM) and / or non-volatile memory such as read-only memory (ROM) or flash memory (flash RAM), and the memory includes at least one storage chip. The memory is an example of computer-readable media.
[0120] Computer-readable media includes permanent and non-permanent, removable and non-removable media that can store information by any method or technology. The information can be computer-readable instructions, data structures, program modules, or other data. Examples of computer storage media include, but are not limited to, phase change memory (PRAM), static random access memory (SRAM), dynamic random access memory (DRAM), other types of random access memory (RAM), read-only memory (ROM), electrically erasable programmable read-only memory (EEPROM), flash memory or other memory technologies, compact disc read-only memory (CD-ROM), digital versatile disc (DVD) or other optical storage, magnetic cassette tapes, magnetic tape disk storage or other magnetic storage devices, or any other non-transitory media that can be used to store information that can be accessed by a computing device. As defined herein, computer-readable media does not include transitory computer-readable media, such as modulated data signals and carrier waves.
[0121] In the description of the present invention, it should be understood that if terms such as "upper", "lower", "front", "rear", "left", and "right" are used to indicate the orientation or positional relationship, they are based on the orientation or positional relationship shown in the drawings, and are only for the convenience of describing the present invention and simplifying the description, rather than indicating or implying that the indicated position or element must have a specific orientation, be constructed and operated in a specific orientation, and therefore should not be construed as a limitation of the present invention.
[0122] It should be noted that, in this document, relational terms such as first and second are only used to distinguish one entity or operation from another entity or operation, and do not necessarily require or imply any actual relationship or order between these entities or operations. It should also be noted that the term "comprising", "including" or any other variant thereof is intended to cover non-exclusive inclusion, so that a process, method, commodity or device comprising a series of elements not only includes those elements, but also includes other elements not expressly listed, or further includes elements inherent to such process, method, commodity or device. Without further limitation, an element defined by the statement "comprising an..." does not exclude the presence of additional identical elements in the process, method, commodity or device comprising the element.
[0123] Those skilled in the art should understand that the embodiments of the present invention can be provided as a method, a system or a computer program product. Therefore, the present invention can take the form of a completely hardware embodiment, a completely software embodiment or an embodiment combining software and hardware aspects. Moreover, the present invention can take the form of a computer program product implemented on one or more computer-usable storage media (including but not limited to disk storage, CD-ROM, optical storage, etc.) containing computer-usable program code.
[0124] The above are only the embodiments of the present invention and are not intended to limit the present invention. For those skilled in the art, the present invention can have various modifications and changes. Any modification, equivalent replacement, improvement, etc. made within the spirit and principle of the present invention shall be included within the scope of the present invention.
Claims
1. A method for generating referential expression understanding data, characterized in that Including: Obtain a target detection image, where at least one target annotation box is marked in the target detection image; Input the target detection image into a multi-modal large language model, and trigger the multi-modal large language model to generate a text description of the specified category target in the target detection image through a preset query statement, so as to obtain the target text description output by the multi-modal large language model; Input the target text description and the target detection image into a referential expression understanding model, so that the referential expression understanding model generates an initial bounding box for the specified category target in the target detection image according to the target text description; Calculate the intersection over union of the initial bounding box and each target annotation box in the target detection image; Select the target annotation box with the highest intersection over union as the target bounding box of the specified category target in the target detection image; Combine the target bounding box and the target text description to generate referential expression understanding data related to the target detection image.
2. The method according to claim 1, characterized in that Before selecting the target annotation box with the highest intersection over union as the target bounding box of the specified category target in the target detection image, the method further includes: Verify whether the highest intersection over union is greater than or equal to a preset intersection over union threshold. If not, discard the initial bounding box. If so, perform the step of selecting the target annotation box with the highest intersection over union as the target bounding box of the specified category target in the target detection image.
3. The method according to claim 2, wherein The step of selecting the target annotation box with the highest intersection over union as the target bounding box of the specified category target in the target detection image includes: In the case where there are more than two target annotation boxes with the same highest intersection over union, calculate the center point distance between the initial bounding box and each target annotation box with the same highest intersection over union, and select the target annotation box with the smallest center point distance as the target bounding box of the specified category target in the target detection image.
4. The method according to claim 2, wherein The step of selecting the target annotation box with the highest intersection over union as the target bounding box of the specified category target in the target detection image includes: In the case where there are more than two target annotation boxes with the same highest intersection over union, calculate the aspect ratio difference between the initial bounding box and each target annotation box with the same highest intersection over union, and select the target annotation box with the smallest aspect ratio difference as the target bounding box of the specified category target in the target detection image.
5. The method according to claim 1, characterized in that, Before inputting the target text description and the target detection image into the referential expression understanding model, the method further includes: Obtain an open-source referential expression understanding dataset, where the open-source referential expression understanding dataset includes multiple referential expression understanding data; Based on a preset loss function and optimization strategy, use each referential expression understanding data in the open-source referential expression understanding dataset to train a pre-constructed referential expression understanding model to obtain the trained referential expression understanding model.
6. The method according to claim 5, characterized in that, After combining the target bounding box and the target text description to generate referential expression understanding data related to the target detection image, the method further includes: Retrain the referential expression understanding model by combining the open-source referential expression understanding dataset and the referential expression understanding data related to the object detection image, to obtain the trained referential expression understanding model.
7. The method according to any one of claims 1 to 6, characterized in that, The preset query statement is a prompt word used to guide the multimodal large language model to generate a text description of the appearance features and / or behavior features of the specified category object in the object detection image.
8. An apparatus for generating referential expression understanding data, characterized in that, Including: An object detection image acquisition unit, a target text description acquisition unit, an initial bounding box generation unit, an intersection over union calculation unit, a target bounding box selection unit, and a referential expression understanding data generation unit. The object detection image acquisition unit is configured to acquire an object detection image, where at least one object annotation box is marked in the object detection image. The target text description acquisition unit is configured to input the object detection image into a multimodal large language model, and trigger the multimodal large language model to generate a text description of the specified category object in the object detection image through a preset query statement, to obtain the target text description output by the multimodal large language model. The initial bounding box generation unit is configured to input the target text description and the object detection image into a referential expression understanding model, so that the referential expression understanding model generates an initial bounding box for the specified category object in the object detection image according to the target text description. The intersection over union calculation unit is configured to calculate the intersection over union of the initial bounding box and each object annotation box in the object detection image. The target bounding box selection unit is configured to select the object annotation box with the highest intersection over union as the target bounding box of the specified category object in the object detection image. The referential expression understanding data generation unit is configured to generate referential expression understanding data related to the object detection image by combining the target bounding box and the target text description.
9. A computer-readable storage medium having a program stored thereon, characterized in that, When the program is executed by a processor, it implements the referential expression understanding data generation method according to any one of claims 1 to 7.
10. An electronic device, characterized in that, The electronic device includes at least one processor, at least one memory connected to the processor, and a bus; wherein, the processor and the memory complete communication with each other through the bus; the processor is configured to call program instructions in the memory to execute the referential expression understanding data generation method according to any one of claims 1 to 7.
Citation Information
Patent Citations
Data annotation box correction method
CN117831046A
Target positioning method and related equipment thereof
CN118568289A