Multi-modal Automatic Annotation Method, Device, Storage Medium and Electronic Device
By introducing cross-modal decoder and cross-attention mechanism in the automatic image annotation system, the problem that the prior art cannot process multimodal data is solved, efficient and accurate image annotation is achieved, and labor costs are reduced.
Patent Information
- Application Number
- CN202411776548.5
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2024-12-05
- Publication Date
- 2025-05-27
- Estimated Expiration
- 2044-12-05
AI Technical Summary
The prior art cannot process multimodal data in automatic image annotation, especially in simultaneously processing image and text information, resulting in low labeling efficiency, high cost, and limiting its application in complex scenarios.
By acquiring the to-process text and images, extracting features using natural language processing and image processing units, inputting them into a cross-modal decoder for feature enhancement and fusion, and calculating the correlation score between text and image features using the cross-attention mechanism, generating query information and inputting it to the multimodal object detection unit to obtain labeling information.
It realizes automatic extraction of information from text description and converting it into image annotation, which significantly improves the efficiency and accuracy of image annotation and reduces labor costs.
Smart Images

Figure CN119273997B_ABST
Abstract
Description
Technical Field
[0001] This application relates to the field of artificial intelligence technology, and in particular, to a multi-modal automatic annotation method, device, storage medium, and electronic device. Background Art
[0002] Artificial Intelligence (AI) is a rapidly developing field that uses computer science and data to imitate human intelligence. The applications of artificial intelligence are very extensive, ranging from virtual assistants and recommendation systems in daily life to professional fields such as healthcare, finance, and manufacturing.
[0003] In the field of automated object detection and recognition, image annotation technology has become an indispensable key link. The image annotation methods in related technologies mainly rely on manual operations, and the names and positions of objects are annotated by manually selecting specific regions in the image.
[0004] However, the annotation methods in related technologies can only automatically annotate fixed object categories. When new categories appear, manual annotation and retraining are required. Moreover, they lack the ability to process multi-modal data, that is, they cannot process image and text information simultaneously. This not only has a high annotation cost but also limits their application in complex scenarios. Summary of the Invention
[0005] The purpose of this application is to provide a multi-modal automatic annotation method, device, storage medium, and electronic device, which can automatically extract information from text descriptions and convert it into image annotations, greatly improving the efficiency and accuracy of image annotation and reducing labor costs.
[0006] This application provides a multi-modal automatic annotation method, including:
[0007] Obtain the text to be processed and the image to be processed, extract features from the text to be processed through the natural language processing unit to obtain the text features to be processed, and process the image to be processed through the image processing unit to obtain the image features to be processed; input the text features to be processed and the image features to be processed into the cross-modal decoder, enhance the features of the text features to be processed and the image features to be processed to obtain enhanced text features and enhanced image features, and use the cross-attention mechanism to perform feature fusion to obtain the correlation scores between the enhanced text features and the enhanced image features corresponding to different image regions; based on the correlation scores between the enhanced text features and the enhanced image features corresponding to different image regions, select the enhanced image feature with the highest correlation with the enhanced text feature, and generate query information corresponding to each selected image feature; input the enhanced text features, the enhanced image features, and the query information corresponding to each selected image feature into the multi-modal object detection unit to obtain the annotation information corresponding to each query information.
[0008] Optionally, the inputting the text features to be processed and the image features to be processed into the cross-modal decoder, enhancing the features of the text features to be processed and the image features to be processed to obtain enhanced text features and enhanced image features, includes: inputting the image features to be processed into multiple attention branch units of the cross-modal decoder to obtain the regional image features output by each attention branch unit; fusing the regional image features output by each attention branch unit to obtain the enhanced image features; wherein, the image to be processed is divided into multiple image regions, and one attention branch unit corresponds to one image region.
[0009] Optionally, the using the cross-attention mechanism to perform feature fusion to obtain the correlation scores between the enhanced text features and the enhanced image features corresponding to different image regions, includes: using the image features to be processed as the query, the text features to be processed as the key and value, and using the cross-attention mechanism to calculate the first correlation scores between the image features and the text features corresponding to different image regions.
[0010] Optionally, the using the cross-attention mechanism to perform feature fusion to obtain the correlation scores between the enhanced text features and the enhanced image features corresponding to different image regions, includes: using the text features to be processed as the query, the image features to be processed as the key and value, and using the cross-attention mechanism to calculate the second correlation scores between the text features and the image features corresponding to different image regions.
[0011] Optionally, the feature fusion using the cross-attention mechanism to obtain the correlation score between the enhanced text feature and the enhanced image features corresponding to different image regions includes: calculating the correlation score between the enhanced text feature and the enhanced image features corresponding to different image regions based on the first correlation score and the second correlation score.
[0012] Optionally, based on the correlation score between the enhanced text feature and the enhanced image features corresponding to different image regions, selecting the enhanced image feature with the highest correlation with the enhanced text feature and generating query information corresponding to each selected image feature includes: screening at least one image region with the highest correlation with the enhanced text feature from the multiple image regions based on the correlation score between the enhanced text feature and the enhanced image features corresponding to different image regions, and generating query information corresponding to each image region in the at least one image region; wherein, the query information includes: feature parameters and location information; the feature parameters are updated by backpropagation during the training phase.
[0013] Optionally, the multi-modal object detection unit includes: a prediction head and multiple decoders; each decoder includes; cross-attention from image to text, cross-attention from text to image; inputting the enhanced text feature, the enhanced image feature, and the query information corresponding to each selected image feature into the multi-modal object detection unit to obtain the annotation information corresponding to each query information includes: sequentially inputting the enhanced text feature, the enhanced image feature, and the query information corresponding to each selected image feature into the multiple decoders to obtain multi-modal features; inputting the multi-modal features into the prediction head to obtain the annotation information corresponding to each query information; wherein, the multiple decoders are connected in sequence, and the output of the previous decoder is used as the input of the subsequent decoder.
[0014] The present application also provides a multi-modal automatic annotation device, including:
[0015] A feature extraction module, configured to obtain a text to be processed and an image to be processed, and perform feature extraction on the text to be processed through the natural language processing unit to obtain text features to be processed, and process the image to be processed through the image processing unit to obtain image features to be processed; a feature fusion module, configured to input the text features to be processed and the image features to be processed into the cross-modal decoder, perform feature enhancement on the text features to be processed and the image features to be processed to obtain enhanced text features and enhanced image features, and use a cross-attention mechanism to perform feature fusion to obtain correlation scores between the enhanced text features and the enhanced image features corresponding to different image regions; a query information generation module, configured to select, based on the correlation scores between the enhanced text features and the enhanced image features corresponding to different image regions, the enhanced image features with the highest correlation with the enhanced text features, and generate query information corresponding to each selected image feature; a labeled information generation module, configured to input the enhanced text features, the enhanced image features, and the query information corresponding to each selected image feature into the multi-modal object detection unit to obtain labeled information corresponding to each query information.
[0016] Optionally, the feature fusion module is specifically configured to input the image features to be processed into multiple attention branch units of the cross-modal decoder to obtain regional image features output by each attention branch unit; the feature fusion module is further specifically configured to fuse the regional image features output by each attention branch unit to obtain the enhanced image features; wherein, the image to be processed is divided into multiple image regions, and one attention branch unit corresponds to one image region.
[0017] Optionally, the feature fusion module is specifically configured to use the image features to be processed as queries, the text features to be processed as keys and values, and use the cross-attention mechanism to calculate the first correlation scores between the image features and the text features corresponding to different image regions.
[0018] Optionally, the feature fusion module is specifically configured to use the text features to be processed as queries, the image features to be processed as keys and values, and use the cross-attention mechanism to calculate the second correlation scores between the text features and the image features corresponding to different image regions.
[0019] Optionally, the feature fusion module is specifically configured to calculate the correlation scores between the enhanced text features and the enhanced image features corresponding to different image regions based on the first correlation scores and the second correlation scores.
[0020] Optionally, the query information generation module is specifically configured to screen out at least one image region with the highest correlation with the enhanced text feature from the multiple image regions based on the correlation score between the enhanced text feature and the enhanced image features corresponding to different image regions, and generate query information corresponding to each image region in the at least one image region; wherein, the query information includes: feature parameters and location information; the feature parameters are updated by backpropagation during the training phase.
[0021] Optionally, the annotation information generation module is specifically configured to sequentially input the enhanced text feature, the enhanced image feature, and the query information corresponding to each selected image feature into the multiple decoders to obtain multimodal features; the annotation information generation module is further specifically configured to input the multimodal features into the prediction head to obtain annotation information corresponding to each query information; wherein, the multiple decoders are connected in sequence, and the output of the previous decoder serves as the input of the subsequent decoder.
[0022] The present application also provides a computer program product, including a computer program / instructions, which when executed by a processor, implements the steps of the multimodal automatic annotation method as described in any one of the above.
[0023] The present application also provides an electronic device, including a memory, a processor, and a computer program stored on the memory and executable on the processor, wherein when the processor executes the program, it implements the steps of the multimodal automatic annotation method as described in any one of the above.
[0024] The present application also provides a computer-readable storage medium, on which a computer program is stored, and when the computer program is executed by a processor, it implements the steps of the multimodal automatic annotation method as described in any one of the above.
[0025] The multi-modal automatic annotation method, device, storage medium and electronic device provided by this application obtain the text to be processed and the image to be processed, extract features from the text to be processed through the natural language processing unit to obtain the text features to be processed, and process the image to be processed through the image processing unit to obtain the image features to be processed; input the text features to be processed and the image features to be processed into the cross-modal decoder to enhance the features of the text features to be processed and the image features to be processed, obtain enhanced text features and enhanced image features, and use the cross-attention mechanism for feature fusion to obtain the correlation scores between the enhanced text features and the enhanced image features corresponding to different image regions; based on the correlation scores between the enhanced text features and the enhanced image features corresponding to different image regions, select the enhanced image feature with the highest correlation with the enhanced text feature, and generate query information corresponding to each selected image feature; input the enhanced text features, the enhanced image features and the query information corresponding to each selected image feature into the multi-modal target detection unit to obtain the annotation information corresponding to each query information. In this way, information can be automatically extracted from the text description and converted into image annotations, which not only greatly improves the efficiency and accuracy of image annotation, but also reduces the labor cost. BRIEF DESCRIPTION OF THE DRAWINGS
[0026] In order to more clearly illustrate the technical solutions in this application or the prior art, the following will briefly introduce the drawings required for the description of the embodiments or the prior art. Obviously, the drawings in the following description are some embodiments of this application. For those of ordinary skill in the art, other drawings can be obtained based on these drawings without creative efforts.
[0027] Figure 1 is one of the flow diagrams of the multi-modal automatic annotation method provided by this application;
[0028] Figure 2 is another flow diagram of the multi-modal automatic annotation method provided by this application;
[0029] Figure 3 is the structural diagram of the multi-modal automatic annotation device provided by this application;
[0030] Figure 4 is the structural diagram of the electronic device provided by this application. DETAILED DESCRIPTION OF THE EMBODIMENTS
[0031] To make the objectives, technical solutions and advantages of the present application clearer, the technical solutions in the present application will be clearly and completely described below with reference to the accompanying drawings in the present application. Apparently, the described embodiments are some, but not all, of the embodiments of the present application. All other embodiments obtained by those of ordinary skill in the art based on the embodiments in the present application without creative efforts shall fall within the protection scope of the present application.
[0032] The terms "first", "second", etc. in the specification and claims of the present application are used to distinguish similar objects, rather than to describe a specific order or sequence. It should be understood that such used data may be interchanged under appropriate circumstances so that the embodiments of the present application can be implemented in an order different from those illustrated or described herein, and the objects distinguished by "first", "second", etc. are generally of the same type, and the number of objects is not limited. For example, the first object may be one or more. In addition, "and / or" in the specification and claims means at least one of the connected objects, and the character " / " generally means an "or" relationship between the associated objects before and after.
[0033] With the rapid development of artificial intelligence technology, especially the continuous progress in the field of computer vision, image annotation technology has become a key link in realizing automatic object detection and recognition. Existing image annotation methods mainly rely on manual operations, and the names and positions of objects are annotated by manually selecting specific regions in the image. Most existing automatic annotation systems focus on fixed objects, and manual annotation of the object must be carried out first, and a target detection model is trained through the annotation data to obtain a pre-trained model with recognition functions, so as to further realize the automatic annotation of the target object.
[0034] The annotation methods in the related art can only automatically annotate fixed object categories. When new categories appear, manual annotation and retraining are required. Moreover, the annotation methods in the related art lack the ability to process multi-modal data, that is, they cannot process image and text information at the same time. This not only has low annotation efficiency and high cost, but also limits its application in complex scenarios.
[0035] In view of the above technical problems existing in the related art, the embodiments of the present application provide a multi-modal automatic annotation system and an automatic annotation method executed by the system, which can automatically extract information from text descriptions and convert it into image annotations, and have important practical significance and application value for improving the efficiency and accuracy of image annotation and reducing labor costs. As Figure 1As shown in the figure, the method includes: 1. The multi-modal input interface receives text data and image data to be annotated; 2. Extract text features from the text data and image features from the image data; 3. The cross-modal decoder fuses the features of the above two modalities, identifies the target object specified in the text in the image, and generates annotation information.
[0036] The multi-modal automatic annotation method provided by the embodiments of the present application will be described in detail below in conjunction with the accompanying drawings through specific embodiments and their application scenarios. The multi-modal automatic annotation system includes: a natural language processing unit, an image processing unit, a cross-modal decoder, and a multi-modal target detection unit.
[0037] As Figure 2 shown, a multi-modal automatic annotation method provided by an embodiment of the present application may include the following steps 201 to 204:
[0038] Step 201, obtain the text to be processed and the image to be processed, and extract the features of the text to be processed through the natural language processing unit to obtain the text features to be processed, and process the image to be processed through the image processing unit to obtain the image features to be processed.
[0039] Exemplarily, the user can input a text description and an input image through the system interface to obtain the above-mentioned text to be processed and the image to be processed. The text description contains the name, attributes, or other relevant information of the object, and multiple objects can be detected simultaneously. The input image is the target scene to be annotated, and multiple image data can be input simultaneously.
[0040] Exemplarily, after obtaining the text to be processed and the image to be processed, the input text description can be analyzed through the natural language processing unit to extract keywords and features, such as attributes such as object name and color. The input image is subjected to feature extraction through the image processing unit to obtain multi-scale image features, that is, the above-mentioned image features to be processed.
[0041] Step 202, input the text features to be processed and the image features to be processed into the cross-modal decoder, enhance the features of the text features to be processed and the image features to be processed to obtain enhanced text features and enhanced image features, and use the cross-attention mechanism to perform feature fusion to obtain the correlation scores between the enhanced text features and the enhanced image features corresponding to different image regions.
[0042] Exemplarily, the multi-modal feature fusion unit of the cross-modal decoder, after receiving the above-mentioned text features to be processed and image features to be processed, uses the self-attention mechanism to enhance the text features and uses the enhanced multi-head self-attention module to enhance the image features. Then, by adding cross-attention from image to text and cross-attention from text to image, feature fusion is achieved.
[0043] Specifically, in the above step 202, the steps of enhancing the image features may include the following steps 202a1 and 202a2:
[0044] Step 202a1: Input the image features to be processed into multiple attention branch units of the cross-modal decoder to obtain regional image features output by each attention branch unit.
[0045] Step 202a2: After fusing the regional image features output by each attention branch unit, obtain the enhanced image features.
[0046] Among them, the image to be processed is divided into multiple image regions, and one attention branch unit corresponds to one image region.
[0047] Exemplarily, the embodiments of the present application provide an enhanced multi-head self-attention module, which combines deformable convolution and channel attention mechanism. With the help of the dynamic adjustment ability of deformable convolution, it can adapt to the specific features of the image content, such as edges, corners or specific texture patterns. Deformable convolution is a special convolution operation that allows the shape of the convolution kernel to be dynamically adjusted according to the input feature map. This adjustment is achieved by adding additional offsets, which can learn how to better align to the boundaries of the target object. The advantage of this is that the model can more flexibly capture the shape and position changes of the target, thereby improving the ability to locate and recognize the target.
[0048] Exemplarily, for each branch of the multi-head self-attention module, the same structure and parameters are used. Each branch only focuses on a local area of the original image features, thereby enhancing the model's perception of the local features of the image and improving the model's recognition ability for small but key differences in the image. Input the multi-scale image features obtained in the above steps into the multi-head self-attention module. The input image features will be divided into multiple local areas, so that the attention can be restricted within the local areas. Each local area enters a branch, and each branch is an enhanced attention module. Finally, the outputs of all branches are fused to form the final enhanced image feature representation.
[0049] Exemplarily, after obtaining the above enhanced text features and enhanced image features, feature fusion is also required. Specifically, the step of fusing the enhanced text features and the enhanced image features in step 202 above may include the following steps 202b1 and 202b2:
[0050] Step 202b1: Using the image features to be processed as queries, the text features to be processed as keys and values, and calculating the first correlation score between the image features corresponding to different image regions and the text features by using the cross-attention mechanism.
[0051] Exemplarily, in the cross-attention from image to text, the image features are used as queries, the text features are used as keys and values, and the correlation score between the image features and the text features is calculated through the cross-attention mechanism.
[0052] Step 202b2: Using the text features to be processed as queries, the image features to be processed as keys and values, and calculating the second correlation score between the text features and the image features corresponding to different image regions by using the cross-attention mechanism.
[0053] Exemplarily, in the cross-attention from text to image, the text features are used as queries, the image features are used as keys and values, and the correlation score between the text features and the image features is calculated through the cross-attention mechanism.
[0054] It can be understood that through the above two cross-attention mechanisms, the model can effectively fuse the image features and the text features, improve the model's understanding of the image content, and enhance its response ability to the text description. This tightly coupled feature fusion method helps the model achieve better performance in the open-set object detection task, especially when dealing with categories not seen during training.
[0055] Specifically, in the above step 202, the step of calculating the correlation score between the enhanced text features and the enhanced image features corresponding to different image regions may further include the following step 202c:
[0056] Step 202c: Based on the first correlation score and the second correlation score, calculating the correlation score between the enhanced text features and the enhanced image features corresponding to different image regions.
[0057] Exemplarily, in the above step, based on the correlation score between the enhanced text features and the enhanced image features corresponding to different image regions, it is possible to select the region most relevant to the input text through the text features. Combining the text features with the dynamic anchor boxes, the query selection guided by the language is used to determine which anchor box regions are most relevant to the input text.
[0058] Step 203: Based on the correlation scores between the enhanced text features and the enhanced image features corresponding to different image regions, select the enhanced image feature with the highest correlation with the enhanced text features, and generate query information corresponding to each selected image feature.
[0059] Exemplarily, according to the correlation scores between the enhanced text features and the enhanced image features corresponding to different image regions, the image features most relevant to the text features can be selected. These image features are considered to be the regions most likely to contain the objects described in the text.
[0060] Specifically, the above step 203 may further include the following step 203a:
[0061] Step 203a: Based on the correlation scores between the enhanced text features and the enhanced image features corresponding to different image regions, filter out at least one image region with the highest correlation with the enhanced text features from the multiple image regions, and generate query information corresponding to each image region in the at least one image region.
[0062] Wherein, the query information includes: feature parameters and location information; the feature parameters are updated in the training stage by means of backpropagation.
[0063] Exemplarily, in the embodiments of the present application, a query information is generated for each selected image feature, and each query information includes the following content: learnable parameters and location information. The learnable parameters are learnable in the training stage of the model, which means that during the training process, the model can learn how to better match the objects described in the input text. At the beginning of the model training, the learnable parameters are randomly initialized and updated through the backpropagation algorithm during the training process. The loss function tells the model the difference between the current learnable parameters and the actual target objects to capture the local image feature representation most relevant to the input text. These learnable parameters can be regarded as the "understanding" or "representation" of the objects described in the input text by the model, and they are continuously optimized as the training data accumulates. Through repeated training and adjustment, these content parts gradually learn how to capture the key features described in the input text, so as to more accurately identify and locate the target objects. The location information is represented by dynamic anchor boxes. The dynamic anchor box is a flexible bounding box representation that allows the model to generate bounding boxes with different positions and sizes for each object, increasing the adaptability of the model to various object shape and size changes.
[0064] Step 204: Input the enhanced text features, the enhanced image features, and the query information corresponding to each selected image feature into the multi-modal target detection unit to obtain the annotation information corresponding to each query information.
[0065] Exemplarily, after obtaining the query information corresponding to each selected image feature, the enhanced text feature, the enhanced image feature, and the query information corresponding to each selected image feature can be input into the multi-modal object detection unit for object detection to generate the annotation information corresponding to each query information.
[0066] Specifically, the multi-modal object detection unit includes: a prediction head and a plurality of decoders; each decoder includes; cross-attention from image to text, cross-attention from text to image, and the above step 204 may further include the following step 204a1 and step 204a2:
[0067] Step 204a1: Input the enhanced text feature, the enhanced image feature, and the query information corresponding to each selected image feature into the plurality of decoders in sequence to obtain multi-modal features.
[0068] Step 204a2: Input the multi-modal features into the prediction head to obtain the target annotation information corresponding to each query information.
[0069] Among them, the plurality of decoders are connected in sequence, and the output of the previous decoder is used as the input of the subsequent decoder.
[0070] Exemplarily, the above multi-modal object detection unit is used for accurate object detection, and is composed of a plurality of decoders and a prediction head, inputs the enhanced text feature and enhanced image feature obtained in step 3, and the query information obtained in the above steps, and outputs the position, size, and category of the final predicted target for each query.
[0071] Exemplarily, each decoder is composed of a self-attention, a cross-attention from image to text, a cross-attention from text to image, and a fully connected layer. Through multiple decoder layers, the model can gradually deepen the fusion of image and text features at different stages. Each layer can add more feature fusion to help the model better align and understand cross-modal information. More decoder layers mean that the model has a stronger ability to express complex relationships and patterns, especially when dealing with complex interactions between vision and language. At the end of the plurality of decoders, the output of the decoder, that is, the final multi-modal feature, is passed to the prediction head, and the prediction head outputs the predicted bounding box, category, and category probability.
[0072] Exemplarily, after obtaining the above annotation information, a YOLO format annotation file and a result image with detection boxes can be generated. These annotation data files can be directly used for the training and verification of the YOLO object detection model.
[0073] The multimodal automatic annotation method provided in the embodiment of the present application obtains the text to be processed and the image to be processed, and extracts the features of the text to be processed by the natural language processing unit to obtain the features of the text to be processed, and processes the image to be processed by the image processing unit to obtain the features of the image to be processed; the features of the text to be processed and the features of the image to be processed are input into the cross-modal decoder, the features of the text to be processed and the features of the image to be processed are enhanced to obtain enhanced text features and enhanced image features, and the cross-attention mechanism is used to perform feature fusion to obtain the correlation score between the enhanced text features and the enhanced image features corresponding to different image regions; based on the correlation score between the enhanced text features and the enhanced image features corresponding to different image regions, the enhanced image features with the highest correlation with the enhanced text features are selected, and query information corresponding to each selected image feature is generated; the enhanced text features, the enhanced image features and the query information corresponding to each selected image feature are input into the multimodal target detection unit to obtain the annotation information corresponding to each query information. In this way, information can be automatically extracted from the text description and converted into image annotation, which not only greatly improves the efficiency and accuracy of image annotation, but also reduces labor costs.
[0074] It should be noted that the multimodal automatic annotation method provided in the embodiment of the present application can be executed by a multimodal automatic annotation device, or a control module in the multimodal automatic annotation device for executing the multimodal automatic annotation method. In the embodiment of the present application, the multimodal automatic annotation device executing the multimodal automatic annotation method is taken as an example to illustrate the multimodal automatic annotation device provided in the embodiment of the present application.
[0075] It should be noted that in the embodiments of the present application, the multimodal automatic annotation methods shown in the above-mentioned method drawings are all illustrated by taking one of the drawings in the embodiments of the present application as an example. In specific implementation, the multimodal automatic annotation methods shown in the above-mentioned method drawings can also be implemented in combination with any other drawings that can be combined as shown in the above-mentioned embodiments, which will not be repeated here.
[0076] The multimodal automatic annotation device provided by the present application is described below, and the multimodal automatic annotation method described below and above can be referenced to each other.
[0077] Figure 3 A schematic diagram of the structure of a multi-modal automatic annotation device provided in an embodiment of the present application is shown in FIG. Figure 3 As shown, specifically including:
[0078] A feature extraction module 301 is configured to obtain a text to be processed and an image to be processed, extract features from the text to be processed through the natural language processing unit to obtain text features to be processed, and process the image to be processed through the image processing unit to obtain image features to be processed; A feature fusion module 302 is configured to input the text features to be processed and the image features to be processed into the cross-modal decoder, enhance the text features to be processed and the image features to be processed to obtain enhanced text features and enhanced image features, and use a cross-attention mechanism for feature fusion to obtain the correlation scores between the enhanced text features and the enhanced image features corresponding to different image regions; A query information generation module 303 is configured to select the enhanced image feature with the highest correlation with the enhanced text feature based on the correlation scores between the enhanced text features and the enhanced image features corresponding to different image regions, and generate query information corresponding to each selected image feature; A labeled information generation module 304 is configured to input the enhanced text features, the enhanced image features, and the query information corresponding to each selected image feature into the multi-modal object detection unit to obtain labeled information corresponding to each query information.
[0079] Optionally, the feature fusion module 302 is specifically configured to input the image features to be processed into multiple attention branch units of the cross-modal decoder to obtain regional image features output by each attention branch unit; The feature fusion module 302 is further specifically configured to fuse the regional image features output by each attention branch unit to obtain the enhanced image features; wherein, the image to be processed is divided into multiple image regions, and one attention branch unit corresponds to one image region.
[0080] Optionally, the feature fusion module 302 is specifically configured to use the image features to be processed as queries, the text features to be processed as keys and values, and use a cross-attention mechanism to calculate the first correlation scores between the image features and the text features corresponding to different image regions.
[0081] Optionally, the feature fusion module 302 is specifically configured to use the text features to be processed as queries, the image features to be processed as keys and values, and use a cross-attention mechanism to calculate the second correlation scores between the text features and the image features corresponding to different image regions.
[0082] Optionally, the feature fusion module 302 is specifically configured to calculate the correlation scores between the enhanced text features and the enhanced image features corresponding to different image regions based on the first correlation scores and the second correlation scores.
[0083] Optionally, the query information generation module 303 is specifically configured to screen out at least one image region with the highest correlation with the enhanced text feature from the multiple image regions based on the correlation score between the enhanced text feature and the enhanced image features corresponding to different image regions, and generate query information corresponding to each image region in the at least one image region; wherein, the query information includes: feature parameters and location information; the feature parameters are updated in a backpropagation manner during the training phase.
[0084] Optionally, the annotation information generation module 304 is specifically configured to input the enhanced text feature, the enhanced image feature, and the query information corresponding to each selected image feature into the multiple decoders in sequence to obtain multimodal features; the annotation information generation module 304 is further specifically configured to input the multimodal features into the prediction head to obtain annotation information corresponding to each query information; wherein, the multiple decoders are connected in sequence, and the output of the previous decoder is used as the input of the subsequent decoder.
[0085] The multimodal automatic annotation device provided in this application acquires a text to be processed and an image to be processed, extracts features of the text to be processed through the natural language processing unit to obtain text features to be processed, and processes the image to be processed through the image processing unit to obtain image features to be processed; inputs the text features to be processed and the image features to be processed into the cross-modal decoder to enhance the features of the text features to be processed and the image features to be processed, obtaining enhanced text features and enhanced image features, and uses a cross-attention mechanism to perform feature fusion to obtain the correlation score between the enhanced text feature and the enhanced image features corresponding to different image regions; based on the correlation score between the enhanced text feature and the enhanced image features corresponding to different image regions, selects the enhanced image feature with the highest correlation with the enhanced text feature, and generates query information corresponding to each selected image feature; inputs the enhanced text feature, the enhanced image feature, and the query information corresponding to each selected image feature into the multimodal object detection unit to obtain annotation information corresponding to each query information. In this way, information can be automatically extracted from the text description and converted into image annotations, which not only greatly improves the efficiency and accuracy of image annotation, but also reduces the labor cost.
[0086] Figure 4 An example of the physical structure diagram of an electronic device is as Figure 4As shown in the figure, the electronic device may include: a processor 410, a communications interface 420, a memory 430, and a communication bus 440. Among them, the processor 410, the communications interface 420, and the memory 430 complete communication with each other through the communication bus 440. The processor 410 may call the logical instructions in the memory 430 to execute a multi-modal automatic annotation method, which includes: obtaining a text to be processed and an image to be processed, and extracting features of the text to be processed through the natural language processing unit to obtain text features to be processed, and processing the image to be processed through the image processing unit to obtain image features to be processed; inputting the text features to be processed and the image features to be processed into the cross-modal decoder to enhance the features of the text features to be processed and the image features to be processed, obtain enhanced text features and enhanced image features, and use the cross-attention mechanism for feature fusion to obtain the correlation scores between the enhanced text features and the enhanced image features corresponding to different image regions; based on the correlation scores between the enhanced text features and the enhanced image features corresponding to different image regions, select the enhanced image feature with the highest correlation with the enhanced text feature, and generate query information corresponding to each selected image feature; input the enhanced text features, the enhanced image features, and the query information corresponding to each selected image feature into the multi-modal object detection unit to obtain annotation information corresponding to each query information.
[0087] In addition, when the logical instructions in the above-mentioned memory 430 are implemented in the form of software functional units and sold or used as independent products, they may be stored in a computer-readable storage medium. Based on such an understanding, the technical solution of the present application, in essence, or the part that contributes to the prior art, or a part of this technical solution, may be embodied in the form of a software product. This computer software product is stored in a storage medium and includes several instructions for causing a computer device (which may be a personal computer, a server, or a network device, etc.) to execute all or part of the steps of the methods described in various embodiments of the present application. The foregoing storage medium includes: various media such as USB flash drives, mobile hard disks, read-only memories (ROM, Read-Only Memory), random access memories (RAM, Random Access Memory), magnetic disks, or optical discs that can store program codes.
[0088] On the other hand, the present application also provides a computer program product, which includes a computer program stored on a computer-readable storage medium. The computer program includes program instructions. When the program instructions are executed by a computer, the computer can execute the multi-modal automatic annotation method provided by the above-mentioned various methods. The method includes: obtaining a text to be processed and an image to be processed, and extracting features of the text to be processed through the natural language processing unit to obtain text features to be processed, and processing the image to be processed through the image processing unit to obtain image features to be processed; inputting the text features to be processed and the image features to be processed into the cross-modal decoder to enhance the features of the text features to be processed and the image features to be processed, obtaining enhanced text features and enhanced image features, and using a cross-attention mechanism for feature fusion to obtain the correlation scores between the enhanced text features and the enhanced image features corresponding to different image regions; based on the correlation scores between the enhanced text features and the enhanced image features corresponding to different image regions, selecting the enhanced image feature with the highest correlation with the enhanced text feature, and generating query information corresponding to each selected image feature; inputting the enhanced text features, the enhanced image features, and the query information corresponding to each selected image feature into the multi-modal target detection unit to obtain annotation information corresponding to each query information.
[0089] On another aspect, the present application also provides a computer-readable storage medium, on which a computer program is stored. When the computer program is executed by a processor, it is configured to execute the multi-modal automatic annotation method provided by the above-mentioned various methods. The method includes: obtaining a text to be processed and an image to be processed, and extracting features of the text to be processed through the natural language processing unit to obtain text features to be processed, and processing the image to be processed through the image processing unit to obtain image features to be processed; inputting the text features to be processed and the image features to be processed into the cross-modal decoder to enhance the features of the text features to be processed and the image features to be processed, obtaining enhanced text features and enhanced image features, and using a cross-attention mechanism for feature fusion to obtain the correlation scores between the enhanced text features and the enhanced image features corresponding to different image regions; based on the correlation scores between the enhanced text features and the enhanced image features corresponding to different image regions, selecting the enhanced image feature with the highest correlation with the enhanced text feature, and generating query information corresponding to each selected image feature; inputting the enhanced text features, the enhanced image features, and the query information corresponding to each selected image feature into the multi-modal target detection unit to obtain annotation information corresponding to each query information.
[0090] The device embodiments described above are merely illustrative. The units described as separate components may or may not be physically separated, and the components shown as units may or may not be physical units, that is, they may be located in one place or distributed to multiple network units. Some or all of the modules can be selected according to actual needs to achieve the purpose of the solution of this embodiment. Those of ordinary skill in the art can understand and implement it without creative efforts.
[0091] Through the description of the above embodiments, those skilled in the art can clearly understand that each embodiment can be implemented by means of software plus a necessary general hardware platform, and of course, it can also be implemented by hardware. Based on such an understanding, the essence of the above technical solution, or the part that contributes to the prior art, can be embodied in the form of a software product. The computer software product can be stored in a computer-readable storage medium, such as ROM / RAM, magnetic disk, optical disk, etc., and includes several instructions for causing a computer device (which can be a personal computer, a server, or a network device, etc.) to execute the methods described in each embodiment or some parts of the embodiments.
[0092] Finally, it should be noted that the above embodiments are only used to illustrate the technical solutions of the present application, rather than to limit them; although the present application has been described in detail with reference to the foregoing embodiments, those of ordinary skill in the art should understand that they can still modify the technical solutions recorded in the foregoing embodiments, or perform equivalent replacements for some of the technical features; and these modifications or replacements do not make the essence of the corresponding technical solutions deviate from the spirit and scope of the technical solutions of the embodiments of the present application.
Claims
1. A multimodal automatic annotation method, characterized in that: Applied to a multimodal automatic annotation system, the automatic annotation system comprises: a natural language processing unit, an image processing unit, a cross-modal decoder and a multimodal target detection unit; The method comprises: Acquire a text to be processed and an image to be processed, and extract features of the text to be processed by the natural language processing unit to obtain features of the text to be processed, and process the image to be processed by the image processing unit to obtain features of the image to be processed; Inputting the to-be-processed text features and the to-be-processed image features into the cross-modal decoder, performing feature enhancement on the to-be-processed text features and the to-be-processed image features to obtain enhanced text features and enhanced image features, and performing feature fusion using a cross-attention mechanism to obtain correlation scores between the enhanced text features and enhanced image features corresponding to different image regions; Based on the correlation scores between the enhanced text features and the enhanced image features corresponding to different image regions, the enhanced image features with the highest correlation with the enhanced text features are selected, and query information corresponding to each selected image feature is generated; Inputting the query information corresponding to the enhanced text feature, the enhanced image feature and each selected image feature into the multimodal object detection unit to obtain annotation information corresponding to each query information; The step of selecting an enhanced image feature with the highest correlation with the enhanced text feature based on the correlation score between the enhanced text feature and the enhanced image features corresponding to different image regions, and generating query information corresponding to each selected image feature includes: Based on the correlation scores between the enhanced text features and the enhanced image features corresponding to different image regions, at least one image region with the highest correlation with the enhanced text features is screened out from multiple image regions, and query information corresponding to each image region in the at least one image region is generated.
2. The method according to claim 1, characterized in that The step of inputting the to-be-processed text features and the to-be-processed image features into the cross-modal decoder, and performing feature enhancement on the to-be-processed text features and the to-be-processed image features to obtain enhanced text features and enhanced image features comprises: Inputting the image features to be processed into multiple attention branch units of the cross-modal decoder to obtain regional image features output by each attention branch unit; After fusing the regional image features output by each attention branch unit, the enhanced image features are obtained; The image to be processed is divided into multiple image regions, and one attention branch unit corresponds to one image region.
3. The method according to claim 1 or 2, characterized in that: The cross-attention mechanism is used to perform feature fusion to obtain a correlation score between the enhanced text feature and the enhanced image features corresponding to different image regions, including: The image features to be processed are used as queries and the text features to be processed are used as keys and values, and a cross-attention mechanism is used to calculate first correlation scores between image features and text features corresponding to different image regions.
4. The method according to claim 3, characterized in that The cross-attention mechanism is used to perform feature fusion to obtain a correlation score between the enhanced text feature and the enhanced image features corresponding to different image regions, including: The text features to be processed are used as queries and the image features to be processed are used as keys and values, and a cross-attention mechanism is used to calculate a second correlation score between the text features and the image features corresponding to different image regions.
5. The method according to claim 4, characterized in that The cross-attention mechanism is used to perform feature fusion to obtain a correlation score between the enhanced text feature and the enhanced image features corresponding to different image regions, including: Based on the first correlation score and the second correlation score, a correlation score between the enhanced text feature and enhanced image features corresponding to different image regions is calculated.
6. The method according to claim 1 or 5, characterized in that: The query information includes: feature parameters and location information; the feature parameters are updated by back propagation during the training phase.
7. The method according to claim 6, characterized in that The multimodal object detection unit includes: a prediction head and a plurality of decoders; each decoder includes: image-to-text cross attention, text-to-image cross attention; The step of inputting the query information corresponding to the enhanced text feature, the enhanced image feature, and each selected image feature into the multimodal target detection unit to obtain annotation information corresponding to each query information includes: Inputting the enhanced text feature, the enhanced image feature and query information corresponding to each selected image feature into the multiple decoders in sequence to obtain a multimodal feature; Inputting the multimodal features into the prediction head to obtain labeling information corresponding to each query information; The multiple decoders are connected in sequence, and the output of the previous decoder serves as the input of the subsequent decoder.
8. A multi-modal automatic annotation device, characterized in that: Applied to a multimodal automatic annotation system, the automatic annotation system comprises: a natural language processing unit, an image processing unit, a cross-modal decoder and a multimodal target detection unit; The device comprises: A feature extraction module, used to obtain a text to be processed and an image to be processed, and to extract features from the text to be processed by the natural language processing unit to obtain features of the text to be processed, and to process the image to be processed by the image processing unit to obtain features of the image to be processed; A feature fusion module, used for inputting the to-be-processed text features and the to-be-processed image features into the cross-modal decoder, performing feature enhancement on the to-be-processed text features and the to-be-processed image features to obtain enhanced text features and enhanced image features, and performing feature fusion using a cross-attention mechanism to obtain a correlation score between the enhanced text features and enhanced image features corresponding to different image regions; A query information generating module, configured to select, based on the correlation scores between the enhanced text features and the enhanced image features corresponding to different image regions, the enhanced image features with the highest correlation with the enhanced text features, and generate query information corresponding to each selected image feature; An annotation information generating module, used for inputting the query information corresponding to the enhanced text feature, the enhanced image feature and each selected image feature into the multimodal object detection unit to obtain annotation information corresponding to each query information; The query information generation module is specifically used to screen out at least one image region with the highest correlation with the enhanced text feature from multiple image regions based on the correlation score between the enhanced text feature and the enhanced image features corresponding to different image regions, and generate query information corresponding to each image region in the at least one image region.
9. An electronic device, characterized in that: The invention comprises a memory, a processor and a computer program stored in the memory and executable on the processor, wherein when the processor executes the program, the steps of the multimodal automatic annotation method according to any one of claims 1 to 7 are implemented.
10. A computer-readable storage medium, characterized in that: A computer program is stored thereon, and when the computer program is executed by a processor, the steps of the multimodal automatic annotation method as claimed in any one of claims 1 to 7 are implemented.
Citation Information
Patent Citations
Multi-modal feature processing method and device, storage medium and electronic equipment
CN116861363A
Multi-modal named entity recognition method based on multi-task cooperative characterization
CN116956920A