Object labeling method and device based on visual language model and area suggestion network

Through the combination of visual language model and regional proposal network, efficient and accurate labeling of unseen objects is achieved, the problem of opening set detection in the prior art is solved, the complexity and calculation overhead of labeling are reduced, and the labeling efficiency and accuracy are improved.

CN120388202APending Publication Date: 2025-07-29TSINGHUA UNIVERSITY
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510310943.2
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-03-17
Publication Date
2025-07-29

AI Technical Summary

Technical Problem

The existing pretrained models cannot accurately label when facing unseen objects and cannot adapt to open set detection, which increases the complexity and calculation overhead of the labeling process and reduces the labeling efficiency and accuracy.

Method used

Using a method based on visual language model and regional suggestion network, object names are generated through multimodal fusion, and regression classification and matching annotation are performed through regional suggestion networks to generate object annotation images.

Benefits of technology

It improves the accuracy of object detection, especially when facing unseen objects, can maintain high detection performance, reduces the cost and calculation overhead of manual labeling, and improves the efficiency and accuracy of the labeling process.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120388202A_ABST
    Figure CN120388202A_ABST
Patent Text Reader

Abstract

The embodiment of the invention provides an object labeling method and device based on a visual language model and a region suggestion network. The method comprises the steps that an image to be labeled and a prompt statement are acquired; performing multi-modal fusion on the to-be-labeled image and the prompt statement through a visual language model to generate an object name; carrying out regression classification on the to-be-labeled image according to the object name through a region suggestion network, and generating an object category candidate region; according to the method, matching labeling is performed according to the object name and the object category candidate area, and an object labeling image is generated, so that the accuracy of object detection can be effectively improved, and especially when an unseen object is faced, relatively high detection performance can still be kept, and open set detection is adapted; the method can significantly reduce the cost of manual labeling, reduces the complexity and calculation overhead of the labeling process, improves the efficiency and accuracy of the labeling process, and has a wide application prospect.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the field of computer technology, in particular to the field of artificial intelligence technology, and more particularly to an object annotation method and device based on a vision-language model and a region proposal network. Background Art

[0002] Automatic annotation technology has made remarkable progress in recent years, especially in the field of computer vision. Image annotation has become the basis for tasks such as object detection, image recognition, and semantic segmentation. Existing annotation methods usually rely on pre-trained models to perform automatic annotation of images. The pre-trained models are usually based on deep neural networks, especially convolutional neural networks (CNNs) and transformers, which can extract features in images and perform classification and localization. However, the pre-trained models can only annotate based on the objects they have encountered during the training phase. When encountering new objects or images containing unknown objects, traditional pre-trained models often cannot accurately annotate, and cannot adapt to the open-vocabulary detection problem; in the case of multiple objects and complex scenes, it is necessary to perform complex matching and reasoning on image and language information, increasing the complexity of the annotation process, with a large computational overhead, and low annotation efficiency and accuracy. Summary of the Invention

[0003] An object of the present invention is to provide an object annotation method based on a vision-language model and a region proposal network, which can effectively improve the accuracy of object detection, especially when facing unseen objects, still maintain high detection performance, and adapt to open-vocabulary detection; it can significantly reduce the cost of manual annotation, reduce the complexity and computational overhead of the annotation process, and at the same time improve the efficiency and accuracy of the annotation process, having broad application prospects. Another object of the present invention is to provide an object annotation device based on a vision-language model and a region proposal network. Still another object of the present invention is to provide a computer-readable medium. Yet another object of the present invention is to provide a computer device.

[0004] To achieve the above object, on the one hand, the present invention discloses an object annotation method based on a vision-language model and a region proposal network, including:

[0005] Obtain an image to be annotated and a prompt statement;

[0006] Through the vision-language model, perform multimodal fusion on the image to be annotated and the prompt statement to generate an object name;

[0007] Through the region proposal network, perform regression classification on the image to be annotated according to the object name to generate candidate regions for object categories;

[0008] Match and label according to the object name and the object category candidate region to generate an object labeled image.

[0009] Preferably, through a vision-language model, perform multimodal fusion on the image to be labeled and the prompt statement to generate an object name, including:

[0010] Extract features from the image to be labeled to generate visual features;

[0011] Extract features from the prompt statement to generate text features;

[0012] Fuse the visual features and the text features to generate an object name.

[0013] Preferably, fuse the visual features and the text features to generate an object name, including:

[0014] Generate the object similarity between the image to be labeled and each preset candidate object name according to the visual features and the text features;

[0015] Determine the object name according to multiple object similarities.

[0016] Preferably, through a region proposal network, perform regression classification on the image to be labeled according to the object name to generate an object category candidate region, including:

[0017] Extract the visual features of the image to be labeled through a convolutional neural network to generate a feature map;

[0018] Traverse the feature map according to a preset sliding window to generate multiple object candidate boxes;

[0019] Classify according to the object candidate region indicated by the object candidate box and the object name to determine the object category candidate region.

[0020] Preferably, classify according to the object candidate region indicated by the object candidate box and the object name to determine the object category candidate region, including:

[0021] Calculate the similarity between the object candidate region indicated by the object candidate box and the object name to generate the matching probability between each object candidate region and the object name;

[0022] Determine the object category candidate region according to multiple matching probabilities.

[0023] Preferably, match and label according to the object name and the object category candidate region to generate an object labeled image, including:

[0024] Judge whether the object name generated by the vision-language model is consistent with the object name included in the object category candidate region;

[0025] If they are consistent, calculate the similarity between the object name and the object category candidate region to generate the matching degree between the object name and the object category candidate region;

[0026] If the matching degree is greater than the preset matching degree threshold, label the object category candidate region as the object name on the image to be labeled to generate an object-labeled image.

[0027] The present invention also discloses an object labeling device based on a vision-language model and a region proposal network, including:

[0028] An acquisition unit for acquiring an image to be labeled and a prompt statement;

[0029] A multimodal fusion unit for performing multimodal fusion on the image to be labeled and the prompt statement through a vision-language model to generate an object name;

[0030] A regression classification unit for performing regression classification on the image to be labeled according to the object name through a region proposal network to generate an object category candidate region;

[0031] An object labeling unit for performing matching labeling according to the object name and the object category candidate region to generate an object-labeled image.

[0032] Preferably, the multimodal fusion unit is specifically configured to extract features from the image to be labeled to generate visual features; extract features from the prompt statement to generate text features; and fuse the visual features and the text features to generate an object name.

[0033] Preferably, the multimodal fusion unit is specifically configured to generate the object similarity between the image to be labeled and each preset candidate object name according to the visual features and the text features; and determine the object name according to multiple object similarities.

[0034] Preferably, the regression classification unit is specifically configured to extract the visual features of the image to be labeled through a convolutional neural network to generate a feature map; traverse the feature map according to a preset sliding window to generate multiple object candidate boxes; and classify according to the object candidate regions and the object name indicated by the object candidate boxes to determine the object category candidate region.

[0035] Preferably, the regression classification unit is specifically configured to calculate the similarity between the object candidate regions and the object name indicated by the object candidate boxes to generate the matching probability between each object candidate region and the object name; and determine the object category candidate region according to multiple matching probabilities.

[0036] Preferably, the object annotation unit is specifically configured to determine whether the object name generated by the vision-language model is consistent with the object name included in the object category candidate region; if they are consistent, calculate the similarity between the object name and the object category candidate region to generate the matching degree between the object name and the object category candidate region; if the matching degree is greater than a preset matching degree threshold, label the object category candidate region as the object name on the image to be annotated to generate an object annotation image.

[0037] The present invention also discloses a computer-readable medium, on which a computer program is stored, and when the program is executed by a processor, the above-mentioned method is implemented.

[0038] The present invention also discloses a computer device, including a memory and a processor, the memory is used to store information including program instructions, the processor is used to control the execution of the program instructions, and when the processor executes the program, the above-mentioned method is implemented.

[0039] The present invention also discloses a computer program product, including computer program / instructions, and when the computer program / instructions are executed by a processor, the above-mentioned method is implemented.

[0040] The present invention obtains an image to be annotated and a prompt statement; through a vision-language model, performs multimodal fusion on the image to be annotated and the prompt statement to generate an object name; through a region proposal network, performs regression classification on the image to be annotated according to the object name to generate an object category candidate region; performs matching annotation according to the object name and the object category candidate region to generate an object annotation image, which can effectively improve the accuracy of object detection, especially when facing unseen objects, can still maintain high detection performance and adapt to open-set detection; can significantly reduce the cost of manual annotation, reduce the complexity and computational overhead of the annotation process, and at the same time improve the efficiency and accuracy of the annotation process, and has broad application prospects. Description of the Drawings

[0041] In order to more clearly illustrate the technical solutions in the embodiments of the present invention or the prior art, the following will briefly introduce the drawings required for the description of the embodiments or the prior art. Obviously, the following drawings are only some embodiments of the present invention, and those of ordinary skill in the art can also obtain other drawings according to these drawings without creative efforts.

[0042] Figure 1 It is a flowchart of an object annotation method based on a vision-language model and a region proposal network provided by an embodiment of the present invention;

[0043] Figure 2 It is a flowchart of another object annotation method based on a vision-language model and a region proposal network provided by an embodiment of the present invention;

[0044] Figure 3 Schematic structural diagram of an object annotation device based on a vision - language model and a region proposal network provided by an embodiment of the present invention;

[0045] Figure 4 Schematic structural diagram of a computer device provided by an embodiment of the present invention. Detailed implementation manners

[0046] Next, the technical solutions in the embodiments of the present invention will be clearly and completely described in conjunction with the accompanying drawings in the embodiments of the present invention. Obviously, the described embodiments are only a part of the embodiments of the present invention, rather than all of the embodiments. All other embodiments obtained by those of ordinary skill in the art based on the embodiments of the present invention without creative efforts shall fall within the protection scope of the present invention.

[0047] It should be noted that a method and device for object annotation based on a vision - language model and a region proposal network disclosed in the present application can be used in the field of artificial intelligence technology, and can also be used in any field other than the field of artificial intelligence technology. The application field of the method and device for object annotation based on a vision - language model and a region proposal network disclosed in the present application is not limited.

[0048] To facilitate the understanding of the technical solution provided by this application, the relevant content of the technical solution of this application will be described first. With the development of deep learning technology, especially the application of visual language models (VLM) and region proposal networks (RPN), automatic annotation methods have gradually become more intelligent and automated. By combining natural language with visual information, VLM can provide richer context information and has certain open-set reasoning capabilities. For example, VLMs such as Grounding Language-Image Pre-training (GLIP), Contrastive Language-Image Pre-training (CLIP), and Bootstrapping Language-Image Pre-training (BLIP) can achieve image annotation through the similarity between images and natural language descriptions. The core advantage of this method is that VLM can generate an understanding of images through natural language prompts. For example, users can input a prompt like "Please tell me what objects are in this picture", and VLM can identify and list the names of the objects in the image. The advantage of this method is that it can handle open-set objects, that is, VLM can identify objects outside the training set, thus avoiding the limitation that traditional methods can only identify object categories in the training data.

[0049] For example, GLIP achieves stronger cross-modal understanding by fusing the features of language and images, and can perform object recognition and localization without explicit labels. GLIP first pre-trains on large-scale data and then uses images and corresponding language prompts for reasoning, thus achieving the recognition of unknown objects.

[0050] RPN is a network used to automatically generate candidate regions for objects and is often used in combination with object detection models (such as Faster R-CNN). RPN generates a series of candidate regions through the feature map in the image, and these candidate regions may contain the objects in the image. Then, the object detection model further classifies and regresses these candidate regions to determine the position and category of the object. The advantage of RPN is that it can effectively narrow down the area to be detected, reduce the computational amount, and improve the accuracy of object detection. In the automatic annotation system of this application, RPN can not only be used to generate candidate regions for objects, but also be combined with visual language models to assist the annotation process. By matching the object names generated by VLM with the candidate regions generated by RPN, the regions in the image can be effectively corresponded to the object names, thus achieving accurate automatic annotation.

[0051] Taking the object annotation device based on the vision - language model and the region proposal network as an example of the execution subject, the implementation process of the object annotation method based on the vision - language model and the region proposal network provided by the embodiments of the present invention is described. It can be understood that the execution subject of the object annotation method based on the vision - language model and the region proposal network provided by the embodiments of the present invention includes, but is not limited to, the object annotation device based on the vision - language model and the region proposal network.

[0052] Figure 1 The following is a flowchart of an object annotation method based on a vision - language model and a region proposal network provided by an embodiment of the present invention. As Figure 1 shown, the method includes:

[0053] Step 101: Obtain the image to be annotated and the prompt statement.

[0054] In the embodiments of the present invention, the image to be annotated is an image that needs to be automatically annotated for objects, and the prompt statement is the content or requirement to be annotated input by the user. The prompt statement is in natural language.

[0055] Step 102: Through the VLM, perform multi - modal fusion on the image to be annotated and the prompt statement to generate the object name.

[0056] In the embodiments of the present invention, given the input image to be annotated and the prompt statement, the VLM fuses the information of the image to be annotated and the prompt statement, and identifies the objects in the image by calculating the similarity between the language and visual features. Specifically, the VLM generates the probability value of the object through the following formula:

[0057]

[0058] where P(Object i |I, Prompt) is the probability value that the object belongs to the i - th candidate object name; f VLM is the feature representation function of the VLM model; I is the image to be annotated; Prompt is the prompt statement; Object i is the i - th candidate object name.

[0059] In the embodiments of the present invention, the VLM calculates the similarity between the image to be annotated and each object name, outputs the probability value of each object, and finally determines the objects existing in the image according to the probability value.

[0060] For example: The image to be annotated is an image containing "apple" and "banana", and the prompt statement is: "Please tell me what objects are in this picture"; The image to be annotated and the prompt statement are input into the VLM, and the VLM outputs the object names through multi-modal fusion: O = {apple, banana}. Among them, O represents the set of object names. These object names will then be used in the subsequent region generation and matching processes.

[0061] Step 103: Through the RPN, perform regression classification on the image to be annotated according to the object names, and generate object category candidate regions.

[0062] In the embodiment of the present invention, the RPN traverses the feature map of the image to be annotated by means of a sliding window, and generates multiple candidate boxes with different scales and aspect ratios at each position. Each candidate box will be regressed with the target position of the actual object, and a target category will be assigned to each box. The goal of the RPN is to generate candidate regions of objects through two tasks of regression and classification. Specifically, the RPN generates the probability of matching each candidate region with the candidate object name through the following formula:

[0063]

[0064] where P(Object i |ROI j ) is the probability of the j-th candidate region matching the i-th candidate object name; f VLM is the feature representation function of the VLM model; f RPN is the feature representation function of the RPN model; ROI j is the j-th candidate region; Object i is the i-th candidate object name.

[0065] In the embodiment of the present invention, the RPN calculates the similarity between each candidate region and the object name, outputs the probability of matching each candidate region with the candidate object name, and finally determines the object category candidate region according to the matching probability, that is: the object name included in this candidate region.

[0066] For example: The image to be annotated is an image containing "apple" and "banana"; The RPN generates multiple candidate regions, and some of these candidate regions may contain apples, while others contain bananas. The RPN will match each candidate region with the candidate object name according to the features of the region, so as to select the regions that may contain the object name.

[0067] Step 104: Perform matching annotation according to the object name and the object category candidate region, and generate an object annotation image.

[0068] In the embodiments of the present invention, by interacting VLM with RPN, the object names generated by VLM are matched with the candidate regions generated by RPN to determine which candidate regions contain the objects output by VLM; by comparing the prompt information of VLM with the candidate regions generated by RPN, the targets in the image can be automatically labeled.

[0069] In the embodiments of the present invention, first, the candidate regions are matched with the object names to screen out the matching pairs where the candidate regions match the object names; through the following conditional probability formula, the matching degree of the matching pairs that match is calculated to generate a matching degree; and the candidate regions are labeled with objects according to this matching degree.

[0070]

[0071] Where P(Match i |ROI j ,Object i ) is the matching degree between the j-th candidate region and the i-th candidate object name; f VLM is the feature representation function of the VLM model; f RPN is the feature representation function of the RPN model; ROI j is the j-th candidate region; Object i is the i-th candidate object name.

[0072] For example: The image to be labeled is an image containing "apple" and "banana"; RPN generates multiple candidate regions, and a certain region is determined to contain "apple"; the matching degree between this candidate region and the object name "apple" will be calculated. If this matching degree is greater than the preset matching degree threshold, this region will be labeled as "apple". Similarly, if the matching degree between another candidate region determined to contain "banana" and the object name "banana" is greater than the preset matching degree threshold, it will be labeled as "banana".

[0073] In the technical solution provided by the embodiments of the present invention, the image to be labeled and the prompt statement are obtained; through the vision-language model, the image to be labeled and the prompt statement are subjected to multimodal fusion to generate object names; through the region proposal network, regression classification is performed on the image to be labeled according to the object names to generate candidate regions of object categories; and matching labeling is performed according to the object names and the candidate regions of object categories to generate an object-labeled image, which can effectively improve the accuracy of object detection. Especially when facing unseen objects, it can still maintain high detection performance and adapt to open-set detection; it can significantly reduce the cost of manual labeling, reduce the complexity and computational overhead of the labeling process, and at the same time improve the efficiency and accuracy of the labeling process, and has a wide range of application prospects.

[0074] Figure 2The flowchart of another object annotation method provided by an embodiment of the present invention based on a vision - language model and a region proposal network is as follows: Figure 2 As shown, the method includes:

[0075] Step 201, obtain the image to be annotated and the prompt statement.

[0076] In an embodiment of the present invention, each step is executed by an object annotation device based on a vision - language model and a region proposal network.

[0077] In an embodiment of the present invention, the image to be annotated is an image that needs to be automatically annotated for objects, and the prompt statement is the content or requirement to be annotated input by the user. The prompt statement is in natural language. For example, the image to be annotated is an image containing "apple" and "banana", and the prompt statement is: "Please tell me what objects are in this picture."

[0078] Step 202, perform feature extraction on the image to be annotated to generate visual features.

[0079] In an embodiment of the present invention, the main task of the VLM is to inform the names of objects existing in the image through natural - language prompts. Through a vision - language model (such as CLIP, BLIP, GLIP, etc.), by combining image and language information for cross - modal reasoning, the names of objects in the image can be effectively identified.

[0080] In an embodiment of the present invention, through the visual feature extraction module in the VLM, feature extraction is performed on the image to be annotated to extract the visual features of the annotated image for subsequent multi - modal information fusion with text features.

[0081] Step 203, perform feature extraction on the prompt statement to generate text features.

[0082] In an embodiment of the present invention, through the text feature extraction module in the VLM, feature extraction is performed on the prompt statement to extract text features for subsequent multi - modal information fusion with visual features.

[0083] Step 204, fuse the visual features and the text features to generate the object name.

[0084] In an embodiment of the present invention, step 204 specifically includes:

[0085] Step 2041, generate the object similarity between the image to be annotated and each preset candidate object name according to the visual features and the text features.

[0086] Specifically, through the methods of dot product and normalization, the visual features of the image to be annotated, the text features of the hint statement, and each candidate object name are used to calculate the similarity, generating the object similarity between the image to be annotated and each preset candidate object name.

[0087] In the embodiments of the present invention, the greater the object similarity, the greater the probability that the candidate object name exists in the image to be annotated; the smaller the object similarity, the smaller the probability that the candidate object name exists in the image to be annotated.

[0088] Step 2042: Determine the object name according to multiple object similarities.

[0089] As an alternative solution, each object similarity is compared with a preset similarity threshold, and the object similarities greater than the similarity threshold are screened out; the candidate object names corresponding to the screened-out object similarities are determined as the object names.

[0090] It should be noted that the similarity threshold can be set according to actual needs, and the embodiments of the present invention do not limit this.

[0091] Step 205: Extract the visual features of the image to be annotated through a convolutional neural network to generate a feature map.

[0092] In the embodiments of the present invention, the main task of the RPN is to generate candidate regions, identifying the regions in the image where objects may exist. The RPN is part of the region proposal network and is usually used in combination with an object detection model (such as Faster R-CNN). The RPN performs convolutional operations on the input image and outputs multiple candidate boxes (bounding boxes), which may contain the positions of objects.

[0093] In the embodiments of the present invention, in the RPN, the generation of the feature map usually depends on the basic convolutional neural network. The convolutional neural network is used to extract the visual features of the image to be annotated to generate a feature map, which is further used to generate candidate regions. These regions can be passed to the subsequent object detection network for precise positioning and classification. The size of the feature map is usually small, but it contains rich high-level semantic information.

[0094] Step 206: Traverse the feature map according to a preset sliding window to generate multiple object candidate boxes.

[0095] In the embodiments of the present invention, the size of the sliding window can be set according to actual needs, and the embodiments of the present invention do not limit this.

[0096] In the embodiments of the present invention, the feature map is traversed in the way of a sliding window, and multiple object candidate boxes with different scales and aspect ratios are generated at each position, indicating the positions of the objects.

[0097] Step 207: Classify according to the object candidate area and object name indicated by the object candidate box to determine the object category candidate area.

[0098] In the embodiment of the present invention, each object candidate box will be regressed to the target position of the actual object, and a target category will be assigned to each object candidate box. The goal of RPN is to generate candidate areas of objects through two tasks: regression and classification.

[0099] In the embodiment of the present invention, step 207 specifically includes:

[0100] Step 2071: Calculate the similarity between the object candidate area and object name indicated by the object candidate box to generate the matching probability between each object candidate area and the object name.

[0101] Specifically, the dot product and normalization methods are used to calculate the similarity between the object candidate area and the object name to generate the matching probability between each object candidate area and the object name.

[0102] In the embodiment of the present invention, the greater the matching probability, the greater the probability that the object candidate area and the object name match; the smaller the matching probability, the smaller the probability that the object candidate area and the object name match.

[0103] Step 2072: Determine the object category candidate area according to multiple matching probabilities.

[0104] In the embodiment of the present invention, multiple matching probabilities are compared to select the maximum matching probability; the object name corresponding to the maximum matching probability is determined as the object name included in the object candidate area, and the object candidate area including the object name is the object category candidate area.

[0105] Step 208: Determine whether the object name generated by VLM is the same as the object name included in the object category candidate area. If they are the same, execute step 209; if they are not the same, end the process.

[0106] In the embodiment of the present invention, if the object name generated by VLM is the same as the object name included in the object category candidate area, it indicates that the candidate area matches the object name, that is, the recognition result of VLM is the same as the regression classification result of RPN, and continue to execute step 209; if the object name generated by VLM is not the same as the object name included in the object category candidate area, it indicates that the candidate area does not match the object name, that is, the recognition result of VLM is the same as the regression classification result of RPN, and continue to determine whether the object name generated by the next group of VLM in the image to be labeled is the same as the object name included in the object category candidate area. If they are not the same, end the process.

[0107] Step 209: Calculate the similarity between the object name and the object category candidate region to generate the matching degree between the object name and the object category candidate region.

[0108] In the embodiment of the present invention, by using the dot product and normalization methods, the similarity between the matching object name and the object category candidate region is calculated to generate the matching degree between the object name and the object category candidate region.

[0109] In the embodiment of the present invention, the larger the matching degree, the higher the matching accuracy between the object name and the object category candidate region; the smaller the matching degree, the lower the matching accuracy between the object name and the object category candidate region.

[0110] Step 210: Determine whether the matching degree is greater than a preset matching degree threshold. If so, execute Step 211; if not, end the process.

[0111] In the embodiment of the present invention, if the matching degree is greater than the matching degree threshold, it indicates that the object included in the object category candidate region is the object name, and continue to execute Step 211; if the matching degree is less than or equal to the matching degree threshold, it indicates that the object included in the object category candidate region is not the object name, and continue to determine whether the matching degree between the next set of object name and the object category candidate region in the image to be labeled is greater than the matching degree threshold. If all are not, end the process.

[0112] It should be noted that the matching degree threshold can be set according to actual needs, and the embodiment of the present invention does not limit this.

[0113] Step 211: Label the object category candidate region as the object name on the image to be labeled to generate an object-labeled image.

[0114] In the embodiment of the present invention, the object category candidate region is labeled as the object name on the image to be labeled in a preset labeling manner to achieve automatic object labeling; the image after all objects are labeled is the object-labeled image.

[0115] It should be noted that the labeling method can be set according to actual needs, and the embodiment of the present invention does not limit this.

[0116] Further, multiple object-labeled images are combined to form a training set, which can be used for subsequent training of a machine learning model, reducing the training cost and improving the training efficiency.

[0117] By combining VLM and RPN, the system can more accurately locate objects and perform efficient matching based on object names and regional features, thereby improving the accuracy of detection. It can automatically generate object annotations, avoiding the high cost and cumbersome process of manual annotation, and significantly improving the accuracy and efficiency of the object detection system. Through the natural language prompt of VLM, the system can recognize objects outside the training set, thus having the ability of open-set detection, being able to handle new and unseen objects, and having strong robustness and generalization ability. This system can adapt to different application scenarios, can be used for the annotation of common objects, and can also handle unseen objects, with strong flexibility and scalability.

[0118] It should be noted that in the technical solutions of this application, the acquisition, storage, use, processing, etc. of data all comply with the relevant regulations of laws and regulations. The user information in the embodiments of this application is obtained through legal and compliant channels, and the acquisition, storage, use, processing, etc. of user information have obtained the authorization and consent of the customers.

[0119] It should be noted that the information collected in this application is information and data authorized by the user or fully authorized by all parties, and the processing of relevant data such as collection, storage, use, processing, transmission, provision, disclosure, and application complies with the relevant laws, regulations, and standards of relevant countries and regions, takes necessary confidentiality measures, does not violate public order and good customs, and provides corresponding operation entrances for users to choose to authorize or refuse.

[0120] It should be noted that the technical solution provided in this application provides corresponding operation entrances for users to choose to agree or refuse the results of automated decision-making; if the user chooses to refuse, the expert decision-making process will be entered.

[0121] In the technical solution of the object annotation method based on a vision language model and a region proposal network provided by the embodiments of the present invention, an image to be annotated and a prompt statement are obtained; through the vision language model, multi-modal fusion of the image to be annotated and the prompt statement is performed to generate an object name; through the region proposal network, regression classification of the image to be annotated is performed according to the object name to generate candidate regions for object categories; matching and annotation are performed according to the object name and the candidate regions for object categories to generate an object annotation image, which can effectively improve the accuracy of object detection, especially when facing unseen objects, still maintaining high detection performance and adapting to open-set detection; it can significantly reduce the cost of manual annotation, reduce the complexity of the annotation process and the computational overhead, and at the same time improve the efficiency and accuracy of the annotation process, and has broad application prospects.

[0122] Figure 3The structural schematic diagram of an object annotation device provided by an embodiment of the present invention. This device is used to execute the above-mentioned object annotation method based on a vision-language model and a region proposal network, as Figure 3 shown. The device includes: an acquisition unit 11, a multimodal fusion unit 12, a regression classification unit 13, and an object annotation unit 14.

[0123] The acquisition unit 11 is used to acquire the image to be annotated and the prompt statement.

[0124] The multimodal fusion unit 12 is used to perform multimodal fusion on the image to be annotated and the prompt statement through a vision-language model to generate the object name.

[0125] The regression classification unit 13 is used to perform regression classification on the image to be annotated according to the object name through a region proposal network to generate object category candidate regions.

[0126] The object annotation unit 14 is used to perform matching annotation based on the object name and the object category candidate regions to generate an object annotation image.

[0127] In an embodiment of the present invention, the multimodal fusion unit 12 is specifically used to extract features from the image to be annotated to generate visual features; extract features from the prompt statement to generate text features; and fuse the visual features and the text features to generate the object name.

[0128] In an embodiment of the present invention, the multimodal fusion unit 12 is specifically used to generate the object similarity between the image to be annotated and each preset candidate object name according to the visual features and the text features; and determine the object name according to multiple object similarities.

[0129] In an embodiment of the present invention, the regression classification unit 13 is specifically used to extract the visual features of the image to be annotated through a convolutional neural network to generate a feature map; traverse the feature map according to a preset sliding window to generate multiple object candidate boxes; and classify according to the object candidate regions indicated by the object candidate boxes and the object name to determine the object category candidate regions.

[0130] In an embodiment of the present invention, the regression classification unit 13 is specifically used to calculate the similarity between the object candidate regions indicated by the object candidate boxes and the object name to generate the matching probability between each object candidate region and the object name; and determine the object category candidate regions according to multiple matching probabilities.

[0131] In an embodiment of the present invention, the object annotation unit 14 is specifically configured to determine whether the object name generated by the vision-language model is consistent with the object name included in the object category candidate region; if they are consistent, calculate the similarity between the object name and the object category candidate region to generate a matching degree between the object name and the object category candidate region; if the matching degree is greater than a preset matching degree threshold, label the object category candidate region as the object name on the image to be annotated to generate an object-annotated image.

[0132] In the solution of the embodiment of the present invention, an image to be annotated and a prompt statement are obtained; through a vision-language model, multi-modal fusion is performed on the image to be annotated and the prompt statement to generate an object name; through a region proposal network, regression classification is performed on the image to be annotated according to the object name to generate an object category candidate region; and matching annotation is performed according to the object name and the object category candidate region to generate an object-annotated image, which can effectively improve the accuracy of object detection. Especially when facing unseen objects, it can still maintain high detection performance and adapt to open-set detection; it can significantly reduce the cost of manual annotation, reduce the complexity and computational overhead of the annotation process, and at the same time improve the efficiency and accuracy of the annotation process, and has a wide range of application prospects.

[0133] The system, apparatus, module or unit illustrated in the above embodiments can be specifically implemented by a computer chip or an entity, or by a product with certain functions. A typical implementation device is a computer device. Specifically, the computer device can be, for example, a personal computer, a laptop computer, a cellular phone, a camera phone, a smart phone, a personal digital assistant, a media player, a navigation device, an email device, a game console, a tablet computer, a wearable device, or any combination of these devices.

[0134] An embodiment of the present invention provides a computer device, including a memory and a processor. The memory is used to store information including program instructions, and the processor is used to control the execution of the program instructions. When the program instructions are loaded and executed by the processor, the steps of the above embodiment of the object annotation method based on a vision-language model and a region proposal network are implemented. For specific descriptions, reference can be made to the embodiment of the object annotation method based on a vision-language model and a region proposal network above.

[0135] Next, refer to Figure 4 , which shows a schematic structural diagram of a computer device 600 suitable for implementing an embodiment of the present application.

[0136] As Figure 4As shown, the computer device 600 includes a central processing unit (CPU) 601, which can perform various appropriate operations and processes according to the programs stored in the read-only memory (ROM) 602 or the programs loaded from the storage section 608 into the random access memory (RAM) 603. In the RAM 603, various programs and data required for the operation of the computer device 600 are also stored. The CPU 601, ROM 602, and RAM 603 are connected to each other via a bus 604. An input / output (I / O) interface 605 is also connected to the bus 604.

[0137] The following components are connected to the I / O interface 605: an input section 606 including a keyboard, a mouse, etc.; an output section 607 including, for example, a cathode ray tube (CRT), a liquid crystal display (LCD), etc. and a speaker, etc.; a storage section 608 including a hard disk, etc.; and a communication section 609 including a network interface card such as a LAN card, a modem, etc. The communication section 609 performs communication processing via a network such as the Internet. A drive 610 is also connected to the I / O interface 605 as needed. A removable medium 611, such as a magnetic disk, an optical disk, a magneto-optical disk, a semiconductor memory, etc., is installed on the drive 610 as needed so that a computer program read from it can be installed in the storage section 608 as needed.

[0138] In particular, according to an embodiment of the present invention, the process described above with reference to the flowchart can be implemented as a computer software program. For example, an embodiment of the present invention includes a computer program product, which includes a computer program tangibly embodied on a machine-readable medium, and the computer program includes program code for performing the method shown in the flowchart. In such an embodiment, the computer program can be downloaded and installed from a network via the communication section 609, and / or installed from the removable medium 611.

[0139] A computer-readable medium includes both permanent and non-permanent, removable and non-removable media and can implement information storage by any method or technology. The information can be computer-readable instructions, data structures, program modules, or other data. Examples of computer storage media include, but are not limited to, phase change memory (PRAM), static random access memory (SRAM), dynamic random access memory (DRAM), other types of random access memory (RAM), read-only memory (ROM), electrically erasable programmable read-only memory (EEPROM), flash memory or other memory technologies, compact disc read-only memory (CD-ROM), digital versatile disc (DVD) or other optical storage, magnetic cassette tapes, magnetic disk storage or other magnetic storage devices, or any other non-transitory medium that can be used to store information that can be accessed by a computing device. As defined herein, a computer-readable medium does not include transitory computer-readable media such as modulated data signals and carrier waves.

[0140] For convenience of description, when describing the above devices, they are described separately as various units according to their functions. Of course, when implementing the present application, the functions of each unit can be implemented in one or more software and / or hardware.

[0141] The present invention is described with reference to the flowcharts and / or block diagrams of methods, apparatuses (systems), and computer program products according to embodiments of the present invention. It should be understood that each process and / or block in the flowchart and / or block diagram, and the combination of processes and / or blocks in the flowchart and / or block diagram, can be implemented by computer program instructions. These computer program instructions can be provided to the processor of a general-purpose computer, a special-purpose computer, an embedded processor, or other programmable data processing device to generate a machine, such that the instructions executed by the processor of the computer or other programmable data processing device generate means for implementing the functions specified in Figure 1 one process or multiple processes and / or blocks Figure 1 one block or multiple blocks.

[0142] These computer program instructions can also be stored in a computer-readable memory that can direct a computer or other programmable data processing device to work in a specific manner, such that the instructions stored in the computer-readable memory generate a manufactured article including instruction means that implement the functions specified in Figure 1 one process or multiple processes and / or blocks Figure 1 one block or multiple blocks.

[0143] These computer program instructions can also be loaded onto a computer or other programmable data processing device, so that a series of operation steps are executed on the computer or other programmable device to generate a computer-implemented process. Thus, the instructions executed on the computer or other programmable device provide steps for implementing the functions specified in one process Figure 1 one process or multiple processes and / or blocks Figure 1 steps of the functions specified in one block or multiple blocks.

[0144] It should also be noted that the term "including", "comprising" or any other variant thereof is intended to cover non-exclusive inclusion, so that a process, method, commodity or device including a series of elements not only includes those elements, but also includes other elements not expressly listed, or also includes elements inherent in such process, method, commodity or device. Without further limitation, an element defined by the statement "including one..." does not exclude the presence of additional identical elements in the process, method, commodity or device including the said element.

[0145] In the technical solution of this application, the acquisition, storage, use, processing, etc. of data all comply with the relevant provisions of national laws and regulations.

[0146] It should be noted that in the embodiments of this application, some industry-existing solutions such as certain software, components, models, etc. may be mentioned. They should be regarded as exemplary. The purpose is only to illustrate the feasibility in the implementation of the technical solution of this application, but it does not mean that the applicant has already or necessarily used this solution.

[0147] Those skilled in the art should understand that the embodiments of this application can be provided as a method, a system or a computer program product. Therefore, this application can take the form of a complete hardware embodiment, a complete software embodiment or an embodiment combining software and hardware aspects. Moreover, this application can take the form of a computer program product implemented on one or more computer-usable storage media (including but not limited to disk storage, CD-ROM, optical storage, etc.) containing computer-usable program code.

[0148] This application can be described in the general context of computer-executable instructions executed by a computer, such as program modules. Generally, program modules include routines, programs, objects, components, data structures, etc. that perform specific tasks or implement specific abstract data types. This application can also be practiced in a distributed computing environment, where tasks are performed by remote processing devices connected through a communication network. In a distributed computing environment, program modules can be located in local and remote computer storage media including storage devices.

[0149] Each embodiment in this specification is described in a progressive manner. For the same or similar parts among the embodiments, reference can be made to each other, and the key point of each embodiment is to illustrate the differences from other embodiments. In particular, for the system embodiment, since it is basically similar to the method embodiment, the description is relatively simple, and for the relevant parts, reference can be made to the partial description of the method embodiment.

[0150] The above are only the embodiments of the present application and are not intended to limit the present application. For those skilled in the art, various modifications and changes can be made to the present application. Any modification, equivalent replacement, improvement, etc. made within the spirit and principle of the present application shall be included within the scope of the claims of the present application.

Claims

1. An object annotation method based on a vision-language model and a region proposal network, characterized in that, The method includes: Obtaining an image to be annotated and a prompt statement; Performing multimodal fusion on the image to be annotated and the prompt statement through a vision-language model to generate an object name; Performing regression classification on the image to be annotated according to the object name through a region proposal network to generate candidate regions for object categories; Performing matching annotation based on the object name and the candidate regions for object categories to generate an object-annotated image.

2. The object annotation method based on a vision-language model and a region proposal network according to claim 1, wherein The performing multimodal fusion on the image to be annotated and the prompt statement through a vision-language model to generate an object name includes: Extracting features from the image to be annotated to generate visual features; Extracting features from the prompt statement to generate text features; Fusing the visual features and the text features to generate an object name.

3. The object annotation method based on a vision language model and a region proposal network according to claim 2, wherein The fusing the visual features and the text features to generate an object name includes: Generating an object similarity between the image to be annotated and each preset candidate object name according to the visual features and the text features; Determining the object name according to multiple such object similarities.

4. The object annotation method based on a vision-language model and a region proposal network according to claim 1, characterized in that, The performing regression classification on the image to be annotated according to the object name through a region proposal network to generate candidate regions for object categories includes: Extracting visual features of the image to be annotated through a convolutional neural network to generate a feature map; Traversing the feature map according to a preset sliding window to generate multiple object candidate boxes; Classifying according to the object candidate regions indicated by the object candidate boxes and the object name to determine the candidate regions for object categories.

5. The object annotation method based on a vision-language model and a region proposal network according to claim 4, characterized in that The classifying according to the object candidate regions indicated by the object candidate boxes and the object name to determine the candidate regions for object categories includes: Calculating a similarity between the object candidate regions indicated by the object candidate boxes and the object name to generate a matching probability between each object candidate region and the object name; Determining the candidate regions for object categories according to multiple such matching probabilities.

6. The object annotation method based on a vision-language model and a region proposal network according to claim 1, wherein, The performing matching annotation based on the object name and the candidate regions for object categories to generate an object-annotated image includes: Judging whether the object name generated by the vision-language model is consistent with the object name included in the candidate regions for object categories; If they are consistent, calculating a similarity between the object name and the candidate regions for object categories to generate a matching degree between the object name and the candidate regions for object categories; If the matching degree is greater than a preset matching degree threshold, annotating the candidate regions for object categories as the object name on the image to be annotated to generate the object-annotated image.

7. An object annotation device based on a vision language model and a region proposal network, characterized in that, The device includes: An acquisition unit, configured to obtain an image to be annotated and a prompt statement; A multimodal fusion unit, configured to perform multimodal fusion on the image to be annotated and the prompt statement through a vision-language model to generate an object name; A regression classification unit, configured to perform regression classification on the image to be annotated according to the object name through a region proposal network to generate candidate regions for object categories; An object annotation unit, configured to perform matching annotation based on the object name and the candidate regions for object categories to generate an object-annotated image.

8. The object annotation device based on the vision language model and the region proposal network according to claim 7, wherein, The multi-modal fusion unit is specifically configured to extract features from the image to be annotated to generate visual features; extract features from the prompt statement to generate text features; and fuse the visual features and the text features to generate an object name.

9. The object annotation device based on a vision language model and a region proposal network according to claim 8, wherein The multi-modal fusion unit is specifically configured to generate an object similarity between the image to be annotated and each preset candidate object name according to the visual features and text features; and determine the object name according to the multiple object similarities.

10. The object annotation device based on a vision language model and a region proposal network according to claim 7, characterized in that The regression classification unit is specifically configured to extract visual features of the image to be annotated through a convolutional neural network to generate a feature map; traverse the feature map according to a preset sliding window to generate a plurality of object candidate boxes; classify according to the object candidate region and the object name indicated by the object candidate box to determine the object category candidate region.

11. The object annotation device based on a vision language model and a region proposal network according to claim 10, wherein The regression classification unit is specifically configured to calculate a similarity between the object candidate region and the object name indicated by the object candidate box to generate a matching probability between each object candidate region and the object name; and determine the object category candidate region according to the multiple matching probabilities.

12. The object annotation device based on a vision-language model and a region proposal network according to claim 7, wherein The object annotation unit is specifically configured to determine whether the object name generated by the vision-language model is consistent with the object name included in the object category candidate region; if they are consistent, calculate a similarity between the object name and the object category candidate region to generate a matching degree between the object name and the object category candidate region; if the matching degree is greater than a preset matching degree threshold, label the object category candidate region as the object name on the image to be annotated to generate the object-annotated image.

13. A computer-readable medium having a computer program stored thereon, characterized in that, When the program is executed by a processor, it implements the object annotation method based on a vision-language model and a region proposal network according to any one of claims 1 to 6.

14. A computer device, comprising a memory and a processor, the memory being used for storing information including program instructions, and the processor being used for controlling the execution of the program instructions, characterized in that, When the program instructions are loaded and executed by a processor, it implements the object annotation method based on a vision-language model and a region proposal network according to any one of claims 1 to 6.

15. A computer program product, comprising computer programs / instructions, characterized in that, When the computer program / instructions are executed by a processor, it implements the object annotation method based on a vision-language model and a region proposal network according to any one of claims 1 to 6.