Image labeling method and related device

Through an image labeling method based on description prompt information and visual question-and-answer questions, the problem of insufficient dependence on sample images and labeling accuracy in the prior art is solved, and efficient and accurate image labeling is achieved.

CN120260040APending Publication Date: 2025-07-04TENCENT TECHNOLOGY (SHENZHEN) CO LTD
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202410011058.X
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2024-01-03
Publication Date
2025-07-04

AI Technical Summary

Technical Problem

Existing image annotation technology relies on training of large-scale sample images and the labeling quality depends on sample quality, resulting in unstable model performance and insufficient labeling accuracy.

Method used

By describing the image content based on preset description prompt information, constructing target description information, and tag summary with visual question-and-answer questions, reducing dependence on sample images and realizing automated annotation.

Benefits of technology

It improves the accuracy and efficiency of image labeling, reduces resource consumption, avoids labeling errors caused by unstable sample labeling quality, and realizes automatic labeling of image labels.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120260040A_ABST
    Figure CN120260040A_ABST
Patent Text Reader

Abstract

The invention relates to the field of image processing, and provides an image annotation method and a related device, which are used for improving the image annotation efficiency, and the method comprises the steps: carrying out the image content description of a to-be-annotated image based on preset description prompt information, obtaining a corresponding content description result, and carrying out the annotation of the to-be-annotated image based on the obtained content description result. Constructing target description information of the to-be-labeled image; based on the target description information, combining a preset summary prompt word to construct a visual question and answer question; and based on the visual question and answer question and the to-be-labeled image, performing label summary on the target description information to obtain a labeling result of the to-be-labeled image. The embodiment of the invention can be applied to various scenes such as cloud technology, artificial intelligence, intelligent traffic, auxiliary driving and the like. According to the method, the image content of the to-be-annotated image is summarized, and the label is extracted according to the summarized content, so that automatic annotation of the to-be-annotated image is realized, computing resource consumption is reduced, and image annotation efficiency is improved.
Need to check novelty before this filing date? Find Prior Art

Description

Background Art

[0002] In image annotation, it is necessary to express as much image content information as possible with a small number of tags, making full use of the semantic relationships between the tags, so that the annotation result is closer to the image content.

[0003] In related technologies, usually based on each labeled sample image, machine learning technology is used to train an image annotation model to obtain a trained image annotation model. Then, the trained image annotation model is used to predict the label set corresponding to the image to be annotated.

[0004] However, first, the model requires a large number of sample images for supervised training. However, in actual business, it is very difficult to collect labeled sample images, which requires a lot of time and resources. And a small number of sample images are difficult to meet the model's needs, which will cause underfitting, thus affecting the model performance and further affecting the accuracy of image annotation. Second, the annotation effect of the model largely depends on the annotation quality of the sample images. If the sample images are marked incorrectly or insufficiently, it will affect the model performance, and the model cannot recognize labels other than the labels of each sample image, making it difficult to ensure the accuracy of image annotation. Summary of the Invention

[0005] Embodiments of the present application provide an image annotation method and related devices to improve the accuracy of image annotation.

[0006] In a first aspect, embodiments of the present application provide an image annotation method, including:

[0007] Obtain an image to be annotated;

[0008] Based on each preset description prompt information, respectively describe the content of the image to be annotated to obtain corresponding content description results, and based on the obtained content description results, construct target description information of the image to be annotated;

[0009] Based on the target description information, combined with a preset summary prompt word, construct a visual question and answer question;

[0010] Based on the visual question and answer question and the image to be annotated, summarize the target description information to obtain an annotation result of the image to be annotated.

[0011] As a possible implementation manner, the constructing the target description information of the image to be annotated based on the obtained content description results includes:

[0012] Based on the obtained content description results, construct initial description information of the image to be annotated;

[0013] Based on each content query question set for each key content, content queries are respectively combined with the image to be annotated to obtain corresponding answer results;

[0014] Based on the obtained answer results of each item, combined with the initial description information, the target description information of the image to be annotated is obtained.

[0015] As a possible implementation manner, before the content queries are respectively combined with the image to be annotated based on each content query question for each key content to obtain corresponding content answer results, it further includes:

[0016] Obtain the key contents set for the target application scenario;

[0017] Based on the information types corresponding to the key contents respectively, combined with the question templates respectively set for each information type, obtain each content query question set for the key contents.

[0018] As a possible implementation manner, the obtaining of each content query question set for the key contents based on the information types corresponding to the key contents respectively and combined with the question templates respectively set for each information type includes:

[0019] If the information type of one key content among the key contents is a noun, then adopt the question template set for nouns, and based on the one key content, obtain the corresponding content query question;

[0020] If the information type of one key content among the key contents is a verb, then adopt the question template set for verbs, and based on the one key content, obtain the corresponding content query question.

[0021] As a possible implementation manner, the obtaining of the annotation result of the image to be annotated by performing label summarization on the target description information based on the visual question and the image to be annotated includes:

[0022] Based on the visual question and the image to be annotated, perform label summarization on the target description information to obtain a candidate label set of the image to be annotated;

[0023] Based on the graphic-text matching degrees between the candidate labels in the candidate label set and the image to be annotated respectively, screen at least one target label that meets the set screening conditions from the candidate labels;

[0024] Take the at least one target label as the annotation result of the image to be annotated.

[0025] As a possible implementation, screening at least one target label that meets the set screening conditions from the candidate labels based on the graphic-text matching degree between each candidate label in the candidate label set and the image to be labeled includes:

[0026] Sorting the candidate labels based on the graphic-text matching degree between each candidate label in the candidate label set and the image to be labeled to obtain a label sorting result;

[0027] Based on the label sorting result, at least one candidate label is sequentially screened from the candidate labels according to the set label screening quantity threshold, and the at least one candidate label screened out is used as at least one target label that meets the set screening conditions.

[0028] As a possible implementation, the graphic-text matching degree between a candidate label and the image to be labeled is obtained in the following manner:

[0029] Inputting the initial label and the image to be labeled into a graphic-text matching model to obtain the graphic-text matching degree between the candidate label and the image to be labeled, where the graphic-text matching model is pre-trained using each first sample image.

[0030] As a possible implementation, respectively describing the content of the image to be labeled based on the preset description prompt messages to obtain corresponding content description results includes:

[0031] Respectively combining the description prompt messages with the image to be labeled and inputting them into an image content description model to obtain corresponding image description content, where the image content description model is pre-trained using each second sample image;

[0032] Summarizing the target description information based on the visual question and the image to be labeled to obtain the labeling result of the image to be labeled includes:

[0033] Combining the visual question with the image to be labeled and inputting them into a visual question answering model to obtain the labeling result of the image to be labeled, where the visual question answering model is pre-trained using each third sample image.

[0034] As a possible implementation, the graphic-text matching model, the image content description model, and the visual question answering model are the same model.

[0035] As a possible implementation, it further includes:

[0036] Based on the annotation result of the image to be annotated and combined with the object information of the target object, a recommendation result of the image to be annotated is obtained, where the recommendation result is used to indicate whether to recommend the image to be annotated to the target object; or,

[0037] Based on the annotation result of the image to be annotated, an anomaly recognition result of the image to be annotated is obtained, where the anomaly recognition result is used to indicate whether the image to be annotated contains abnormal content.

[0038] In a second aspect, an embodiment of the present application provides an image annotation device, including:

[0039] An image acquisition unit, configured to acquire an image to be annotated;

[0040] A content description unit, configured to respectively perform image content description on the image to be annotated based on each preset description prompt information, obtain corresponding content description results, and construct target description information of the image to be annotated based on the obtained content description results;

[0041] A question construction unit, configured to construct a visual question and answer question based on the target description information and in combination with a preset summary prompt word;

[0042] A content summary unit, configured to perform label summary on the target description information based on the visual question and answer question and the image to be annotated, and obtain an annotation result of the image to be annotated.

[0043] As a possible implementation manner, when constructing the target description information of the image to be annotated based on the obtained content description results, the content description unit is specifically configured to:

[0044] Construct initial description information of the image to be annotated based on the obtained content description results;

[0045] Based on each content query question set for each key content, respectively perform content query in combination with the image to be annotated, and obtain corresponding answer results;

[0046] Based on the obtained answer results and in combination with the initial description information, obtain the target description information of the image to be annotated.

[0047] As a possible implementation manner, before respectively performing content query in combination with the image to be annotated based on each content query question of each key content to obtain corresponding content answer results, the content description unit is further configured to:

[0048] Obtain each key content set for a target application scenario;

[0049] Based on the information types corresponding to the respective key contents, in combination with the question templates respectively set for each information type, content inquiry questions set for the respective key contents are obtained.

[0050] As a possible implementation manner, when obtaining the content inquiry questions set for the respective key contents based on the information types corresponding to the respective key contents and in combination with the question templates respectively set for each information type, the content description unit is specifically configured to:

[0051] If the information type of one of the key contents among the respective key contents is a noun, then the question template set for nouns is adopted, and based on the one key content, the corresponding content inquiry question is obtained;

[0052] If the information type of one of the key contents among the respective key contents is a verb, then the question template set for verbs is adopted, and based on the one key content, the corresponding content inquiry question is obtained.

[0053] As a possible implementation manner, when summarizing the target description information based on the visual question and answer question and the image to be annotated to obtain the annotation result of the image to be annotated, the content summarization unit is specifically configured to:

[0054] Based on the visual question and answer question and the image to be annotated, summarize the target description information to obtain a candidate label set of the image to be annotated;

[0055] Based on the graphic-text matching degree between each candidate label in the candidate label set and the image to be annotated, at least one target label that meets the set screening conditions is screened out from the respective candidate labels;

[0056] Use the at least one target label as the annotation result of the image to be annotated.

[0057] As a possible implementation manner, when screening at least one target label that meets the set screening conditions from the respective candidate labels based on the graphic-text matching degree between each candidate label in the candidate label set and the image to be annotated, the content summarization unit is specifically configured to:

[0058] Based on the graphic-text matching degree between each candidate label in the candidate label set and the image to be annotated, sort the respective candidate labels to obtain a label sorting result;

[0059] Based on the label sorting result, according to the set label screening quantity threshold, at least one candidate label is sequentially screened out from the respective candidate labels, and the at least one candidate label screened out is used as at least one target label that meets the set screening conditions.

[0060] As a possible implementation, the content summarization unit is further configured to:

[0061] Input the one candidate tag and the image to be annotated into a text-image matching model to obtain the text-image matching degree between the one candidate tag and the image to be annotated, where the text-image matching model is obtained by pre-training with each first sample image.

[0062] As a possible implementation, when the content description unit respectively performs image content descriptions on the image to be annotated based on each preset description prompt information to obtain corresponding content description results, the content description unit specifically is configured to:

[0063] Input each of the description prompt information and the image to be annotated into an image content description model to obtain corresponding image description content, where the image content description model is obtained by pre-training with each second sample image;

[0064] When the content summarization unit performs tag summarization on the target description information based on the visual question and the image to be annotated to obtain the annotation result of the image to be annotated, the content summarization unit specifically is configured to:

[0065] Input the visual question and the image to be annotated into a visual question answering model to obtain the annotation result of the image to be annotated, where the visual question answering model is obtained by pre-training with each third sample image.

[0066] As a possible implementation, the image content description model and the visual question answering model are the same model.

[0067] As a possible implementation, the content summarization unit is further configured to:

[0068] Based on the annotation result of the image to be annotated and in combination with the object information of the target object, obtain a recommendation result of the image to be annotated, where the recommendation result is used to represent whether to recommend the image to be annotated to the target object; or,

[0069] Based on the annotation result of the image to be annotated, obtain an anomaly recognition result of the image to be annotated, where the anomaly recognition result is used to represent whether the image to be annotated contains abnormal content.

[0070] In a third aspect, an embodiment of the present application provides an electronic device, including a processor and a memory, where the memory stores a computer program, and when the computer program is executed by the processor, the processor is caused to execute the steps of the above method.

[0071] Fourthly, an embodiment of the present application provides a computer-readable storage medium, which includes a computer program. When the computer program runs on an electronic device, the computer program is used to cause the electronic device to execute the steps of the method in any of the above aspects.

[0072] Fifthly, an embodiment of the present application provides a computer program product. The program product includes a computer program. The computer program is stored in a computer-readable storage medium. A processor of an electronic device reads and executes the computer program, so that the electronic device executes the steps of the method in any of the above aspects.

[0073] In the embodiment of the present application, based on each preset description prompt information, the acquired images to be labeled are respectively described in terms of image content to obtain corresponding content description results. Then, based on the obtained content description results, the target description information of the images to be labeled is constructed. Next, based on the target description information and in combination with the preset summary prompt words, visual question and answer questions are constructed. Then, based on the visual question and answer questions and the images to be labeled, the target description information is summarized by tags to obtain the labeling results of the images to be labeled.

[0074] By establishing the image content description of the images to be labeled and summarizing the image content description, the labeling results are obtained. There is no need for manual labeling and model training, which reduces resource consumption, improves the labeling efficiency, effectively reduces costs, and at the same time does not depend on the label annotation quality of sample images, ensuring the accuracy of image annotation and realizing the automatic annotation of image labels, thereby improving the labeling efficiency of image labels. In addition, in the case of performing multiple image content descriptions on the images to be labeled, compared with performing one image content description, it can also avoid to a certain extent the problem that the target description information is inaccurate due to the randomness and instability of the description results, thereby improving the description accuracy of the target description information. Furthermore, when summarizing the tags based on the target description information subsequently, the quality of image annotation is ensured.

[0075] Other features and advantages of the present application will be described in the subsequent description. And, some of them will become obvious from the description, or be understood by implementing the present application. The objectives and other advantages of the present application can be realized and obtained by the structures specifically pointed out in the written description, claims, and drawings. BRIEF DESCRIPTION OF THE DRAWINGS

[0076] The drawings described herein are used to provide a further understanding of the present application and constitute a part of the present application. The schematic embodiments of the present application and their descriptions are used to explain the present application and do not constitute an improper limitation to the present application. In the drawings:

[0077] Figure 1 It is a schematic diagram of an application scenario provided in an embodiment of the present application;

[0078] Figure 2 It is a schematic flowchart of an image annotation method provided in an embodiment of the present application;

[0079] Figure 3 It is a schematic logical diagram of an image content description process provided in an embodiment of the present application;

[0080] Figure 4 It is a schematic logical diagram of an image information supplement process provided in an embodiment of the present application;

[0081] Figure 5 It is a schematic logical diagram of an image content summary process provided in an embodiment of the present application;

[0082] Figure 6 It is a schematic logical diagram of a target label screening process provided in an embodiment of the present application;

[0083] Figure 7 It is a schematic diagram of a visual question answering task provided in an embodiment of the present application;

[0084] Figure 8 It is a schematic diagram of an image content description task provided in an embodiment of the present application;

[0085] Figure 9 It is a schematic diagram of a graphic-text matching task provided in an embodiment of the present application;

[0086] Figure 10 It is a schematic logical diagram of an image annotation process provided in an embodiment of the present application;

[0087] Figure 11 It is a schematic structural diagram of an image annotation device provided in an embodiment of the present application;

[0088] Figure 12 It is a schematic structural diagram of an electronic device provided in an embodiment of the present application. Detailed implementation manners

[0089] To make the objectives, technical solutions, and advantages of the embodiments of the present application clearer, the technical solutions of the present application will be clearly and completely described below in conjunction with the accompanying drawings in the embodiments of the present application. Apparently, the described embodiments are some, rather than all, of the embodiments of the present application. All other embodiments obtained by those of ordinary skill in the art based on the embodiments recorded in this application document without creative efforts shall fall within the scope of protection of the technical solutions of the present application.

[0090] In the embodiments of the present application, the term "module" or "unit" refers to a computer program with a predetermined function or a part of a computer program, which works together with other related parts to achieve a predetermined goal, and can be fully or partially implemented by using software, hardware (such as a processing circuit or a memory), or a combination thereof. Similarly, one processor (or multiple processors or memories) can be used to implement one or more modules or units. In addition, each module or unit can be a part of the overall module or unit that includes the function of the module or unit.

[0091] It can be understood that in the specific implementation of the present application, when it comes to data such as the image to be annotated, when the above embodiments of the present application are applied to specific products or technologies, user permission or consent needs to be obtained, and the collection, use, and processing of relevant data need to comply with the relevant laws, regulations, and standards of relevant countries and regions.

[0092] Artificial Intelligence (AI) is the theory, method, technology, and application system that uses digital computers or machines controlled by digital computers to simulate, extend, and expand human intelligence, perceive the environment, acquire knowledge, and use knowledge to obtain the best results. In other words, artificial intelligence is a comprehensive technology in computer science. It attempts to understand the essence of intelligence and produce a new intelligent machine that can react in a way similar to human intelligence. Artificial intelligence also studies the design principles and implementation methods of various intelligent machines, enabling the machines to have the functions of perception, reasoning, and decision-making.

[0093] Artificial intelligence technology is an interdisciplinary subject, covering a wide range of fields, including both hardware-level technologies and software-level technologies. The basic technologies of artificial intelligence generally include sensors, dedicated artificial intelligence chips, cloud computing, distributed storage, big data processing technology, pre-trained model technology, operation / interaction systems, mechatronics, etc. Among them, the pre-trained model, also known as the large model or the foundation model, can be widely applied to downstream tasks in various directions of artificial intelligence after fine-tuning. The artificial intelligence software technology mainly includes several major directions such as computer vision technology, speech processing technology, natural language processing technology, and machine learning / deep learning.

[0094] Computer Vision Technology (CV) is a science that studies how to enable machines to "see". Further speaking, it refers to machine vision that uses cameras and computers to replace human eyes for object recognition, monitoring, and measurement, and further performs graphics processing to make the computer-processed images more suitable for human eyes to observe or be transmitted to instruments for detection. As a scientific discipline, computer vision studies related theories and technologies and attempts to build artificial intelligence systems that can obtain information from images or multi-dimensional data. Large model technology has brought important changes to the development of computer vision technology. Pretrained models in the field of vision such as Swin Transformer, Vision Transformer (ViT), Vision MoE (V-MoE) based on Mixture of Experts (MoE), and Masked Autoencoders (MAE) can be quickly and widely applied to specific downstream tasks after fine-tuning. Computer vision technology usually includes image processing, image recognition, image semantic understanding, image retrieval, Optical Character Recognition (OCR), video processing, video semantic understanding, video content / behavior recognition, 3D object reconstruction, 3D technology, virtual reality, augmented reality, simultaneous localization and mapping, etc. It also includes common biometric recognition technologies such as face recognition and fingerprint recognition.

[0095] The key technologies of Speech Technology include Automatic Speech Recognition (ASR), Text-to-Speech (TTS), and voiceprint recognition technology. Enabling computers to listen, see, speak, and feel is the future development direction of human-computer interaction. Among them, speech has become one of the most promising human-computer interaction methods. Large model technology has brought changes to the development of speech technology. Pretrained models that follow the Transformer architecture such as WavLM and UniSpeech have strong generalization and versatility and can excellently complete speech processing tasks in various directions.

[0096] Natural Language Processing (NLP) is an important direction in the fields of computer science and artificial intelligence. It studies various theories and methods that can enable effective communication between humans and computers in natural language. Natural language processing involves natural language, that is, the language people use in daily life, and is closely related to linguistic research; at the same time, it involves disciplines such as computer science and mathematics. The pre-trained model, an important technology for model training in the field of artificial intelligence, has evolved from the large language model in the NLP field. After fine-tuning, the large language model can be widely applied to downstream tasks. Natural language processing technologies usually include text processing, semantic understanding, machine translation, robot question answering, knowledge graphs and other technologies.

[0097] Machine Learning (ML) is an interdisciplinary subject that involves multiple disciplines such as probability theory, statistics, approximation theory, convex analysis, and algorithm complexity theory. It specifically studies how computers simulate or implement human learning behaviors to acquire new knowledge or skills and reorganize the existing knowledge structure to continuously improve their own performance. Machine learning is the core of artificial intelligence and the fundamental way to make computers intelligent. Its applications cover all fields of artificial intelligence. Machine learning and deep learning usually include technologies such as artificial neural networks, belief networks, reinforcement learning, transfer learning, inductive learning, and rote learning. The pre-trained model is the latest development result of deep learning, integrating the above technologies.

[0098] Autopilot technology refers to the vehicle's ability to drive itself without driver operation. It usually includes technologies such as high-precision maps, environmental perception, computer vision, behavior decision-making, path planning, and motion control. Autopilot includes various development paths such as single-vehicle intelligence, vehicle-road cooperation, and networked cloud control. Autopilot technology has broad application prospects. Currently, in addition to the fields of logistics, public transportation, taxis, and intelligent transportation, it will be further developed in the future.

[0099] With the research and progress of artificial intelligence technology, artificial intelligence technology has been studied and applied in multiple fields. For example, common ones include smart homes, smart wearable devices, virtual assistants, smart speakers, smart marketing, driverless, autopilot, drones, digital twins, virtual humans, robots, artificial intelligence-generated content (AIGC), conversational interactions, intelligent healthcare, intelligent customer service, game AI, etc. It is believed that with the development of technology, artificial intelligence technology will be applied in more fields and play an increasingly important role.

[0100] Pre-training model, also known as cornerstone model or big model, refers to a deep neural network (DNN) with large parameters. It is trained on massive unlabeled data. The function approximation ability of large-parameter DNN is used to enable PTM to extract common features from the data. After fine tuning, parameter efficient fine tuning (PEFT), prompt-tuning and other technologies, it is suitable for downstream tasks. Therefore, the pre-training model can achieve ideal results in few-shot or zero-shot scenarios. PTM can be divided into language models (ELMO, BERT, GPT), visual models (swin-transformer, ViT, V-MOE), speech models (VALL-E), multimodal models (ViBERT, CLIP, Flamingo, Gato), etc. according to the data modality processed. Among them, the multimodal model refers to a model that establishes two or more data modality feature representations. The pre-training model is an important tool for outputting artificial intelligence generated content (AIGC), and can also be used as a general interface to connect multiple specific task models.

[0101] The solution provided in the embodiment of the present application involves the application of a large model, specifically: using a pre-trained large visual language cross-modal model to achieve image annotation. Among them, the large visual language cross-modal model is a model that can support visual question-answering tasks, image content description tasks, and image-text matching tasks, has large-scale parameters, and is pre-trained on large-scale samples. The specific model application process is described below and will not be repeated here.

[0102] In addition, the embodiments of the present application also involve prompt engineering, which is a technical and engineering method for designing and optimizing the input prompt words of the large model in order to efficiently and appropriately apply specific tasks to the large model. In the embodiments of the present application, the specific design is based on the input prompt words such as the description prompt information (i.e., the description prompt words), indicating what actions the visual language cross-modal large model should take or what output it should generate when performing specific tasks (such as visual question-answering tasks, image content description tasks, and image-text matching tasks).

[0103] See also Figure 1 As shown, it is a schematic diagram of an application scenario provided in an embodiment of the present application. The application scenario includes a terminal device 110 and a server 120. The number of terminal devices 110 can be one or more. The number of servers 120 can also be one or more. The present application does not specifically limit the number of terminal devices 110 and servers 120.

[0104] In the embodiments of the present application, the terminal device 110 may be a smart phone, a tablet computer, a laptop computer, a desktop computer, a smart speaker, a smart watch, an Internet of Things device, a smart voice interaction device, a smart home appliance, a vehicle terminal, an aircraft, etc., but is not limited thereto. The terminal device 110 supports a client for content recommendation, and the server 120 is the background server corresponding to the client.

[0105] The server 120 may be an independent physical server, or a server cluster or a distributed system composed of multiple physical servers, or a cloud server that provides basic cloud computing services such as cloud services, cloud databases, cloud computing, cloud functions, cloud storage, network services, cloud communications, middleware services, domain name services, security services, Content Delivery Network (CDN), and big data and artificial intelligence platforms.

[0106] The terminal device 110 and the server 120 may be directly or indirectly connected through wired or wireless communication methods, and the present application does not limit this.

[0107] In some embodiments, the server 120 obtains an image to be annotated; based on each preset description prompt information, respectively describes the content of the image to be annotated to obtain corresponding content description results, and based on the obtained content description results, constructs target description information of the image to be annotated; based on the target description information, combines preset summary prompt words to construct a visual question and answer question; based on the visual question and answer question and the image to be annotated, summarizes the target description information to obtain an annotation result of the image to be annotated.

[0108] It should be noted that the image annotation method mentioned in the embodiments of the present application may be executed by the server or the terminal device, or jointly executed by the server and the terminal device, and this is not limited. This article only takes the server as an example for illustration.

[0109] The embodiments of the present application can be applied to various scenarios, including but not limited to cloud technology, artificial intelligence, intelligent transportation, assisted driving, recommendation systems, network security detection, etc.

[0110] Refer to Figure 2 As shown, it is a schematic flowchart of an image annotation method provided in the embodiments of the present application, and the process includes:

[0111] S201. Obtain an image to be annotated.

[0112] In the embodiments of the present application, the annotation result of the image to be annotated is used to describe the content such as entities, behaviors, and scenarios in the image to be annotated, but is not limited thereto.

[0113] In the embodiments of the present application, Chinese labels can be marked for the image to be labeled, and English labels can also be marked for the image to be labeled. Of course, labels in other languages can also be marked. The marking process of the labels in each language is similar. In the following, the process of marking English labels for the image to be labeled will be taken as an example for illustration.

[0114] S202. Based on each preset description prompt information, respectively perform image content description on the image to be labeled to obtain corresponding content description results, and based on the obtained content description results, construct the target description information of the image to be labeled.

[0115] Among them, each description prompt information is used to indicate performing image content description on the image to be labeled. The description prompt information can also be called a description prompt word.

[0116] In the embodiments of the present application, for the marking of labels in different languages, corresponding language description prompt information is adopted. Taking the marking of English labels as an example, the description prompt information is English description prompt information.

[0117] For example, the description prompt information includes but is not limited to one or more of: "Image shows", "A photo of", "Describe image", "Picture context describe", "Image context captioning", "Theme of picture is".

[0118] In the embodiments of the present application, the content description result refers to the description of the entities, behaviors, scenes, etc. in the image to be labeled using natural language.

[0119] As a possible implementation manner, image content description can be performed on the image to be labeled based on one description prompt information to obtain corresponding content description results. However, considering that in the case of a single description prompt information, there are randomness and instability in the description of the image to be labeled, resulting in possible phenomena such as description errors, omissions, hallucinations, etc. Therefore, image content description can also be performed on the image to be labeled by respectively combining multiple preset description prompt information to obtain corresponding content description results, and the accuracy of the content description can be improved through multiple rounds of image description.

[0120] In some embodiments, the content description result can be obtained by using an image content description model. Specifically, each description prompt information is respectively combined with the image to be labeled and input into the image content description model to obtain corresponding image description content. Among them, the image content description model is obtained by pre-training using each second sample image.

[0121] For example, referring to Figure 3 as shown, the description prompt "Image shows" and the image to be annotated are input into the image content description model to obtain content description result 1, and content description result 1 is "A zebra is eating grass on afield"; the description prompt "Aphoto of" and the image to be annotated are input into the image content description model to obtain content description result 2, and content description result 2 is "A zebra is on a field"; the description prompt "Describeimage" and the image to be annotated are input into the image content description model to obtain content description result 3, and content description result 3 is "A zebra is eating grass"; the description prompt "Picture context describe" and the image to be annotated are input into the image content description model to obtain content description result 4, and content description result 4 is "An animal iseating on something"; the description prompt "Image context captioning" and the image to be annotated are input into the image content description model to obtain content description result 5, and content description result 5 is "A zebra is eatinggrass on a field"; the description prompt "Theme of picture is" and the image to be annotated are input into the image content description model to obtain content description result 6, and content description result 6 is "An anima is on a field".

[0122] After obtaining each content description result, the target description information can be constructed in the following ways, but is not limited to:

[0123] Method 1: Directly splice each content description result to construct the target description information of the image to be annotated. It should be noted that direct splicing can be to splice each content description result using some connection characters (such as characters with a connection relationship like ";", "and", "&", etc.), or to perform semantic fusion on each content description result, and this is not limited.

[0124] For example, directly splice the content description results 1, 2, 3, 4, 5, and 6 to construct the target description information of the image to be annotated. The target description information is "Content description result 1 (A zebra is eating grass on a field); Content description result 2 (A zebra is on a field); Content description result 3 (A zebra is eating grass); Content description result 4 (An animal is eating on something); Content description result 5 (A zebra is eating grass on a field); Content description result 6 (An animal is on a field).

[0125] Method 2: In the process of multi-round image description, due to the randomness and instability of the output content description results, the desired label information may not be obtained. At this time, to ensure the comprehensiveness and accuracy of the description of specific information, image information supplementation can also be performed. Specifically, in the embodiments of the present application, the image information supplementation can be implemented in but not limited to the following ways:

[0126] Based on the obtained content description results, construct the initial description information of the image to be annotated;

[0127] Based on each content query question set for each key content, respectively combine it with the image to be annotated to perform content query and obtain the corresponding answer results;

[0128] Based on the obtained answer results, combine with the initial description information to obtain the target description information of the image to be annotated.

[0129] Among them, a content query question is used to query whether the image to be annotated contains one or a category of key content. The key content can be an entity, such as a cat, a dog, an airplane, etc.; the key content can also be an action, such as eating grass, having dinner, playing ball, etc.; the key content can also be a scene, such as a party, indoor, outdoor, etc., but is not limited thereto.

[0130] In the embodiments of the present application, as an example, for different key contents, corresponding content query questions can be preset in advance, so that directly according to the preset content query questions, combined with the image to be annotated, the corresponding answer results can be obtained; as another example, after constructing the initial description information of the image to be annotated, content query questions can be set for each key content, and then according to the set content query questions, combined with the image to be annotated, the corresponding answer results can be obtained.

[0131] Among them, each content inquiry question can be set in the following ways, but not limited to:

[0132] Obtain each key content set for the target application scenario;

[0133] Based on the information types corresponding to each key content obtained, combined with the question templates respectively set for each information type, obtain each content inquiry question set for each key content.

[0134] The target application scenario includes but is not limited to: content recommendation, abnormal image recognition, etc. For the content recommendation scenario, it can be further divided into

[0135] To improve the construction efficiency of content inquiry questions, in the embodiments of the present application, according to the information types of key content, using the question templates set for different information types, content inquiry questions can be constructed. Among them, the information type can be a noun or a verb, but is not limited thereto.

[0136] Specifically, when constructing content inquiry questions, according to the information type corresponding to the key content, there are but not limited to the following two situations:

[0137] Situation 1, if the information type of one of the key contents is a noun, then use the question template set for nouns, and based on this key content, obtain the corresponding content inquiry question.

[0138] It should be noted that in the embodiments of the present application, one or more question templates can be preset for a class of key contents (i.e., one information type). In the following, only one question template set for a class of key contents is taken as an example for illustration, but it is not limited thereto. If there are multiple question templates set for a class of key contents, for one key content, one question template can be arbitrarily selected from multiple question templates, and using the selected one question target, obtain one content inquiry question corresponding to this key content, or multiple content inquiry questions corresponding to this key content can be obtained using multiple question templates.

[0139] For example, referring to Figure 4 As shown, the key contents include "cat", "dog", "airplane", "zebra", "drawingpicture", "driving car", "playing basketball", "cooking dinner", among which, "cat", "dog", "airplane", "zebra" belong to nouns, and "drawing picture", "driving car", "playingbasketball", "cooking dinner" belong to verbs.

[0140] Suppose the question template for nouns (i.e., the noun question template) is "Does this image contain______?". Using the question template for nouns, based on the key content "cat", the content inquiry question 1 is obtained, and the content inquiry question 1 is "Does this image contain cat?"; using the question template for nouns, based on the key content "dog", the content inquiry question 2 is obtained, and the content inquiry question 2 is "Does this image contain dog?"; using the question template for nouns, based on the key content "airplane", the content inquiry question 3 is obtained, and the content inquiry question 3 is "Does this image contain airplane?"; using the question template for nouns, based on the content key information "zebra", the content inquiry question 4 is obtained, and the content inquiry question 4 is "Does this image contain zebra?".

[0141] As another possible case, if the information type of one of the content key information is verb key information, then the question template set for the verb key information is used, and based on the content key information, the corresponding content inquiry question is obtained.

[0142] Still referring to Figure 4As shown, "drawing picture", "driving car", "playing basketball", and "cooking dinner" are verbs. Suppose the question template set for verbs (i.e., the verb question template) is "Is there someone ______ in image?", then, using the question template set for verbs, based on the key content "drawing picture", the content inquiry question 5 is obtained, and the content inquiry question 5 is "Is there someone drawing picture in image?"; using the question template set for verbs, based on the key content "driving car", the content inquiry question 6 is obtained, and the content inquiry question 6 is "Is there someone driving car in image?"; using the question template set for verbs, based on the key content "playing basketball", the content inquiry question 7 is obtained, and the content inquiry question 7 is "Is there someone playing basketball in image?"; using the question template set for verbs, based on the key content "cooking dinner", the content inquiry question 8 is obtained, and the content inquiry question 8 is "Is there someone cooking dinner in image?".

[0143] After obtaining each content inquiry question, as a possible implementation, a visual question answering model can be used to process the visual question answering task. Exemplarily, each content inquiry question is combined with the image to be annotated and input into the visual question answering model to obtain the corresponding answer result.

[0144] For example, still referring to Figure 4As shown in the figure, the content inquiry question 1 and the image to be annotated are input into the visual question answering model to obtain the answer result 1, and the answer result 1 indicates that the image to be annotated contains a cat; the content inquiry question 2 and the image to be annotated are input into the visual question answering model to obtain the answer result 2, and the answer result 2 indicates that the image to be annotated does not contain a dog; the content inquiry question 3 and the image to be annotated are input into the visual question answering model to obtain the answer result 3, and the answer result 3 indicates that the image to be annotated does not contain an airplane; the content inquiry question 4 and the image to be annotated are input into the visual question answering model to obtain the answer result 4, and the answer result 4 indicates that the image to be annotated contains a zebra; the content inquiry question 5 and the image to be annotated are input into the visual question answering model to obtain the answer result 5, and the answer result 5 indicates that there is no drawing behavior in the image to be annotated; the content inquiry question 6 and the image to be annotated are input into the visual question answering model to obtain the answer result 6, and the answer result 6 indicates that there is no driving behavior in the image to be annotated; the content inquiry question 7 and the image to be annotated are input into the visual question answering model to obtain the answer result 7, and the answer result 7 indicates that there is no ball-playing behavior in the image to be annotated; the content inquiry question 8 and the image to be annotated are input into the visual question answering model to obtain the answer result 8, and the answer result 8 indicates that there is no behavior of cooking dinner in the image to be annotated.

[0145] After obtaining each answer result, based on each answer result and combined with the initial description information, the target description information of the image to be annotated is obtained. Exemplarily, each answer result and the initial description information can be spliced to obtain the target description information of the image to be annotated.

[0146] It should be noted that direct splicing can be to splice each answer result and the initial description information using some connection characters (such as characters with connection relationships like ";", "and", "&", etc.), or to perform semantic fusion on each answer result and the initial description information, and this is not limited.

[0147] For example, the initial description information is "A zebra is eating grass on a field". Based on the answer results 1 to 8 and combined with the initial description information for splicing, the target description information is obtained, and the target description information is "A zebraand a cat are eating grass on a field".

[0148] S203. Based on the target description information and combined with the preset summary prompt words, construct a visual question answering question.

[0149] In the embodiment of the present application, by adding summary prompt words to the target description information, the summary ability for the target description information and the image to be annotated is established.

[0150] As a possible implementation, the target description information and the summary prompt are concatenated to obtain a visual question and answer question. The front-to-back concatenation order of the target description information and the summary prompt can be determined according to the semantics of the set summary prompt.

[0151] For example, refer to Figure 5 As shown, the summary prompt is "Based on the image content, summarize the following text into a few keywords", and the target description information is "A zebra and a cat are eating grass on a field". The summary prompt and the target description information are concatenated to obtain a visual question and answer question: "Based on the image content, summarize the following text into a few keywords: A zebra and a cat are eating grass on a field".

[0152] S204. Based on the visual question and answer question and the image to be labeled, perform label summarization on the target description information to obtain the annotation result of the image to be labeled.

[0153] Specifically, in the embodiments of the present application, when executing S204, the following implementation manners can be adopted but are not limited to:

[0154] Method 1. Based on the visual question and answer question and the image to be labeled, perform label summarization on the target description information to obtain a candidate label set of the image to be labeled, and directly use the candidate label set of the image to be labeled as the annotation result of the image to be labeled.

[0155] As a possible implementation, the visual question and answer question and the image to be labeled are input into a visual question and answer model to obtain a candidate label set of the image to be labeled. The visual question and answer model is obtained by pre-training with each third sample image.

[0156] For example, still refer to Figure 5 As shown, the visual question and answer question and the image to be labeled are input into the visual question and answer model to obtain a candidate label set of the image to be labeled. The candidate label set contains the following candidate labels: "zebra", "eating grass", "field", "cat". Then, "zebra", "eating grass", "field", "cat" are used as the annotation result of the image to be labeled.

[0157] Method 2: Considering that in the actual application process, multiple candidate words are often obtained, and only a fixed number of tags are required in the actual application process, screening is performed through keyword similarity matching. Specifically, in Method 2, the following steps are included:

[0158] Based on the visual question and the image to be annotated, summarize the tags of the target description information to obtain the candidate tag set of the image to be annotated;

[0159] Based on the graphic-text matching degree between each candidate tag in the candidate tag set and the image to be annotated, screen at least one target tag that meets the set screening conditions from each candidate tag;

[0160] Use at least one target tag as the annotation result of the image to be annotated.

[0161] Among them, in the process of summarizing the tags of the target description information based on the visual question and the image to be annotated to obtain the candidate tag set of the image to be annotated, a visual question answering model can be used to obtain the candidate tag set of the image to be annotated. Specifically, input the visual question and the image to be annotated into the visual question answering model to obtain the candidate tag set of the image to be annotated.

[0162] For example, refer to Figure 6 As shown, input the visual question and the image to be annotated into the visual question answering model to obtain the candidate tag set of the image to be annotated. The candidate tag set contains the following candidate tags: candidate tag 1 ("zebra"), candidate tag 2 ("eating grass"), candidate tag 3 ("field"), candidate tag 4 ("cat").

[0163] After obtaining the candidate tag set of the image to be annotated, calculate the graphic-text matching degree between each candidate tag in the candidate tag set and the image to be annotated. As a possible implementation, the graphic-text matching degree can be obtained by using a graphic-text matching model. Specifically, for any one candidate tag in each candidate tag, input the candidate tag and the image to be annotated into the graphic-text matching model to obtain the graphic-text matching degree between the candidate tag and the image to be annotated.

[0164] Among them, the graphic-text matching degree between a candidate tag and the image to be annotated is used to represent: the probability that the image to be annotated contains the content indicated by the candidate tag. The graphic-text matching model is pre-trained using each first sample image.

[0165] For example, still refer to Figure 6As shown, the candidate label "zebra" (candidate label 1) and the image to be labeled are input into the image-text matching model to obtain the image-text matching degree 1 between the candidate label "zebra" and the image to be labeled, and the value of the image-text matching degree 1 is 0.95; the candidate label "eating grass" (candidate label 2) and the image to be labeled are input into the image-text matching model to obtain the image-text matching degree 2 between the candidate label "eating grass" and the image to be labeled, and the value of the image-text matching degree 2 is 0.9; the candidate label "field" (candidate label 3) and the image to be labeled are input into the image-text matching model to obtain the image-text matching degree 3 between the candidate label "field" and the image to be labeled, and the value of the image-text matching degree 3 is 0.92; the candidate label "cat" (candidate label 4) and the image to be labeled are input into the image-text matching model to obtain the image-text matching degree 4 between the candidate label "cat" and the image to be labeled, and the value of the image-text matching degree 4 is 0.85.

[0166] As a possible implementation manner, when screening at least one target label that meets the set screening conditions from each candidate label based on the image-text matching degree between each candidate label in the candidate label set and the image to be labeled, the following manner can be adopted but is not limited to:

[0167] Based on the image-text matching degree between each candidate label in the candidate label set and the image to be labeled, sort each candidate label to obtain a label sorting result; based on the label sorting result, according to the set label screening quantity threshold, screen out at least one candidate label from each candidate label in turn, and use the at least one candidate label screened out as at least one target label that meets the set screening conditions.

[0168] As an example, assume that the higher the value of the image-text matching degree between a candidate label and the image to be labeled, the higher the labeling accuracy of the candidate label. Based on the image-text matching degree between each candidate label in the candidate label set and the image to be labeled, sort each candidate label in descending order to obtain a label sorting result, and then based on the label sorting result, according to the set label screening quantity threshold k, screen out the first N candidate labels (i.e., the top N candidate labels) from each candidate label, and use the first N candidate labels screened out as N target labels that meet the set screening conditions.

[0169] For example, still referring to Figure 6As shown, based on the graphic-text matching degrees between the candidate labels "zebra", "eating grass", "field", and "cat" and the image to be annotated respectively, the candidate labels are sorted in descending order to obtain a label sorting result. The label sorting result is as follows: candidate label "zebra", candidate label "field", candidate label "eating grass", candidate label "cat". Assume that the value of the label screening quantity threshold is 3, that is, N = 3. Based on the label sorting result, according to the set label screening quantity threshold, the first 3 candidate labels are screened out from each candidate label: candidate label "zebra", candidate label "field", candidate label "eating grass", and the screened candidate labels "zebra", "field", and "eating grass" are used as the target labels that meet the set screening conditions.

[0170] It should be noted that in the embodiments of the present application, there is no limitation on the sorting method. It can also be in ascending order, which is not limited here and will not be elaborated further.

[0171] Furthermore, in some embodiments, after obtaining the annotation result of the image to be annotated, content recommendation or anomaly recognition can also be performed using the image to be annotated.

[0172] Specifically, as a possible implementation manner, based on the annotation result of the image to be annotated and combined with the object information of the target object, a recommendation result of the image to be annotated is obtained, and the recommendation result is used to represent whether to recommend the image to be annotated to the target object.

[0173] Among them, the target object can be one or a class of objects. The object information of the target object is used to represent the content preferences of one or a class of objects.

[0174] For example, the annotation result of the image to be annotated is: "zebra", "field", "eating grass", and the object information of the target object represents that the target object prefers animal images. Based on the annotation result of the image to be annotated and combined with the object information of the target object, a recommendation result of the image to be annotated is obtained, and the recommendation result represents recommending the image to be annotated to the target object.

[0175] Through the above implementation manner, during image annotation, personalized content recommendation can be realized according to the annotation result, improving the accuracy and efficiency of content recommendation.

[0176] As another possible implementation manner, based on the annotation result of the image to be annotated, an anomaly recognition result of the image to be annotated is obtained, and the anomaly recognition result is used to represent whether the image to be annotated contains abnormal content.

[0177] Among them, anomaly recognition may refer to the recognition of content containing sensitive or illegal content. Through anomaly recognition, it is possible to help prevent the spread of abnormal content.

[0178] For example, the child protection function in Internet products needs to identify and block images that may contain sensitive content to protect the healthy growth of minors. This method can help identify and prevent children from accessing inappropriate content.

[0179] Another example is that an advertising platform needs to review advertising content to ensure that the advertisements do not contain illegal content and are consistent with the actual products, thereby helping to automatically review advertising images and improve the review efficiency and accuracy.

[0180] In some embodiments, the image-text matching model, visual question answering model, and image content description model may be the same model. Correspondingly, the first sample image set (including each first sample image), the second sample image set (including each first sample image), and the third sample image set (including each first sample image) may be the same sample image set.

[0181] A model with the functions of image-text matching, visual question answering, and image content description is called a visual language cross-modal large model (which can be abbreviated as the large model). The visual language cross-modal large model is a model that can support visual question answering tasks, image content description, and image-text matching capabilities, has a large number of parameters, and is pre-trained on a large-scale sample. The specific three ability forms of the visual language cross-modal large model are as follows:

[0182] (1) Visual question answering task: This task requires the model to be given an image and an open-ended natural language question, and requires the model to output a natural language answer. That is to say, the input data of the visual language cross-modal large model is an image and a natural language question, and the output data is a natural language answer. For example, referring to Figure 7 As shown, the natural language question is "What is a mustache made of?" The natural language question and Image 1 are input into the visual language cross-modal large model to obtain the natural language answer "banana".

[0183] (2) Image content description task: This task requires the model to be given an image and requires the model to output a natural language description text that can describe the image. That is to say, the input data of the visual language cross-modal large model is an image, and the output data is a natural language description text. For example, referring to Figure 8 As shown, Image 2 is input into the visual language cross-modal large model to obtain the natural language description text "The baseball player is waiting for the serve"; Image 3 is input into the visual language cross-modal large model to obtain the natural language description text "A bus is parked beside a tall building".

[0184] (3) Image-Text Matching Ability: This task requires the given model to be provided with an image and a natural language description of the image, and the model is required to output the image-text matching degree between the text and the image. That is to say, the input data of the vision-language cross-modal large model is an image and a natural language text, and the output data is the image-text matching degree. For example, as shown in Figure 9 Input the image 4 and the text "A zebra and a cat are eating grass on a field" into the vision-language cross-modal large model to obtain the image-text matching degree between the text and the image.

[0185] Next, the present application will be described in conjunction with a specific embodiment.

[0186] As shown in Figure 10 It is a schematic diagram of the image annotation process for Chinese labels provided in the embodiment of the present application. The image annotation process can be divided into four stages: multi-round image description, image information supplementation, image information summary, and keyword matching and screening.

[0187] In the multi-round matching stage, based on each preset description prompt information, the image content is described in combination with the image to be annotated respectively to obtain corresponding content description results, and based on the obtained content description results, the initial description information of the image to be annotated is constructed. Specifically, each description prompt information includes: description prompt information 1 "The image content includes", description prompt information 2 "Overview of the image content", description prompt information 3 "Summary of the image content", description prompt information 4 "Scenes included in the image". The image to be annotated is input into the vision-language cross-modal large model using different description prompt information, and using the image-text matching ability of the vision-language cross-modal large model, 4 content description results are obtained. Among the 4 content description results, content description result 1 is "A car", content description result 2 is "A tall building", content description result 3 is "A car and a tall building", and content description result 4 is "A car is parked beside the tall building". Then, the 4 content description results are spliced to obtain the initial description information of the image to be annotated, and the initial description information is "A car is parked beside the tall building".

[0188] During the multi-round image description process, due to the randomness and instability of the output of the vision-language cross-modal large model, the desired label information may not be obtained. At this time, to ensure the comprehensiveness and accuracy of the description of specific information, the visual question-answering ability of the large model is used to supplement the image information.

[0189] In the image information supplement stage, obtain each key content set for the target application scenario, and based on the information type corresponding to each key content obtained, combine the question templates respectively set for each information type to obtain each content query question set for each key content. Then, based on each content query question, respectively combine with the image to be annotated for content query, obtain the corresponding answer results, and based on the obtained answer results, combine with the initial description information to obtain the target description information of the image to be annotated.

[0190] Specifically, each key content set for the target application scenario includes "bus", "private car", "taxi", "office building", "driving", "waiting for traffic lights", "walking on the sidewalk", "dancing". Among them, "bus", "private car", "taxi", "office building" belong to nouns, and "driving", "waiting for traffic lights", "walking on the sidewalk" belong to verbs.

[0191] Suppose the question template set for nouns is "Does the image contain ___?", and using the question template for nouns, based on the key content "bus", obtain the content query question 1 "Does the image contain a bus?"; based on the key content "private car", obtain the content query question 2 "Does the image contain a private car?"; based on the key content "taxi", obtain the content query question 3 "Does the image contain a taxi?"; based on the content key information "office building", obtain the content query question 4 "Does the image contain an office building?"

[0192] The question template set for verbs is "Does there exist a person or object ___ in the image?", and using the question template for nouns, based on the key content "driving", obtain the content query question 5 "Does there exist a person or object driving in the image?"; based on the key content "waiting for traffic lights", obtain the content query question 6 "Does there exist a person or object waiting for traffic lights in the image?"; based on the key content "walking on the sidewalk", obtain the content query question 7 "Does there exist a person or object walking on the sidewalk in the image?"; based on the key content "dancing", obtain the content query question 8 "Does there exist a person or object dancing in the image?"

[0193] After obtaining each content query question, input the image to be annotated and content query questions 1, 2,..., 8 into the visual language cross-modal large model respectively, and use the visual question answering ability of the visual language cross-modal large model to obtain 8 answer results. The 8 answer results indicate that the image to be annotated contains a bus, does not contain a private car, a taxi, or an office building, and there are no behaviors such as driving, waiting for traffic lights, walking on the sidewalk, or dancing.

[0194] Based on 8 response results and combined with the initial description information, obtain the target description information of the image to be labeled: "A bus is parked beside a tall building."

[0195] In the image information summarization stage, considering that there may be a large amount of key information for prompting, the new description text is often very long. Especially for the key information that does not exist in the image, it may become a large amount of noise. Therefore, the question-answering ability of the cross-modal large model is used to further summarize the description text, so as to obtain the candidate label keywords.

[0196] Specifically, the preset summarization prompt word is "Based on the image content, summarize some keywords from the following text". Based on the target description information and combined with the preset summarization prompt word, construct a visual question-answering question: "Based on the image content, summarize some keywords from the following text: A bus is parked beside a tall building."

[0197] Then, input the visual question-answering question and the image to be labeled into the visual-language cross-modal large model, and use the visual question-answering ability of the visual-language cross-modal large model to obtain the candidate label set of the image to be labeled {Label 1, Label 2, Label 3,...}. Exemplarily, Label 1 is "bus", Label 2 is "tall building", and Label 3 is "parked".

[0198] In the keyword matching and screening stage, input the image to be labeled and each candidate label in the candidate label set into the visual-language cross-modal large model respectively, use the image-text matching ability of the visual-language cross-modal large model to obtain the image-text matching degree between each candidate label and the image to be labeled, and then, based on the image-text matching degree between each candidate label and the image to be labeled, screen out the top N (such as the top 3) labels from each candidate label as the target labels of the image to be labeled.

[0199] Based on the same inventive concept, an embodiment of the present application provides an image annotation device. As Figure 11 shown, it is a schematic structural diagram of the image annotation device 1100, and may include:

[0200] An image acquisition unit 1101, configured to acquire an image to be labeled;

[0201] A content description unit 1102, configured to respectively perform image content description on the image to be labeled based on each preset description prompt information, obtain corresponding content description results, and construct target description information of the image to be labeled based on the obtained content description results;

[0202] A question construction unit 1103, configured to construct a visual question-answering question based on the target description information and combined with a preset summarization prompt word;

[0203] A content summarization unit 1104, configured to perform label summarization on the target description information based on the visual question and the image to be annotated, so as to obtain an annotation result of the image to be annotated.

[0204] As a possible implementation manner, when constructing the target description information of the image to be annotated based on the obtained content description results, the content description unit 1102 is specifically configured to:

[0205] Construct initial description information of the image to be annotated based on the obtained content description results;

[0206] Perform content query in combination with the image to be annotated respectively based on the content inquiry questions set for each key content, so as to obtain corresponding answer results;

[0207] Based on the obtained answer results and in combination with the initial description information, obtain the target description information of the image to be annotated.

[0208] As a possible implementation manner, before performing content query in combination with the image to be annotated respectively based on the content inquiry questions for each key content to obtain corresponding content answer results, the content description unit 1102 is further configured to:

[0209] Obtain the key contents set for the target application scenario;

[0210] Based on the information types corresponding to the key contents respectively, in combination with the question templates set for each information type, obtain the content inquiry questions set for the key contents.

[0211] As a possible implementation manner, when obtaining the content inquiry questions set for the key contents based on the information types corresponding to the key contents respectively and in combination with the question templates set for each information type, the content description unit 1102 is specifically configured to:

[0212] If the information type of a key content among the key contents is a noun, then adopt the question template set for nouns, and based on the key content, obtain the corresponding content inquiry question;

[0213] If the information type of a key content among the key contents is a verb, then adopt the question template set for verbs, and based on the key content, obtain the corresponding content inquiry question.

[0214] As a possible implementation manner, when performing label summarization on the target description information based on the visual question and the image to be annotated to obtain an annotation result of the image to be annotated, the content summarization unit 1104 is specifically configured to:

[0215] Based on the visual question and the image to be annotated, perform label summarization on the target description information to obtain a candidate label set for the image to be annotated;

[0216] Based on the graphic-text matching degree between each candidate label in the candidate label set and the image to be annotated, screen at least one target label that meets the set screening conditions from the candidate labels;

[0217] Use the at least one target label as the annotation result of the image to be annotated.

[0218] As a possible implementation manner, when screening at least one target label that meets the set screening conditions from the candidate labels based on the graphic-text matching degree between each candidate label in the candidate label set and the image to be annotated, the content summarization unit 1104 is specifically configured to:

[0219] Based on the graphic-text matching degree between each candidate label in the candidate label set and the image to be annotated, sort the candidate labels to obtain a label sorting result;

[0220] Based on the label sorting result, sequentially screen at least one candidate label from the candidate labels according to a set label screening quantity threshold, and use the at least one candidate label screened out as at least one target label that meets the set screening conditions.

[0221] As a possible implementation manner, the content summarization unit 1104 is further configured to:

[0222] Input the candidate label and the image to be annotated into a graphic-text matching model to obtain the graphic-text matching degree between the candidate label and the image to be annotated, where the graphic-text matching model is pre-trained using each first sample image.

[0223] As a possible implementation manner, when respectively performing image content description on the image to be annotated based on each preset description prompt information to obtain corresponding content description results, the content description unit 1102 is specifically configured to:

[0224] Respectively combine each description prompt information with the image to be annotated and input them into an image content description model to obtain corresponding image description content, where the image content description model is pre-trained using each second sample image;

[0225] When performing label summarization on the target description information based on the visual question and the image to be annotated to obtain the annotation result of the image to be annotated, the content summarization unit 1104 is specifically configured to:

[0226] Input the visual question - answering question and the image to be annotated into a visual question - answering model to obtain the annotation result of the image to be annotated. The visual question - answering model is obtained by pre - training with each third - sample image.

[0227] As a possible implementation, the image content description model and the visual question - answering model are the same model.

[0228] As a possible implementation, the content summarization unit 1104 is further configured to:

[0229] Based on the annotation result of the image to be annotated and combined with the object information of the target object, obtain a recommendation result for the image to be annotated. The recommendation result is used to represent whether to recommend the image to be annotated to the target object; or,

[0230] Based on the annotation result of the image to be annotated, obtain an anomaly recognition result for the image to be annotated. The anomaly recognition result is used to represent whether the image to be annotated contains abnormal content.

[0231] For the convenience of description, the above - mentioned parts are divided into various modules (or units) according to functions and described separately. Of course, when implementing this application, the functions of each module (or unit) can be implemented in the same or multiple software or hardware.

[0232] Regarding the device in the above - mentioned embodiments, the specific manner in which each unit executes the request has been described in detail in the embodiments related to the method, and will not be elaborated here.

[0233] Those skilled in the art can understand that various aspects of the present application can be implemented as a system, a method, or a program product. Therefore, various aspects of the present application can be specifically implemented in the following forms: a complete hardware implementation, a complete software implementation (including firmware, micro - code, etc.), or an implementation combining hardware and software aspects, which can be collectively referred to as "circuit", "module", or "system" here.

[0234] Based on the same inventive concept, the embodiments of the present application also provide an electronic device. In one embodiment, the electronic device can be a server or a terminal device. Refer to Figure 12 As shown, it is a schematic structural diagram of a possible electronic device provided in the embodiments of the present application. Figure 12 In it, the electronic device 1200 includes: a processor 1210 and a memory 1220.

[0235] Among them, the memory 1220 stores a computer program that can be executed by the processor 1210. The processor 1210 can execute the steps of the above - mentioned image annotation method by executing the instructions stored in the memory 1220.

[0236] The memory 1220 can be a volatile memory, such as a random-access memory (RAM); the memory 1220 can also be a non-volatile memory, such as a Read-Only Memory (ROM), a flash memory, a hard disk drive (HDD) or a solid-state drive (SSD); or the memory 1220 is any other medium that can be used to carry or store the desired program code in the form of instructions or data structures and can be accessed by a computer, but is not limited thereto. The memory 1220 can also be a combination of the above memories.

[0237] The processor 1210 can include one or more central processing units (CPUs) or be a digital processing unit, etc. When the processor 1210 executes the computer program stored in the memory 1220, the above image annotation method is implemented.

[0238] In some embodiments, the processor 1210 and the memory 1220 can be implemented on the same chip. In some embodiments, they can also be separately implemented on independent chips.

[0239] In the embodiments of the present application, the specific connection medium between the above-mentioned processor 1210 and the memory 1220 is not limited. In the embodiments of the present application, taking the connection between the processor 1210 and the memory 1220 through a bus as an example, the bus is Figure 12 described by a thick line in. The connection manners between other components are only for illustrative purposes and are not to be construed as limiting. The bus can be divided into an address bus, a data bus, a control bus, etc. For the sake of convenience of description, Figure 12 only a thick line is used to describe it in, but it does not describe that there is only one bus or one type of bus.

[0240] Based on the same inventive concept, the embodiments of the present application provide a computer-readable storage medium, which includes a computer program. When the computer program runs on an electronic device, the computer program is used to cause the electronic device to execute the steps of the above image annotation method. In some possible implementation manners, various aspects of the image annotation method provided in the present application can also be implemented in the form of a program product, which includes a computer program. When the program product runs on an electronic device, the computer program is used to cause the electronic device to execute the steps in the above image annotation method. For example, the electronic device can execute as Figure 2 the steps shown in.

[0241] The program product can adopt any combination of one or more readable media. The readable media can be a readable signal medium or a readable storage medium. The readable storage medium can be, for example, but not limited to, an electrical, magnetic, optical, electromagnetic, infrared, or semiconductor system, apparatus, or device, or any combination of the above. More specific examples (non-exhaustive list) of the readable storage medium include: an electrical connection having one or more wires, a portable disk, a hard disk, a RAM, a ROM, an erasable programmable read-only memory (EPROM or flash memory), an optical fiber, a portable compact disk read-only memory (CD-ROM), an optical storage device, a magnetic storage device, or any suitable combination of the above.

[0242] The program product of the embodiments of the present application can adopt a CD-ROM and include a computer program, and can run on an electronic device. However, the program product of the present application is not limited thereto. In this document, the readable storage medium can be any tangible medium that contains or stores a computer program, and this computer program can be used by or in combination with a command execution system, apparatus, or device.

[0243] The readable signal medium can include a data signal propagated in a baseband or as part of a carrier wave, which carries a readable computer program. Such a propagated data signal can take various forms, including but not limited to electromagnetic signals, optical signals, or any suitable combination of the above. The readable signal medium can also be any readable medium other than the readable storage medium, and this readable medium can send, propagate, or transmit a computer program for use by or in combination with a command execution system, apparatus, or device.

[0244] Although the preferred embodiments of the present application have been described, those skilled in the art can make additional changes and modifications to these embodiments once they know the basic creative concept. Therefore, the appended claims are intended to be construed as including the preferred embodiments and all changes and modifications falling within the scope of the present application.

[0245] Obviously, those skilled in the art can make various changes and modifications to the present application without departing from the spirit and scope of the present application. Thus, if these modifications and variations of the present application fall within the scope of the claims of the present application and their equivalent technologies, the present application also intends to include these modifications and variations.

Claims

1. An image annotation method, characterized in that, Including: Obtain the image to be labeled; Based on each preset description prompt information, respectively describe the content of the image to be labeled, obtain corresponding content description results, and based on the obtained content description results, construct the target description information of the image to be labeled; Based on the target description information, combined with the preset summary prompt words, construct a visual question and answer question; Based on the visual question and answer question and the image to be labeled, summarize the target description information to obtain the labeling result of the image to be labeled.

2. The method according to claim 1, wherein The constructing the target description information of the image to be labeled based on the obtained content description results includes: Based on the obtained content description results, construct the initial description information of the image to be labeled; Based on each content query question set for each key content, respectively combine with the image to be labeled to perform content query, and obtain corresponding answer results; Based on the obtained answer results, combined with the initial description information, obtain the target description information of the image to be labeled.

3. The method according to claim 2, characterized in that Before the step of respectively combining with the image to be labeled based on each content query question for each key content to perform content query and obtain corresponding content answer results, it further includes: Obtain each key content set for the target application scenario; Based on the information type corresponding to each key content, combined with the question templates respectively set for each information type, obtain each content query question set for each key content.

4. The method according to claim 3, characterized in that, The obtaining each content query question set for each key content based on the information type corresponding to each key content, combined with the question templates respectively set for each information type includes: If the information type of a key content among each key content is a noun, then adopt the question template set for nouns, and based on the one key content, obtain the corresponding content query question; If the information type of a key content among each key content is a verb, then adopt the question template set for verbs, and based on the one key content, obtain the corresponding content query question.

5. The method according to any one of claims 1 to 4, characterized in that, The summarizing the target description information to obtain the labeling result of the image to be labeled based on the visual question and answer question and the image to be labeled includes: Based on the visual question and answer question and the image to be labeled, summarize the target description information to obtain a candidate label set of the image to be labeled; Based on the graphic-text matching degree between each candidate label in the candidate label set and the image to be labeled, screen at least one target label that meets the set screening conditions from the candidate labels; Use the at least one target label as the labeling result of the image to be labeled.

6. The method according to claim 5, characterized in that, The screening at least one target label that meets the set screening conditions from the candidate labels based on the graphic-text matching degree between each candidate label in the candidate label set and the image to be labeled includes: Based on the graphic-text matching degree between each candidate label in the candidate label set and the image to be labeled, sort the candidate labels to obtain a label sorting result; Based on the tag sorting result, at least one candidate tag is sequentially selected from each of the candidate tags according to a set threshold of the number of tags to be screened, and the at least one candidate tag selected is used as at least one target tag that meets the set screening conditions.

7. The method according to claim 5, wherein The graphic-text matching degree between a candidate tag and the image to be annotated is obtained in the following manner: The initial tag and the image to be annotated are input into a graphic-text matching model to obtain the graphic-text matching degree between the candidate tag and the image to be annotated. The graphic-text matching model is pre-trained using each first sample image.

8. The method according to claim 5, characterized in that Based on each preset description prompt information, the image content of the image to be annotated is respectively described to obtain corresponding content description results, including: Each description prompt information is respectively combined with the image to be annotated and input into an image content description model to obtain corresponding image description content. The image content description model is pre-trained using each second sample image. Based on the visual question and the image to be annotated, the target description information is summarized by tags to obtain the annotation result of the image to be annotated, including: The visual question is combined with the image to be annotated and input into a visual question answering model to obtain the annotation result of the image to be annotated. The visual question answering model is pre-trained using each third sample image.

9. The method according to claim 8, wherein The graphic-text matching model, the image content description model, and the visual question answering model are the same model.

10. The method according to any one of claims 1 to 4, characterized in that, It further includes: Based on the annotation result of the image to be annotated and combined with the object information of the target object, a recommendation result of the image to be annotated is obtained. The recommendation result is used to represent whether to recommend the image to be annotated to the target object. Or, Based on the annotation result of the image to be annotated, an anomaly recognition result of the image to be annotated is obtained. The anomaly recognition result is used to represent whether the image to be annotated contains abnormal content.

11. An image annotation device, characterized in that, It includes: An image acquisition unit, configured to acquire an image to be annotated; A content description unit, configured to respectively describe the image content of the image to be annotated based on each preset description prompt information to obtain corresponding content description results, and construct target description information of the image to be annotated based on the obtained content description results; A question construction unit, configured to construct a visual question based on the target description information and in combination with a preset summary prompt word; A content summary unit, configured to summarize the target description information by tags based on the visual question and the image to be annotated to obtain the annotation result of the image to be annotated.

12. An electronic device, characterized in that, It includes a processor and a memory. The memory stores a computer program. When the computer program is executed by the processor, the processor executes the steps of the method according to any one of claims 1 to 10.

13. A computer-readable storage medium, characterized in that, It includes a computer program. When the computer program runs on an electronic device, the computer program is used to cause the electronic device to execute the steps of the method according to any one of claims 1 to 10.

14. A computer program product, characterized in that, It includes a computer program stored in a computer-readable storage medium, and a processor of an electronic device reads and executes the computer program from the computer-readable storage medium, so that the electronic device executes the steps of any one of claims 1 to 10.