Image label labeling method and device, electronic equipment and storage medium
By performing entity recognition and style feature extraction of images, and combining clustering technology, image label annotation is automated, the problem of inaccurate manual labeling is solved, and the accuracy and automation of image labels are improved.
Patent Information
- Application Number
- CN202510499590.5
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-04-18
- Publication Date
- 2025-07-18
AI Technical Summary
The existing image generation model needs to manually label image tags during training, resulting in inaccurate labels and affects the accuracy of the generated image.
By realizing the target image, extracting image style features and clustering, using entity description text for labeling, establishing a mapping relationship between image content and style, and reducing manual annotation dependence.
It improves the accuracy and automation of image labels, avoids inefficiency and errors of manual labels, and ensures the unified expression and style consistency of image semantic information.
Smart Images

Figure CN120339661A_ABST
Abstract
Description
Technical Field
[0001] This application relates to the field of image processing technologies, and in particular, to a method and device for annotating image tags, an electronic device, and a storage medium. Background Art
[0002] The style of an image refers to the design features of the image, and images of different styles are suitable for different application scenarios. For example, during the insurance marketing activities during the Spring Festival, the elements in the image should more reflect the characteristics of the Spring Festival. In the related technologies of image generation, relevant image generation models are usually used to generate images of the required style. However, this image generation model needs to be pre-trained, and most of the image tags required for training this image generation model are manually annotated, and there are situations where the annotation is inaccurate or even incorrect, which in turn affects the accuracy of the image generated by this image generation model. Therefore, how to improve the annotation accuracy of image tags has become an urgent problem to be solved. Summary of the Invention
[0003] The main purpose of the embodiments of this application is to propose a method and device for annotating image tags, an electronic device, and a storage medium, aiming to improve the annotation accuracy of image tags.
[0004] To achieve the above object, the first aspect of the embodiments of this application proposes a method for annotating image tags, and the method includes:
[0005] Obtain a target image;
[0006] Perform entity recognition on the target image to obtain an entity description text;
[0007] Based on the entity description text, perform image style extraction on the target image to obtain the image style feature of the target image;
[0008] Based on the image style feature, perform image clustering on the target image to obtain a clustering region;
[0009] Obtain the target description text by obtaining the entity description text of the same clustering region, and obtain the current image by obtaining the target images belonging to the same clustering region;
[0010] Perform label annotation on the current image based on the target description text.
[0011] In some embodiments, perform text generation on the target image through a preset text description generation model to obtain an image description text; wherein, the image description text is used to describe the target image;
[0012] Perform entity recognition on the image description text to obtain the image entity of the target image;
[0013] Based on the image entity, perform entity description on the target image to obtain the entity description text.
[0014] In some embodiments, perform target detection on the target image to obtain detection entities;
[0015] Obtain the union of the detection entities and the semantic entities to get the union entity of the target image;
[0016] Perform entity description on the union entity to obtain the entity description text of the target image.
[0017] In some embodiments, perform image encoding on the target image to obtain image encoding;
[0018] Perform text encoding on the entity description text to obtain text encoding;
[0019] Determine the image style feature based on the difference between the image encoding and the text encoding.
[0020] In some embodiments, calculate the cosine similarity between any two of the target images based on the image style feature to obtain image similarity;
[0021] Perform image clustering on the target images based on a preset threshold and the image similarity to obtain the clustering region.
[0022] In some embodiments, perform extraction of common description labels based on the target description text to obtain candidate common labels;
[0023] Count the frequencies of the candidate common labels to obtain candidate label frequencies;
[0024] Perform label annotation on the current image based on the target description text, the candidate common labels, and the candidate label frequencies of the candidate common labels.
[0025] In some embodiments, determine the importance of the candidate common labels based on the entity description text through an attention model;
[0026] Determine the confidence of the candidate common labels based on the importance and the candidate label frequencies;
[0027] Filter the candidate common labels based on the confidence to obtain target common labels;
[0028] Perform label annotation on the current image based on the target common labels.
[0029] To achieve the above object, a second aspect of the embodiments of the present application proposes an annotation device for image tags, the device comprising:
[0030] A first acquisition module, configured to acquire a target image;
[0031] An entity recognition module, configured to perform entity recognition on the target image to obtain an entity description text;
[0032] A style extraction module, configured to perform image style extraction on the target image based on the entity description text to obtain an image style feature of the target image;
[0033] An image clustering module, configured to perform image clustering on the target image based on the image style feature to obtain a clustering region;
[0034] A second acquisition module, configured to acquire the entity description text of the same clustering region to obtain a target description text, and acquire the target images belonging to the same clustering region to obtain a current image;
[0035] A tag annotation module, configured to perform tag annotation on the current image based on the target description text.
[0036] To achieve the above object, a third aspect of the embodiments of the present application proposes an electronic device, the electronic device comprising a memory and a processor, the memory storing a computer program, and the processor implementing the method described in the first aspect when executing the computer program.
[0037] To achieve the above object, a fourth aspect of the embodiments of the present application proposes a computer-readable storage medium, the computer-readable storage medium storing a computer program, and the computer program implementing the method described in the first aspect when executed by a processor.
[0038] A method and device for annotating image tags, an electronic device, and a storage medium proposed in the present application. By performing entity recognition on a target image, an entity description text that can reflect the semantic information of the image is obtained. On this basis, the style features of the image are further extracted, and a mapping relationship between the image content and the style is established, thereby improving the accuracy of image style feature extraction. Secondly, image clustering is performed using the image style features, and images with similar styles are divided into the same clustering region. Then, the entity description texts within the clustering region are used as target description texts, and the images within the clustering region are used as target images and current images, respectively, to achieve the construction of an image set with a consistent style semantic background. Finally, based on the target description texts, the current images in the same clustering region are annotated with tags, which can avoid the inefficiency and errors caused by completely relying on manual annotation of each image one by one, and realize the automatic classification of image semantic information while maintaining style consistency, thereby significantly improving the accuracy and automation of image tag annotation. BRIEF DESCRIPTION OF THE DRAWINGS
[0039] Figure 1 is a flowchart of the method for annotating image tags provided in an embodiment of the present application;
[0040] Figure 2 is Figure 1 a flowchart of step S102 in
[0041] Figure 3 is Figure 2 a flowchart of step S203 in
[0042] Figure 4 is Figure 1 a flowchart of step S103 in
[0043] Figure 5 is Figure 1 a flowchart of step S104 in
[0044] Figure 6 is Figure 1 a flowchart of step S106 in
[0045] Figure 7 is Figure 3 a flowchart of step S603 in
[0046] Figure 8 is a schematic structural diagram of the device for annotating image tags provided in an embodiment of the present application;
[0047] Figure 9 is a schematic hardware structure diagram of the electronic device provided in an embodiment of the present application. DETAILED DESCRIPTION OF THE EMBODIMENTS
[0048] In order to make the objectives, technical solutions and advantages of the present application more clear and understandable, the present application will be further described in detail below with reference to the accompanying drawings and embodiments. It should be understood that the specific embodiments described herein are only used to explain the present application and are not used to limit the present application.
[0049] It should be noted that although the functional modules are divided in the device schematic diagram and the logical order is shown in the flowchart, in some cases, the steps shown or described can be executed in a different order from the module division in the device or the order in the flowchart. The terms "first", "second", etc. in the description, claims and the above-mentioned drawings are used to distinguish similar objects and do not necessarily need to describe a specific order or sequence.
[0050] Unless otherwise defined, all technical and scientific terms used herein have the same meaning as commonly understood by those skilled in the technical field to which this application belongs. The terms used herein are only for the purpose of describing the embodiments of this application and are not intended to limit this application.
[0051] First, several terms involved in the present application are parsed:
[0052] Attention model: The attention model is a technology widely used in the field of artificial intelligence, especially in machine learning and deep learning, to enhance the model's ability to selectively focus on information. This model was initially inspired by the human visual attention mechanism and can simulate the ability to focus on certain parts when processing a large amount of information. In the fields of computer science and artificial intelligence, especially in applications such as natural language processing (NLP), image recognition, and speech recognition, the attention model has become one of the key technologies to improve model performance. The attention model helps the model more effectively extract important features by assigning different importance weights to different parts of the data, thereby improving the information processing process. For example, in machine translation and text summarization tasks, the attention model can help the system more accurately capture the keywords and phrases in the sentence, thereby generating more natural and accurate translations and summaries. In addition, the attention mechanism is also used to enhance recurrent neural networks (RNNs) and Transformer models. In this way, it changes the way the model processes sequential data and greatly improves the processing efficiency and effect. The attention model is not limited to language processing and also shows its strong adaptability and effectiveness in many other technical fields such as video processing and music generation.
[0053] Confidence: Confidence is a metric used in the fields of statistics and artificial intelligence to measure the accuracy of prediction results. It expresses the degree of certainty of the model regarding the prediction results it gives. In machine learning and data science, confidence is commonly used in classification and regression tasks to help researchers and engineers evaluate the reliability of model predictions. For example, in a machine learning model, confidence can be expressed as a probability value, and the closer this value is to 1, the more confident the model is in the prediction result. The calculation of confidence is not limited to traditional statistical methods and is also widely applied in fields such as deep learning, pattern recognition, and natural language processing. In natural language processing (NLP), confidence assessment is commonly used in tasks such as speech recognition, sentiment analysis, and text classification to determine the degree of certainty of the model regarding its recognition or classification results.
[0054] The style of an image refers to the design features of the image, and images of different styles are suitable for different application scenarios. For example, during the insurance marketing activities during the Spring Festival, the elements in the image should more prominently reflect the characteristics of the Spring Festival. In related technologies for image generation, relevant image generation models are usually used to generate images of the required style. However, this image generation model needs to be pre-trained, and most of the image labels required for training this image generation model are manually annotated. This annotation method has situations of inaccurate or even incorrect annotation, which in turn affects the accuracy of the images generated by this image generation model. Therefore, how to improve the annotation accuracy of image labels has become an urgent problem to be solved.
[0055] Based on this, the embodiments of the present application provide a method and device for annotating image labels, an electronic device, and a storage medium, aiming to improve the annotation accuracy of image labels.
[0056] The method and device for annotating image labels, an electronic device, and a storage medium provided by the embodiments of the present application are specifically described through the following embodiments. First, the method for annotating image labels in the embodiments of the present application is described.
[0057] The embodiments of the present application can acquire and process relevant data based on artificial intelligence technology. Among them, Artificial Intelligence (AI) is the theory, method, technology, and application system that uses digital computers or machines controlled by digital computers to simulate, extend, and expand human intelligence, perceive the environment, acquire knowledge, and use knowledge to obtain the best results.
[0058] Artificial intelligence basic technologies generally include technologies such as sensors, dedicated artificial intelligence chips, cloud computing, distributed storage, big data processing technology, operation / interaction systems, and mechatronics. Artificial intelligence software technologies mainly include several major directions such as computer vision technology, robotics, biometric technology, speech processing technology, natural language processing technology, and machine learning / deep learning.
[0059] The image label annotation method provided by the embodiments of this application relates to the field of image processing technology. The image label annotation method provided by the embodiments of this application can be applied to a terminal, or to a server, or to software running on a terminal or a server. In some embodiments, the terminal can be a smart phone, a tablet computer, a laptop computer, a desktop computer, etc.; the server can be configured as an independent physical server, or as a server cluster or distributed system composed of multiple physical servers, or as a cloud server providing basic cloud computing services such as cloud services, cloud databases, cloud computing, cloud functions, cloud storage, network services, cloud communications, middleware services, domain name services, security services, CDN, and big data and artificial intelligence platforms; the software can be an application that implements the image label annotation method, etc., but is not limited to the above forms.
[0060] This application can be used in many general or specific computer system environments or configurations. For example: personal computers, server computers, handheld or portable devices, tablet devices, multi-processor systems, microprocessor-based systems, set-top boxes, programmable consumer electronics, network PCs, minicomputers, mainframe computers, distributed computing environments including any of the above systems or devices, and so on. This application can be described in the general context of computer-executable instructions executed by a computer, such as program modules. Generally, program modules include routines, programs, objects, components, data structures, etc. that perform specific tasks or implement specific abstract data types. This application can also be practiced in a distributed computing environment where tasks are performed by remote processing devices connected through a communication network. In a distributed computing environment, program modules can be located in local and remote computer storage media including storage devices.
[0061] It should be noted that in each specific embodiment of this application, when it comes to performing relevant processing based on data related to the user's identity or characteristics, such as user information, user behavior data, user historical data, and user location information, the user's permission or consent will be obtained first. Moreover, the collection, use, and processing of these data will comply with relevant laws, regulations, and standards. In addition, when the embodiments of this application need to obtain the user's sensitive personal information, the user's separate permission or separate consent will be obtained through methods such as pop-up windows or redirecting to a confirmation page. After clearly obtaining the user's separate permission or separate consent, the necessary user-related data for the normal operation of the embodiments of this application will be obtained.
[0062] Figure 1 is an alternative flowchart of the image label annotation method provided by the embodiments of this application. Figure 1The method in [it] may include but is not limited to steps S101 to S106.
[0063] Step S101: Obtain a target image;
[0064] Step S102: Perform entity recognition on the target image to obtain an entity description text;
[0065] Step S103: Based on the entity description text, extract the image style of the target image to obtain the image style feature of the target image;
[0066] Step S104: Based on the image style feature, perform image clustering on the target image to obtain a clustering region;
[0067] Step S105: Obtain the entity description text of the same clustering region to obtain a target description text, and obtain the target images belonging to the same clustering region to obtain the current image;
[0068] Step S106: Perform label annotation on the current image based on the target description text.
[0069] Steps S101 to S106 illustrated in the embodiments of the present application, by performing entity recognition on the target image, obtain an entity description text that can reflect the semantic information of the image. On this basis, further extract the style feature of the image, establish a mapping relationship between the image content and the style, so as to improve the accuracy of image style feature extraction. Secondly, use the image style feature for image clustering, divide the images with similar styles into the same clustering region, and then use the entity description texts within the clustering region as the target description texts, and use the images within the clustering region as the target images and the current images, so as to realize the construction of an image set with a consistent style semantic background. Finally, perform label annotation on the current images in the same clustering region based on the target description text, which can avoid the inefficiency and error problems caused by completely relying on manual annotation of each image one by one, and realize the automatic classification of image semantic information while maintaining style consistency, thus significantly improving the accuracy and automation of image label annotation.
[0070] In step S101 of some embodiments, a target image is obtained. The target image is the original image to be subjected to label annotation processing, usually existing in the form of a static image, and includes entity information that can be used to characterize the semantic content of the image and visual elements reflecting the design intention of the image. For example, during the insurance marketing activities during the Spring Festival, the target image can be the activity poster image used for publicity, and such an image may contain elements with obvious festival symbolic meanings such as "lanterns", "firecrackers", "fortune characters", or "red main tone". The ways to obtain the target image can include, but are not limited to: obtaining it by shooting with an image acquisition device, such as using the camera of a smart terminal to obtain a static image; or calling historical image data from a preset image database; or scraping image data that conforms to the predetermined keyword rules from publicly available Internet resources, etc.
[0071] Please refer to Figure 2 , in some embodiments, step S102 may include, but is not limited to, steps S201 to S203:
[0072] Step S201, generate text for the target image through a preset text description generation model to obtain an image description text; wherein, the image description text is used to describe the target image;
[0073] Step S202, perform entity recognition on the image description text to obtain the image entities of the target image;
[0074] Step S203, based on the image entities, perform entity description on the target image to obtain an entity description text.
[0075] In steps S201 to S203 illustrated in the embodiments of the present application, the preset text description generation model processes the target image, and can automatically generate an image description text according to the visual content of the image, thereby converting unstructured image data into semantic information that can be expressed in natural language, avoiding the traditional method of relying on manual writing of description texts, and improving the consistency and objectivity of image semantic expression. Secondly, based on the generated image description text, entity recognition processing is performed to extract the key entities with semantic meanings in the target image. Finally, entity description is performed on the target image based on the image entities of the target image, converting the target image from an unquantifiable text into a quantifiable entity description text with entities as the carrier, thereby effectively overcoming problems such as the dependence on manual processing, unclear structure, and inconsistent expression in image semantic expression in related technologies.
[0076] In step S201 of some embodiments, the text description generation model is CLIP interrogator. CLIP interrogator is a pre-trained model based on the image-text alignment mechanism, which combines the joint embedding space of the image encoder and the text encoder. After encoding the input image, it retrieves the text description that is most semantically relevant to the image in the large-scale text space, so as to generate the image description text for expressing the semantic content of the image. The image description text is a semantic expression text in natural language form, which is used to reflect the scene background, main objects, object categories, action features, etc. contained in the target image. It is usually expressed in a complete sentence pattern and can accurately match the key information contained in the image at the lexical, word order, and semantic levels. For example, "A woman in a red dress stands under a cherry tree."
[0077] In step S202 of some embodiments, the entity recognition process can be implemented based on a variety of natural language processing technologies. The specific methods include Named Entity Recognition (NER), Dependency Parsing, Semantic Role Labeling, etc. Image entities are keywords or phrases with specific semantic attributes extracted from the image description text, usually including person names, locations, object categories, action behaviors, scene elements, etc. Their forms can be noun phrases or structured fields, such as "woman", "red dress", "cherry tree", etc.
[0078] Please refer to Figure 3 , in some embodiments, step S203 may include but is not limited to steps S301 to S303:
[0079] Step S301, perform object detection on the target image to obtain detected entities;
[0080] Step S302, obtain the union between the detected entities and the semantic entities to obtain the union entities of the target image;
[0081] Step S303, perform entity description on the union entities to obtain the entity description text of the target image.
[0082] Steps S301 to S303 shown in the embodiments of the present application can detect the target entity and the category information of the target entity from the image pixel level through target detection operations on the target image, that is, the detected entity is obtained. Then, a set operation is performed on the detected entity and the semantic entity obtained based on text processing to construct a union entity set containing visual content and text semantics. The union entity covers the explicit targets that can be recognized by the model in the image and the potential semantic units reflected in the text generation process, enhancing the integrity of entity information and semantic diversity. Finally, an entity description text is generated based on the union entity to express various semantic elements contained in the image in the form of natural language, thereby realizing the unified expression of image content and semantic information. In this way, while improving the label coverage, the semantic accuracy and context consistency of the label expression are ensured, thus solving the problems of missing annotation information and incomplete semantic expression in image labels.
[0083] In step S301 of some embodiments, target detection is a technology in which a user identifies and locates a target in an image and outputs the corresponding target category information. Target detection can be implemented through various deep neural network models. For example, Faster R-CNN, YOLO series, RetinaNet, and DETR, etc. These models can output detection results containing category labels, bounding box coordinates, and confidence scores in a structured form. The detected entity is a set of semantic objects output by the target detection model, used to represent the entity units that have been detected in the image.
[0084] In step S302 of some embodiments, the union of the detected entity and the semantic entity is obtained because the detected entity is only obtained based on image pixel-level feature extraction and may not cover the abstract elements that are not explicitly presented in the image but have significant semantics. The semantic entity supplements the high-level semantic information that is difficult to directly obtain through visual means in the image. Therefore, performing a union operation on the detected entity and the semantic entity can fuse two dimensions of semantic sources in the entity set, thereby forming a union entity set with a more complete structure and richer semantics.
[0085] In step S303 of some embodiments, entity description processing is performed on the union entity to generate an entity description text for the target image, which can be achieved by the following method: connecting the union entity in sequence according to the natural language expression method and separating them with commas to form the entity description text.
[0086] Please refer to Figure 4 , in some embodiments, step S103 may include but is not limited to steps S401 to S403:
[0087] Step S401, perform image encoding on the target image to obtain an image encoding;
[0088] Step S402: Perform text encoding on the entity description text to obtain a text encoding.
[0089] Step S403: Determine the image style feature based on the difference between the image encoding and the text encoding.
[0090] Steps S401 to S403 illustrated in the embodiments of the present application first perform image encoding on the target image to obtain the corresponding image encoding, then perform text encoding on the entity description text to obtain the corresponding text encoding, and further determine the image style feature of the target image based on the difference between the above-mentioned image encoding and text encoding. The image encoding can effectively capture the visual content information of the target image, while the text encoding can fully reflect the semantic features in the entity description text. Through the difference analysis between the two, the style difference between the image and the text entity content is accurately characterized, and the style difference is used as the image style. Compared with directly extracting the image style in the related art, subtracting the image encoding from the text encoding to obtain the difference as the image style feature reflects the difference in the style dimension between the image visual information and the text semantic information, effectively removing the interference factors generated by the inherent entity features in the image content itself and the text description, avoiding the repeated expression of the image content and the semantic entity, and thus more accurately highlighting the unique style elements in the image, significantly improving the accuracy and distinctiveness of the image style extraction.
[0091] In step S401 of some embodiments, the image encoding refers to characterizing visual features such as pixel information, texture structure, edge contour, and spatial distribution in the target image, and constructing a vector representation for characterizing the image content in numerical form. The image encoding can be implemented by an image encoding model based on a convolutional neural network (CNN), such as models like ResNet (Residual Neural Network), EfficientNet (Efficient Neural Network), or Vision Transformer.
[0092] In step S402 of some embodiments, the text encoding refers to characterizing the word order, context dependence, grammatical structure, and semantic information in the entity description text, and constructing a vector representation for characterizing the text semantics in numerical form. The text encoding can be implemented by a text encoding model based on deep learning, such as the text encoder in BERT (Bidirectional Encoder Representations from Transformers), RoBERTa, ERNIE, or CLIP. By encoding the text, a text encoding result with context semantic recognition ability is output, providing a semantic basis for the difference calculation of the image style feature.
[0093] In step S403 of some embodiments, the reason for determining the image style feature based on the difference between the image encoding and the text encoding is that the image encoding can effectively capture the visual content information of the target image, while the text encoding can fully reflect the semantic features in the entity description text. Subtracting the image encoding from the text encoding to obtain the difference as the image style feature reflects the difference between the image visual information and the text semantic information in the style dimension, effectively removing the interference factors generated by the inherent entity features in the image content itself and the text description, avoiding the repeated expression of the image content and the semantic entity, and thus more accurately highlighting the unique style elements in the image, significantly improving the accuracy and distinctiveness of image style extraction.
[0094] Please refer to Figure 5 , in some embodiments, step S104 includes but is not limited to steps S501 to S502:
[0095] Step S501, based on the image style feature, calculate the cosine similarity between any two target images to obtain the image similarity;
[0096] Step S502, based on the preset threshold and the image similarity, perform image clustering on the target images to obtain the clustering region.
[0097] Steps S501 to S502 illustrated in the embodiments of the present application, by calculating the cosine similarity between any two target images, are used to quantify the similarity relationship between the image style features, and then based on the preset threshold and the image similarity, perform image clustering on the target images to obtain a clustering region with a relatively high style consistency. Thus, the images are classified according to their styles to obtain the clustering region.
[0098] In step S501 of some embodiments, the cosine similarity is a numerical index used to measure the similarity degree between two target images in the image style feature space, specifically represented as the cosine value of the angle between two image style feature vectors. The value range of the cosine similarity is usually [-1, 1], where the closer the value is to 1, the higher the similarity degree between the two target images in the style feature dimension.
[0099] In step S502 of some embodiments, the preset threshold refers to the similarity boundary parameter preset for image clustering, which is used to determine whether two target images should be classified into the same clustering region. Specifically, if the cosine similarity between any two target images is greater than or equal to the preset threshold, it is considered that the two target images have sufficient similarity in style features and can be classified into the same category. The preset threshold can be set according to the distribution law of image style features on the training set or by experimental tuning. The clustering region refers to the image set formed by dividing target images with similar image style features into the same category. The image clustering method can be a density-based clustering algorithm, a hierarchical clustering algorithm, a partitioning-based clustering algorithm, or a graph model-based clustering algorithm.
[0100] Please refer to Figure 6 , in some embodiments, step S106 includes but is not limited to steps S601 to S603:
[0101] Step S601, based on the target description text, extract common description labels to obtain candidate common labels;
[0102] Step S602, count the frequencies of the candidate common labels to obtain candidate label frequencies;
[0103] Step S603, based on the target description text, candidate common labels, and candidate label frequencies of the candidate common labels, label the current image.
[0104] Steps S601 to S603 illustrated in the embodiments of the present application extract common description labels based on the target description text to obtain candidate common labels that can reflect the common semantic meanings of multiple image styles. Then, further count the occurrence frequencies of each candidate common label in the target description text to obtain candidate label frequencies for characterizing the common intensity. The candidate label frequencies reflect the stability and representativeness of different labels in the clustering region. Finally, combine the target description text, candidate common labels, and candidate label frequencies to label the current image, enhancing the generality of the labels and the matching degree with the image content.
[0105] In step S601 of some embodiments, the candidate common label refers to the set of label texts extracted from the target description text that can reflect the common descriptive content of multiple target images at the semantic level. As described above, the target description text is the union entity connected in sequence according to the natural language expression method and separated by commas to form the description text. The extraction of the common description label can be achieved in the following way: divide each target description text according to the comma to obtain the entity corresponding to each target description text, and then take the union of all target description texts to obtain the candidate common label.
[0106] In step S602 of some embodiments, the candidate label frequency refers to the statistical number of occurrences of candidate common labels in the target description text set, which is used to quantify the frequency of occurrence of each candidate common label in the entire text corpus. The candidate label frequency can be represented in the form of an integer value. The larger the value, the more times the label appears in the text within the clustering region.
[0107] Please refer to Figure 7 , in some embodiments, step S603 may include but is not limited to steps S701 to S704:
[0108] Step S701, based on the entity description text, determine the importance of the candidate common label through an attention model;
[0109] Step S702, based on the importance and the candidate label frequency, determine the confidence of the candidate common label;
[0110] Step S703, based on the confidence, screen the candidate common labels to obtain the target common labels;
[0111] Step S704, label the current image based on the target common labels.
[0112] Steps S701 to S704 illustrated in the embodiments of the present application, through an attention model, combined with the entity description text, determine the strength of the semantic association between the candidate common label and the entity description text, and obtain the importance that can reflect the expression weight at the semantic level. Then, combine the importance of the candidate common label with the candidate label frequency to calculate the confidence of the candidate common label, which is used to comprehensively represent the representative degree of the candidate common label in two dimensions of significance and generality. Subsequently, based on the confidence, screen the candidate common labels to obtain the target common labels that can accurately represent the common semantics within the image clustering region. Finally, use the target common labels to label the current image to complete the label output with accurate semantics and unified expression.
[0113] In step S701 of some embodiments, the attention model refers to a neural network model used to calculate the semantic association strength between the entity description text and the candidate common label. By assigning attention weights to each label, it highlights the information content with higher semantic relevance. This attention model can be implemented through various deep learning structures. For example, the self-attention mechanism in the Transformer structure, the multi-head attention mechanism in the Bidirectional Encoder Representations from Transformers (BERT), or an attention network based on the encoder-decoder structure can be used. The importance is a value between 0 and 1. The larger the value, the stronger the semantic association between the candidate common label and the entity description text.
[0114] In step S702 of some embodiments, the confidence can be determined by multiplication or by multiplying according to a preset weight, which is not specifically limited in this application. The confidence is a numerical value, such as 5.3, and is used to comprehensively represent the representative degree of the candidate common label in the two dimensions of significance and generality.
[0115] It should be noted that the reason for determining the confidence of the candidate common label based on the importance and the frequency of the candidate label is that the semantic representativeness of the candidate common label in the entity description text and the stability of its appearance in the target description text set are two core factors for measuring the effectiveness and generality of the label. Among them, the importance reflects the semantic association strength between the candidate common label and the entity description text and can represent the prominence degree of the label in semantic expression; while the candidate label frequency reflects from a statistical perspective the degree of repeated appearance of the label in multiple image descriptions within the clustering region and has the ability to measure the commonness and universality of the label. If the confidence judgment is only based on a certain dimension, it is easy to cause the label screening result to be biased towards labels that are semantically related but rare, or labels that appear frequently but have insufficient semantic weight, thereby affecting the accuracy of the final label annotation. Therefore, by jointly modeling the importance and the candidate label frequency, a comprehensive evaluation of the candidate common label is achieved in the two dimensions of semantic strength and appearance frequency, which can more comprehensively reflect the representative degree of the label in the current image clustering region, thereby improving the reliability of the confidence evaluation and providing a reasonable basis for the subsequent screening of the target common label.
[0116] In step S703 of some embodiments, the screening method can be to set a screening threshold based on the confidence, specifically including: when the confidence of the candidate common label is greater than or equal to the threshold, the corresponding label is retained as the target common label.
[0117] The screening method can also be a ranking screening method, specifically including: sorting all candidate common labels in descending order according to the confidence value, and selecting the top several labels to form the target common label set.
[0118] Please refer to Figure 8 , this application embodiment also provides an image label annotation device, which can implement the above image label annotation method. The device includes:
[0119] The first acquisition module 801 is used to acquire a target image;
[0120] The entity recognition module 802 is used to perform entity recognition on the target image to obtain an entity description text;
[0121] The style extraction module 803 is used to perform image style extraction on the target image based on the entity description text to obtain the image style feature of the target image;
[0122] An image clustering module 804, configured to perform image clustering on a target image based on image style features to obtain a clustering region;
[0123] A second acquisition module 805, configured to acquire an entity description text of the same clustering region to obtain a target description text, and acquire target images belonging to the same clustering region to obtain a current image;
[0124] A label annotation module 806, configured to perform label annotation on the current image based on the target description text.
[0125] The specific implementation manner of the image label annotation device is basically the same as the specific embodiments of the above image label annotation method, and will not be described herein again.
[0126] An embodiment of the present application further provides an electronic device. The electronic device includes a memory and a processor. The memory stores a computer program, and when the processor executes the computer program, the above image label annotation method is implemented. The electronic device may be any intelligent terminal including a tablet computer, a vehicle-mounted computer, etc.
[0127] Please refer to Figure 9 , Figure 9 , which schematically shows the hardware structure of an electronic device according to another embodiment. The electronic device includes:
[0128] A processor 901, which may be implemented by using a general-purpose CPU (Central Processing Unit), a microprocessor, an application-specific integrated circuit (ASIC), or one or more integrated circuits, etc., and is configured to execute relevant programs to implement the technical solutions provided by the embodiments of the present application;
[0129] A memory 902, which may be implemented in the form of a read-only memory (ROM), a static storage device, a dynamic storage device, or a random access memory (RAM), etc. The memory 902 may store an operating system and other application programs. When implementing the technical solutions provided by the embodiments of this specification through software or firmware, the relevant program codes are stored in the memory 902, and the processor 901 is called to execute the image label annotation method of the embodiments of the present application;
[0130] An input / output interface 903, configured to implement information input and output;
[0131] A communication interface 904, which is used to implement communication and interaction between this device and other devices, can achieve communication through wired means (such as USB, network cable, etc.), and can also achieve communication through wireless means (such as mobile network, WIFI, Bluetooth, etc.);
[0132] A bus 905, which transmits information between various components of the device (such as a processor 901, a memory 902, an input / output interface 903, and a communication interface 904);
[0133] Among them, the processor 901, the memory 902, the input / output interface 903, and the communication interface 904 achieve communication connections with each other inside the device through the bus 905.
[0134] The embodiment of the present application also provides a computer-readable storage medium, which stores a computer program, and when the computer program is executed by a processor, it implements the above-mentioned method for annotating image tags.
[0135] As a non-transitory computer-readable storage medium, the memory can be used to store non-transitory software programs and non-transitory computer-executable programs. In addition, the memory can include high-speed random access memory, and can also include non-transitory memory, such as at least one disk storage device, a flash memory device, or other non-transitory solid-state storage devices. In some embodiments, the memory optionally includes a memory remotely set relative to the processor, and these remote memories can be connected to the processor through a network. Examples of the above network include but are not limited to the Internet, an enterprise intranet, a local area network, a mobile communication network, and combinations thereof.
[0136] The method for annotating image tags, the device for annotating image tags, the electronic device, and the storage medium provided by the embodiment of the present application, through entity recognition of the target image, obtain an entity description text that can reflect the semantic information of the image. On this basis, further extract the style features of the image, establish a mapping relationship between the image content and the style, so as to improve the accuracy of image style feature extraction. Secondly, use the image style features for image clustering, divide images with similar styles into the same clustering area, and then use the entity description texts in the clustering area as the target description texts, and use the images in the clustering area as the target images as the current image, so as to realize the construction of an image set with a consistent style semantic background. Finally, based on the target description text, label the current images in the same clustering area, which can avoid the inefficiency and errors caused by completely relying on manual annotation of each image one by one, and realize the automatic classification of image semantic information while maintaining style consistency, thus significantly improving the accuracy and automation degree of image tag annotation.
[0137] The embodiments described in the embodiments of the present application are to more clearly illustrate the technical solutions of the embodiments of the present application, and do not constitute a limitation on the technical solutions provided by the embodiments of the present application. Those skilled in the art will know that with the evolution of technology and the emergence of new application scenarios, the technical solutions provided by the embodiments of the present application are equally applicable to similar technical problems.
[0138] Those skilled in the art can understand that the technical solutions shown in the figures do not constitute a limitation on the embodiments of the present application, and may include more or fewer steps than those shown in the figures, or combine certain steps, or different steps.
[0139] The device embodiments described above are merely illustrative. The units described as separate components may or may not be physically separated, that is, they may be located in one place, or may be distributed to multiple network units. Some or all of the modules can be selected according to actual needs to achieve the purpose of the solution of this embodiment.
[0140] Those of ordinary skill in the art can understand that all or some of the steps in the methods disclosed above, and the functional modules / units in the systems and devices, can be implemented as software, firmware, hardware, and their appropriate combinations.
[0141] The terms "first", "second", "third", "fourth", etc. (if any) in the specification of the present application and the above-mentioned drawings are used to distinguish similar objects, and do not have to be used to describe a specific order or sequence. It should be understood that the data used in this way can be interchanged under appropriate circumstances, so that the embodiments of the present application described here can be implemented in an order other than those illustrated or described here. In addition, the terms "including" and "having" and any variations thereof are intended to cover non-exclusive inclusion. For example, a process, method, system, product, or device that includes a series of steps or units does not have to be limited to those steps or units clearly listed, but may include other steps or units that are not clearly listed or are inherent to these processes, methods, products, or devices.
[0142] It should be understood that in this application, "at least one (item)" means one or more, and "a plurality" means two or more. "And / or" is used to describe the association relationship of associated objects, indicating that there can be three relationships. For example, "A and / or B" can mean: only A exists, only B exists, and both A and B exist at the same time. Among them, A and B can be singular or plural. The character " / " generally indicates that the associated objects before and after are in an "or" relationship. "At least one (item) of the following" or its similar expression refers to any combination of these items, including any combination of single item (item) or plural items (items). For example, at least one (item) of a, b, or c can mean: a, b, c, "a and b", "a and c", "b and c", or "a and b and c", where a, b, c can be single or multiple.
[0143] In several embodiments provided in this application, it should be understood that the disclosed devices and methods can be implemented in other ways. For example, the device embodiments described above are merely illustrative. For example, the above division of units is only a logical function division. In actual implementation, there can be other division methods. For example, multiple units or components can be combined or integrated into another system, or some features can be ignored or not executed. Another point is that the displayed or discussed couplings or direct couplings or communication connections to each other can be through some interfaces. The indirect couplings or communication connections of devices or units can be in electrical, mechanical or other forms.
[0144] The units described above as separate components may or may not be physically separated. The components displayed as units may or may not be physical units, that is, they can be located in one place, or they can be distributed to multiple network units. Some or all of the units can be selected according to actual needs to achieve the purpose of the solution of this embodiment.
[0145] In addition, in each embodiment of this application, the functional units can be integrated in a processing unit, or each unit can exist physically alone, or two or more units can be integrated in one unit. The above integrated units can be implemented in the form of hardware or in the form of software functional units.
[0146] When the integrated unit is implemented in the form of a software functional unit and sold or used as an independent product, it can be stored in a computer-readable storage medium. Based on this understanding, the technical solution of this application, in essence, or the part that contributes to the prior art, or all or part of this technical solution, can be embodied in the form of a software product. This computer software product is stored in a storage medium and includes multiple instructions for causing a computer device (which may be a personal computer, a server, or a network device, etc.) to execute all or part of the steps of the methods of the various embodiments of this application. The aforementioned storage medium includes: various media that can store programs, such as USB flash drives, mobile hard disks, read-only memory (ROM), random access memory (RAM), magnetic disks, or optical discs.
[0147] The preferred embodiments of the embodiments of this application have been described above with reference to the accompanying drawings, which does not limit the scope of the rights of the embodiments of this application. Any modifications, equivalent replacements, and improvements made by those skilled in the art without departing from the scope and essence of the embodiments of this application shall be within the scope of the rights of the embodiments of this application.
Claims
1. A method for annotating image tags, characterized in that, The method includes: Obtain a target image; Perform entity recognition on the target image to obtain an entity description text; Based on the entity description text, perform image style extraction on the target image to obtain the image style features of the target image; Based on the image style features, perform image clustering on the target image to obtain a clustering region; Obtain the entity description text of the same clustering region to obtain a target description text, and obtain the current image of the target images belonging to the same clustering region; Perform label annotation on the current image based on the target description text.
2. The method according to claim 1, wherein The performing entity recognition on the target image to obtain an entity description text includes: Through a preset text description generation model, perform text generation on the target image to obtain an image description text; wherein, the image description text is used to describe the target image; Perform entity recognition on the image description text to obtain the image entities of the target image; Based on the image entities, perform entity description on the target image to obtain the entity description text.
3. The method according to claim 2, wherein The performing entity description on the target image based on the image entities to obtain the entity description text includes: Perform target detection on the target image to obtain detection entities; Obtain the union between the detection entities and the semantic entities to obtain the union entities of the target image; Perform entity description on the union entities to obtain the entity description text of the target image.
4. The method according to claim 1, wherein The performing image style extraction on the target image based on the entity description text to obtain the image style features of the target image includes: Perform image encoding on the target image to obtain an image encoding; Perform text encoding on the entity description text to obtain a text encoding; Based on the difference between the image encoding and the text encoding, determine the image style features.
5. The method according to claim 1, wherein The performing image clustering on the target image based on the image style features to obtain a clustering region includes: Based on the image style features, calculate the cosine similarity between any two of the target images to obtain an image similarity; Based on a preset threshold and the image similarity, perform image clustering on the target images to obtain the clustering region.
6. The method according to any one of claims 1 to 5, characterized in that The performing label annotation on the current image based on the target description text includes: Based on the target description text, perform extraction of common description labels to obtain candidate common labels; Count the frequencies of the candidate common labels to obtain candidate label frequencies; Based on the target description text, the candidate common labels, and the candidate label frequencies of the candidate common labels, perform label annotation on the current image.
7. The method according to claim 6, wherein The performing label annotation on the current image based on the target description text, the candidate common labels, and the candidate label frequencies of the candidate common labels includes: Through an attention model, based on the entity description text, determine the importance of the candidate common labels; Based on the importance and the candidate label frequencies, determine the confidence of the candidate common labels; Based on the confidence, screen the candidate common labels to obtain target common labels; Perform label annotation on the current image based on the target common label.
8. An annotation device for image tags, characterized in that, The device includes: A first acquisition module, configured to acquire a target image; An entity recognition module, configured to perform entity recognition on the target image to obtain an entity description text; A style extraction module, configured to perform image style extraction on the target image based on the entity description text to obtain an image style feature of the target image; An image clustering module, configured to perform image clustering on the target image based on the image style feature to obtain a clustering region; A second acquisition module, configured to acquire the entity description text of the same clustering region to obtain a target description text, and acquire the target images belonging to the same clustering region to obtain the current image; A label annotation module, configured to perform label annotation on the current image based on the target description text.
9. An electronic device, characterized in that, The electronic device includes a memory and a processor. The memory stores a computer program, and when the processor executes the computer program, it implements the image label annotation method according to any one of claims 1 to 7.
10. A computer-readable storage medium storing a computer program, characterized in that, When the computer program is executed by the processor, it implements the image label annotation method according to any one of claims 1 to 7.