Image content annotation method, device, storage medium and computer equipment
By performing basic annotation and secondary annotation of clustering processing on the image, more accurate target annotation is generated, which solves the problems of time-consuming and labor-intensive image annotation and low retrieval efficiency in the existing technology, and realizes more efficient image management and retrieval.
Patent Information
- Application Number
- CN202311114752.6
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2023-08-29
- Publication Date
- 2025-10-03
- Estimated Expiration
- 2043-08-29
AI Technical Summary
In the existing technology, image annotation methods are time-consuming and labor-intensive, and automatic annotation methods can only identify content elements within the image, resulting in low efficiency in image management and retrieval, and difficulty in locking the target image.
By performing basic annotation on the image, generating a feature map, and combining it with clustering processing in the image library, secondary reasoning annotation is performed to generate more accurate target annotations.
The standardization and retrieval efficiency of image classification are improved, and more accurate annotation information is achieved.
Smart Images

Figure CN117079052B_ABST
Abstract
Description
Technical Field
[0001] The present application relates to the field of image processing technology, and in particular to an image content annotation method, apparatus, storage medium, and computer equipment. Background Art
[0002] In the era of digital photography, people take and store a large number of photos, which are generally compiled into photo albums. In order to make photo management more standardized, the content in the photos is identified and annotated, and the content in the photos is described based on the annotations.
[0003] However, many images captured by users often lack annotations that fully describe their content. For example, a user might capture and save a landscape photo without a detailed caption. The core task of image annotation is to leverage these complex, unannotated images to provide relevant services to internet users.
[0004] Traditional annotation methods are all based on manual annotation, which is time-consuming and labor-intensive. Automatic annotation methods can only identify and annotate content elements within an image. For example, if there are elements of people and tables in a photo, only the content in the photo is annotated. When managing and searching images, the annotated information is generally used as a search element or clustering element. However, this method is relatively rough in image management and can only retrieve and manage based on specific content elements in the image. In devices that store massive amounts of image data, there are significant limitations in retrieving images using the annotation information formed by this method, making it difficult to lock on to the target image, resulting in low efficiency. Summary of the Invention
[0005] The present application provides an image content annotation method, apparatus, storage medium, and computer equipment. By performing basic annotation on images and then performing secondary inference annotation based on the results of image clustering processing and the basic annotation, more accurate annotation information can be obtained, thereby making image classification more standardized and greatly improving the efficiency of subsequent retrieval.
[0006] This application provides an image content annotation method, including:
[0007] Extracting element information from the original image and obtaining basic annotations corresponding to the element information through image recognition;
[0008] generating a feature map of the original image based on the correlation relationship between the basic annotations in the original image;
[0009] Clustering the feature map of the original image with the feature maps of other images in the image library so that the original image and other images in the image library that meet the clustering conditions form a cluster set;
[0010] An inferred annotation is generated according to the basic annotations of all images in the cluster set, and the inferred annotation is determined as the target annotation of the original image.
[0011] Optionally, extracting element information from the original image and obtaining basic annotations corresponding to the element information through image recognition includes:
[0012] Positioning the target in the original image using a target detection algorithm;
[0013] Selecting the target according to the positioning result, and determining the selected image content as element information;
[0014] Identify the element information to obtain the corresponding basic annotation.
[0015] Optionally, extracting element information from the original image and obtaining basic annotations corresponding to the element information through image recognition includes:
[0016] Perform pixel classification on the original image using a semantic segmentation algorithm;
[0017] Determine each divided pixel set as corresponding element information according to the classification result;
[0018] Identify the element information to obtain the corresponding basic annotation.
[0019] Optionally, clustering the feature map of the original image with the feature maps of other images in the image gallery so that the original image and other images in the image gallery that meet the clustering conditions form a cluster set, including:
[0020] generating a semantic description according to a graph structure corresponding to the feature graph of the original image, and using the semantic description as semantic information of the original image;
[0021] Calculating the similarity between the semantic information of the original image and the semantic information corresponding to the feature maps of other images in the image library;
[0022] The image whose similarity exceeds a preset value and the original image are determined as a cluster set.
[0023] Optionally, generating an inferred annotation based on the basic annotations of all images in the cluster set, and determining the inferred annotation as the target annotation of the original image includes:
[0024] Obtaining basic annotations for all images in the cluster set;
[0025] All basic annotations are input into a preset neural network model to output inference annotations of the original image, and the inference annotations are determined as target annotations of the original image.
[0026] Optionally, the method of generating a semantic description according to the graph structure corresponding to the feature graph of the original image includes:
[0027] Performing feature representation on the graph structure of the feature graph through a graph neural network, and generating the semantic description based on the feature representation result; or,
[0028] The graph structure of the feature graph is semantically represented through a knowledge graph, and the semantic description is generated based on the semantic representation.
[0029] Optionally, obtaining basic annotations of all images in the cluster set includes:
[0030] Extracting time information and position information of the original image;
[0031] Based on the time information and the location information, remove images that do not meet corresponding time conditions and location conditions from all images in the cluster set;
[0032] Get the basic annotations of all remaining images in the cluster set.
[0033] Optionally, the method further includes:
[0034] Acquiring ambient sound data and / or human voice data recorded when the original image is captured;
[0035] The environmental sound data and / or human voice data are analyzed, and the analysis results are added to the basic annotation or target annotation of the original image.
[0036] This application also provides an image content annotation device, comprising:
[0037] A first recognition unit is used to extract element information from the original image and obtain basic annotations corresponding to the element information through image recognition;
[0038] A generating unit, configured to generate a feature map of the original image based on the correlation relationship between the basic annotations in the original image;
[0039] a clustering unit, configured to perform clustering processing on the feature map of the original image and the feature maps of other images in the image library, so that the original image and other images in the image library that meet the clustering conditions form a cluster set;
[0040] The second recognition unit is configured to generate an inferred annotation based on the basic annotations of all images in the cluster set, and determine the inferred annotation as the target annotation of the original image.
[0041] The present application also provides a computer device, comprising a memory and a processor, wherein the processor executes the above-mentioned image content annotation method by calling a computer program stored in the memory.
[0042] The image content annotation method, device, storage medium and computer equipment provided by the present application can extract element information from the original image, and obtain the basic annotation corresponding to the element information through image recognition, generate a feature map of the original image based on the correlation between the basic annotations in the original image, cluster the feature map of the original image with the feature maps of other images in the gallery, so that the original image and other images in the gallery that meet the clustering conditions form a cluster set, generate inference annotations based on the basic annotations of all images in the cluster set, and determine the inference annotations as the target annotations of the original image. The solution provided by the present application can obtain more accurate annotation information by performing basic annotations on the image and then performing secondary annotations based on the results of clustering processing, making image classification more standardized and greatly improving the efficiency of subsequent retrieval. BRIEF DESCRIPTION OF THE DRAWINGS
[0043] In order to more clearly illustrate the technical solutions in the embodiments of the present application, the following briefly introduces the drawings required for use in the description of the embodiments. Obviously, the drawings described below are only some embodiments of the present application. For those skilled in the art, other drawings can be obtained based on these drawings without creative work.
[0044] Figure 1 A flowchart of the image content annotation method provided in an embodiment of the present application.
[0045] Figure 2 Another flowchart of the image content annotation method provided in an embodiment of the present application.
[0046] Figure 3 A schematic diagram of an original image provided by an embodiment of the present invention.
[0047] Figure 4 A schematic diagram of the characteristic graph structure provided by an embodiment of the present invention.
[0048] Figure 5 A schematic diagram of a flow chart of an image content annotation apparatus provided in an embodiment of the present application.
[0049] Figure 6 Another flowchart of the image content annotation device provided in an embodiment of the present application is shown.
[0050] Figure 7 A schematic diagram of the computer device structure provided in an embodiment of the present invention. DETAILED DESCRIPTION
[0051] The following will be combined with the drawings in the embodiments of this application to clearly and completely describe the technical solutions in the embodiments of this application. Obviously, the embodiments described are only part of the embodiments of this application, not all of the embodiments. Based on the embodiments in this application, all other embodiments obtained by those skilled in the art without making creative efforts are within the scope of protection of this application.
[0052] It should be noted that, in this article, the terms "include", "comprises" or any other variations thereof are intended to cover non-exclusive inclusion, so that a process, method, article or device that includes a series of elements includes not only those elements, but also includes other elements that are not explicitly listed, or also includes elements inherent to such process, method, article or device. In the absence of further restrictions, an element defined by the sentence "includes a..." does not exclude the presence of other identical elements in the process, method, article or device that includes the element. In addition, parts, features, and elements with the same name in different embodiments of the present application may have the same meaning or different meanings, and their specific meanings need to be determined by their explanation in the specific embodiment or further combined with the context of the specific embodiment.
[0053] The embodiments of the present application provide an image content annotation method, apparatus, storage medium, and computer device. Specifically, the image content annotation method of the embodiments of the present application can be executed by a computer device or server, wherein the computer device can be a terminal. The terminal can be a smartphone, tablet computer, laptop computer, touch screen, game console, personal computer (PC), personal digital assistant (PDA), smart home device, etc. The terminal can also include a client, which can be a media player client or instant messaging client, etc.
[0054] See also Figure 1 , the specific process of the image content annotation method can be as follows:
[0055] Step 101: extract at least one element information from an original image, and obtain basic annotations corresponding to the at least one element information through image recognition.
[0056] In one embodiment, it is necessary to first obtain an original image. A specific implementation method for obtaining the original image may be to obtain a photograph in a digital image format (e.g., BMP, JPG, etc.), such as by using a digital camera or mobile phone to instantly take a photo of a person. Specifically, the electronic device may receive an image acquisition instruction, which may activate the electronic device's camera and capture the current scene to obtain the original image. It will be readily apparent to those skilled in the art that the original image may also be obtained through methods such as network downloading, video screenshots, and photo scanning, and the embodiments of the present invention are not limited thereto.
[0057] After obtaining the original image, at least one element information in the original image can be further extracted. Specifically, the original image data can be processed to extract key features in the image data, so that the various elements in the image can be divided into regions based on these features, such as people, animals, objects, background, etc. in a photo. Then, these elements are further image recognized, and the recognition results are annotated, such as through AI image recognition tools or pre-trained recognition models to obtain specific annotation information of the above elements in the original image. For example, the recognized content elements can be annotated as "cat", "dog", "sofa", etc. as the basic annotation of the original image.
[0058] In one embodiment, when extracting at least one element information from the original image, target detection can also be performed on the original image data using a target detection algorithm, such as using a YOLO algorithm to generate a bounding box, select each content element in the original image, and identify the selected content to perform basic annotation. Specifically, the steps of extracting at least one element information from the original image and obtaining basic annotations corresponding to the at least one element information through image recognition can include: locating at least one target in the original image using a target detection algorithm, selecting the target based on the positioning result, determining the selected image content as element information, and identifying the element information to obtain corresponding basic annotations.
[0059] In another embodiment, when extracting at least one element information from the original image, the original image data may be pixel-classified using a semantic segmentation algorithm, and then the pixels in the original image may be clustered to identify the element content in the image data, thereby performing basic annotation. Specifically, the steps of extracting at least one element information from the original image and obtaining the basic annotation corresponding to the at least one element information through image recognition may include: performing pixel classification on the original image using a semantic segmentation algorithm, determining each divided pixel set as corresponding element information based on the classification results, and identifying the element information to obtain the corresponding basic annotation.
[0060] In one embodiment, the above-mentioned target detection algorithm and semantic segmentation algorithm can also be used simultaneously to perform element extraction and basic annotation. Specifically, the target detection algorithm, such as YOLO, Faster R-CNN, etc., can be used first to locate and classify the target objects in the original image, and obtain information such as the bounding box and category label of each target object. Then, for the area within each bounding box, a semantic segmentation algorithm, such as U-Net, DeepLab, etc., is used to classify each pixel to obtain the pixel-level segmentation result of each target. Finally, the results of target detection and semantic segmentation are fused to obtain an image with both target detection and semantic segmentation information. The semantic segmentation information in the image can be used as the basic annotation of the original image.
[0061] Step 102: Generate a feature map of the original image based on the correlation between the basic annotations in the original image.
[0062] In one embodiment, a feature map generation model can be used to generate a feature map of the original image. Specifically, the basic annotations of the elements in the original image can be used as nodes, and the relationships can be used as edges to form a directed graph. Specifically, the original image can be first subjected to object detection (Object detection) to identify different targets in the original image and give their positions and categories. The above-mentioned target detection can be performed by a traditional detection algorithm or a target detection algorithm based on deep learning. The traditional detection algorithm can be detected by a VJ (Viola Jones) detector, a HOG (Histogram of Oriented Gradients) detector, or a DPM (Deformable Parts Model) detector. The target detection algorithm based on deep learning can include two-stage target detection algorithms such as R-CNN, SPP-Net, Fast R-CNN, and Faster R-CNN, as well as one-stage target detection algorithms such as OverFeat, YOLO series, SSD, and RetinaNet.
[0063] It can be understood that the feature graph generation model can be a scene graph generation model, and the scene graph generation model can output a semantic graph structure.
[0064] After obtaining the position and category information of different targets in the original image, attribute recognition can be performed on each element information to identify their attributes such as color, shape, size, and give their values, which are also the attribute values of at least one of the above-mentioned element information. Next, relationship recognition is required between each target to identify the spatial, semantic or causal relationships between them, which are also the association relationships between at least one of the above-mentioned element information, and give their types. Finally, the above information can be integrated into a directed graph, called a scene graph, which is also the above-mentioned feature graph, in which each node represents a target and its attributes, and each edge represents the relationship between a pair of targets and their types. The scene graph can be used to describe the content in the picture, and can also be used to generate a description of the picture or answer questions related to the picture.
[0065] Step 103 : clustering the feature map of the original image and the feature maps of other images in the image library, so that the original image and other images in the image library that meet the clustering conditions form a cluster set.
[0066] In one embodiment, the image library can be further searched for images with high similarity to the feature graph of the above-mentioned original image, and clustering can be performed based on these images and the original image to obtain a cluster set. In particular, when performing similarity comparison of feature graphs, the similarity of the nodes in the feature graphs and the similarity of the edges can be compared respectively. For example, when the first similarity between the nodes in the feature graph of a sample image in the image library and the nodes in the original image exceeds a first preset value, and the second similarity of the edges in the feature graphs of the two exceeds a second preset value, it can be determined that the current sample image is a similar image and clustered with the original image. In other embodiments, it is also possible to calculate the sum of the first similarity and the second similarity and compare it with a third preset value. When it exceeds the third preset value, it can be determined that the current sample image is a similar image and clustered with the original image.
[0067] It is understood that a large number of sample images are pre-stored in the image library, and these sample images are processed to generate corresponding feature maps. In addition, calculating the similarity between the feature map of the original image and the feature map of the sample image can be achieved through a similarity evaluation model.
[0068] Step 104 : Generate inferred annotations based on the basic annotations of all images in the cluster set, and determine the inferred annotations as target annotations of the original image.
[0069] In one embodiment, the original image is annotated again. Specifically, based on the clustering results, the basic annotations of all images in the clustering results are input into a pre-trained neural network model to output an inferred annotation of the original image, which serves as the target annotation for the original image. It should be noted that the inferred annotation information can be multiple. That is, when elements within an image can be summarized into multiple inferred annotations, all of these inferred annotations can be used to annotate the image. For example, if a birthday party is held in a restaurant, two inferred annotations can be annotated: the restaurant name and the birthday party.
[0070] Among them, the above-mentioned neural network model can be obtained through training with a large number of training sets, and the training results can be adjusted through human intervention, so that it can accurately summarize and analyze the basic annotations and output correct inference annotations. Its training set can also include a geographic information database, which is used to determine location information based on scene elements in image data.
[0071] As can be seen from the above, the image content annotation method provided by the embodiment of the present application can extract element information from the original image, and obtain the basic annotation corresponding to the element information through image recognition, generate a feature map of the original image based on the correlation between the basic annotations in the original image, cluster the feature map of the original image with the feature maps of other images in the gallery, so that the original image and other images in the gallery that meet the clustering conditions form a cluster set, generate inference annotations based on the basic annotations of all images in the cluster set, and determine the inference annotations as the target annotations of the original image. The solution provided by the embodiment of the present application can obtain more accurate annotation information by performing basic annotations on the image and then performing secondary annotations based on the results of clustering processing, making image classification more standardized and greatly improving the efficiency of subsequent retrieval.
[0072] See also Figure 2 , is another flow chart of the image content annotation method provided in an embodiment of the present application. The specific flow of the method may be as follows:
[0073] Step 201: extract at least one element information from the original image, and obtain basic annotations corresponding to the at least one element information through image recognition.
[0074] In one embodiment, after acquiring the original image, at least one element information in the original image can be further extracted. Specifically, the original image data can be processed to extract key features in the image data, and then the elements in the image can be divided based on these features, such as people, animals, objects, background, etc. in a photo. Then, these elements are further image recognized, and the recognition results are annotated, such as output through AI image recognition tools or pre-trained recognition models to obtain specific annotation information of the above elements in the original image. For example, Figure 3 As shown, the recognized content elements can be labeled as “girl”, “cake”, “candle”, “balloon”, and “table” as basic labels for the original image.
[0075] Step 202 : Acquire the ambient sound data and / or human voice data recorded when the original image is captured, analyze the ambient sound data and / or human voice data, and add the analysis results to the basic annotation of the original image.
[0076] This embodiment can also support multimodal data annotation. For example, when image data is generated, the corresponding audio data can be obtained to analyze the environmental sounds or human voices in the photo and add richer annotation information to the photo.
[0077] The embodiments of the present application utilize artificial intelligence technology and the association of multiple information to achieve automatic annotation of photo content. Furthermore, secondary annotation can be performed based on basic annotations to achieve classification of massive image data, which is beneficial for users to retrieve target images. Furthermore, secondary annotation based on inferences from basic annotations can make the annotation results more detailed and comprehensive, providing users with more useful information. By influencing the annotation results of the AI model on image data by associating image data, annotation accuracy and efficiency can be improved.
[0078] Step 203: Generate a feature map of the original image based on the correlation between the basic annotations in the original image.
[0079] In one embodiment, a feature map generation model can be used to generate a feature map of the original image. Specifically, the elements and attributes of the elements in the original image can be used as nodes, and the relationships between the basic annotations in the original image can be used as edges to form a directed graph. Figure 4 As shown, the elements displayed in the original image are converted into nodes of girl, cake, candle, balloon, and table, and the edges are girl-blowing-candle, candle-on-cake, and balloon-behind-table.
[0080] Step 204 : Generate a semantic description based on the graph structure corresponding to the feature graph of the original image, and use the semantic description as the semantic information of the original image.
[0081] In one embodiment, a semantic description can be generated through the graph structure generated above, and the semantic description can be used as the semantic information of each graph. Specifically, it can be achieved by using methods such as a graph neural network or a knowledge graph. A graph neural network is a neural network model that can process graph structured data. It can learn the feature representation of each node and edge through neighbor aggregation or spectral analysis, and is used for tasks such as classification, clustering, and generation. A knowledge graph is a graph structured data that can represent structured knowledge. It can learn the semantic representation of each entity and relationship through embedding or logical reasoning, and is used for tasks such as completion, alignment, and inference. This embodiment will not go into further detail on this. Therefore, the above-mentioned method of generating a semantic description based on the graph structure corresponding to the feature graph of the original image includes: performing feature representation on the graph structure of the feature graph through a graph neural network, and generating a semantic description based on the feature representation result; or, performing semantic representation on the graph structure of the feature graph through a knowledge graph, and generating a semantic description based on the semantic representation.
[0082] Step 205 , calculating the similarity between the semantic information of the original image and the semantic information corresponding to the feature maps of other images in the image library, and determining the images with similarities exceeding a preset value and the original image as a cluster set.
[0083] After obtaining the semantic information of the original image, the semantic information of each image in the image library is further obtained using the method in the previous step. Semantic similarity can then be calculated to determine the degree of elemental similarity between the original image and each image in the image library. The images can then be clustered using a clustering model to obtain a set of clusters. Specifically, image clustering can be achieved using the K-means clustering model. Alternatively, the DSSM semantic similarity model can be used to analyze the semantic similarity between images.
[0084] Step 206: Obtain basic annotations for all images in the cluster set.
[0085] In one embodiment, the image data of the original image may also include time information and location information. That is, when the original image data is generated, the current timestamp information and GPS data may be recorded simultaneously. When secondary annotation of the original image is required, the basic annotations of all images in the cluster set may be filtered based on the time information and location information. Because the relevant themes of image data taken within a certain period of time are generally consistent, their image content is often related. For example, if a photo is labeled as a company party, it can be inferred that other photos taken at a similar time are also party photos, and the same applies to the same location. For images in the cluster set that have significantly different time information or location information from the original image, or that fall outside the corresponding time threshold range or location threshold range, their corresponding basic annotations may be removed, thereby improving the accuracy of the final secondary annotation. In other words, in this embodiment, the step of obtaining basic annotations for all images in the cluster set may include extracting the time information and location information of the original image, and based on the time information and location information, removing images that do not meet the corresponding time and location conditions from all images in the cluster set, thereby obtaining basic annotations for all remaining images in the cluster set.
[0086] In step 207 , all basic annotations are input into a preset neural network model to output inference annotations of the original image, and the inference annotations are determined as target annotations of the original image.
[0087] In one embodiment, the original image is annotated a second time. Specifically, this can be based on the aforementioned cluster set, and the basic annotations of all images in the cluster set are input into a pre-trained neural network model to output the inferred annotations of the original image. During the image data annotation process, this embodiment also allows users to provide feedback and corrections to the basic or inferred annotation results to improve the accuracy of the final target annotation. Through user participation, the artificial intelligence model is continuously optimized and trained to provide more accurate automatic annotation results.
[0088] In one embodiment, these basic annotations can be used as neuron inputs of the corresponding AI model. Through the processing of the AI model, the probability value of the corresponding inference annotation information is calculated, and the corresponding inference annotation information is output to annotate the image. For example, the content elements in the image data include birthday cakes, balloons, or other birthday-related elements. The basic annotation information of these element contents is used as the input of the AI model, and then the inference annotation information indicating that the image is a birthday party can be output, and this can be used as the target annotation of the original image. In addition, when the content elements in the image data include scene elements such as landmarks, buildings, beaches, mountain views, etc., the AI model can output the location information of the corresponding image as the target annotation.
[0089] All of the above technical solutions can be combined in any way to form optional embodiments of the present application, and will not be described in detail here.
[0090] As can be seen from the above, the image content annotation method provided by the embodiment of the present application can extract at least one element information in the original image, and obtain the basic annotation corresponding to the at least one element information through image recognition, obtain the ambient sound data and / or human voice data recorded when the original image is shot, analyze the ambient sound data and / or human voice data, and add the analysis results to the basic annotation of the original image, generate a feature map of the original image based on the correlation between the basic annotations in the original image, generate a semantic description according to the graph structure corresponding to the feature map of the original image, and use the semantic description as the semantic information of the original image, calculate the similarity between the semantic information of the original image and the semantic information corresponding to the feature maps of other images in the image library, determine the images with similarity exceeding a preset value and the original image as a cluster set, obtain the basic annotations of all images in the cluster set, input all the basic annotations into a preset neural network model to output the inference annotation of the original image, and determine the inference annotation as the target annotation of the original image. The solution provided by the embodiment of the present application can obtain more accurate annotation information by performing basic annotation on the image and then performing secondary annotation in combination with the results of clustering processing, making image classification more standardized and greatly improving the efficiency of subsequent retrieval.
[0091] In order to implement the above method, an embodiment of the present invention further provides an image content annotation device, which can be integrated into a computer device such as a mobile phone, a personal computer, a tablet computer, and the like.
[0092] For example, Figure 5 FIG. 1 is a schematic diagram of a first structural embodiment of an image content annotation apparatus provided by an embodiment of the present invention. The image content annotation apparatus may include:
[0093] The first recognition unit 301 is used to extract element information from the original image and obtain basic annotations corresponding to the element information through image recognition;
[0094] A generating unit 302 is configured to generate a feature map of the original image based on the correlation relationship between the basic annotations in the original image;
[0095] A clustering unit 303 is configured to perform clustering processing on the feature map of the original image and the feature maps of other images in the image library, so that the original image and other images in the image library that meet the clustering conditions form a cluster set;
[0096] The second recognition unit 304 is configured to generate an inferred annotation based on the basic annotations of all images in the cluster set, and determine the inferred annotation as the target annotation of the original image.
[0097] In one embodiment, see Figure 6 , Figure 6 2 is a schematic diagram of a second structure of the image content tagging apparatus provided by an embodiment of the present invention, wherein the clustering unit 303 specifically includes:
[0098] A generating subunit 3031 is configured to generate a semantic description according to a graph structure corresponding to the feature graph of the original image, and use the semantic description as semantic information of the original image;
[0099] A calculation subunit 3032 is configured to calculate the similarity between the semantic information of the original image and the semantic information corresponding to the feature maps of other images in the image library;
[0100] The determining subunit 3033 is configured to determine the image with a similarity exceeding a preset value and the original image as a cluster set.
[0101] In one embodiment, the second identification unit 304 specifically includes:
[0102] An acquisition subunit 3041 is used to acquire basic annotations of all images in the cluster set;
[0103] The processing subunit 3042 is configured to input all basic annotations into a preset neural network model to output an inferred annotation of the original image, and determine the inferred annotation as a target annotation of the original image.
[0104] The image content annotation device proposed in the embodiment of the present invention can extract element information from the original image, and obtain the basic annotation corresponding to the element information through image recognition, generate a feature map of the original image based on the correlation between the basic annotations in the original image, cluster the feature map of the original image with the feature maps of other images in the gallery, so that the original image and other images in the gallery that meet the clustering conditions form a cluster set, generate inference annotations based on the basic annotations of all images in the cluster set, and determine the inference annotations as the target annotations of the original image. The solution provided in the embodiment of the present application can obtain more accurate annotation information by performing basic annotation on the image and then performing secondary annotation in combination with the results of clustering processing, making image classification more standardized and greatly improving the efficiency of subsequent retrieval.
[0105] All of the above technical solutions can be combined in any way to form optional embodiments of the present application, and will not be described in detail here.
[0106] An embodiment of the present application further provides a computer device, comprising a memory and a processor, wherein the processor is configured to execute the process of the image content annotation method provided in this embodiment by calling a computer program stored in the memory.
[0107] For example, the computer device mentioned above can be a terminal device with corresponding functions such as a mobile phone, tablet computer, personal computer, cloud computer, etc. Figure 7 , Figure 7 A schematic diagram of the structure of a computer provided in an embodiment of the present application.
[0108] The computer device 300 may include components such as a memory 301 and a processor 302. Those skilled in the art will appreciate that Figure 7 The computer device structure shown in the figure does not constitute a limitation to the computer device, and may include more or fewer components than shown in the figure, or combine certain components, or arrange the components differently.
[0109] Memory 301 can be used to store applications and data. The applications stored in memory 301 include executable code. Applications can be composed of various functional modules. Processor 302 executes various functional applications and data processing by running the applications stored in memory 301.
[0110] The processor 302 is the control center of the computer device. It uses various interfaces and lines to connect the various parts of the entire computer device. By running or executing applications stored in the memory 301 and calling data stored in the memory 301, it performs various functions of the computer device and processes data, thereby monitoring the computer device as a whole.
[0111] In this embodiment, the processor 302 in the computer device loads the executable code corresponding to one or more application processes into the memory 301 according to the following instructions, and the processor 302 runs the application stored in the memory 301 to execute:
[0112] Extracting element information from the original image and obtaining basic annotations corresponding to the element information through image recognition;
[0113] generating a feature map of the original image based on the correlation relationship between the basic annotations in the original image;
[0114] Clustering the feature map of the original image with the feature maps of other images in the image library so that the original image and other images in the image library that meet the clustering conditions form a cluster set;
[0115] An inferred annotation is generated according to the basic annotations of all images in the cluster set, and the inferred annotation is determined as the target annotation of the original image.
[0116] Those skilled in the art will appreciate that all or part of the steps in the various methods of the above embodiments may be accomplished by instructions, or by controlling related hardware through instructions. The instructions may be stored in a computer-readable storage medium and loaded and executed by a processor.
[0117] To this end, an embodiment of the present application provides a storage medium storing a plurality of instructions that can be loaded by a processor to execute the steps of any of the data calculation methods provided in the embodiments of the present application. For example, the instructions can execute the following steps:
[0118] Extracting element information from the original image and obtaining basic annotations corresponding to the element information through image recognition;
[0119] generating a feature map of the original image based on the correlation relationship between the basic annotations in the original image;
[0120] Clustering the feature map of the original image with the feature maps of other images in the image library so that the original image and other images in the image library that meet the clustering conditions form a cluster set;
[0121] An inferred annotation is generated according to the basic annotations of all images in the cluster set, and the inferred annotation is determined as the target annotation of the original image.
[0122] The specific implementation of the above operations can be found in the previous embodiments and will not be repeated here.
[0123] The storage medium may include a read-only memory (ROM), a random access memory (RAM), a magnetic disk or an optical disk, etc.
[0124] Since the instructions stored in the storage medium can execute the steps in any data calculation method provided in the embodiments of the present application, the beneficial effects that can be achieved by any data calculation method provided in the embodiments of the present application can be achieved. Please refer to the previous embodiments for details and will not be repeated here.
[0125] In the embodiments of the computer device and readable storage medium provided in this application, all technical features of the embodiments of the above-mentioned method are included. The expanded and explained contents of the specification are applicable to the embodiments of the above-mentioned positioning method in the same way and will not be repeated here.
[0126] An embodiment of the present application also provides a chip, including a memory and a processor, wherein the memory is used to store programs, and the processor is used to call and run programs from the memory, so that a device equipped with the chip executes the methods in the various possible embodiments above.
[0127] In the above embodiments, the description of each embodiment has its own focus. For parts that are not described in detail in a certain embodiment, please refer to the detailed description of the image content tagging device above, and will not be repeated here.
[0128] The image content annotation method provided in the embodiment of the present application and the image content annotation device in the above embodiment have the same concept. The specific implementation process is detailed in the embodiment of the image content annotation device and will not be repeated here.
[0129] It should be noted that, with respect to the image content annotation method described in the embodiments of the present application, those skilled in the art will understand that all or part of the process of implementing the image content annotation method described in the embodiments of the present application can be accomplished by controlling related hardware through a computer program. The computer program can be stored in a computer-readable storage medium, such as a memory, and executed by at least one processor. During execution, the process may include the process of the embodiment of the image content annotation method. The storage medium can be a magnetic disk, an optical disk, a read-only memory, a random access memory (RAM), etc.
[0130] The above is a detailed introduction to the image content annotation method, device, storage medium and computer equipment provided in the embodiments of the present application. Specific examples are used in this article to illustrate the principles and implementation methods of the present application. The description of the above embodiments is only used to help understand the method of the present application and its core ideas. At the same time, for technical personnel in this field, based on the ideas of the present application, there will be changes in the specific implementation methods and application scope. In summary, the contents of this specification should not be understood as limiting the present application.
[0131] The above description is merely an embodiment of the present application and does not limit the patent scope of the present application. Any equivalent structure or equivalent process transformation made using the contents of the description and drawings of this application, such as the mutual combination of technical features between the embodiments, or direct or indirect application in other related technical fields, are also included in the patent protection scope of the present application.
Claims
1. A method for labeling image content, characterized in that: include: Extracting element information from the original image and obtaining basic annotations corresponding to the element information through image recognition; generating a feature map of the original image based on the correlation relationship between the basic annotations in the original image; generating a semantic description according to a graph structure corresponding to the feature graph of the original image, and using the semantic description as semantic information of the original image; Calculating the similarity between the semantic information of the original image and the semantic information corresponding to the feature maps of other images in the image library; Determine the image with a similarity exceeding a preset value and the original image as a cluster set; An inferred annotation is generated according to the basic annotations of all images in the cluster set, and the inferred annotation is determined as the target annotation of the original image.
2. The image content annotation method according to claim 1, wherein: The step of extracting element information from the original image and obtaining basic annotations corresponding to the element information through image recognition includes: Positioning the target in the original image using a target detection algorithm; Selecting the target according to the positioning result, and determining the selected image content as element information; Identify the element information to obtain the corresponding basic annotation.
3. The image content annotation method according to claim 1, wherein: The step of extracting element information from the original image and obtaining basic annotations corresponding to the element information through image recognition includes: Perform pixel classification on the original image using a semantic segmentation algorithm; Determine each divided pixel set as corresponding element information according to the classification result; Identify the element information to obtain the corresponding basic annotation.
4. The image content annotation method according to claim 1, wherein: Generating an inferred annotation based on the basic annotations of all images in the cluster set, and determining the inferred annotation as the target annotation of the original image, includes: Obtaining basic annotations for all images in the cluster set; All basic annotations are input into a preset neural network model to output inference annotations of the original image, and the inference annotations are determined as target annotations of the original image.
5. The image content annotation method according to claim 1, wherein: The method of generating a semantic description according to the graph structure corresponding to the feature graph of the original image includes: Performing feature representation on the graph structure of the feature graph through a graph neural network, and generating the semantic description based on the feature representation result; or, The graph structure of the feature graph is semantically represented through a knowledge graph, and the semantic description is generated based on the semantic representation.
6. The image content annotation method according to claim 4, wherein: The obtaining of basic annotations for all images in the cluster set includes: Extracting time information and position information of the original image; Based on the time information and the location information, remove images that do not meet corresponding time conditions and location conditions from all images in the cluster set; Get the basic annotations of all remaining images in the cluster set.
7. The image content annotation method according to claim 1, wherein: The method further comprises: Acquiring ambient sound data and / or human voice data recorded when the original image is captured; The environmental sound data and / or human voice data are analyzed, and the analysis results are added to the basic annotations of the original image.
8. An image content annotation device, characterized in that: include: A first recognition unit is used to extract element information from the original image and obtain basic annotations corresponding to the element information through image recognition; A generating unit, configured to generate a feature map of the original image based on the correlation relationship between the basic annotations in the original image; a clustering unit configured to generate a semantic description based on the graph structure corresponding to the feature graph of the original image, and use the semantic description as the semantic information of the original image; calculate the similarity between the semantic information of the original image and the semantic information corresponding to the feature graphs of other images in the image library; and determine the images whose similarity exceeds a preset value and the original image as a cluster set; The second recognition unit is configured to generate an inferred annotation based on the basic annotations of all images in the cluster set, and determine the inferred annotation as the target annotation of the original image.
9. A computer-readable storage medium having a computer program stored thereon, characterized in that: When the computer program is executed on a computer, the computer is caused to execute the image content tagging method according to any one of claims 1 to 7.
10. A computer device comprising a memory and a processor, characterized in that: The processor executes the image content tagging method according to any one of claims 1 to 7 by calling the computer program stored in the memory.