Image tagging method and device

The method addresses the challenge of connecting image tags by using area distribution maps to determine relevance and generate combined tags, enhancing search accuracy in image tagging.

WO2025135625A1PCT designated stage expired Publication Date: 2025-06-26SAMSUNG ELECTRONICS CO LTD
View PDF 4 Cites 0 Cited by

Patent Information

Application Number
PCT/KR2024/019748
Authority / Receiving Office
WO · WO
Patent Type
Applications
Current Assignee / Owner
Priority Date
2023-12-22
Filing Date
2024-12-04
Publication Date
2025-06-26

AI Technical Summary

Technical Problem

Existing image tagging methods struggle to effectively connect image tags to accurately represent the content of images, leading to incorrect search results when users input specific queries.

Method used

A method and device for performing image tagging using image tags obtained from an image tag generation model, where the device determines the degree of relevance between image tags based on area distribution maps and generates combined tags to improve search accuracy.

Benefits of technology

The proposed method enhances image tagging by generating combined tags that better represent the image content, thereby improving search accuracy and relevance.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure KR2024019748_26062025_PF_FP_ABST
    Figure KR2024019748_26062025_PF_FP_ABST
Patent Text Reader

Abstract

This image tagging method involves: acquiring a plurality of image tags from an image by using an image tag generation model; determining the degree of relevance between the plurality of image tags on the basis of a plurality of area distribution maps corresponding to the plurality of image tags, respectively; generating, on the basis of the determined degree of relevance, a combined tag in which the image tags are linked; and performing image tagging of the image by using the plurality of image tags and the combined tag.
Need to check novelty before this filing date? Find Prior Art

Description

Method and device for performing image tagging

[0001] It relates to a method and device for performing image tagging.

[0002] With the advancement of the Internet and mobile devices, the supply and demand for content is rapidly increasing. In particular, content uploading through social networking services and registering keywords to search for and access content has made content search and sharing easier.

[0003] For visual content such as photos and videos, image tags can be used as a search tool. Users can search for content using search queries that include image tags. As content sharing via the Internet increases, the need for search methods that enable users to effectively access desired content is growing.

[0004] The present invention provides a method and device for performing image tagging using image tags obtained using an image tag generation model and a combined tag connected to image tags.

[0005] One embodiment of the present disclosure will be set forth in part in the description which follows, and in part, as will be apparent from the description or may be learned by practice of the described embodiments.

[0006] According to one embodiment of the present disclosure, a method for performing image tagging is provided. The method for performing image tagging includes a step of obtaining a plurality of image tags from an image using an image tag generation model. Furthermore, the method for performing image tagging includes a step of determining a degree of relevance between the plurality of image tags based on a plurality of area distribution maps each corresponding to the plurality of image tags. Furthermore, the method for performing image tagging includes a step of generating a combined tag in which the image tags are linked based on the determined degree of relevance. Furthermore, the method for performing image tagging includes a step of performing image tagging on the image using the plurality of image tags and the combined tag.

[0007] According to one embodiment of the present disclosure, a device for performing image tagging is provided. The device for performing image tagging includes a memory storing at least one instruction and at least one processor operably connected to the memory. The at least one processor executes the at least one instruction to obtain a plurality of image tags from an image using an image tag generation model. Furthermore, the at least one processor executes the at least one instruction to determine a degree of correlation between the plurality of image tags based on a plurality of area distribution maps each corresponding to the plurality of image tags. Furthermore, the at least one processor executes the at least one instruction to generate a combined tag to which the image tags are linked based on the determined degree of correlation. Furthermore, the at least one processor executes the at least one instruction to perform image tagging on the image using the plurality of image tags and the generated combined tag.

[0008] According to one embodiment of the present disclosure, a computer-readable recording medium having recorded thereon a program for executing the above-described method can be provided.

[0009] The above and other aspects, features, and advantages of one embodiment of the present disclosure will become more apparent from the following description taken in conjunction with the accompanying drawings.

[0010] FIG. 1 is a diagram for explaining a process of obtaining an image tag and performing image tagging using an image tag generation model according to one embodiment of the present disclosure.

[0011] FIG. 2 is a diagram for explaining an image tag generation model according to one embodiment of the present disclosure.

[0012] FIG. 3 is a flowchart illustrating a method for performing image tagging according to one embodiment of the present disclosure.

[0013] FIG. 4 is a detailed flowchart of a step for determining a degree of relevance between image tags according to one embodiment of the present disclosure.

[0014] FIG. 5 is a diagram illustrating a method for generating a combination of image tags according to one embodiment of the present disclosure.

[0015] FIG. 6 is a diagram for explaining a process of generating an area distribution map corresponding to an image tag according to one embodiment of the present disclosure.

[0016] FIG. 7 is a diagram illustrating a process for determining a degree of relevance between image tags corresponding to a combination of image tags according to one embodiment of the present disclosure.

[0017] FIG. 8 is a diagram illustrating an example of using an image tag generation model based on a transformer decoder according to one embodiment of the present disclosure.

[0018] FIG. 9 is a detailed flowchart of a step for generating a combined tag based on the degree of association between image tags according to one embodiment of the present disclosure.

[0019] FIG. 10 is a block diagram illustrating an electronic device performing image tagging according to one embodiment of the present disclosure.

[0020] FIG. 11 is a block diagram illustrating the configuration and operation of an electronic device that performs image tagging according to one embodiment of the present disclosure.

[0021] FIG. 12 is a block diagram illustrating the operation of a relevance analysis module and a combined tag generation module according to one embodiment of the present disclosure.

[0022] FIG. 13 is a block diagram illustrating the operation of a relevance analysis module and a combined tag generation module according to one embodiment of the present disclosure.

[0023] Hereinafter, embodiments of the present invention will be described in detail with reference to the accompanying drawings. In the drawings, identical components are designated by the same reference numerals, and redundant descriptions are omitted. The embodiments described herein are merely exemplary, and the present invention is not limited thereto and may be implemented in various other forms. It should be understood that the singular form "a" includes plural references unless the context clearly dictates otherwise.

[0024] Hereinafter, terms used in this specification will be briefly described, and the present disclosure will be described in detail. In this disclosure, the expression "at least one of a, b, or c" may refer to "a," "b," "c," "a and b," "a and c," "b and c," "all of a, b, and c," or variations thereof.

[0025] The terms used in this disclosure are selected from widely used, common terms, taking into account the functions of the disclosure. However, these terms may vary depending on the intentions of those skilled in the art, precedents, the emergence of new technologies, etc. Furthermore, in certain cases, terms may be arbitrarily selected by the applicant, in which case their meanings will be described in detail in the relevant description. Therefore, the terms used in this disclosure should not be defined simply as names, but rather based on the meanings of the terms and the overall content of the disclosure.

[0026] Singular expressions may include plural expressions unless the context clearly indicates otherwise. Terms used herein, including technical or scientific terms, have the same meaning as commonly understood by a person of ordinary skill in the art described herein. Furthermore, terms containing ordinal numbers, such as "first" or "second," used herein may be used to describe various components, but such components should not be limited by such terms. Such terms are used solely to distinguish one component from another.

[0027] When a part of the specification is said to "include" a component, unless otherwise specifically stated, this does not exclude other components but rather implies the inclusion of other components. Furthermore, terms such as "part" and "module" used in the specification refer to a unit that processes at least one function or operation, which may be implemented in hardware, software, or a combination of hardware and software.

[0028] The artificial intelligence-related function according to one embodiment of the present disclosure operates via a processor and memory. The processor may be composed of one or more processors. In this case, one or more processors may be a general-purpose processor such as a CPU, an AP, a Digital Signal Processor (DSP), a graphics-only processor such as a GPU, a Vision Processing Unit (VPU), or an artificial intelligence-only processor such as an NPU. The one or more processors control the processing of input data according to predefined operation rules or artificial intelligence models stored in the memory. Alternatively, if one or more processors are artificial intelligence-only processors, the artificial intelligence-only processor may be designed with a hardware structure specialized for processing a specific artificial intelligence model.

[0029] The predefined operation rules or artificial intelligence models are characterized by being created through learning. Here, being created through learning means that the basic artificial intelligence model is trained using a learning algorithm using a plurality of learning data, thereby creating a predefined operation rules or artificial intelligence model set to perform a desired characteristic (or purpose). This learning may be performed on the device itself on which the artificial intelligence according to the present disclosure is performed, or may be performed through a separate server and / or system. Examples of the learning algorithm include, but are not limited to, supervised learning, unsupervised learning, semi-supervised learning, or reinforcement learning.

[0030] An artificial intelligence model may be composed of multiple neural network layers. Each of the multiple neural network layers has multiple weight values, and performs neural network operations through operations between the operation results of the previous layer and the multiple weights. The multiple weights of the multiple neural network layers may be optimized based on the learning results of the artificial intelligence model. For example, the multiple weights may be updated so that the loss value or cost value obtained from the artificial intelligence model is reduced or minimized during the learning process. The artificial neural network may include a deep neural network (DNN), such as a convolutional neural network (CNN), a recurrent neural network (RNN), a restricted boltzmann machine (RBM), a deep belief network (DBN), a bidirectional recurrent deep neural network (BRDNN), or deep Q-networks, but is not limited to the examples described above.

[0031] Below, embodiments of the present disclosure are described in detail with reference to the attached drawings so that those skilled in the art can practice the present disclosure. However, the present disclosure may be implemented in various different forms and is not limited to the embodiments described herein.

[0032] FIG. 1 is a diagram illustrating a process of obtaining an image tag and performing image tagging using an image tag generation model according to one embodiment of the present disclosure.

[0033] In this disclosure, 'image' means visualized information such as each frame that constitutes a photograph, picture, illustration, graph, map, or video.

[0034] In this disclosure, an "image tag" may include information about visual aspects of an image, such as color or shape, or information about what the image is about. For example, an image tag may include information about an object or scene within the image. Alternatively, an image tag may include information about a specific object or event with a unique name, or a specific location or time. Alternatively, an image tag may include information about an abstract emotion or existence.

[0035] According to one embodiment of the present disclosure, image tags can be obtained using an image tag generation model. The image tag generation model can take an image as input and output at least one image tag. As illustrated in Fig. 1, the image tag generation model can take an image of a "boy wearing a red shirt on a white background" as input and output various image tags such as "young," "boy," "wear," "red," "shirt," and "white." The operation of the image tag generation model will be described below with reference to Fig. 2.

[0036] In this disclosure, "image tagging" refers to the process of storing an image by associating a tag for the image with the image. Image tagging can be performed by obtaining a tag for the image from the image, or by storing the image in association with the tag so that the image can be obtained from the tag for the image. For example, image tagging can be performed by storing an image and a tag for the image so that one of them references the other. Alternatively, image tagging can be performed by additionally storing tag data for the image in a space where image data is stored. Image tagging can be performed manually by a person or automatically by a computer or machine. In this disclosure, image tagging is assumed to be automatically performed by a computer or machine, but manual tagging may be applied in some processes.

[0037] According to one embodiment of the present disclosure, image tagging can be performed by storing an image by associating one or more tags for the image. Referring to FIG. 1, image tagging is illustrated by associating and storing the image tags 'young', 'boy', 'wear', 'red', 'shirt', and 'white' for the image of 'a boy wearing a red shirt on a white background'. In this way, in the case of an image for which image tagging has been performed, the image can be searched by using the image tags used for the image tagging as an image search query. In the example illustrated in FIG. 1, when 'shirt' is entered as a search query for a database in which images are stored, all images tagged with 'shirt', including the image in FIG. 1, can be retrieved.

[0038] FIG. 2 is a diagram for explaining an image tag generation model according to one embodiment of the present disclosure.

[0039] An image tag generation model can recognize an image and extract an image tag for the image. A neural network model that generates an image tag, trained using a training dataset, can recognize the image and generate at least one image tag for the image. A device that performs image tagging can generate an image tag based on feature information extracted from the image using the image tag generation model. The image tag generation model can generate at least one image tag based on feature information extracted from the image.

[0040] An image tag generation model can be trained using a training dataset. The training dataset can include labeled input data sufficient to learn to output image tags from images. For the image tag generation model, a loss corresponding to the difference between the actual image tags corresponding to the input images and the image tags output from the image tag generation model can be determined. An image tagging device can train the image tag generation model in a direction that reduces the loss of the loss function through backpropagation. The image tagging device can consider the training of the image tag generation model to be complete when the loss determined during the training process of the image tag generation model is minimized.

[0041] As illustrated in FIG. 2, the image tag generation model may be in the form of a multi-label recognition model. The multi-label recognition model may be in the form of a neural network composed of multiple layers. The multi-label recognition model may be in the form of shared backbone layers and additional layers that generate tags, each of which takes the output of the shared backbone as input. The multi-label recognition model may include a feature extractor that extracts feature information and a classifier that generates image tags using the feature information extracted from the feature extractor. For example, the image tag generation model may be a multi-label recognition model in the form of connecting multiple classification heads that generate image tags to a feature extractor that extracts feature information from an image, as illustrated in FIG. 2.

[0042] A feature extractor can extract information about the components that make up an image. For example, a feature extractor can extract information about the edges of each image. A feature extractor can extract information about the colors of each image. A feature extractor can extract information about the brightness or contrast of each image. A feature extractor can include an image encoder.

[0043] A classifier can perform the task of generating different types of image tags. The classifier can be a computational unit performing a given operation or a neural network composed of layers. For example, in Figure 2, Tag-A can be an image tag relating to an object in the image. Tag-B can be an image tag relating to an attribute or action in the image. Tag-C can be an image tag that infers a scene in the image.

[0044] A predictor can generate tags corresponding to different tag domains. For example, a predictor might generate "boy" or "shirt" as Tag-A. A predictor might generate "young," "red," or "white" as Tag-B. A predictor might generate "wear" as Tag-C.

[0045] The image tag generation model can independently generate image tags containing objects, attributes, or actions. This is because the training dataset used to train the image tag generation model independently labels objects, attributes, or actions corresponding to the classes recognized in each image. The image tag generation model only independently generates image tags and does not create image tags by associating related image tags. Therefore, the image tag generation model can generate image tags related to attributes or actions within the image for an input image, but the image tags do not indicate which object within the image they belong to or which object action they represent.

[0046] For example, among the image tags corresponding to the image in Fig. 1, 'red' and 'white' are image tags related to colors contained in the image, but do not indicate which object in the image they are related to. Therefore, when a user enters a search query 'white shirt' to search for a picture of 'a person wearing a white shirt' for images tagged with image tags, there is a problem that the image shown in Fig. 1, which has both 'white' and 'shirt' as image tags, is also retrieved as a search result.

[0047] To address this issue, the classes associated with attributes or actions must be recognizable for the objects they represent. To achieve this, the training dataset used to train the image tag generation model must be labeled with the objects corresponding to the recognized classes and their attributes or actions, in a form that correlates with each other. However, creating a training dataset with additional labels that relate objects to their attributes or actions is costly and time-consuming, and the number of labels (e.g., "white shirt," "red shirt," etc.) can increase exponentially, making training the image tag generation model practically challenging.

[0048] Below, a method of performing image tagging using image tags generated from an image tag generation model and image tags extended from the generated image tags is described in detail.

[0049] FIG. 3 is a flowchart illustrating a method for performing image tagging according to one embodiment of the present disclosure.

[0050] Referring to FIG. 3, in step S310, a device performing image tagging can obtain a plurality of image tags from an image using an image tag generation model.

[0051] The device performing image tagging can be an electronic device or server capable of capturing images, or an electronic device equipped with a camera that captures images. Examples include image processing devices such as smartphones, smart glasses, wearable devices, digital cameras, laptops, augmented reality (AR) devices, and virtual reality (VR) devices.

[0052] Devices performing image tagging can be equipped with various types of neural network models. For example, a device performing image tagging can be equipped with at least one model such as a CNN, DNN, RNN, or BRDNN, and can also utilize these models in combination.

[0053] The image may be captured by a device that performs image tagging, or may be received from an external device. The device that performs image tagging may generate multiple image tags from the image. The device that performs image tagging may utilize an image tag generation model. The image tag generation model may classify the class corresponding to the image tag based on features extracted from the input image.

[0054] In step S320, the device for performing image tagging can determine the degree of relevance between the acquired image tags based on an area distribution map corresponding to each image tag within the image. The device for performing image tagging can generate an area distribution map corresponding to each image tag. In other words, the device for performing image tagging can generate a plurality of area distribution maps each corresponding to a plurality of image tags. The device for performing image tagging can obtain the degree of relevance between the acquired image tags by comparing the generated area distribution maps.

[0055] The region distribution corresponding to an image tag contains information about which part of the image influenced the classification into the corresponding class when classifying the image into a class and generating an image tag. The region distribution corresponding to an image tag indicates the part of the image utilized in the process of generating the image tag, and is not limited to its name. The region distribution may be a visualization of the part of the image related to the image tag. For example, the region distribution may be based on an activation map that specifies the part that played a role (activation) in generating the image tag (classifying it into a class corresponding to the image tag). Alternatively, the region distribution may be based on an attention map that can identify which part of the image was focused on when generating the image tag. Alternatively, the region distribution may be based on a feature map that can identify which features were extracted from the image when generating the image tag. Therefore, the area distribution corresponding to the image tag can also be called an activation map, a class activation map, an attention map, a feature map, etc.

[0056] The degree of relevance between image tags refers to the degree to which multiple image tags are related to each other. The degree of relevance between image tags can be determined for two or more image tags. A relationship between image tags indicates a high degree of relevance. For example, a first image tag representing a first object and a second image tag representing an attribute of the first object can be said to have a high degree of relevance. Conversely, a first image tag representing a first object and a third image tag representing an attribute of a second object different from the first object can be said to have a low degree of relevance.

[0057] FIG. 4 is a detailed flowchart of a step for determining a degree of relevance between image tags according to one embodiment of the present disclosure.

[0058] In step S410, the device for performing image tagging can determine a combination of image tags from among the acquired image tags. The device for performing image tagging can determine a combination of image tags by selecting image tags from among the acquired image tags to determine a correlation between the image tags. When the number of acquired image tags is small, the device for performing image tagging can determine a combination of all image tags and determine a correlation between the image tags for each combination of image tags. As the number of acquired image tags increases, determining a combination of all image tags and determining a correlation between the image tags for each combination of image tags can be a burden in terms of computational amount.

[0059] A device performing image tagging can determine two or more image tags from among a plurality of image tags as a combination of image tags. The device performing image tagging can determine a combination of image tags from the selected image tags based on the parts of speech of the acquired image tags or whether the acquired image tags correspond to a relationship between objects and characteristics of the objects.

[0060] FIG. 5 is a diagram illustrating a method for generating a combination of image tags according to one embodiment of the present disclosure.

[0061] Referring to Fig. 5, an example is shown in which a total of four image tags are obtained from one image, but the number of image tags is not limited thereto. Tag-A is an image tag called 'people', is a noun, and corresponds to an object within the image. Tag-B is an image tag called 'car', is a noun, and corresponds to an object within the image. Tag-C is an image tag called 'red', is an adjective, and corresponds to an attribute of an object within the image. Tag-D is an image tag called 'running', is a verb, and corresponds to an action of an object within the image.

[0062] According to one embodiment of the present disclosure, when the task of selecting image tags to determine the correlation between image tags among the acquired image tags is not performed, the device performing image tagging can generate all possible combinations of image tags. In the case of the example of FIG. 5, the device performing image tagging can generate combinations of image tags composed of two image tags, three image tags, and four image tags among four image tags. The combinations of image tags composed of two image tags are (Tag-A, Tag-B), (Tag-A, Tag-C), (Tag-A, Tag-D), (Tag-B, Tag-C), (Tag-B, Tag-D), (Tag-C, Tag-D) (i.e., a total of six combinations). The combinations of image tags consisting of three image tags are (Tag-A, Tag-B, Tag-C), (Tag-A, Tag-B, Tag-D), (Tag-B, Tag-C, Tag-D), (Tag-A, Tag-C, Tag-D) (i.e., a total of four combinations). The combinations of image tags consisting of four image tags are (Tag-A, Tag-B, Tag-C, Tag-D) (i.e., a total of one combination). There may be combinations of image tags that are not appropriate as search queries or that seem unrelated in general concepts. For example, in the example of Fig. 5, the combination of (Tag-C, Tag-D) may be a combination of image tags of (red, running) or (running, red), which may be unrelated to each other.

[0063] According to one embodiment of the present disclosure, a device for performing image tagging can select image tags from among acquired image tags according to the parts of speech of the image tags, and generate a combination of image tags. The device for performing image tagging can determine a combination of image tags by checking the parts of speech of each image tag. For example, referring to FIG. 5, the device for performing image tagging can determine a combination of image tags of (Tag-B, Tag-C) according to the form of 'adjective + noun', such as, for example, 'red car'. In addition, the device for performing image tagging can determine a combination of image tags of (Tag-A, Tag-D) according to the form of 'verb + noun', such as, for example, 'running people'. In addition, the device for performing image tagging can determine a combination of image tags of (Tag-B, Tag-C, Tag-D) according to the form of 'verb + adjective + noun', such as, for example, 'running red car'.

[0064] According to one embodiment of the present disclosure, a device for performing image tagging can select image tags from among acquired image tags based on whether the image tags have a relationship corresponding to an object and a characteristic of the object, and generate a combination of image tags. The device for performing image tagging can determine a combination of image tags based on whether each image tag relates to an object or to an attribute or action of the object. For example, referring to FIG. 5, the device for performing image tagging can determine a combination of image tags of (Tag-B, Tag-C) based on the form of 'object attribute + object', such as, for example, 'red car'. In addition, the device for performing image tagging can determine a combination of image tags of (Tag-A, Tag-D) based on the form of 'object action + object', such as, for example, 'running people'. Additionally, the device performing image tagging can determine a combination of image tags (Tag-B, Tag-C, Tag-D) based on the form of 'object action + object property + object', for example, 'running red car'.

[0065] According to one embodiment of the present disclosure, a device for performing image tagging may determine a combination of image tags by checking both the parts of speech of the acquired image tags and whether the acquired image tags have a relationship corresponding to an object and a characteristic of the object.

[0066] A device performing image tagging can determine combinations of image tags that are likely to be search queries. Conversely, the device performing image tagging can exclude combinations of image tags that are inappropriate or cannot be search queries.

[0067] According to one embodiment of the present disclosure, a device for performing image tagging may use a lookup table pre-specified by a user based on general definitions or meanings of words corresponding to image tags to determine a combination of image tags.

[0068] According to one embodiment of the present disclosure, a device for performing image tagging may utilize an ontology database of words corresponding to image tags to determine a combination of image tags. An ontology database expresses concepts, characteristics, and relationships between objects in a form that can be processed by a computer. A second image tag corresponding to another word retrieved from the ontology database of a word corresponding to a first image tag may be a combination of image tags with the first image tag.

[0069] According to one embodiment of the present disclosure, a device for performing image tagging may utilize a word association analysis model for words corresponding to image tags to determine a combination of image tags. The word association analysis model may be a model that can input two or more words and output a numerical value representing a general association between the input words. When a word corresponding to a first image tag and a word corresponding to a second image tag are input to the word association analysis model and a numerical value exceeding a predetermined threshold is output, the first image tag and the second image tag may be a combination of image tags.

[0070] Referring back to FIG. 4, at step S420, the device performing image tagging can generate an area distribution map for each image tag. When the device performing image tagging determines a combination of image tags from among the acquired image tags, it can generate an area distribution map for each image tag belonging to the combination of image tags.

[0071] Since the area distribution corresponding to an image tag represents the portion of the image used to generate the image tag, each image tag may have a different area distribution. If image tags have similar area distributions, it may indicate that the portions of the image used to generate the image tags are similar.

[0072] FIG. 6 is a diagram for explaining a process of generating an area distribution map corresponding to an image tag according to one embodiment of the present disclosure.

[0073] An image tag generation model according to one embodiment of the present disclosure may be, but is not limited to, a CNN (Convolution Neural Network) model including a plurality of convolutional layers and a fully-connected layer connected to a feature map, as illustrated in FIG. 6. Referring to FIG. 6, the image tag generation model may recognize a class called "car" by inputting an image, and generate an image tag called "car" accordingly.

[0074] The area distribution map illustrated in Fig. 6 is an example of an activation map that specifies a part that played a role in generating an image tag called 'car' (i.e., classifying it into a class corresponding to the image tag), and is based on Grad-CAM (Class Activation Map), but is not limited thereto. The formula for determining the Grad-CAM score for class 'C' is as shown in Mathematical Expression 1.

[0075] [Mathematical Formula 1]

[0076]

[0077] In mathematical expression 1, C is a class, is the kth feature map, is the weight (influence of feature map) for the kth feature map for classification into class 'C', is the value at the (i,j) location in the kth feature map, and Z represents the sum for each feature map. In Grad-CAM, the weight for each feature map is determined through partial differentiation rather than learning.

[0078] Referring to Figure 6, the pixel-wise average value of the gradient for the k-th feature map for classification into a class corresponding to the image tag called 'car' is Each feature map You can create a heatmap by multiplying it, take a pixel-wise sum, and then apply the ReLU function to obtain a class activation map.

[0079] Referring back to FIG. 4, at step S430, the device performing image tagging can generate a mask or bounding box corresponding to each region distribution generated for each image tag. The mask or bounding box is the result of processing that identifies the region used to distinguish the portion of the image used to generate the image tag, and the method of processing that identifies the region is not limited thereto.

[0080] According to one embodiment of the present disclosure, a device for performing image tagging can generate a mask corresponding to each area distribution map generated for each image tag. That is, the device can generate a plurality of masks corresponding to each of a plurality of image tags. The masks corresponding to the area distribution maps may be binary masks. For example, the binary masks corresponding to the area distribution maps can represent a portion utilized in generating the image tags by performing binarization processing (setting a score equal to or greater than a threshold value to 1, and setting the rest to 0) on the area distribution maps.

[0081] According to one embodiment of the present disclosure, a device for performing image tagging can generate a bounding box corresponding to each area distribution map generated for each image tag. That is, the device can generate a plurality of bounding boxes corresponding to each of a plurality of image tags. The bounding box corresponding to the area distribution map can represent a predetermined area including a portion used to generate an image tag by surrounding a predetermined area including a portion used to generate an image tag within the area distribution map with a shape such as a rectangle.

[0082] At step S440, the device performing image tagging can determine the degree of relevance between image tags based on the degree of overlap between masks or bounding boxes corresponding to each area distribution map generated for each image tag.

[0083] The relevance between image tags refers to the general relevance between words corresponding to image tags and the relevance between image tags in a specific image. A device performing image tagging can compare the similarity between masks or bounding boxes corresponding to each area distribution map generated for each image tag.

[0084] According to one embodiment of the present disclosure, a device for performing image tagging can determine a degree of correlation between acquired image tags based on the degree of overlap between the generated masks when generating masks corresponding to each area distribution map generated for each image tag. The device for performing image tagging can determine an intersection of unions (IoU) between the masks corresponding to the image tags.

[0085] According to one embodiment of the present disclosure, a device for performing image tagging can determine a degree of correlation between acquired image tags based on the degree of overlap between the generated bounding boxes when generating bounding boxes corresponding to each area distribution map generated for each image tag. The device for performing image tagging can determine an intersection of unions (IoU) between bounding boxes corresponding to the image tags.

[0086] FIG. 7 is a diagram illustrating a process for determining a degree of relevance between image tags corresponding to a combination of image tags according to one embodiment of the present disclosure.

[0087] Referring to Fig. 7, multiple image tags can be obtained from an image of a red car. The multiple image tags can be 'car', 'red', 'people', 'street', 'old', 'stand', 'park', etc. Fig. 7 illustrates a process of determining the degree of correlation between the first image tag 'car' and the second image tag 'red' among the multiple image tags.

[0088] Referring to FIG. 7, the first region distribution corresponding to the first image tag 'car' represents information about a portion of the image that influenced classification into a class corresponding to the first image tag within the image. The first region distribution corresponding to the first image tag 'car' can be visualized so that a region recognized as a car within the image is distinguished from a region other than a car. The first region distribution can be displayed in different colors depending on the degree to which the portion of the image influenced classification into a class corresponding to the first image tag. The first mask corresponding to the first region distribution (or the first class activation map) is obtained through binarization processing (setting a score equal to or greater than a threshold value to 1, and setting the rest to 0) on the class activation map, and can distinguish a portion utilized to generate the first image tag.

[0089] The second region distribution corresponding to the second image tag 'red' represents information about the part of the image that influenced classification into the class corresponding to the second image tag within the image. The second region distribution corresponding to the second image tag 'red' can be visualized so that the area recognized as red within the image is distinguished from other areas. The second region distribution can be displayed in different colors depending on the degree to which the part of the image influenced classification into the class corresponding to the second image tag. The second mask corresponding to the second region distribution (or the second class activation map) is obtained by binarizing the class activation map and can distinguish the part used to generate the second image tag.

[0090] A device for performing image tagging can obtain an overlap between a first mask corresponding to a first area distribution and a second mask corresponding to a second area distribution. Referring to FIG. 7, the overlap between the first mask and the second mask can be obtained by determining an Intersection of Union (IoU) value between an unmasked area of ​​the first mask and an unmasked area of ​​the second mask. In FIG. 7, since the IoU between the first mask and the second mask is '0.64', it can be determined that the first image tag 'car' and the second image tag 'red' have a relevance of 64%.

[0091] FIG. 8 is a diagram illustrating an example of using an image tag generation model based on a transformer decoder according to one embodiment of the present disclosure.

[0092] Referring to FIG. 8, the image tag generation model based on the transformer decoder may include backbone layers that extract spatial features from an input image, a multi-layer transformer decoder that inputs spatial features output from the backbone layers, and a linear projection layer. The multi-layer transformer decoder may perform query updating and adaptive feature pooling, and the linear projection layer may perform logit determination of each image tag.

[0093] The transformer decoder in each layer can perform cross-attention on the spatial features output from the backbone layers and the label embeddings corresponding to each image tag used as a query. Furthermore, the transformer decoder in each layer can perform self-attention based on the label embeddings corresponding to each image tag output from the transformer decoder in the previous layer, thereby updating the label embeddings corresponding to each image tag used as a query.

[0094] Label embeddings corresponding to each image tag are updated at each layer of a multi-layer transformer decoder and can reflect spatial features extracted from the image through cross-attention. A region distribution map corresponding to the image tag can be generated based on a cross-attention map that determines whether the label embedding corresponding to each image tag is influenced by a specific spatial feature extracted from the image. A device that performs image tagging can determine the degree of relevance between image tags based on the region distribution map derived from the cross-attention map corresponding to each image tag.

[0095] Referring back to FIG. 3, at step S330, a device for performing image tagging may generate a combined tag in which image tags are connected based on the determined relevance between the image tags. In the present disclosure, a 'combined tag' refers to an image tag generated by connecting image tags. The combined tag may be connected by concatenating image tags in a predetermined order. The combined tag may be connected to other image tags by changing the endings of some of the image tags. The connected image tags may be separated by a space.

[0096] A device performing image tagging can determine the connection order of image tags in order to connect them. For example, if there is a first image tag and a second image tag, the device performing image tagging can connect them as "first image tag + second image tag" by having the first image tag come first, or can connect them as "second image tag + second image tag" by having the second image tag come first.

[0097] A device performing image tagging can determine whether to change the suffix when a verb, such as a verb or adjective, is used to connect image tags. For example, if there is a first image tag that is a verb and a second image tag that is a noun, the device performing image tagging can determine whether to change the suffix when connecting the first image tag so that it comes first, if the first image tag is a verb infinitive.

[0098] FIG. 9 is a detailed flowchart of a step for generating a combined tag based on the degree of association between image tags according to one embodiment of the present disclosure.

[0099] In step S910, a device performing image tagging can obtain a relevance corresponding to a combination of image tags. The device performing image tagging can obtain a relevance between image tags belonging to a combination of image tags.

[0100] At step S920, the device performing image tagging can determine whether the relevance corresponding to the combination of image tags is greater than or equal to a predetermined threshold. The device performing image tagging can determine whether the relevance between image tags belonging to the combination of image tags is greater than or equal to a predetermined threshold.

[0101] If the degree of relevance is less than the threshold, the device may proceed to step S940.

[0102] In step S930, the device for performing image tagging can generate a combined tag in which the image tags are connected based on whether the degree of relevance corresponding to the combination of image tags is greater than or equal to a predetermined threshold, in other words, whether the degree of relevance between image tags belonging to the combination of image tags is greater than or equal to a predetermined threshold. The device for performing image tagging can generate a combined tag by connecting the image tags by determining the connection order of the image tags based on the probability of being input as a search query, and determining whether to change the ending of the image tag.

[0103] In the example of Fig. 7, if the threshold for determining the degree of relevance between two image tags is 60%, the degree of relevance between the first image tag 'car' and the second image tag 'red' is '64%', so a new image tag called 'red car' can be created.

[0104] At step 940, the device performing image tagging determines whether there are more combinations of image tags, and if there are more combinations of image tags, the process can proceed again from step 910, and if there are no more combinations of image tags, the process of generating a combined tag can be terminated.

[0105] Referring back to FIG. 3, at step S340, the device performing image tagging can perform image tagging on an image using the acquired image tags and the generated combined tags. The device performing image tagging can perform image tagging on an image by adding the generated combined tags as well as the image tags acquired using the image tag generation model. The device performing image tagging can associate and store the generated combined tags with the image as well as the image tags for the corresponding image.

[0106] FIG. 10 is a block diagram illustrating an electronic device (100) that performs image tagging according to one embodiment of the present disclosure.

[0107] Referring to FIG. 10, an electronic device (100) performing image tagging according to one embodiment may include a memory (110), a processor (120), a communication interface (130), an input / output interface (140), and a camera (150).

[0108] The memory (110) may store instructions, data structures, and program codes that can be read by the processor (120). In one embodiment of the present disclosure, operations performed by the processor (120) may be implemented by executing instructions or codes of a program stored in the memory (110).

[0109] The memory (110) may include a flash memory type, a hard disk type, a multimedia card micro type, a card type memory (e.g., secure digital (SD) or XD memory, etc.), and may include a non-volatile memory including at least one of a ROM (Read-Only Memory), an Electrically Erasable Programmable ROM (EEPROM), a PROM (Programmable ROM), a magnetic memory, a magnetic disk, and an optical disk, and a volatile memory such as a dynamic RAM (DRAM) or a static RAM (SRAM).

[0110] According to one embodiment, the memory (110) may store one or more instructions and / or programs that control an electronic device (100) performing image tagging to process an image. For example, the memory (110) may store an image tag generation module, a relevance analysis module, a combined tag generation module, and an image tagging module. If learning of an artificial intelligence model is required within the electronic device (100) performing image tagging, a model learning module may be additionally installed.

[0111] The processor (120) can control the operations or functions performed by the electronic device (100) performing image tagging by executing instructions or programmed software modules stored in the memory (110). The processor (120) can be configured with hardware components that perform arithmetic, logic, and input / output operations and signal processing. The processor (120) can control the overall operations performed by the electronic device (100) performing image tagging by executing one or more instructions stored in the memory (110).

[0112] The processor (120) may be configured with at least one of, for example, a Central Processing Unit, a microprocessor, a GPU, an Application Specific Integrated Circuits (ASICs), a Digital Signal Processors (DSPs), a Digital Signal Processing Devices (DSPDs), a Programmable Logic Devices (PLDs), a Field Programmable Gate Arrays (FPGAs), an Application Processor, a Neural Processing Unit, or an artificial intelligence processor designed with a hardware structure specialized for processing artificial intelligence models, but is not limited thereto. Each processor constituting the processor (120) may be a dedicated processor for performing a predetermined function.

[0113] The communication interface (130) can perform wired or wireless communication with other devices or networks. The communication interface (130) can include a communication circuit or a communication module that supports at least one of various wired or wireless communication methods. For example, the communication interface (130) can perform data communication between the electronic device (100) that performs image tagging and other devices by using at least one of data communication methods including wired LAN, wireless LAN, Wi-Fi, Bluetooth, ZigBee, Wi-Fi Direct (WFD), infrared Data Association (IrDA), Bluetooth Low Energy (BLE), Near Field Communication (NFC), Wireless Broadband Internet (Wibro), World Interoperability for Microwave Access (WiMAX), Shared Wireless Access Protocol (SWAP), Wireless Gigabit Alliances (WiGig), and RF communication.

[0114] According to one embodiment of the present disclosure, the communication interface (130) can receive an artificial intelligence model or image used for performing image tagging from an external device. For example, the communication interface (130) can receive an artificial intelligence model or database trained on a server from the server. The communication interface (130) can receive an image from another electronic device. The communication interface (130) can share an image generated by the electronic device (100) performing image tagging with another user or transmit the image to an external device so that it can be displayed on the external device.

[0115] A communication interface (130) according to one embodiment of the present disclosure can transmit an image or information about an image acquired by an electronic device (100) performing image tagging to a server, or receive an image from the server. For example, the communication interface (130) can transmit an image captured by the electronic device (100) performing image tagging or an image received from another electronic device to the server. The communication interface (130) can transmit information about an image input through the input / output interface (140) of the electronic device (100) performing image tagging to the server. The communication interface (130) can receive an image stored in the server (200) from the server (200).

[0116] The input / output interface (140) includes an output unit that provides information or images, and may further include an input unit that receives input. The output unit may include a display panel and a controller that controls the display panel, and may be implemented in various ways, such as an OLED (Organic Light Emitting Diodes) display, an AM-OLED (Active-Matrix Organic Light-Emitting Diode) display, an LCD (Liquid Crystal Display), etc. The input unit may receive various types of input from a user, and may include at least one of a touch panel, a keypad, and a pen recognition panel. The input / output interface (140) may be provided in the form of a touch screen in which a display panel and a touch panel are combined, and may be implemented in a flexible or foldable manner.

[0117] An input / output interface (140) according to one embodiment of the present disclosure can obtain information about an image from a user. The user can select an image for image tagging from among multiple images via the input / output interface (140). The input / output interface (140) can receive user input regarding information or control commands required during the image tagging process.

[0118] The camera (150) is a hardware module that captures images. The camera (150) can capture images. The camera (150) may include at least one camera module, and may support functions such as depth, telephoto, wide-angle, and ultra-wide-angle, depending on the specifications of the electronic device (100) that performs image tagging. Unless the electronic device (100) that performs image tagging is used to directly capture images, the camera (150) may be excluded from the electronic device (100) that performs image tagging.

[0119] FIG. 11 is a block diagram illustrating the configuration and operation of an electronic device (100) that performs image tagging according to one embodiment of the present disclosure.

[0120] Referring to FIG. 11, an electronic device (100) for performing image tagging includes a memory (110) storing at least one instruction and at least one processor (120) operatively connected to the memory (110) to execute at least one instruction. The processor (120) may execute at least one instruction to load and execute commands or codes for an image tag generation module (121), a relevance analysis module (123), a combination tag generation module (125), and an image tagging module (127).

[0121] The processor (120) can obtain a plurality of image tags from an image by executing at least one instruction, using an image tag generation model. The processor (120) can use the image tag generation model, which is a multi-label recognition model, through the image tag generation module (121). The processor (120) can classify a class corresponding to an image tag based on features extracted from the image through the image tag generation module (121), and obtain a plurality of image tags based on the classification result.

[0122] The processor (120) may execute at least one instruction to determine the degree of relevance between the acquired image tags based on the area distribution map corresponding to each image tag within the image. The processor (120) may determine a combination of image tags among the image tags acquired by the image tag generation module (121) through the degree of relevance analysis module (123). The processor (120) may generate an area distribution map corresponding to each image tag through the degree of relevance analysis module (123), and may obtain the degree of relevance between the image tags acquired by the image tag generation module (121) by comparing the generated area distribution maps. The area distribution map may be based on an activation map or an attention map corresponding to each image tag. The processor (120) may generate a mask corresponding to each of the generated area distribution maps through the degree of relevance analysis module (123), and may determine the degree of relevance between the image tags acquired by the image tag generation module (121) based on the degree of overlap between the generated masks. The processor (120) can generate a bounding box corresponding to each generated area distribution through the relevance analysis module (123), and determine the relevance between image tags obtained from the image tag generation module (121) based on the degree of overlap between the generated bounding boxes.

[0123] The processor (120) can determine a combination of image tags from image tags selected from among the image tags obtained from the image tag generation module (121) through the relevance analysis module (123). The processor (120) can determine a combination of image tags from the selected image tags through the relevance analysis module (123) based on the part-of-speech of the image tags obtained from the image tag generation module (121) or whether the obtained image tags have a relationship corresponding to the characteristics of objects. The processor (120) can determine a relevance between image tags belonging to a combination of image tags through the relevance analysis module (123) based on the comparison result of the area distribution for each combination of image tags.

[0124] The processor (120) may execute at least one instruction to generate a combined tag in which image tags are connected based on the relevance determined by the relevance analysis module (123). The processor (120) may generate a combined tag in which the acquired image tags are connected if the relevance between the image tags acquired by the image tag generation module (121) is greater than a predetermined threshold through the combined tag generation module (125). The processor (120) may determine the connection order of the image tags based on the probability of being input as a search query through the combined tag generation module (125), and determine whether to change the ending of the image tag, thereby connecting the image tags acquired by the image tag generation module (121).

[0125] The processor (120) may perform image tagging on an image by executing at least one instruction using the image tags obtained from the image tag generation module (121) and the combined tag generated from the combined tag generation module (125). The processor (120) may perform image tagging on an image by adding the combined tag generated from the combined tag generation module (125) as well as the image tags obtained using the image tag generation model from the image tag generation module (121) through the image tagging module (127). The processor (120) may associate the combined tag with the image tag as well as the image tag for the corresponding image through the image tagging module (127) and store the tagged image in the memory (110).

[0126] FIG. 12 is a block diagram for explaining the operation of a relevance analysis module (123) and a combined tag generation module (125) according to one embodiment of the present disclosure.

[0127] According to one embodiment of the present disclosure, the relevance analysis module (123) may include an image tag combination unit, an area distribution generation unit, a mask generation unit, a mask comparison unit, and a relevance determination unit. The relevance analysis module (123) may receive a plurality of image tags as input. Referring to FIG. 12, a plurality of image tags (Tag-A, Tag-B, Tag-C, Tag-D) may be acquired for an image.

[0128] The image tag combination unit can generate all possible combinations of image tags. As illustrated in Fig. 12, when a total of four image tags are input to the relevance analysis module, the image tag combination unit can generate image tag combinations consisting of two image tags, three image tags, and four image tags.

[0129] The region distribution generation unit can generate a region distribution corresponding to each image tag. The region distribution generation unit can generate a region distribution for each image tag, which indicates a part of the image that has influenced classification into a class corresponding to each image tag. For example, the region distribution corresponding to Tag-A can indicate an image region that has influenced classification into a class corresponding to Tag-A within the image. The region distribution corresponding to Tag-B can indicate an image region that has influenced classification into a class corresponding to Tag-B within the image.

[0130] The mask generation unit can generate a mask corresponding to each area distribution generated for each image tag. The mask generation unit can indicate a area used to generate an image tag by masking the remaining portion except for the portion used to generate the image tag within the area distribution. For example, by masking the remaining portion except for the portion used to generate Tag-A within the area distribution corresponding to Tag-A, a mask of an area distribution corresponding to Tag-A can be generated. For example, by masking the remaining portion except for the portion used to generate Tag-B within the area distribution corresponding to Tag-B, a mask of an area distribution corresponding to Tag-B can be generated.

[0131] The mask comparison unit can compare the similarity between masks corresponding to each area distribution map generated for each image tag for a specific image. For example, the mask comparison unit can determine the degree of overlap between masks corresponding to each area distribution map generated for each image tag. The mask comparison unit can determine the degree of overlap between masks corresponding to image tags for all combinations of image tags generated by the image tag combination unit.

[0132] According to one embodiment of the present disclosure, the mask generation unit may be replaced with a bounding box generation unit, and the mask comparison unit may be replaced with a bounding box comparison unit. The bounding box generation unit may generate a bounding box corresponding to each area distribution map generated for each image tag. The bounding box generation unit may represent a predetermined area including a portion used to generate an image tag by surrounding a predetermined area including a portion used to generate an image tag within the area distribution map with a shape such as a rectangle. The bounding box comparison unit may compare the similarity between the bounding boxes corresponding to each area distribution map generated for each image tag for a specific image. For example, the bounding box comparison unit may determine the degree of overlap between the bounding boxes corresponding to each area distribution map generated for each image tag. The bounding box comparison unit may determine the degree of overlap between the bounding boxes corresponding to the image tags for all combinations of image tags generated by the image tag combination unit.

[0133] The relevance judgment unit can obtain the relevance between image tags based on the determined overlap. The relevance judgment unit can determine that the higher the overlap between the masks or bounding boxes corresponding to each area distribution map generated for each image tag, the higher the relevance between the image tags. The relevance judgment unit can output, to the combined tag generation module (125), a combination of image tags having a relevance between the image tags greater than or equal to a predetermined threshold based on the determined overlap. Referring to FIG. 12, it can be seen that the relevance judgment unit determines whether the relevance between the image tags is greater than or equal to a predetermined threshold, and outputs (Tag-A, Tag-C), (Tag-A, Tag-D), (Tag-B, Tag-C), (Tag-B, Tag-D) among the combinations of image tags generated by the image tag combination unit to the combined tag generation module (125).

[0134] The combined tag generation module (125) can generate combined tags by connecting image tags belonging to a combination of image tags input from the relevance analysis module (123). The combined tag generation module (125) can determine the connection order of image tags based on the probability of being input as a search query, and determine whether to change the ending of the image tag, thereby connecting image tags belonging to a combination of image tags.

[0135] Referring to FIG. 12, it can be seen that the combined tag generation module (125) generates a combined tag, Tag-CA, from the combination of image tags of (Tag-A, Tag-C), a combined tag, Tag-DA, from the combination of image tags of (Tag-A, Tag-D), a combined tag, Tag-CB, from the combination of image tags of (Tag-B, Tag-C), and a combined tag, Tag-DB, from the combination of image tags of (Tag-B, Tag-D), thereby generating a plurality of combined tags. Applying the example of FIG. 5, the combined tag generation module (125) can generate a combined tag, 'red people', from the combination of image tags of (people, red) by determining whether to change the connection order of the image tags and the ending of the image tag. The combined tag generation module (125) can generate a combined tag, 'running people', from the combination of image tags of (people, running). The combined tag generation module (125) can generate a combined tag, 'red car', from a combination of image tags, 'car, red'. The combined tag generation module (125) can generate a combined tag, 'running car', from a combination of image tags, 'car, running'.

[0136] FIG. 13 is a block diagram for explaining the operation of a relevance analysis module (123) and a combined tag generation module (125) according to one embodiment of the present disclosure.

[0137] According to one embodiment of the present disclosure, the relevance analysis module (123) may include an image tag selection unit, an image tag combination unit, an area distribution generation unit, a mask generation unit, a mask comparison unit, and a relevance determination unit. The same details as those described in detail in FIG. 12 for components with the same names are omitted below.

[0138] The image tag selection unit can select image tags based on the parts of speech of the image tags among the image tags. The image tag selection unit can select image tags by checking the parts of speech of each image tag. For example, the image tag selection unit can select image tags to be used in generating image tag combinations based on common combinations such as 'adjective + noun', 'verb + noun', 'verb + adjective + noun', and 'noun + noun'.

[0139] The image tag selection unit can select image tags based on whether the image tags correspond to objects and their characteristics. The image tag selection unit can select image tags to be used in generating a combination of image tags based on whether each image tag relates to an object or to an attribute or action of the object.

[0140] The image tag selection unit can select image tags by considering combinations of image tags likely to be search queries. Conversely, it may not select image tags if the combination is inappropriate or ineligible for a search query. The image tag selection unit can utilize lookup tables, word ontology databases, and word association analysis models to guide image tag selection.

[0141] The image tag combination unit can generate a combination of image tags based on the image tags selected by the image tag selection unit. The image tag combination unit can only generate a combination of image tags composed of the image tags selected by the image tag selection unit, rather than generating a combination of image tags based on all image tags input to the relevance analysis module (123).

[0142] The area distribution generation unit can generate an area distribution corresponding to each image tag belonging to the image tag combination based on the combination of image tags generated by the image tag combination unit. The mask generation unit can generate a mask corresponding to each area distribution generated for each image tag. The mask comparison unit can determine the degree of overlap between masks corresponding to the image tags for the combination of image tags generated by the image tag combination unit.

[0143] According to one embodiment of the present disclosure, the mask generation unit may be replaced with a bounding box generation unit, and the mask comparison unit may be replaced with a bounding box comparison unit. The bounding box generation unit may generate a bounding box corresponding to each area distribution map generated for each image tag. The bounding box comparison unit may determine the degree of overlap between bounding boxes corresponding to image tags for a combination of image tags generated by the image tag combination unit.

[0144] The relevance judgment unit can output combinations of image tags whose relevance between the image tags is greater than or equal to a predetermined threshold based on the determined overlap, to the combined tag generation module (125). Referring to FIG. 13, it can be seen that the relevance judgment unit determines whether the relevance between the image tags is greater than or equal to a predetermined threshold, and outputs (Tag-A, Tag-D), (Tag-B, Tag-C), and (Tag-B, Tag-D) among the combinations of image tags generated by the image tag combination unit to the combined tag generation module (125).

[0145] The combined tag generation module (125) can generate a combined tag by connecting image tags belonging to a combination of image tags input from the relevance analysis module (123).

[0146] Referring to FIG. 13, it can be seen that the combined tag generation module (125) generates a combined tag, Tag-DA, from a combination of image tags of (Tag-A, Tag-D), a combined tag, Tag-CB, from a combination of image tags of (Tag-B, Tag-C), and a combined tag, Tag-DB, from a combination of image tags of (Tag-B, Tag-D), thereby generating multiple combined tags. Applying the example of FIG. 5, the combined tag generation module (125) can generate a combined tag, 'running people', from a combination of image tags of (people, running) by determining whether to change the connection order of the image tags and the ending of the image tag. The combined tag generation module (125) can generate a combined tag, 'red car', from a combination of image tags of (car, red). The combined tag generation module (125) can generate a combined tag, 'running car', from a combination of image tags of (car, running). However, unlike the case of Fig. 12, in Fig. 13, the combination of image tags (people, red) is excluded from being generated, so the combined tag 'red people' is not generated.

[0147] Meanwhile, embodiments of the present disclosure may also be implemented in the form of a recording medium containing computer-executable instructions, such as program modules, executed by a computer. Computer-readable media may be any available media that can be accessed by a computer, and include both volatile and nonvolatile media, removable and non-removable media. Furthermore, computer-readable media may include computer storage media and communication media. Computer storage media include both volatile and nonvolatile, removable and non-removable media implemented in any method or technology for storing information, such as computer-readable instructions, data structures, program modules, or other data. Communication media may typically include computer-readable instructions, data structures, or other data in a modulated data signal, such as program modules.

[0148] Additionally, a computer-readable storage medium may be provided in the form of a non-transitory storage medium. Here, the term "non-transitory storage medium" simply means a tangible device that does not contain signals (e.g., electromagnetic waves). This term does not distinguish between cases where data is permanently stored in the storage medium and cases where data is temporarily stored. For example, a "non-transitory storage medium" may include a buffer in which data is temporarily stored.

[0149] According to one embodiment, the method according to various embodiments disclosed in the present document may be provided as included in a computer program product. The computer program product may be traded as a product between a seller and a buyer. The computer program product may be distributed in the form of a machine-readable storage medium (e.g., compact disc read-only memory (CD-ROM)), or may be distributed online (e.g., downloaded or uploaded) through an application store or directly between two user devices (e.g., smartphones). In the case of online distribution, at least a portion of the computer program product (e.g., a downloadable app) may be temporarily stored or temporarily generated in a machine-readable storage medium, such as the memory of a manufacturer's server, an application store's server, or an intermediary server.

[0150] According to one embodiment of the present disclosure, a method for performing image tagging is provided. The method for performing image tagging may include a step (S310) of obtaining a plurality of image tags from an image using an image tag generation model. The method for performing image tagging may include a step (S320) of determining a degree of relevance between the obtained image tags based on an area distribution map corresponding to each image tag within the image. The method for performing image tagging may include a step (S330) of generating a combined tag in which the image tags are linked based on the determined degree of relevance. The method for performing image tagging may include a step (S340) of performing image tagging on the image using the obtained image tags and the generated combined tag.

[0151] According to one embodiment of the present disclosure, the step of determining relevance (S320) may include a step of generating a region distribution map corresponding to each image tag (S410, S420). The step of determining relevance (S320) may include a step of obtaining a relevance between the obtained image tags by comparing the generated region distribution maps (S430, S440).

[0152] The steps (S430, S440) for obtaining a degree of relevance may include a step (S430) of generating a mask corresponding to each generated area distribution map. The steps (S430, S440) for obtaining a degree of relevance may include a step (S440) of determining a degree of relevance between acquired image tags based on the degree of overlap between the generated masks.

[0153] The steps (S430, S440) for obtaining a degree of relevance may include a step (S430) of generating a bounding box corresponding to each generated area distribution map. The steps (S430, S440) for obtaining a degree of relevance may include a step (S440) of determining a degree of relevance between the acquired image tags based on the degree of overlap between the generated bounding boxes.

[0154] According to one embodiment of the present disclosure, the step of determining the degree of relevance (S320) may include a step of determining a combination of image tags from image tags selected from among the acquired image tags (S410). The step of determining the degree of relevance (S320) may include a step of determining the degree of relevance between image tags belonging to a combination of image tags based on a comparison result of the area distribution for each combination of image tags (S420, S430, S440).

[0155] The step (S410) of determining a combination of image tags may include determining a combination of image tags from selected image tags according to the parts of speech of the acquired image tags or according to whether the acquired image tags have a relationship corresponding to an object and a characteristic of the object.

[0156] According to one embodiment of the present disclosure, the area distribution may be based on an activation map or an attention map corresponding to each image tag.

[0157] According to one embodiment of the present disclosure, the step (S330) of generating a combined tag may include generating a combined tag in which the acquired image tags are connected when the degree of correlation between the acquired image tags is greater than or equal to a predetermined threshold.

[0158] The step (S330) of generating a combined tag may include determining the connection order of image tags based on the probability of being entered as a search query, determining whether to change the ending of the image tag, and connecting the obtained image tags.

[0159] According to one embodiment of the present disclosure, the step (S310) of obtaining a plurality of image tags may include classifying a class corresponding to an image tag based on features extracted from an image using the image tag generation model, which is a multi-label recognition model, and obtaining a plurality of image tags based on the classification result.

[0160] According to one embodiment of the present disclosure, a computer-readable recording medium having recorded thereon a program for executing a method for performing the above image tagging is provided.

[0161] According to one embodiment of the present disclosure, a device for performing image tagging is provided. The device for performing image tagging may include a memory (110) storing at least one instruction and at least one processor (120) operatively connected to the memory (110) and executing at least one instruction. The at least one processor (120) may execute at least one instruction to obtain a plurality of image tags from an image using an image tag generation model. The at least one processor (120) may execute at least one instruction to determine a degree of correlation between the obtained image tags based on a distribution map of areas corresponding to each image tag within the image. The at least one processor (120) may execute at least one instruction to generate a combined tag in which the image tags are connected based on the determined degree of correlation. The at least one processor (120) may execute at least one instruction to perform image tagging on an image using the obtained image tags and the generated combined tag.

[0162] According to one embodiment of the present disclosure, at least one processor (120) may execute at least one instruction to generate an area distribution map corresponding to each image tag, and compare the generated area distribution maps to obtain a correlation between the obtained image tags.

[0163] At least one processor (120) can execute at least one instruction to generate a mask corresponding to each generated area distribution, and determine a degree of correlation between acquired image tags based on the degree of overlap between the generated masks.

[0164] At least one processor (120) may execute at least one instruction to generate a bounding box corresponding to each generated area distribution, and determine a degree of correlation between acquired image tags based on the degree of overlap between the generated bounding boxes.

[0165] According to one embodiment of the present disclosure, at least one processor (120) may execute at least one instruction to determine a combination of image tags from image tags selected from among acquired image tags, and, based on a comparison result of area distribution for each combination of image tags, determine a degree of correlation between image tags belonging to the combination of image tags.

[0166] At least one processor (120) may execute at least one instruction to determine a combination of image tags from the selected image tags based on the parts of speech of the acquired image tags or based on whether the acquired image tags have a relationship corresponding to an object and a characteristic of the object.

[0167] According to one embodiment of the present disclosure, the area distribution may be based on an activation map or an attention map corresponding to each image tag.

[0168] According to one embodiment of the present disclosure, at least one processor (120) may execute at least one instruction to generate a combined tag in which the acquired image tags are connected when the degree of correlation between the acquired image tags is greater than or equal to a predetermined threshold.

[0169] Additionally, at least one processor (120) may execute at least one instruction to determine the connection order of image tags based on the probability of being entered as a search query, and to determine whether to change the ending of the image tag, thereby connecting the acquired image tags.

[0170] The above description of the present disclosure is provided for illustrative purposes only, and those skilled in the art will readily appreciate that modifications to other specific forms can be made without altering the technical spirit or essential features of the present disclosure. Therefore, the embodiments described above should be understood as illustrative in all respects and not restrictive. For example, components described as being single may be implemented in a distributed manner, and similarly, components described as being distributed may be implemented in a combined manner.

[0171] The scope of the present disclosure is indicated by the claims described below rather than the detailed description above, and all changes or modifications derived from the meaning and scope of the claims and their equivalent concepts should be interpreted as being included in the scope of the present disclosure.

Claims

1. A step (S310) of obtaining multiple image tags from an image using an image tag generation model; A step (S320) of determining the degree of relevance between the plurality of image tags based on a plurality of area distribution maps each corresponding to the plurality of image tags; Based on the determined relevance, a step (S330) of generating a combined tag in which image tags are connected; and A method for performing image tagging, comprising a step (S340) of performing image tagging on the image using the plurality of image tags and the combined tag.

2. In paragraph 1, The step of determining the above relevance (S320) is Step of generating the above multiple area distribution maps (S410, S420); and A method for performing image tagging, comprising a step (S430, S440) of obtaining a correlation between the plurality of image tags by comparing the plurality of area distribution maps.

3. In paragraph 1 or 2, The step (S430, S440) of obtaining the above relevance is: Step (S430) of generating multiple masks corresponding to each of the multiple area distribution maps; and A method for performing image tagging, comprising a step (S440) of determining a degree of relevance between the plurality of image tags based on the degree of overlap between the plurality of masks.

4. In any one of paragraphs 1 to 3, The step (S430, S440) of obtaining the above relevance is: Step (S430) of generating multiple bounding boxes corresponding to each of the multiple area distributions above; and A method for performing image tagging, comprising a step (S440) of determining a degree of relevance between the plurality of image tags based on the degree of overlap between the plurality of bounding boxes.

5. In any one of paragraphs 1 to 4, The step of determining the above relevance (S320) is Step (S410) of determining a combination of image tags from image tags selected from among the above plurality of image tags; and A method for performing image tagging, comprising a step (S420, S430, S440) of determining a degree of relevance between image tags of a combination of image tags based on a comparison result of the area distribution of each image tag of the combination of image tags.

6. In any one of paragraphs 1 to 5, The step (S410) of determining the combination of the above image tags is A method for performing image tagging, wherein a combination of image tags is determined from the selected image tags based on a part of speech of the plurality of image tags or a relationship between the plurality of image tags and an object and a characteristic of the object.

7. In any one of paragraphs 1 to 6, Each area distribution of the above multiple area distributions is, A method for performing image tagging based on an activation map or an attention map each corresponding to the plurality of image tags.

8. In any one of paragraphs 1 to 7, The step (S330) of generating the above combination tag is: A method for performing image tagging, wherein the combined tag is generated based on the degree of correlation between the plurality of image tags being greater than a predefined threshold.

9. In any one of paragraphs 1 to 8, The step (S330) of generating the above combination tag is: A method for performing image tagging, wherein the connection order of image tags is determined based on the probability of being entered as a search query, and whether to change the ending of at least one image tag of the plurality of image tags is determined, thereby connecting the plurality of image tags.

10. In any one of paragraphs 1 to 9, The step (S310) of obtaining the above multiple image tags is A method for performing image tagging, wherein the image tag generation model, which is a multi-label recognition model, is used to classify a class corresponding to an image tag based on features extracted from the image, and the plurality of image tags are obtained based on the classification result.

11. A memory (110) storing at least one instruction; and comprising at least one processor (120) operatively connected to the memory (110); The at least one processor (120) executes the at least one instruction, An image tagging device which obtains a plurality of image tags from an image using an image tag generation model, determines a degree of relevance between the plurality of image tags based on a plurality of area distribution maps each corresponding to the plurality of image tags, generates a combined tag in which the image tags are connected based on the determined degree of relevance, and performs image tagging on the image using the plurality of image tags and the generated combined tag.

12. In paragraph 11, The at least one processor (120) executes the at least one instruction, A device for performing image tagging, which generates a plurality of area distribution maps and compares the plurality of area distribution maps to obtain a correlation between the plurality of image tags.

13. In clause 11 or 12, The at least one processor (120) executes the at least one instruction, A device for performing image tagging, which generates a plurality of masks each corresponding to a plurality of area distribution maps, and determines a degree of correlation between the plurality of image tags based on the degree of overlap between the plurality of masks.

14. In any one of paragraphs 11 to 13, The at least one processor (120) executes the at least one instruction, A device for performing image tagging, which determines a combination of image tags from image tags selected from among the plurality of image tags, and determines a degree of correlation between image tags of the combination of image tags based on a comparison result of the area distribution of each image tag of the combination of image tags.

15. In a computer-readable recording medium on which instructions are recorded, The above instructions, when executed by at least one processor, cause the at least one processor to: Using the image tag generation model, multiple image tags are obtained from an image, Based on a plurality of area distribution maps each corresponding to the plurality of image tags, the degree of relevance between the plurality of image tags is determined, Based on the above judged relevance, a combined tag is created in which image tags are linked, A recording medium that performs image tagging on the image using the plurality of image tags and the combination tag.

Citation Information

Patent Citations

  • Content similarity degree calculating apparatus, content similarity degree calculating method, and program

    JP2014186582A

  • Related tag group generating apparatus and method

    KR1020100132763A

  • A line light connector

    KR1020230009627A

  • Information provision device, information provision method, and program

    WO2014027415A1