Dataset generation device utilized in image generation artificial intelligence model
The dataset generation device addresses the challenge of efficiently generating prompt data and image-text pair datasets for image generation AI models by utilizing modules for caption data generation and merging, resulting in reduced learning time and costs.
Patent Information
- Application Number
- PCT/KR2024/018485
- Authority / Receiving Office
- WO · WO
- Patent Type
- Applications
- Current Assignee / Owner
- Priority Date
- 2023-11-24
- Filing Date
- 2024-11-21
- Publication Date
- 2025-05-30
AI Technical Summary
Existing image generation artificial intelligence models face challenges in efficiently generating prompt data by merging caption data describing image situations and features, and creating image-text pair datasets.
A dataset generation device that includes modules for generating first and second caption data, merging these to create prompt data, and matching image data with prompt data to generate an image-text pair dataset.
This solution minimizes the time and cost required in the learning process of image generation AI models by effectively generating prompt data and creating image-text pair datasets.
Smart Images

Figure KR2024018485_30052025_PF_FP_ABST
Abstract
Description
Dataset generation device used in image generation artificial intelligence models
[0001] The present invention relates to a dataset generation device utilized in an image generation artificial intelligence model, and more particularly, to a dataset generation device utilized in an image generation artificial intelligence model capable of generating prompt data by merging first caption data describing a situation depicted by an image and second caption data describing a feature of the image, and generating an image-text pair dataset by matching image data representing the image with the prompt data.
[0002] Image captioning using deep neural networks involves outputting natural language sentences, or sequences of words, that describe the input image. This process largely consists of a feature extraction step, which extracts features that describe the input image, and a caption generation step, which generates a caption describing the input image based on the extracted features.
[0003] Conventional techniques for generating captions from input images have adopted the encoder-decoder framework, a framework primarily used in machine translation. This framework first performs a feature extraction step using an encoder, then uses the extracted features as input to a decoder, which then outputs the word sequence of the corresponding caption.
[0004] Some studies have proposed applying soft attention to features representing input images. Applying soft attention to input images means not using the image features extracted from a CNN model as is, but applying different weights based on the content of the image features. This allows for more effective features representing the input image, thereby improving caption generation performance. There are two types of soft attention applied to input images: temporal attention and spatial attention. Temporal attention refers to attention that indicates which frames in an image frame sequence should be focused on. In contrast, spatial attention refers to attention that indicates which space within a single frame should be focused on.
[0005] The problem to be solved by the present invention is to provide a dataset generation device utilized in an image generation artificial intelligence model capable of generating prompt data by merging first caption data describing a situation depicted by an image and second caption data describing the characteristics of the image, and generating an image-text pair dataset by matching the image data representing the image with the prompt data.
[0006] The purposes of the present invention are not limited to those mentioned above, and other unmentioned purposes and advantages of the present invention can be understood through the following description and will be more clearly understood through embodiments of the present invention. Furthermore, it will be readily apparent that the purposes and advantages of the present invention can be realized by the means and combinations thereof set forth in the claims.
[0007] According to one aspect of the present invention for solving the above-described problem, a dataset generation device utilized in an image generation artificial intelligence model may include a processor including a first caption generation module that generates first caption data describing a situation depicted by an image; a second caption generation module that generates second caption data describing a feature of the image; a prompt generation module that generates prompt data by merging the first caption data and the second caption data; and a dataset generation module that generates an image-text pair dataset by matching image data representing the image with the prompt data; and a dataset database that stores the image-text pair dataset.
[0008] The above first caption generation module can input the image data as input data to a first visual-language model, and output the first caption data as output data from the first visual-language model.
[0009] The second caption generation module may input the image data as input data to the second visual-language model, and output image embedding as output data from the second visual-language model.
[0010] The dataset generation device utilized in the image generation artificial intelligence model according to the present invention further includes a feature modifier database in which a plurality of candidate caption data that can be selected to be generated as the second caption data are classified and stored by a plurality of type categories according to the type of the feature; and the second caption generation module
[0011] Each of the plurality of candidate caption data can be input as input data to the second visual-language model, and a text embedding for each of the plurality of candidate caption data can be output as output data from the second visual-language model.
[0012] The second caption generation module may calculate a similarity between each of the plurality of text embeddings and the image embedding, select one or more text embeddings from among the plurality of text embeddings based on the similarity, and generate candidate caption data corresponding to the selected one or more text embeddings as the second caption data.
[0013] The second caption generation module may select text embeddings included in the reference ranking in order of high similarity for each of the plurality of type categories, and generate candidate caption data corresponding to the selected text embeddings as the second caption data.
[0014] Other specific details of the present invention are included in the detailed description and drawings.
[0015] A dataset generation device utilized in an image generation artificial intelligence model according to the present invention generates prompt data by merging first caption data describing a situation depicted by an image and second caption data describing a feature of the image, and matches image data representing the image with the prompt data to generate an image-text pair dataset, thereby minimizing the time and cost required in the learning process of the image generation artificial intelligence model.
[0016] The effects of the present invention are not limited to the effects mentioned above, and other effects not mentioned will be clearly understood by those skilled in the art from the description below.
[0017] FIG. 1 is a connection diagram between a dataset generation device used in an image generation artificial intelligence model according to one embodiment of the present disclosure and another device.
[0018] FIG. 2 is a block diagram of a dataset generation device utilized in an image generation artificial intelligence model according to an embodiment of the present disclosure.
[0019] FIG. 3 is a block diagram of a first visual-language model used by a dataset generation device utilized in an image generation artificial intelligence model according to one embodiment of the present disclosure.
[0020] FIG. 4 is a diagram for explaining a process in which a dataset generation device utilized in an image generation artificial intelligence model according to one embodiment of the present disclosure generates first caption data.
[0021] FIG. 5 is a block diagram of a second visual-language model used by a dataset generation device utilized in an image generation artificial intelligence model according to one embodiment of the present disclosure.
[0022] FIG. 6 is a diagram illustrating a process in which a dataset generation device utilized in an image generation artificial intelligence model according to one embodiment of the present disclosure generates second caption data.
[0023] FIG. 7 is a diagram illustrating a process of generating prompt data by a dataset generation device utilized in an image generation artificial intelligence model according to an embodiment of the present disclosure.
[0024] The advantages and features of the present invention, and the methods for achieving them, will become clearer with reference to the embodiments described in detail below together with the accompanying drawings. However, the present invention is not limited to the embodiments disclosed below and may be implemented in various different forms. These embodiments are provided solely to ensure that the disclosure of the present invention is complete and to fully inform those skilled in the art of the scope of the present invention, and the present invention is defined solely by the scope of the claims.
[0025] The terminology used herein is for the purpose of describing embodiments only and is not intended to limit the present invention. In this specification, the singular also includes the plural unless specifically stated otherwise. As used herein, the terms "comprises" and / or "comprising" do not exclude the presence or addition of one or more other components in addition to the mentioned components. Like reference numerals refer to like components throughout the specification, and "and / or" includes each and any combination of one or more of the mentioned components. Although "first", "second", etc. are used to describe various components, these components are not limited by these terms. These terms are only used to distinguish one component from another. Therefore, it should be understood that a first component mentioned below may also be a second component within the technical spirit of the present invention.
[0026] Unless otherwise defined, all terms (including technical and scientific terms) used herein may be used in their common sense to those skilled in the art to which the present invention pertains. Furthermore, terms defined in commonly used dictionaries are not to be interpreted ideally or excessively unless explicitly and specifically defined otherwise.
[0027] The term "part" or "module" as used herein refers to a software or hardware component such as an FPGA or ASIC, and the "part" or "module" performs certain functions. However, the "part" or "module" is not limited to software or hardware. The "part" or "module" may be configured to reside on an addressable storage medium and may be configured to execute one or more processors. Thus, by way of example, the "part" or "module" includes components such as software components, object-oriented software components, class components, and task components, as well as processes, functions, attributes, procedures, subroutines, segments of program code, drivers, firmware, microcode, circuitry, data, databases, data structures, tables, arrays, and variables. The functionality provided within the components and "parts" or "modules" may be combined into fewer components and "parts" or "modules" or further separated into additional components and "parts" or "modules."
[0028] Spatially relative terms such as "below," "beneath," "lower," "above," and "upper" can be used to easily describe the relationship between one component and other components as depicted in the drawings. Spatially relative terms should be understood to include different orientations of the components during use or operation in addition to the orientations depicted in the drawings. For example, if a component depicted in the drawings were flipped over, a component described as "below" or "beneath" another component could instead be located "above" the other component. Thus, the exemplary term "below" can include both the above and below orientations. Components can also be oriented in other directions, and thus spatially relative terms can be interpreted accordingly.
[0029] In this specification, the term "computer" refers to any type of hardware device including at least one processor, and may also be understood to encompass software components operating on the hardware device, depending on the embodiment. For example, the term "computer" may be understood to encompass, but is not limited to, servers, smartphones, tablet PCs, desktops, laptops, and all user clients and applications running on each device.
[0030]
[0031] Hereinafter, embodiments of the present invention will be described in detail with reference to the attached drawings.
[0032] FIG. 1 is a connection configuration diagram between a dataset generation device used in an image generation artificial intelligence model according to an embodiment of the present disclosure and another device, and FIG. 2 is a block diagram of a dataset generation device used in an image generation artificial intelligence model according to an embodiment of the present disclosure.
[0033] Referring to FIGS. 1 and 2, a dataset generation device (100) utilized in an image generation artificial intelligence model according to an embodiment of the present disclosure can receive an image that serves as the basis for generating an image-text pair dataset from an external database (DB), generate prompt data from the image, and then generate an image-text pair dataset in which the image and the prompt data are matched multiple times.
[0034] Meanwhile, a dataset generation device (100) utilized in an image generation artificial intelligence model according to an embodiment of the present disclosure may receive another image from a user device (200), generate prompt data from the other image in the manner described above, and then transmit the prompt data to the user device (200).
[0035] Next, the dataset generation device (100) utilized in the image generation artificial intelligence model can receive modified prompt data from the user device (200), generate a new image based on the received prompt data, and then transmit the new image to the user device (200).
[0036] The dataset generation device (100) utilized in the image generation artificial intelligence model according to one embodiment of the present disclosure can generate first caption data and second caption data from an image, and merge the first caption data and second caption data to generate prompt data.
[0037] Thereafter, a dataset generation device (100) utilized in an image generation artificial intelligence model according to an embodiment of the present disclosure can generate an image-text pair dataset by matching image data representing an image with prompt data.
[0038] To this end, a dataset generation device (100) utilized in an image generation artificial intelligence model according to an embodiment of the present disclosure may include a processor (110), a communication unit (120), a memory (130), a dataset database (DB1), and a feature modifier database (DB2).
[0039] The processor (110) can control the overall operation of the dataset generation device (100) utilized in the image generation artificial intelligence model according to one embodiment of the present disclosure.
[0040] The processor (110) may include a first caption generation module (111), a second caption generation module (112), a prompt generation module (113), and a dataset generation module (114).
[0041] Below, the process of generating the first caption data will be described.
[0042] FIG. 3 is a block diagram of a first visual-language model used by a dataset generation device used in an image generation artificial intelligence model according to an embodiment of the present disclosure, and FIG. 4 is a diagram for explaining a process in which a dataset generation device used in an image generation artificial intelligence model according to an embodiment of the present disclosure generates first caption data.
[0043] Referring further to FIGS. 3 and 4, the first caption generation module (111) can generate first caption data (CD1) that describes the situation depicted by the image.
[0044] This first caption generation module (111) can generate image data (ID), which is a matrix representing an image, and obtain first caption data (CD1) from the image data (ID) using the first visual-language model (M1).
[0045] To this end, the first caption generation module (111) can convert the size and format of the image data (ID) so that the image data (ID) can be input into the first time-language model (M1).
[0046] Specifically, the first caption generation module (111) can convert the size of image data (ID) as a matrix, interpolate the converted image data (ID), and normalize the interpolated image data (ID).
[0047] At this time, the size conversion, interpolation, and normalization of the image data (ID) may differ depending on the type of input terminal of the first visual-language model (M1).
[0048] The first caption generation module (111) can input normalized image data (ID) as input data to the first visual-language model (M1) and receive first caption data (CD1) as output data from the first visual-language model (M1).
[0049] Here, the first visual-language model (M1) may be a BLIP (Bootstrapping Language-Image Pre-training) model, which is an artificial intelligence model composed of a first image encoder (M1-1), a cue-former (M1-2), and a large language model (M1-3).
[0050] Specifically, when the first caption generation module (111) inputs normalized image data (ID) into the first visual-language model (M1), the normalized image data (ID) passes through the first image encoder (M1-1) and the queue-former (M1-2) and is input as a query into the large language model (M1-3), and the large language model (M1-3) can output the first caption data (CD1).
[0051] Meanwhile, the first visual-language model (M1) is not limited in its type as long as it can output first caption data (CD1), which is text that describes the situation depicted by the image represented by the image data (ID).
[0052] Through this, the first caption generation module (111) can obtain first caption data (CD1) “a man taking a picture of another man sitting on a wall”, which is a text describing the situation depicted by the image (I), from the image (I), as illustrated in FIG. 4.
[0053] Below, the process of generating the second caption data will be described.
[0054] FIG. 5 is a block diagram of a second visual-language model used by a dataset generation device used in an image generation artificial intelligence model according to an embodiment of the present disclosure, and FIG. 6 is a diagram for explaining a process in which a dataset generation device used in an image generation artificial intelligence model according to an embodiment of the present disclosure generates second caption data.
[0055] Referring further to FIGS. 5 and 6, the second caption generation module (112) can generate second caption data (CD2) describing the characteristics of the image.
[0056] This second caption generation module (112) can generate image data (ID), which is a matrix representing an image, and obtain second caption data (CD2) from the image data (ID) using a second visual-language model (M2).
[0057] To this end, the second caption generation module (112) can convert the size and format of the image data (ID) so that the image data (ID) can be input into the second time-language model (M2).
[0058] Specifically, the second caption generation module (112) can convert the size of the image data (ID) as a matrix, interpolate the converted image data (ID), and normalize the interpolated image data (ID).
[0059] At this time, the size conversion, interpolation, and normalization of the image data (ID) may differ depending on the type of input terminal of the second visual-language model (M2).
[0060] The second caption generation module (112) can input image data (ID) as input data to the second visual-language model (M2) and receive image embedding (IE) as output data from the second visual-language model (M2).
[0061] Here, the second visual-language model (M2) may be a CLIP (Contrastive Language-Image Pre-training) model, which is an artificial intelligence model composed of a second image encoder (M2-1) and a text encoder (M2-1).
[0062] Specifically, when the second caption generation module (112) inputs normalized image data (ID) into the second image encoder (M2-1) of the second visual-language model (M2), the second image encoder (M2-1) can output an image embedding (IE).
[0063] Additionally, when the second caption generation module (112) inputs candidate caption data (CCD) into the text encoder (M2-2) of the second visual-language model (M2), the text encoder (M2-2) can output a text embedding (TE).
[0064] Meanwhile, the second visual-language model (M2) is not limited in its type as long as it can output image embeddings (IE) from image data (ID) and text embeddings (TE) from candidate caption data (CCD).
[0065] Here, the candidate caption data (CCD) may be data stored in a feature modifier database (DB2).
[0066] Specifically, the feature modifier database (DB2) can store a plurality of candidate caption data (CCD) that can be selected to be generated as second caption data (CD2) classified by a plurality of type categories according to the type of feature.
[0067] For example, multiple type categories might consist of Medium (DB2-1), Artist (DB2-2), Platform (DB2-3), Movement (DB2-4), and Flavors (DB2-5).
[0068] Among the multiple type categories, the type category Medium (DB2-1) includes multiple candidate caption data (CCD), which is a modifier text that modifies the type of the image. Among the multiple type categories, the type category Artist (DB2-2) includes multiple candidate caption data (CCD), which is a modifier text that modifies a specific artist. Among the multiple type categories, the type category Platform (DB2-3) includes multiple candidate caption data (CCD), which is a modifier text that modifies an online platform on which the image is mainly displayed. Among the multiple type categories, the type category Movement (DB2-4) includes multiple candidate caption data (CCD), which is a modifier text that modifies a period in which the style of the image was popular. Among the multiple type categories, the type category Flavors (DB2-5) may include multiple candidate caption data (CCD), which is a modifier text that modifies the style expressed by the image.
[0069] For example, the type category Medium (DB2-1) may contain candidate caption data (CCD) “film still from dune 2020”, the type category Artist (DB2-2) may contain candidate caption data (CCD) “by Saul Yaffie”, “by Leonardo da Vinci”, the type category Platform (DB2-3) may contain candidate caption data (CCD) “unsplash”, “pixabay”, the type category Movement (DB2-4) may contain candidate caption data (CCD) “2010y”, “2000y”, and the type category Flavors (DB2-5) may contain candidate caption data (CCD) “maluma”, “desert mirage”, “hasselblad camera”, “one is blond having fun in the sun”, “alexa mini”, “sitting on temple stairs”, “5 feet distance from the camera”, “photo for a magazine”, “4k shot”.
[0070] In this feature modifier database (DB2), the plurality of candidate caption data (CCD) described above can be input to the text encoder (M2-2) and the output text embedding (TE) can be further stored.
[0071] The generation of such text embeddings (TE) can be performed in advance before performing the process of generating the second caption data (CD2).
[0072] That is, before performing the process of generating the second caption data (CD2), the feature modifier database (DB2) may store a plurality of candidate caption data (CCD) classified by a plurality of type categories and a plurality of text embeddings (TE) generated from each of the plurality of candidate caption data (CCD).
[0073] The second caption generation module (112) inputs normalized image data (ID) into the second image encoder (M2-1) of the second visual-language model (M2) to output an image embedding (IE), and then can calculate the similarity between the image embedding (IE) and each of all multiple text embeddings (TE) stored in the feature modifier database (DB2).
[0074] Next, the second caption generation module (112) can select one or more text embeddings (TE) from among a plurality of text embeddings (TE) based on similarity, and generate candidate caption data (CCD) corresponding to the selected one or more text embeddings (TE) as second caption data (CD2).
[0075] To this end, the second caption generation module (112) normalizes all of the image embeddings (IE) and all of the multiple text embeddings (TE) stored in the feature modifier database (DB2), and can calculate the similarity of each result value obtained by taking the inner product of each of the multiple text embeddings (TE) stored in the feature modifier database (DB2) and the image embeddings (IE).
[0076] At this time, when selecting one or more text embeddings (TE) from among a plurality of text embeddings (TE) based on similarity, the second caption generation module (112) may select text embeddings (TE) included in the reference ranking in order of high similarity for each of a plurality of type categories, and generate candidate caption data (CCD) corresponding to the selected text embeddings (TE) as second caption data (CD2).
[0077] For example, the second caption generation module (112) may select one text embedding (TE) included in the standard ranking “1st” in order of high similarity in each of the type categories Medium (DB2-1), Artist (DB2-2), Platform (DB2-3), and Movement (DB2-4), select 15 text embeddings (TE) included in the standard ranking “15th” in order of high similarity in each of the type categories Flavors (DB2-5), and generate candidate caption data (CCD) corresponding to each of the 19 selected text embeddings (TE) as second caption data (CD2).
[0078] At this time, the second caption generation module (112) can generate second caption data (CD2) by adding truncation data (e.g., “,”) between candidate caption data (CCD) corresponding to each text embedding (TE).
[0079] Through this, the second caption generation module (112) can obtain second caption data (CD2) “film still from dune 2020, by Saul Yaffie, unsplash, 2010y, maluma, desert mirage, hasselblad camera, one is blond having fun in the sun, alexa mini, sitting on temple stairs, 5 feet distance from the camera, photo for a magazine, 4k shot”, which is a text describing the features of the image (I), from the image (I), as shown in FIG. 6.
[0080] Below, we will explain the process of generating prompt data.
[0081] FIG. 7 is a diagram illustrating a process of generating prompt data by a dataset generation device utilized in an image generation artificial intelligence model according to an embodiment of the present disclosure.
[0082] Referring further to FIG. 7, the prompt generation module (113) can merge the first caption data (CD1) and the second caption data (CD2) to generate prompt data (PD).
[0083] Through this, the prompt generation module (113) can generate prompt data (PD) consisting of text describing the situation depicted by the image and text describing the characteristics of the image.
[0084] Thereafter, the dataset creation module (114) can create an image-text pair dataset by matching image data (ID) and prompt data (PD).
[0085] At this time, the dataset creation module (114) can create an image-text pair dataset by matching a plurality of image data (ID) and each of the prompt data (PD) of the plurality of image data (ID).
[0086] Next, the image-text pair dataset generated from the dataset creation module (114) can be stored in the dataset database (DB1).
[0087] Meanwhile, a processor (110) according to another embodiment may control a communication unit (120) to generate prompt data for a new image through the above-described prompt data generation process when a new image is received from a user device (100) and transmit the generated prompt data to the user device (100).
[0088] When the user device (100) receives prompt data in response to transmission of a new image, it can receive a modification input for the prompt data from the user.
[0089] Here, the modified input may mean an input that modifies, deletes, or adds text describing the situation depicted by the new image included in the prompt data and text describing the characteristics of the new image.
[0090] Thereafter, the user device (100) can modify the prompt data in response to the modification input to generate modified prompt data, and input the modified prompt data into an image generation artificial intelligence model to generate a modified image reflecting the modified prompt data.
[0091] Here, the generative artificial intelligence model may be an artificial intelligence model that receives prompt data consisting of text describing the situation depicted by the image and text describing the characteristics of the image as input and generates an image corresponding to the prompt data, and the type thereof is not limited.
[0092] Meanwhile, when the processor (110) according to another embodiment receives the modified prompt data, it can compare it with the prompt data generated from the new image and check the additional text added to the modified prompt data against the prompt data.
[0093] Thereafter, a processor (110) according to another embodiment may classify additional text into one of a plurality of type categories and store it as candidate caption data (CCD) in a feature modifier database (DB2).
[0094] Through this, the processor (110) according to another embodiment can diversify the candidate caption data (CCD) by adding text related to an image input from a user to the candidate caption data (CCD).
[0095] The memory (130) can store various programs and data required for the operation of the dataset generation device (100) utilized in the image generation artificial intelligence model. The memory (130) can be implemented as a non-volatile memory (130), a volatile memory (130), a flash memory (130), a hard disk drive (HDD), or a solid state drive (SSD).
[0096] The processor (110) can control the overall operation of the dataset generation device (100) utilized in the image generation artificial intelligence model using various programs stored in the memory (130). The processor (110) can be composed of a RAM, a ROM, a graphic processing unit, a main CPU, first to n interfaces, and a bus. At this time, the RAM, ROM, graphic processing unit, main CPU, first to n interfaces, etc. can be connected to each other via a bus.
[0097] RAM stores the O / S and application programs. Specifically, when the dataset generation device (100) utilized in the image generation artificial intelligence model is booted, the O / S is stored in RAM, and various application data selected by the user can be stored in RAM.
[0098] ROM stores a set of commands for system booting, etc. When a turn-on command is input and power is supplied, the main CPU copies the O / S stored in the memory (130) to RAM according to the commands stored in the ROM, executes the O / S, and boots the system. When booting is complete, the main CPU copies various application programs stored in the memory (130) to RAM, and executes the application programs copied to RAM to perform various operations.
[0099] The main CPU accesses the memory (130) and performs booting using the OS stored in the memory (130). Then, the main CPU performs various operations using various programs, contents, data, etc. stored in the memory (130).
[0100] The first to nth interfaces are connected to the various components described above. One of the first to nth interfaces may be a network interface that connects to an external device via a network.
[0101] Meanwhile, furthermore, the processor (110) can control an artificial intelligence model. In this case, it goes without saying that the processor (110) can include a graphics-only processor (e.g., GPU) for controlling the artificial intelligence model.
[0102] Meanwhile, the artificial intelligence model (the first visual-language model (M1) and the second visual-language model (M2)) according to the present invention can use a decoder based on a transformer model that overcomes the long-term dependency limitations of a multi-view encoder and a recurrent neural network to obtain relationship information between features of an input image.
[0103] Specifically, the artificial intelligence model according to the present invention, unlike a general image caption generation model, can extract information from images from various viewpoints by using multiple feature extractors rather than a single feature extractor.
[0104] Additionally, the decoder included in the AI model according to the present invention may utilize a self-correcting transformer that enhances the role of the language model, unlike conventional transformers. Furthermore, the image caption generation model according to some embodiments of the present invention can be optimized through SCST, thereby achieving even higher performance.
[0105] The artificial intelligence model according to the present invention can use the structure of a transformer model, which is a machine translation model trained only using an attention mechanism technique, unlike the traditional method using a recurrent neural network.
[0106] The transformer model included in the artificial intelligence model according to the present invention can be implemented using an encoder and a decoder. For example, the transformer model can be implemented using one or more encoders and one or more decoders.
[0107] The transformer model can be composed of four core networks: (1) position encoding, (2) scale-independent attention mechanism, (3) multi-head attention mechanism, and (4) position-wise forward network.
[0108] This explains positional encoding. The process of mapping natural language into real-valued vectors that a model can understand is called word embedding. In natural language processing using traditional recurrent neural network models, word embedding can be used to map natural language into vectors.
[0109] Transformer models also require word embeddings, but because they are parallel methods that utilize attention mechanisms rather than sequential or convolutional methods, they can struggle to maintain the order of words in the input sentence. Therefore, additional word positional information needs to be provided to the word embeddings. To address this issue, Transformer models can incorporate information about the relative and absolute positions of words into their word embeddings through positional encoding.
[0110] Meanwhile, positional encoding can be omitted. For example, if the input is an image for which positional information is not important, the encoder may not apply positional encoding. In this case, the decoder that uses sentences as input may apply positional encoding.
[0111] Describes the internal attention mechanism of scale.
[0112] The transformer model, which is based on the attention mechanism technique, can be used by modifying the internal attention mechanism technique (Dot-Product Attention).
[0113] The formula used in the internal attention mechanism technique is as follows: mathematical formula (1).
[0114]
[0115] In mathematical equation (1), the given input vectors are Q (Query), K (Key), V (Value), and dk (the size of the Key). The inner product of Query and Key is calculated, and then divided by the square root of the size of the Key to reduce the size of the vector. The Softmax function is used to transform the factors in the vector into a distribution that sums to 1, and the weight for Value is obtained by multiplying it by Value.
[0116] The Transformer model can additionally utilize a scaled dot-product attention mechanism, which scales the vector by the size of the key, in addition to the inner product attention technique. Since the size of the key increases, the size of the inner product result increases, and thus the gradient value decreases rapidly when the softmax function is applied. Therefore, it is necessary to scale the vector before applying the softmax function.
[0117] Describes the multi-head attention mechanism.
[0118] A typical attention mechanism technique obtains the weight of Value by calculating the above mathematical expression (1) once.
[0119] However, the transformer model according to the present invention can use a methodology to obtain multiple different phenotypes from a single data and then combine them to obtain multi-angle information, rather than completing the calculation in a single operation.
[0120] Specifically, in the multi-head attention mechanism technique, h scale inner product attention mechanisms can be calculated by linearly projecting Query, Key, and Value h times, as in mathematical equations (2) and (3) below.
[0121]
[0122] Mathematical expression (2) is the formula for calculating each head in the multi-head attention mechanism technique. In mathematical expression (2), each head (i.e., headi) can be calculated by applying the scaled inner product attention mechanism based on the query, key, and value.
[0123] Mathematical expression (3) represents a formula for obtaining a final value by combining all heads. In mathematical expression (3), the final value of the multi-head attention mechanism technique can be calculated by combining the h heads calculated in mathematical expression (2).
[0124] In equations (2) and (3), W is a learnable parameter matrix.
[0125] Through mathematical equations (2) and (3), various attention mechanism information about words can be obtained.
[0126] Describes position-wise forward networks.
[0127] The layers of the transformer model according to the present invention may include independent position-wise feed-forward networks. The position-wise feed-forward network can be calculated using two linear transformations and the ReLU activation function, as shown in the following mathematical equation (4).
[0128]
[0129] In formula (4), W1 and W2 are as follows.
[0130]
[0131] Meanwhile, the artificial intelligence model (the first visual-language model (M1) and the second visual-language model (M2)) according to the present invention may be a model based on supervised learning or unsupervised learning. Furthermore, the artificial intelligence model according to the present invention may include a support vector machine (SVM), a decision tree, a neural network, etc., and a method using these.
[0132] As an example, the artificial intelligence model according to the present invention may be an artificial intelligence model based on a convolutional deep neural network (CNN) trained by inputting training data. However, the present invention is not limited thereto, and it goes without saying that various artificial intelligence models can be applied to the present invention. For example, models such as a deep neural network (DNN), a recurrent neural network (RNN), and a bidirectional recurrent deep neural network (BRDNN) can be used as the artificial intelligence model, but are not limited thereto.
[0133] Here, convolutional deep neural networks (CNNs) are a type of multilayer perceptron designed to utilize minimal preprocessing. A CNN consists of one or more convolutional layers and regular artificial neural network layers layered on top of them, with additional weight and pooling layers. This structure allows a CNN to fully utilize two-dimensional input data. Furthermore, a CNN can be trained using standard backpropagation. Compared to other feedforward artificial neural network techniques, a CNN is easier to train and uses fewer parameters.
[0134] Additionally, deep neural networks (DNN) are artificial neural networks (ANN) that consist of multiple hidden layers between the input layer and the output layer.
[0135] At this time, the structure of a deep neural network can be composed of perceptrons. A perceptron consists of multiple inputs, a processor, and a single output. The processor multiplies each input by a weight and then sums all the weighted inputs. The processor then inputs the summed value into an activation function to produce a single output. If a specific value is desired as the output of the activation function, the weights multiplied by each input can be modified and the output can be recalculated using the modified weights. Each perceptron can use a different activation function. Furthermore, each perceptron receives the outputs from the previous layer as input and uses the activation function to derive the output. The derived output is then passed on to the input of the next layer. Through the above process, several output values can be ultimately obtained.
[0136] A recurrent neural network (RNN) is a neural network in which the connections between the units forming the artificial neural network form directed cycles. Unlike feed-forward neural networks, RNNs can utilize internal memory to process arbitrary inputs.
[0137] Deep Belief Networks (DBNs) are a generative graphical model used in machine learning. In deep learning, they refer to deep neural networks comprised of multiple layers of latent variables. They are characterized by connections between layers, but no connections between units within a layer.
[0138] Deep belief neural networks, due to their generative nature, can be used for pre-training. After learning initial weights through pre-training, they can be fine-tuned using backpropagation or other discriminative algorithms. This characteristic is extremely useful when training data is limited, as the influence of initial weight values on the resulting model becomes stronger with less training data. Pre-trained initial weight values are closer to the optimal weights than randomly set initial weight values, which enables improved performance and speed in the fine-tuning stage.
[0139] The above-described artificial intelligence and its learning methods are provided for illustrative purposes only, and the artificial intelligence and its learning methods utilized in the above-described embodiments are not limited. For example, any type of artificial intelligence technology and its learning methods applicable to the same task by a person skilled in the art can be utilized to implement the system according to the disclosed embodiments.
[0140] Meanwhile, the processor (110) may include one or more cores (not shown) and a graphics processing unit (not shown) and / or a connection path (e.g., a bus) for transmitting and receiving signals with other components.
[0141] In one embodiment, a processor (110) performs a method described in connection with the present invention by executing one or more instructions stored in a memory (130).
[0142] For example, the processor (110) may acquire new learning data by executing one or more instructions stored in the memory (130), perform a test on the acquired new learning data using a learned model, extract first learning data for which labeled information is acquired with an accuracy higher than a predetermined first reference value as a result of the test, delete the extracted first learning data from the new learning data, and retrain the learned model using the new learning data from which the extracted learning data has been deleted.
[0143] Meanwhile, the processor (110) may further include a RAM (Random Access Memory, not shown) and a ROM (Read-Only Memory, not shown) that temporarily and / or permanently store signals (or data) processed within the processor (110). In addition, the processor (110) may be implemented in the form of a system on chip (SoC) that includes at least one of a graphics processing unit, RAM, and ROM.
[0144] The memory (130) can store programs (one or more instructions) for processing and controlling the processor (110). The programs stored in the memory (130) can be divided into multiple modules according to function.
[0145] The communication unit (120) can perform communication. In particular, the communication unit (120) can include various communication chips such as a Wi-Fi chip, a Bluetooth chip, a wireless communication chip, an NFC chip, a low-power Bluetooth chip (BLE chip), etc. At this time, the Wi-Fi chip, the Bluetooth chip, and the NFC chip perform communication in the LAN method, the Wi-Fi method, the Bluetooth method, and the NFC method, respectively. When using a Wi-Fi chip or a Bluetooth chip, various connection information such as an SSID and a session key are first transmitted and received, and after establishing a communication connection using this, various information can be transmitted and received. The wireless communication chip refers to a chip that performs communication according to various communication standards such as IEEE, Zigbee, 3G (3rd Generation), 3GPP (3rd Generation Partnership Project), LTE (Long Term Evolution), 5G (5th Generation), etc.
[0146]
[0147] The steps of a method or algorithm described in connection with an embodiment of the present invention may be implemented directly in hardware, implemented as a software module executed by hardware, or implemented by a combination thereof. The software module may reside in a random access memory (RAM), a read only memory (ROM), an erasable programmable ROM (EPROM), an electrically erasable programmable ROM (EEPROM), a flash memory, a hard disk, a removable disk, a CD-ROM, or any other form of computer-readable recording medium well known in the art to which the present invention pertains.
[0148] Additionally, different embodiments of the present invention may be complementary or combined with each other.
[0149] The components of the present invention may be implemented as a program (or application) and stored on a medium to be executed in conjunction with a computer, which is hardware. The components of the present invention may be implemented as software programs or software elements. Similarly, the embodiments may be implemented in a programming or scripting language such as C, C++, Java, assembler, Python, etc., including various algorithms implemented as a combination of data structures, processes, routines, or other programming components. Functional aspects may be implemented as algorithms executed on one or more processors.
[0150] While the embodiments of the present invention have been described with reference to the attached drawings, those skilled in the art will appreciate that the present invention can be implemented in other specific forms without altering the technical concept or essential features thereof. Therefore, the embodiments described above should be understood to be illustrative in all respects and not restrictive.
Claims
1. A first caption generation module that generates first caption data describing a situation depicted by an image; A second caption generation module for generating second caption data describing the features of the above image; A prompt generation module that generates prompt data by merging the first caption data and the second caption data; and A processor including a dataset generation module that generates an image-text pair dataset by matching image data representing the image and the prompt data; and A dataset database storing an image-text pair dataset; characterized by including: A data set generation device used in image generation artificial intelligence models.
2. In paragraph 1, The above first caption generation module The above image data is input as input data to the first visual-language model, and the first caption data is output as output data from the first visual-language model. A data set generation device used in image generation artificial intelligence models.
3. In paragraph 1, The above second caption generation module The above image data is input as input data to the second visual-language model, and the image embedding is output as output data from the second visual-language model. A data set generation device used in image generation artificial intelligence models.
4. In paragraph 3, A feature modifier database in which a plurality of candidate caption data that can be selected to be generated as the second caption data is stored classified by a plurality of type categories according to the type of the feature; The above second caption generation module Each of the plurality of candidate caption data is input as input data to the second visual-language model, and text embeddings for each of the plurality of candidate caption data are output as output data from the second visual-language model. A data set generation device used in image generation artificial intelligence models.
5. In paragraph 4, The above second caption generation module It is characterized in that the similarity with the image embedding is calculated for each of the plurality of text embeddings, one or more text embeddings are selected from the plurality of text embeddings based on the similarity, and candidate caption data corresponding to the one or more selected text embeddings are generated as the second caption data. A data set generation device used in image generation artificial intelligence models.
6. In paragraph 5, The above second caption generation module It is characterized in that text embeddings included in the criterion ranking are selected in the order of high similarity for each of the above multiple type categories, and candidate caption data corresponding to the selected text embeddings are generated as the second caption data. A data set generation device used in image generation artificial intelligence models.
Citation Information
Patent Citations
System and method for creating caption for image and computer program for the same
KR101996371B1
Bed to prevent patient from pushing
KR1020210113114A
Regenerable electret filter, electret filter regenration system and air cleaner using the same
KR1020220025430A
Apparatus for generating data set used in image generation artificial intelligence model
KR102692052B1
KR20210130980A