Information Processing Apparatus and Information Processing Program

The information processing apparatus and program enhance image caption generation by reinforcing the model with a reward value based on CIDEr score and sentence length, addressing the issue of short captions and maintaining information richness in generated content.

JP7705126B1Active Publication Date: 2025-07-09SOFTBANK CORPORATION +1
View PDF 6 Cites 0 Cited by

Patent Information

Application Number
JP2024033342
Authority / Receiving Office
JP · JP
Patent Type
Patents
Current Assignee / Owner
Filing Date
2024-03-05
Publication Date
2025-07-09
Estimated Expiration
2044-03-05

AI Technical Summary

Technical Problem

Existing image caption generation techniques using machine learning models often result in captions that are too short, leading to a decrease in the amount of information included while maintaining accurate content.

Method used

An information processing apparatus and program that includes an acquisition unit for generating captions using a machine learning model, a calculation unit for evaluating caption accuracy and length, and a model generation unit that reinforces the model with a reward value based on both CIDEr score and sentence length to prevent short captions.

Benefits of technology

Prevents a decrease in the amount of information included in image captions while ensuring accurate content by maximizing the CIDEr score and maintaining an appropriate caption length.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure 0007705126000001_ABST
    Figure 0007705126000001_ABST
Patent Text Reader

Abstract

It is possible to generate an image caption with accurate content while preventing a decrease in the amount of information contained in the image caption. 【Solution means】 The information processing apparatus according to the present application includes an acquisition unit that acquires a target caption, which is a sentence describing the content of the input target image, using a machine learning model, a calculation unit that calculates an evaluation value for evaluating the accuracy of the target caption and a sentence length value indicating the length of the target caption, and a model generation unit that generates a machine learning model that is reinforced by setting a reward value based on the evaluation value and the sentence length value.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to an information processing apparatus and an information processing program.

Background Art

[0002] Conventionally, a technique for generating an image caption, which is a sentence describing the content of an image from the image, is known. For example, a technique for generating an image caption from an image using a machine learning model for generating an image caption (hereinafter, may be referred to as a "caption generation model") is known.

[0003] Also, a CIDEr (Consensus-based Image Description Evaluation) score is known as an evaluation value for evaluating the accuracy of an image caption generated by a caption generation model. Further, as a learning method for improving the CIDEr score, SCST (Self-critical Sequence Training for Image Captioning), which is a type of reinforcement learning, is known.

Prior Art Documents

Non-Patent Documents

[0004]

Non-Patent Document 1

Summary of the Invention

Problems to be Solved by the Invention

[0005] However, in the above prior art, the length of the image caption tends to be short. Therefore, in the above prior art, it is not always possible to prevent a decrease in the amount of information included in the image caption while generating an image caption with accurate content.

[0006] An object of the present application is to provide an information processing apparatus and an information processing program capable of preventing a decrease in the amount of information included in an image caption while generating an image caption with accurate content.

Means for Solving the Problems

[0007] The information processing apparatus according to the present application includes an acquisition unit that acquires a target caption, which is a sentence explaining the content of the input target image, using a machine learning model, a calculation unit that calculates an evaluation value for evaluating the accuracy of the target caption and a sentence length value indicating the length of the target caption, and a model generation unit that generates the machine learning model that is reinforced by setting a reward value based on the evaluation value and the sentence length value.

Effects of the Invention

[0008] According to one aspect of the embodiment, it is possible to prevent a decrease in the amount of information included in the image caption while generating an image caption with accurate content.

Brief Description of the Drawings

[0009]

Figure 1

Figure 2

Figure 3

Figure 4

Figure 5

Figure 6

Figure 7

Mode for Carrying Out the Invention

[0010] Hereinafter, embodiments for implementing the information processing apparatus and the information processing program according to the present application (hereinafter referred to as "embodiments") will be described in detail with reference to the drawings. Note that the information processing apparatus and the information processing program according to the present application are not limited by this embodiment. Also, in the following embodiments, the same parts are denoted by the same reference numerals, and redundant descriptions are omitted.

[0011] (Embodiment) [1. Introduction] Conventionally, a technique related to a machine learning model (hereinafter, may be referred to as a "caption generation model") for generating an image caption, which is a sentence describing the content of an image from an image, has been known. In recent years, as a caption generation model, a machine learning model called a visual language model (VLM) has become mainstream. Hereinafter, the case where the caption generation model is a visual language model will be described.

[0012] A vision-language model is a machine learning model that has been pre-trained based on training data containing pairs of images and corresponding image captions. For example, a vision-language model is a machine learning model trained to output an image caption corresponding to an input image when the image included in the training data is input. For example, a vision-language model is trained to predict the correct next token from an image and a given (partial) token. For example, a vision-language model is trained to estimate the next token from a sequence of tokens being generated. For example, the vision-language model may be CoCa (Contrastive Captioners are Image-Text Foundation Models), BLIP (Bootstrapping Language-Image Pre-training), BLIP2, GIT (Generative Image to Text Transformer), etc.

[0013] Also, the CIDEr (Consensus-based Image Description Evaluation) score is known as an evaluation value for evaluating the accuracy of an image caption generated by a vision-language model. Also, SCST (Self-critical Sequence Training for Image Captioning) is known as a learning method for improving the CIDEr score. SCST is a method of reinforcement learning for a vision-language model using the CIDEr score as a reward value. Specifically, there are two types of methods for a vision-language model to generate an image caption: a greedy method and a search method. A general vision-language model can switch between the greedy method and the search method to generate an image caption when generating an image caption. In SCST, the vision-language model is reinforced by setting, as a reward value, the value obtained by subtracting the CIDEr score of the image caption generated by the greedy method from the CIDEr score of the image caption generated by the search method. The difference between the greedy method and the search method will be described in detail with reference to FIGS. 1 and 2.

[0014] FIG. 1 is a diagram for explaining the generation process of an image caption by a greedy method. In FIG. 1, a vision-language model 1 generates an image caption 31 corresponding to an image 20 from the image 20 by the greedy method. The greedy method is a method in which when the vision-language model 1 outputs the next token of each token, it selects and outputs the token with the highest probability of appearance (also referred to as the appearance probability) from among a plurality of candidate next tokens.

[0015] In FIG. 1, first, the vision-language model 1 outputs a token 311. Subsequently, when the vision-language model 1 outputs the next token of the token 311, it outputs a token 312 with the highest appearance probability among a plurality of candidate next tokens. Subsequently, when the vision-language model 1 outputs the next token of the token 312, it outputs a token 313 with the highest appearance probability among a plurality of candidate next tokens. In this way, the vision-language model 1 generates an image caption 31 by the greedy method by selecting and outputting the token with the highest probability of appearance from among a plurality of candidate next tokens when outputting the next token of each token.

[0016] Also, in FIG. 1, the CIDEr score of the image caption 31 by the greedy method is indicated by a CIDEr score (greedy method) 41. As explained in FIG. 1, the characteristic of the greedy method is that the vision-language model 1 continuously selects the token with the highest appearance probability. In other words, the characteristic of the greedy method is that the vision-language model 1 continuously outputs only the token with the highest appearance probability. Also, it can be said that the image caption generated by the vision-language model 1 continuously selecting the token with the highest appearance probability is the image caption generated as a result of the vision-language model 1 continuously making the best choice at each point in time. Therefore, the image caption generated by the greedy method serves as a reference indicating the current strength of the vision-language model 1. For example, the CIDEr score of the image caption 31 by the greedy method (referred to as "CIDEr score (greedy method) 41" in FIG. 1) serves as a reference score indicating the current strength of the vision-language model 1.

[0017] FIG. 2 is a diagram for explaining the generation process of an image caption by a search method. In FIG. 2, a visual language model 1 generates an image caption 32 corresponding to an image 20 from the image 20 by a search method. The search method is a method in which when the visual language model 1 outputs the next token of each token, two or more tokens are selected from among the candidates of the next tokens in descending order of the probability of appearance next, and one token selected according to a predetermined selection rule is output from among the two or more selected tokens. For example, the predetermined selection rule may be a rule of selecting one token according to some probabilistic rule from among the two or more selected tokens. For example, the predetermined selection rule may be a rule of randomly selecting one token from among the two or more selected tokens.

[0018] In FIG. 2, first, the visual language model 1 selects two tokens 321 and 322 from among a token 321 with a 30% probability of appearance next, a token 322 with a 50% probability of appearance next, and a token 323 with a 20% probability of appearance next, in descending order of the probability of appearance next. Subsequently, the visual language model 1 outputs one token 321 randomly selected from among the two selected tokens 321 and 322. Subsequently, the visual language model 1 selects two tokens 324 and 325 from among a token 324 with a 40% probability of appearance next, a token 325 with a 50% probability of appearance next, and a token 326 with a 10% probability of appearance next, in descending order of the probability of appearance next. Subsequently, the visual language model 1 outputs one token 325 randomly selected from among the two selected tokens 324 and 325. In this way, when the visual language model 1 outputs the next token of each token, two or more tokens are selected from among the candidates of the next tokens in descending order of the probability of appearance next, and one token selected according to a predetermined selection rule is output from among the two or more selected tokens, thereby generating an image caption 32 by a search method.

[0019] Also, in FIG. 2, the CIDEr score of the image caption 32 by the search method is indicated by the CIDEr score (search method) 42. As described with reference to FIG. 2, the search method is different from the greedy method in that the visual language model 1 does not always continue to output only the tokens with the highest appearance probability. In other words, the search method outputs tokens probabilistically selected from among the tokens having an appearance probability equal to or higher than a certain level by the visual language model 1, and thus may generate an image caption of a quality exceeding the current strength of the visual language model 1. That is, the image caption generated by the search method may exceed the current strength of the visual language model 1. For example, the CIDEr score of the image caption 32 by the search method (referred to as "CIDEr score (search method) 42" in FIG. 2) may exceed the reference score (corresponding to the "CIDEr score (greedy method) 41" described in FIG. 1) indicating the current strength of the visual language model 1.

[0020] Therefore, in SCST, the visual language model 1 is reinforced to improve the quality of the image caption generated by the search method. That is, in SCST, the visual language model 1 is reinforced to improve the CIDEr score of the image caption generated by the search method. Specifically, in SCST, the visual language model 1 is reinforced by setting the CIDEr score of the image caption generated by the search method as the reward value. More specifically, in SCST, the visual language model 1 is reinforced by setting, as the reward value, the value obtained by subtracting the CIDEr score of the image caption generated by the greedy method (the "CIDEr score (greedy method)") from the CIDEr score of the image caption generated by the search method (the "CIDEr score (search method)"). The reward value of SCST is represented by the following formula (1).

[0021]

Equation

[0022] Also, for each of the plurality of image captions generated by the search method, the loss function of SCST is represented by an expression obtained by taking the logarithm of the appearance probability of each token included in the image caption, multiplying the logarithms of the appearance probabilities of the respective tokens, and multiplying the result by the reward value represented by the above formula (1). The loss function of SCST is represented by the following formula (2). Note that Σ in the following formula (2) indicates taking the sum for each of the plurality of image captions.

[0023] [Number]

[0024] In SCST, the visual language model 1 is reinforced to minimize the value of the loss function represented by the above formula (2). In SCST, the values of the parameters of the visual language model 1 are learned to minimize the value of the loss function represented by the above formula (2).

[0025] Also, when SCST is applied to the visual language model 1 using the CIDEr score as the reward value, the length of the image caption (hereinafter sometimes referred to as the "generated caption") generated by the learned visual language model 1 tends to be short. This is because the CIDEr score takes a higher value as the similarity between the correct caption of the image and the generated caption is higher. Specifically, as the length of the generated caption increases, the possibility of including incorrect information not included in the correct caption of the image also increases, so the CIDEr score is also likely to decrease. To avoid this, it is considered that the visual language model 1 generates an image caption with a shorter length. However, the fact that the length of the generated caption is short corresponds to the fact that the amount of information included in the generated caption is small. Also, the fact that the amount of information included in the generated caption is small corresponds to the fact that the information that should originally be included in the generated caption is not included. Therefore, a technique for preventing the length of the generated caption from becoming short is desired. That is, a technique for preventing a decrease in the amount of information included in the generated caption is desired.

[0026] In contrast, the information processing apparatus according to the embodiment generates a vision-language model that is reinforced by setting a reward value based on the CIDEr score and the sentence length value indicating the length of the generated caption. FIG. 3 is a diagram for explaining a method of learning the vision-language model according to the embodiment. In FIG. 3, it is different from FIG. 2 in that the sentence length value indicating the length of the generated caption is set as the reward value of reinforcement learning. The reward value according to the embodiment is represented by the following formula (3). The following formula (3) is obtained by adding a term corresponding to the sentence length value indicating the length of the generated caption to the above formula (1).

[0027] [Number]

[0028] In this way, the information processing apparatus according to the embodiment sets a reward value based on the CIDEr score and the sentence length value indicating the length of the generated caption to reinforce the vision-language model, thereby maximizing the CIDEr score and preventing the length of the generated caption from becoming short. As a result, the information processing apparatus can generate an image caption with accurate content and a length that is not extremely short. Therefore, the information processing apparatus can prevent a decrease in the amount of information included in the image caption while generating an image caption with accurate content.

[0029] [2. Configuration of Information Processing Apparatus] With reference to FIG. 4, a configuration example of the information processing apparatus 100 according to the embodiment will be described. FIG. 4 is a diagram showing a configuration example of the information processing apparatus 100 according to the embodiment. The information processing apparatus 100 includes a communication unit 110, a storage unit 120, and a control unit 130.

[0030] (Communication Unit 110) The communication unit 110 is realized by a NIC (Network Interface Card), an antenna, or the like. The communication unit 110 is connected to various networks, either wired or wirelessly, and performs information transmission and reception, for example, with other information processing devices other than the information processing device 100.

[0031] (Memory unit 120) The memory unit 120 is realized by, for example, semiconductor memory elements such as RAM (Random Access Memory) and flash memory, or storage devices such as hard disks and optical disks. Specifically, the memory unit 120 stores various data. For example, the memory unit 120 stores various programs. For example, the memory unit 120 stores the information processing program according to the embodiment. Also, the memory unit 120 may store information regarding a machine learning model that generates an image caption, which is a sentence explaining the content of an image from the image. For example, the memory unit 120 may store information regarding a vision-language model, which is a machine learning model that generates an image caption from an image. Also, the memory unit 120 may store information regarding the target image acquired by the acquisition unit 131. Also, the memory unit 120 may store information regarding the machine learning model generated by the model generation unit 133. For example, the memory unit 120 may store information regarding the vision-language model generated by the model generation unit 133.

[0032] (Control unit 130) The control unit 130 is a controller and is realized, for example, when various programs stored in the storage device inside the information processing device 100 are executed with the RAM as a work area by a CPU (Central Processing Unit), an MPU (Micro Processing Unit), or the like. Also, the control unit 130 is a controller and is realized by an integrated circuit such as an ASIC (Application Specific Integrated Circuit) or an FPGA (Field Programmable Gate Array).

[0033] The control unit 130 may have an acquisition unit 131, a calculation unit 132, a model generation unit 133, and a sentence generation unit 134 as functional units, and may realize or execute the operations of information processing described below. Note that the internal configuration of the control unit 130 is not limited to the configuration shown in FIG. 4, and may be any other configuration as long as it performs the information processing described later. Also, each functional unit represents the function of the control unit 130, and does not necessarily have to be physically distinguished.

[0034] (Acquisition Unit 131) The acquisition unit 131 may acquire a target image that is a target for generating an image caption. Further, the target image may include not only still images but also moving images. Specifically, the acquisition unit 131 may acquire the target image from an external information processing device via the communication unit 110. Also, when the acquisition unit 131 acquires the target image, it may store information regarding the target image in the storage unit 120.

[0035] Further, the acquisition unit 131 may acquire a machine learning model that generates an image caption, which is a sentence describing the content of the image, from the image. Specifically, the acquisition unit 131 may acquire a vision-language model, which is a machine learning model that generates an image caption, which is a sentence describing the content of the image, from the image. More specifically, the acquisition unit 131 may refer to the storage unit 120 and acquire the vision-language model. For example, when acquiring a target image, the acquisition unit 131 may acquire the vision-language model. For example, the acquisition unit 131 may acquire a machine learning model that has been pre-trained based on training data including pairs of images and the image captions corresponding to the images. For example, when an image included in the training data is input, the acquisition unit 131 may acquire a machine learning model that has been trained to output the image caption corresponding to the image. For example, the acquisition unit 131 may acquire a machine learning model that has been trained to predict the correct next token from an image and a given (partial) token. For example, the acquisition unit 131 may acquire a machine learning model that has been trained to estimate the next token from a sequence of tokens being generated. For example, the acquisition unit 131 may acquire a machine learning model that has been trained to estimate and output the next token from the input image and token sequence. For example, the acquisition unit 131 may acquire a vision-language model such as CoCa, BLIP, BLIP2, GIT, etc.

[0036] In addition, the acquisition unit 131 acquires a target caption, which is a sentence describing the content of the input target image, using a machine learning model. Specifically, the acquisition unit 131 generates a target caption from the target image using a machine learning model that generates an image caption, which is a sentence describing the content of the image, from the image. For example, the acquisition unit 131 may generate a target caption, which is an image caption corresponding to the target image, using a vision-language model. For example, when the acquisition unit 131 acquires the target image and the vision-language model, the acquisition unit 131 may generate a target caption, which is an image caption corresponding to the target image, using the vision-language model. More specifically, the acquisition unit 131 may input the target image into a machine learning model that generates an image caption, which is a sentence describing the content of the image, from the image, and cause the machine learning model to generate the target caption. The acquisition unit 131 may generate the target caption by causing the machine learning model to generate the target caption. Note that hereinafter, the "target caption" may be described as the "generated caption" in some cases. For example, the acquisition unit 131 may generate an image caption corresponding to the target image by inputting the target image into the vision-language model. For example, the acquisition unit 131 may acquire a generated caption, which is an image caption output from the vision-language model, by inputting the target image into the vision-language model. In this way, the acquisition unit 131 may acquire a generated caption, which is an image caption generated by the vision-language model, which is a machine learning model that generates an image caption, which is a sentence describing the content of the image, from the image. Further, when the acquisition unit 131 acquires the generated caption, the acquisition unit 131 may output the acquired generated caption to the calculation unit 132.

[0037] For example, the acquisition unit 131 may acquire a greedy caption, which is a generated caption generated by the greedy algorithm. Specifically, the acquisition unit 131 may acquire the greedy caption output from the vision-language model by inputting the target image into the vision-language model in a state where the vision-language model is switched to the greedy algorithm mode. For example, when the vision-language model outputs the next token of each token, the acquisition unit 131 may select and output the token with the highest probability of appearing next from among a plurality of candidate next tokens, thereby generating a greedy caption. Also, the acquisition unit 131 may acquire the generated greedy caption. Further, when the acquisition unit 131 acquires the greedy caption, it may output the acquired greedy caption to the calculation unit 132.

[0038] Also, the acquisition unit 131 may acquire a search caption, which is a target caption generated by the search algorithm. The acquisition unit 131 may acquire a search caption, which is a generated caption generated by the search algorithm. Specifically, the acquisition unit 131 may acquire the search caption output from the vision-language model by inputting the target image into the vision-language model in a state where the vision-language model is switched to the search algorithm mode. For example, when the vision-language model outputs the next token of each token, the acquisition unit 131 may select two or more tokens from among a plurality of candidate next tokens in descending order of the probability of appearing next, and output one token selected according to a predetermined selection rule from among the selected two or more tokens, thereby generating a search caption. Also, the acquisition unit 131 may acquire the generated search caption. Further, when the acquisition unit 131 acquires the search caption, it may output the acquired search caption to the calculation unit 132.

[0039] (Calculation unit 132) The calculation unit 132 may acquire the target caption from the acquisition unit 131. The calculation unit 132 may acquire the generated caption from the acquisition unit 131. Further, when the calculation unit 132 acquires the generated caption from the acquisition unit 131, it may calculate an evaluation value for evaluating the accuracy of the acquired generated caption. In other words, the evaluation value for evaluating the accuracy of the generated caption is the evaluation value for evaluating the quality of the generated caption. Generally, the evaluation value for evaluating the generated caption is calculated based on how similar the generated caption is to the correct caption. Here, the correct caption is a text prepared in advance as the correct answer for the generated caption. The correct caption is a text that describes the content of the target image and is the correct text prepared in advance as the correct answer for the generated caption. In other words, the more similar the generated caption is to the correct caption, the more accurate the generated caption is. Also, the more similar the generated caption is to the correct caption, the higher the quality of the generated caption.

[0040] Further, the calculation unit 132 may calculate an evaluation value based on the similarity between the correct caption and the target caption. The calculation unit 132 may calculate an evaluation value based on the similarity between the correct caption and the generated caption. For example, the calculation unit 132 may calculate an evaluation value that takes a larger value as the similarity between the correct caption and the generated caption is higher. Also, a high degree of coincidence of the n-gram (combination of n consecutive tokens) between the correct caption and the generated caption corresponds to the fact that the correct caption and the generated caption are similar. For example, the calculation unit 132 may calculate an evaluation value based on the degree of coincidence of the n-gram between the correct caption and the generated caption. For example, the calculation unit 132 may calculate an evaluation value that takes a larger value as the degree of coincidence of the n-gram between the correct caption and the generated caption is higher.

[0041] Further, the calculation unit 132 may calculate an evaluation value that is the CIDEr (Consensus-based Image Description Evaluation) score of the target caption. The calculation unit 132 may calculate an evaluation value that is the CIDEr score of the generated caption. CIDEr is an evaluation index used for quality evaluation of the generated caption. Also, the evaluation value corresponding to CIDEr is called the CIDEr score. The CIDEr score takes a larger value as the degree of coincidence of n-grams between the correct caption of the image and the generated caption is higher. Also, the CIDEr score evaluates the degree of coincidence of n-grams between the correct caption of the image and the generated caption in consideration of TF-IDF, which is an evaluation index for evaluating the rarity of words. Specifically, the CIDEr score calculates the weights of TF-IDF so as to reduce the weights for n-grams that appear in many images and increase the weights for n-grams that appear only in specific images.

[0042] Further, the calculation unit 132 may acquire the greedy caption from the acquisition unit 131. Also, when the calculation unit 132 acquires the greedy caption from the acquisition unit 131, the calculation unit 132 may calculate a greedy evaluation value for evaluating the accuracy of the acquired greedy caption. For example, the calculation unit 132 may calculate a greedy evaluation value that is the CIDEr score of the greedy caption. Also, when the calculation unit 132 calculates the greedy evaluation value, the calculation unit 132 may output the greedy evaluation value to the model generation unit 133.

[0043] Further, the calculation unit 132 may acquire the search caption from the acquisition unit 131. Also, when the calculation unit 132 acquires the search caption from the acquisition unit 131, the calculation unit 132 may calculate a search evaluation value for evaluating the accuracy of the acquired search caption. For example, the calculation unit 132 may calculate a search evaluation value that is the CIDEr score of the search caption. Also, when the calculation unit 132 calculates the search evaluation value, the calculation unit 132 may output the search evaluation value to the model generation unit 133.

[0044] Further, when the calculation unit 132 acquires the generated caption from the acquisition unit 131, it may calculate a sentence length value indicating the length of the acquired generated caption. FIG. 5 is a diagram for explaining a method of calculating the sentence length value according to the embodiment. FIG. 5 shows a sequence of tokens included in the generated caption acquired by the calculation unit 132 from the acquisition unit 131. In FIG. 5, the generated caption is the start token <bos>(BOS is the abbreviation of "Beggining of Sequence"), the first token A1, the second token A2, ···, the (N-1)th (N is a natural number) token A(N-1) (not shown in the figure), and the end token <eos>(EOS is an abbreviation for "End of Sequence") and is composed of a sequence of tokens arranged in order from left to right.

[0045] Also, the calculation unit 132 may assign an index to each token included in the generated caption. For example, the calculation unit 132 may assign, as an index, a number that increases by 1 in the order in which the tokens are arranged. In FIG. 5, the calculation unit 132 uses the number 0 as the start token <bos>is assigned as the index. Further, the calculation unit 132 assigns the number 1 as the index of the first token A1 following the start token. Further, the calculation unit 132 assigns the number 2 as the index of the second token A2 following the first token A1. Similarly, the calculation unit 132 assigns the number (N - 1) as the index of the (N - 1)th token. Further, the calculation unit 132 assigns the number N as the end token <eos>It is assigned as an index of

[0046] Further, the calculation unit 132 may calculate a sentence length value based on the length of the sequence of tokens included in the target caption. The calculation unit 132 may calculate a sentence length value based on the length of the sequence of tokens included in the generated caption. The calculation unit 132 may calculate a sentence length value based on the length of the sequence of tokens included in the generated caption based on the index of each token included in the generated caption. For example, the calculation unit 132 may calculate a sentence length value based on the length of the sequence of tokens excluding the start token and the end token from the sequence of tokens included in the generated caption. For example, the calculation unit 132 may calculate a value obtained by subtracting the index of the start token from the index of the end token. Subsequently, the calculation unit 132 may calculate, as the sentence length value of the generated caption, a value obtained by further subtracting 1 from the value obtained by subtracting the index of the start token from the index of the end token. In FIG. 5, the calculation unit 132 may calculate N, which is a value obtained by subtracting 0, which is the index of the start token, from N, which is the index of the end token. Subsequently, the calculation unit 132 may calculate (N−1), which is a value obtained by subtracting 1 from N, as the sentence length value. In this way, the calculation unit 132 may calculate a sentence length value based on the length of the sequence of tokens included in the generated caption.

[0047] Further, when the calculation unit 132 calculates a sentence length value, the calculated sentence length value may be output to the model generation unit 133. For example, when the acquisition unit 131 acquires the search caption, the calculation unit 132 may calculate a search sentence length value, which is the sentence length value of the acquired search caption. Further, when the calculation unit 132 calculates the search sentence length value, the calculated search sentence length value may be output to the model generation unit 133.

[0048] As described above, the calculation unit 132 calculates an evaluation value for evaluating the accuracy of the target caption and a sentence length value indicating the length of the target caption. For example, the calculation unit 132 calculates a search evaluation value, which is the evaluation value of the search caption, and a search sentence length value, which is the sentence length value of the search caption.

[0049] (Model generation unit 133) The model generation unit 133 generates a machine learning model that is reinforced by setting a reward value based on the evaluation value and the sentence length value calculated by the calculation unit 132. For example, the model generation unit 133 generates a vision-language model that is reinforced by setting a reward value based on the evaluation value and the sentence length value calculated by the calculation unit 132. For example, the model generation unit 133 generates a machine learning model that is reinforced by setting a reward value based on the search evaluation value and the search sentence length value calculated by the calculation unit 132. For example, the model generation unit 133 generates a vision-language model that is reinforced by setting a reward value based on the search evaluation value and the search sentence length value calculated by the calculation unit 132. For example, the model generation unit 133 obtains the greedy evaluation value, the search evaluation value, and the search sentence length value from the calculation unit 132. When the model generation unit 133 obtains the greedy evaluation value, the search evaluation value, and the search sentence length value from the calculation unit 132, the model generation unit 133 may set, as the reward value for the reinforcement learning of the vision-language model, the value obtained by adding the value obtained by subtracting the greedy evaluation value from the search evaluation value and the search sentence length value. For example, the model generation unit 133 may set, as the reward value for the reinforcement learning of the vision-language model, the value obtained by adding the value obtained by subtracting the CIDEr score of the greedy caption (hereinafter referred to as the "CIDEr score (greedy method)") from the CIDEr score of the search caption (hereinafter referred to as the "CIDEr score (search method)") and the search sentence length value. For example, the model generation unit 133 may generate a vision-language model that is reinforced by setting the reward value represented by the above formula (3). Here, the search sentence length value corresponds to the "sentence length value" in the above formula (3). The model generation unit 133 may generate a vision-language model that is reinforced so as to maximize the reward value represented by the above formula (3).

[0050] In addition, for each of the search captions, the model generation unit 133 may reinforce the learning of the vision-language model so as to minimize the value of the loss function represented by the sum of the logarithm of the appearance probability of each token included in the search caption and the reward value represented by the above formula (3). More specifically, the model generation unit 133 may reinforce the learning of the vision-language model so as to minimize the value of the loss function represented by the same formula as the above formula (2). The model generation unit 133 may reinforce the learning of the vision-language model so as to minimize the value of the loss function obtained by replacing the value of the reward value in the above formula (2) with the reward value represented by the above formula (3). The model generation unit 133 learns the value of the parameters of the vision-language model so as to minimize the value of the loss function obtained by replacing the value of the reward value in the above formula (2) with the reward value represented by the above formula (3).

[0051] In addition, when the model generation unit 133 generates a vision-language model that has been reinforced based on the reward value represented by the above formula (3), the model generation unit 133 may store information regarding the generated vision-language model in the storage unit 120. In addition, when the model generation unit 133 generates a vision-language model that has been reinforced based on the reward value represented by the above formula (3), the model generation unit 133 may output the generated vision-language model to the text generation unit 134.

[0052] (Text Generation Unit 134) The text generation unit 134 may acquire the machine learning model generated by the model generation unit 133. For example, the text generation unit 134 may acquire the vision-language model generated by the model generation unit 133. For example, the text generation unit 134 may refer to the storage unit 120 to acquire the vision-language model generated by the model generation unit 133. Alternatively, the text generation unit 134 may acquire the vision-language model from the model generation unit 133. In addition, when the text generation unit 134 acquires the vision-language model generated by the model generation unit 133, the text generation unit 134 may acquire the image to be processed. For example, the text generation unit 134 may acquire the image to be processed from an external information processing device via the communication unit 110.

[0053] Further, the text generation unit 134 may generate an image caption corresponding to the image to be processed from the image to be processed using the machine learning model generated by the model generation unit 133. For example, an image caption corresponding to the image to be processed may be generated from the image to be processed using the vision-language model generated by the model generation unit 133. For example, when the text generation unit 134 acquires the vision-language model generated by the model generation unit 133 and the image to be processed, an image caption corresponding to the image to be processed may be generated from the image to be processed using the vision-language model generated by the model generation unit 133. Specifically, the text generation unit 134 may generate an image caption corresponding to the image to be processed by inputting the image to be processed into the vision-language model generated by the model generation unit 133. For example, the text generation unit 134 may acquire an image caption corresponding to the image to be processed output from the vision-language model by inputting the image to be processed into the vision-language model generated by the model generation unit 133.

[0054] 〔3. Processing Procedure〕 FIG. 6 is a flowchart showing a procedure of information processing by the information processing apparatus 100 according to the embodiment. In FIG. 6, the acquisition unit 131 of the information processing apparatus 100 acquires a generated caption, which is an image caption generated by the vision-language model (step S11). Further, the calculation unit 132 of the information processing apparatus 100 calculates an evaluation value for evaluating the accuracy of the generated caption acquired by the acquisition unit 131 and a sentence length value indicating the length of the generated caption (step S12). Further, the model generation unit 133 of the information processing apparatus 100 generates a vision-language model that is reinforced by setting a reward value based on the evaluation value and the sentence length value calculated by the calculation unit 132 (step S13). Further, the text generation unit 134 of the information processing apparatus 100 generates an image caption corresponding to the image to be processed from the image to be processed using the vision-language model generated by the model generation unit 133 (step S14).

[0055] 〔4. Modification Example〕 The processing according to the above-described embodiment may be implemented in various different forms other than the above embodiment.

[0056] In the above-described embodiment, the case where the model generation unit 133 sets, as the reward value for the reinforcement learning of the vision-language model, the value obtained by adding the value obtained by subtracting the greedy evaluation value from the search evaluation value and the search sentence length value has been described. However, the reward value for the reinforcement learning of the vision-language model is not limited to this. For example, the model generation unit 133 may set, as the reward value for the reinforcement learning of the vision-language model, the value obtained by adding the search evaluation value and the search sentence length value. For example, the model generation unit 133 may set, as the reward value for the reinforcement learning of the vision-language model, the value obtained by adding the search evaluation value, which is the CIDEr score of the search caption, and the search sentence length value.

[0057] Also, in the above-described embodiment, the case where the CIDEr score is used as the evaluation value for evaluating the accuracy of the generated caption has been described. However, the evaluation value for evaluating the accuracy of the generated caption is not limited to the CIDEr score. For example, the evaluation value for evaluating the accuracy of the generated caption may be a score corresponding to evaluation metrics such as BLEU (Bilingual Evaluation Understudy), METEOR (Metric for Evaluation of Translation with Explicit ORdering), SPICE (Semantic Propositional Image Caption Evaluation), etc. Also, the evaluation value for evaluating the accuracy of the generated caption may be an evaluation value calculated by combining a plurality of evaluation metrics.

[0058] 〔5. Effects〕 As described above, the information processing apparatus 100 according to the embodiment includes an acquisition unit 131, a calculation unit 132, and a model generation unit 133. The acquisition unit 131 acquires a target caption, which is a sentence describing the content of the input target image, using a machine learning model. The calculation unit 132 calculates an evaluation value for evaluating the accuracy of the target caption and a sentence length value indicating the length of the target caption. The model generation unit 133 generates a machine learning model that is reinforced by setting a reward value based on the evaluation value and the sentence length value.

[0059] In this way, the information processing apparatus 100 sets a reward value based on an evaluation value for evaluating the accuracy of an image caption, which is a sentence describing the content of an image, and a sentence length value indicating the length of the image caption, and reinforces the machine learning model, thereby maximizing the evaluation value for evaluating the accuracy of the image caption and preventing the length of the image caption from becoming short. As a result, the information processing apparatus 100 can generate an image caption with accurate content and a length that is not extremely short. Therefore, the information processing apparatus 100 can prevent a decrease in the amount of information included in the image caption while generating an image caption with accurate content. In addition, since the information processing apparatus 100 can prevent a decrease in the amount of information included in the image caption while generating an image caption with accurate content, it can contribute to the achievement of Goal 9, "Build the infrastructure for industry and innovation," of the Sustainable Development Goals (SDGs).

[0060] In addition, the acquisition unit 131 acquires a search caption, which is a target caption generated by a search method. The calculation unit 132 calculates a search evaluation value, which is an evaluation value of the search caption, and a search sentence length value, which is a sentence length value of the search caption. The model generation unit 133 generates a machine learning model that is reinforced by setting a reward value based on the search evaluation value and the search sentence length value.

[0061] As a result, the information processing apparatus 100 can prevent a decrease in the amount of information included in the search caption while generating a search caption with accurate content by setting a reward value based on an evaluation value for evaluating the accuracy of the search caption and a sentence length value indicating the length of the search caption, and performing reinforcement learning on the machine learning model.

[0062] In addition, the calculation unit 132 calculates an evaluation value based on the similarity between the correct caption and the target caption.

[0063] As a result, the information processing apparatus 100 can prevent a decrease in the amount of information included in the image caption while generating an image caption with accurate content that maximizes the evaluation value based on the similarity between the correct caption and the image caption generated by the machine learning model.

[0064] In addition, the calculation unit 132 calculates an evaluation value that is the CIDEr (Consensus-based Image Description Evaluation) score of the target caption.

[0065] As a result, the information processing apparatus 100 can prevent a decrease in the amount of information included in the image caption while generating an image caption with accurate content that maximizes the CIDEr score.

[0066] In addition, the calculation unit 132 calculates a sentence length value based on the length of the sequence of tokens included in the target caption.

[0067] As a result, the information processing apparatus 100 can prevent the length of the sequence of tokens included in the image caption from becoming short by using the length of the sequence of tokens included in the image caption as a reward value. Therefore, the information processing apparatus 100 can prevent a decrease in the amount of information included in the image caption.

[0068] Further, the information processing apparatus 100 further includes a sentence generation unit 134. The sentence generation unit 134 uses the machine learning model generated by the model generation unit 133 to generate an image caption corresponding to the image to be processed from the image to be processed.

[0069] Thereby, the information processing apparatus 100 can generate an image caption with accurate content while preventing a decrease in the amount of information included in the image caption.

[0070] [6. Hardware Configuration] Further, the information processing apparatus 100 according to the above-described embodiment is realized by a computer 1000 having a configuration as shown in FIG. 7, for example. FIG. 7 is a hardware configuration diagram showing an example of a computer that realizes the functions of the information processing apparatus 100. The computer 1000 includes a CPU 1100, a RAM 1200, a ROM 1300, an HDD 1400, a communication interface (I / F) 1500, an input / output interface (I / F) 1600, and a media interface (I / F) 1700.

[0071] The CPU 1100 operates based on a program stored in the ROM 1300 or the HDD 1400 and controls each part. The ROM 1300 stores a boot program executed by the CPU 1100 when the computer 1000 is started up, a program depending on the hardware of the computer 1000, and the like.

[0072] The HDD 1400 stores a program executed by the CPU 1100, data used by such a program, and the like. The communication interface 1500 receives data from other devices via a predetermined communication network and sends it to the CPU 1100, and sends data generated by the CPU 1100 to other devices via a predetermined communication network.

[0073] The CPU 1100 controls output devices such as displays and printers, and input devices such as keyboards and mice via the input / output interface 1600. The CPU 1100 acquires data from the input device via the input / output interface 1600. Further, the CPU 1100 outputs the generated data to the output device via the input / output interface 1600.

[0074] The media interface 1700 reads a program or data stored in the recording medium 1800 and provides it to the CPU 1100 via the RAM 1200. The CPU 1100 loads such a program from the recording medium 1800 onto the RAM 1200 via the media interface 1700 and executes the loaded program. The recording medium 1800 is, for example, an optical recording medium such as a DVD (Digital Versatile Disc) or a PD (Phase change rewritable Disk), a magneto-optical recording medium such as an MO (Magneto-Optical disk), a tape medium, a magnetic recording medium, or a semiconductor memory.

[0075] For example, when the computer 1000 functions as the information processing apparatus 100 according to the embodiment, the CPU 1100 of the computer 1000 realizes the functions of the control unit 130 by executing the program loaded onto the RAM 1200. The CPU 1100 of the computer 1000 reads and executes these programs from the recording medium 1800. As another example, these programs may be acquired from other devices via a predetermined communication network.

[0076] As described above in detail some of the embodiments of the present application with reference to the drawings, these are examples, and the present invention can be implemented in other forms with various modifications and improvements based on the knowledge of those skilled in the art, starting from the aspects described in the column of the disclosure of the invention.

[0077] 〔7. Others〕 Also, among the respective processes described in the above embodiments and modified examples, all or part of the processes described as being automatically performed can be manually performed, or all or part of the processes described as being manually performed can be automatically performed by a known method. In addition, regarding the processing procedures, specific names, and information including various data and parameters shown in the above documents and drawings, they can be arbitrarily changed unless otherwise specified. For example, the various information shown in each figure is not limited to the illustrated information.

[0078] Also, each component of each illustrated device is conceptually functional and does not necessarily need to be physically configured as shown in the figure. That is, the specific form of the distribution and integration of each device is not limited to that shown in the figure, and all or part of it can be functionally or physically distributed and integrated in any unit according to various loads and usage situations.

[0079] Also, the above-described embodiments and modified examples can be appropriately combined within a range that does not conflict with the processing content.

Explanation of Reference Numerals

[0080] 100 Information Processing Device 110 Communication Unit 120 Storage Unit 130 Control Unit 131 Acquisition Unit 132 Calculation Unit 133 Model Generation Unit 134 Text Generation Unit< / eos> < / bos> < / eos> < / bos>

Claims

1. An acquisition unit that acquires a target caption, which is a sentence describing the content of the input target image, using a machine learning model; A calculation unit that calculates an evaluation value for evaluating the accuracy of the target caption and a sentence length value indicating the length of the target caption; A model generation unit that generates the machine learning model that is reinforced by setting a reward value based on the evaluation value and the sentence length value; An information processing apparatus comprising the above.

2. The acquisition unit: Acquires a search caption, which is the target caption generated by a search method; The calculation unit: Calculates a search evaluation value, which is the evaluation value of the search caption, and a search sentence length value, which is the sentence length value of the search caption; The model generation unit: Generates the machine learning model that is reinforced by setting the reward value based on the search evaluation value and the search sentence length value. The information processing apparatus according to Claim 1.

3. The calculation unit: Calculates the evaluation value based on the similarity between the correct caption and the target caption. The information processing apparatus according to Claim 1.

4. The calculation unit: Calculates the evaluation value, which is the CIDEr (Consensus-based Image Description Evaluation) score of the target caption. The information processing apparatus according to Claim 1.

5. The calculation unit: Calculates the sentence length value based on the length of the sequence of tokens included in the target caption. The information processing apparatus according to Claim 1.

6. The information processing apparatus according to Claim 1, further comprising a sentence generation unit that generates an image caption corresponding to the image to be processed from the image to be processed using the machine learning model generated by the model generation unit. The information processing apparatus according to Claim 1.

7. An acquisition procedure for acquiring a target caption, which is a sentence describing the content of the input target image, using a machine learning model; A calculation procedure for calculating an evaluation value for evaluating the accuracy of the target caption and a sentence length value indicating the length of the target caption; A model generation procedure for generating the machine learning model that is reinforced by setting a reward value based on the evaluation value and the sentence length value; An information processing program for causing a computer to execute the above.

Citation Information

Patent Citations

  • Training method and training device for image description model

    CN114090815A

  • Image description text generation method and device, electronic equipment and storage medium

    CN116844011A

  • Turbidity sensors for smart aquaculture

    KR1020240165630A

  • Method, apparatus, device and medium for generating captioning information of multimedia data

    US20220014807A1

  • Image caption generation model training device, image caption generation device, image caption generation model training method, image caption generation method, and program

    WO2024023884A1