Machine learning device, machine learning method, machine learning program, and inference device
The machine learning device efficiently learns a statistical model for the VQA task by converting non-VQA format samples into VQA format, thereby reducing costs and enhancing accuracy through increased training data diversity and quantity.
Patent Information
- Application Number
- JP2022019858
- Authority / Receiving Office
- JP · JP
- Patent Type
- Patents
- Current Assignee / Owner
- Filing Date
- 2022-02-10
- Publication Date
- 2025-05-26
- Estimated Expiration
- 2042-02-10
AI Technical Summary
The challenge is to efficiently learn a statistical model for the Visual Question Answering (VQA) task at a low cost, as existing methods require vast and varied datasets to achieve high accuracy.
A machine learning device and method that includes a conversion unit to generate VQA format learning samples from non-VQA format samples, and a learning unit to train a statistical model using these generated samples, thereby increasing the diversity and quantity of training data.
This approach allows for the training of a highly accurate statistical model for the VQA task with reduced costs by leveraging non-VQA samples and automatically generating question and answer pairs, thereby enhancing the efficiency of the learning process.
Smart Images

Figure 0007682818000001 
Figure 0007682818000002 
Figure 0007682818000003
Abstract
Description
Technical Field
[0001] Embodiments of the present invention relate to a machine learning device, a machine learning method, a machine learning program, and an inference device.
Background Art
[0002] In the field of machine learning, a task of inputting an image and a text-form question regarding the image and outputting a text-form answer to the question is known. This task is called VQA (Visual Question Answering). A statistical model of the VQA task is trained based on a learning dataset given as a combination (tuple) of an image, a question, and an answer. Since there are enormous variations in the combination of an image and a question regarding the image, in the VQA learning dataset called VQAv2, variations are ensured by preparing hundreds of thousands of questions for tens of thousands of images. For example, when trying to generate a statistical model capable of corresponding to specific animals, plants, and vehicles, it is necessary to prepare images regarding those specific objects and questions and answers with all variations regarding them. Thus, it requires an enormous cost to prepare a learning dataset consisting of a combination of an image, a question, and an answer with various variations. To reduce the cost, even if a statistical model is trained with a learning dataset with few variations, it is impossible to generate a highly accurate statistical model. Efficient learning that can generate a highly accurate statistical model at low cost is desired.
Prior Art Documents
Non-Patent Documents
[0003]
Non-Patent Document 1
Summary of the Invention
Problems to be Solved by the Invention
[0004] The problem to be solved by the present invention is to provide a machine learning device, a machine learning method, a machine learning program, and an inference device that can efficiently learn a statistical model for the VQA task.
Means for Solving the Problems
[0005] The machine learning device according to the embodiment includes a conversion unit and a learning unit. The conversion unit generates a learning sample in the VQA format related to the VQA task based on a sample in a non-VQA format. The learning sample in the VQA format has, as elements, a combination of an object, a question sentence regarding the object, and an answer sentence for the question sentence, and the sample in the non-VQA format has, as elements, a combination of an object and a label related to the object. The learning unit trains a statistical model for the VQA task based on the learning sample in the VQA format generated by the conversion unit.
Brief Description of the Drawings
[0006]
Figure 1
Figure 2
Figure 3
Figure 4
Figure 5
Figure 6
Figure 7
Figure 8
Figure 9
Figure 10
Figure 11
Figure 12
Figure 13
Figure 14
Figure 15
Figure 16
Figure 17
Figure 18
Mode for Carrying Out the Invention
[0007] Hereinafter, a machine learning device, a machine learning method, a machine learning program, and an inference device according to the present embodiment will be described with reference to the drawings.
[0008] FIG. 1 is a diagram showing a configuration example of the machine learning device 1 according to the present embodiment. As shown in FIG. 1, the machine learning device 1 is a computer having a processing circuit 11, a storage device 12, an input device 13, a communication device 14, and a display device 15. Data communication between the processing circuit 11, the storage device 12, the input device 13, the communication device 14, and the display device 15 is performed via a bus.
[0009] The processing circuit 11 includes a processor such as a CPU (Central Processing Unit) and a memory such as a RAM (Random Access Memory). The processing circuit 11 includes an acquisition unit 111, a conversion unit 112, a learning unit 113, and a display control unit 114. The processing circuit 11 realizes the functions of the respective units 111 to 114 by executing a machine learning program. The machine learning program is stored in a non-transitory computer-readable recording medium such as the storage device 12. The machine learning program may be implemented as a single program that describes all the functions of the respective units 111 to 114, or may be implemented as a plurality of modules divided into several functional units. Further, the respective units 111 to 114 may be implemented by an integrated circuit such as an application specific integrated circuit (ASIC) or an FPGA (Field Programmable Gate Array). In this case, it may be implemented by a single integrated circuit or individually implemented by a plurality of integrated circuits.
[0010] The acquisition unit 111 acquires VQA format learning samples and non-VQA format data samples related to the VQA task for training the statistical model of the VQA task. The learning samples of the VQA task have a format as samples for training the statistical model of the VQA task. Specifically, the learning samples of the VQA task have, as its elements, a combination (tuple) of an object, a question sentence regarding the object, and an answer sentence for the question. The object means the data to be processed. Specifically, an image or a video is used as the object. Note that, in addition to images or videos, data obtained by various modalities such as audio, measuring instrument outputs, and / or 3D point clouds may be used as the object according to this embodiment. The VQA format learning samples are acquired from a database storing a large amount of learning samples in the VQA format. The non-VQA format means a format different from the VQA format. The data samples in the non-VQA format have, as its elements, a combination of an object and a label related to the object. The label is text data related to the semantic content of the object. The data samples in the non-VQA format may be learning samples of a task different from the VQA task (non-VQA task), or may not be learning samples. The data samples in the non-VQA format are acquired from a database storing a large amount of data samples in the non-VQA format.
[0011] The conversion unit 112 generates learning samples in the VQA format related to the VQA task based on data samples in a non-VQA format. The learning samples generated by the conversion unit 112 are also used for training the statistical model of the VQA task. Hereinafter, the learning samples in the VQA format obtained from the database of VQA samples by the acquisition unit 111 are referred to as VQA samples, and the learning samples generated by the conversion unit 112 are referred to as additional samples. Also, the data samples in the non-VQA format obtained by the acquisition unit 111 are referred to as non-VQA samples. Also, when not distinguishing VQA samples, non-VQA samples, and additional samples, they may simply be referred to as samples.
[0012] The learning unit 113 trains the statistical model of the VQA task based on the additional samples generated by the conversion unit 112. Note that the learning unit 113 may train the statistical model of the VQA task based on the additional samples generated by the conversion unit 112 and the VQA samples obtained by the acquisition unit 111.
[0013] The display control unit 114 displays various information on the display device 15. For example, the display control unit 114 displays VQA samples, additional samples, prediction results of the VQA task by the statistical model, and the like.
[0014] The storage device 12 is composed of a ROM (Read Only Memory), HDD (Hard Disk Drive), SSD (Solid State Drive), integrated circuit memory device, etc. The storage device 12 stores a machine learning processing program and the like.
[0015] The input device 13 inputs various commands from the operator. As the input device 13, a keyboard, mouse, various switches, touch pad, touch panel display, etc. can be used. The output signal from the input device 13 is supplied to the processing circuit 11. Note that the input device 13 may be an input device of a computer connected to the processing circuit 11 via wire or wireless.
[0016] The communication device 14 is an interface for performing data communication between the machine learning device 1 and an external device connected via a network. As an example, the external device is a database of VQA samples, samples in a non-VQA format, or the like.
[0017] The display device 15 displays various information under the control of the display control unit 114. As the display device 15, a CRT (Cathode-Ray Tube) display, a liquid crystal display, an organic EL (Electro Luminescence) display, an LED (Light-Emitting Diode) display, a plasma display, or any other display known in the art can be appropriately used. Further, the display device 15 may be a projector.
[0018] Hereinafter, the machine learning device 1 will be described in detail. In the following description, it is assumed that a non-VGA sample is a learning sample in a format related to a non-VQA task. The non-VQA task refers to a task different from the VQA task. The non-VQA task is a task of recognizing, understanding, and inferring the relationship between an object and a label related to the object. As an example, an image classification task, an object detection task, a visual grounding task, or an image search task can be applied as the non-VQA task. In the following description, it is assumed that the object is an image.
[0019] FIG. 2 is a diagram illustrating a processing procedure of machine learning processing by the machine learning apparatus 1. FIG. 3 is a diagram schematically showing the machine learning processing illustrated in FIG. 2. As shown in FIGS. 2 and 3, the acquisition unit 111 acquires a VQA sample 31 or a non-VQA sample 32 in order to train a statistical model of the VQA task (step S201). As an example, the acquisition unit 111 acquires a dataset of VQA samples and a dataset of non-VQA samples from the storage device 12 or an external database, and selects samples for one mini-batch from among these datasets of VQA samples and non-VQA samples. In one mini-batch according to the present embodiment, VQA samples and non-VQA samples may be mixed, or only one of VQA samples and non-VQA samples may be included.
[0020] The VQA sample 31 includes a combination of an image, a question sentence regarding the content of the image, and a correct answer sentence regarding the question sentence. The non-VQA sample 32 includes a combination of an image and a correct label for the image. The non-VQA sample 32 does not include a question sentence and a correct answer sentence. For example, when the non-VQA task is an image classification task, an image classification sample is used as the non-VQA sample 32. The image classification sample includes, as its elements, an image and a correct label for the image. The correct label means a class label of an object shown in the image. As another example, when the non-VQA task is an object detection task, an object detection sample is used as the non-VQA sample 32. The object detection sample includes, as its elements, an image and a correct label for the image. The correct label includes a class label of an object shown in the image and parameters of a rectangle (bounding box) surrounding the object.
[0021] When step S201 is performed, the learning unit 113 determines whether the sample obtained in step S201 is a VQA sample (step S202). If the sample is obtained artificially, step S202 may not be executed. If the sample is obtained randomly, when the obtained sample includes a question text and a correct answer text, the learning unit 113 determines that the VQA sample 31 has been obtained, and when the question text and the correct answer text are not included, it may be determined that the VQA sample 31 has not been obtained, that is, the non-VQA sample 32 has been obtained. Alternatively, when an identifier indicating the type of the sample is associated with each sample, the learning unit 113 may determine whether it is a VQA sample based on the identifier.
[0022] If it is determined in step S202 that the VQA sample 31 has not been obtained, that is, the non-VQA sample 32 has been obtained (step S202: NO), the conversion unit 112 generates a question text and a correct answer text based on the correct label of the non-VQA sample 32 (step S203). The generation process of the question text and the correct answer text will be described in detail later.
[0023] If step S203 is performed or if it is determined in step S202 that the VQA sample 31 has been obtained (step S202: YES), the learning unit 113 uses the statistical model M1 of the VQA model to predict an answer text based on the image and the question text (step S204).
[0024] FIG. 4 is a diagram showing an example of the network configuration of the statistical model M1. As shown in FIG. 4, the statistical model M1 is a neural network trained to input an image and a question sentence and output an answer sentence. The statistical model M1 includes an image encoder M11, a text encoder M12, a fuser M13, and an answer sentence converter M14. The image encoder M11 is an encoding network layer that outputs a feature amount as image data (hereinafter, image feature amount) of the input image. The text encoder M12 is an encoding network layer that outputs a feature amount as text data (hereinafter, text feature amount) of the input question sentence. More specifically, the text encoder M12 divides the question sentence into a series of words by a tokenizer, assigns a unique word ID to each word, and thus converts it into a series of word IDs. The word ID is an identifier uniquely assigned in advance for each word. Then, the text encoder M12 encodes (encodes) the series of word IDs and converts them into text feature amounts. The fuser M13 outputs a fused feature amount of the image feature amount output from the image encoder M11 and the text feature amount output from the text encoder M12. The fused feature amount corresponds to the vector representation of the predicted answer sentence. The fuser M13 may be a module that performs an addition operation or concatenation, or may be an encoding network layer.
[0025] The answer sentence converter M14 is a decoding network layer called a string decoder (sequence decoder) that converts the fused feature amount output from the fuser M13 into a string of natural language representing the answer sentence. Specifically, the answer sentence converter M14 converts the fused feature amount into a plurality of answer word vectors respectively corresponding to a plurality of words constituting the predicted answer sentence. The answer sentence converter M14 converts each answer word vector into a series of relative values (logits) representing the occurrence probability of each word. The logit series corresponds to the predicted answer sentence. The answer word vector is a multi-dimensional vector having a dimension corresponding to the number of words (hereinafter referred to as registered words) registered in the tokenizer dictionary, and each element is assigned the logit of each registered word for the target word. The number of registered words is not particularly limited, but illustratively, it is about tens of thousands to hundreds of thousands. Note that the predicted answer sentence may be a numerical sequence or a symbol sequence instead of a character string. The tokenizer dictionary is common to all types of samples without depending on the type of samples such as VQA samples and image classification samples. Note that in the learning stage, it is not necessary to convert the logit series into a language series representing the predicted answer sentence. Note that at the time of inference, the answer sentence converter M14 generates a character string representing the inference answer sentence by converting each of the plurality of logit series into words with reference to the tokenizer dictionary.
[0026] The image encoder M11, the text encoder M12, the fuser M13, and the answer sentence converter M14 are typically composed of multi-layer neural networks. However, the present embodiment is not limited to this, and random forest, recursive partitioning regression tree (such as CART), bagging, boosting, support vector machine, etc. may be used.
[0027] When step S204 is performed, the learning unit 113 calculates the loss between the correct answer sentence and the predicted answer sentence (step S205). The loss calculation is performed for the purpose of feeding back the loss between the predicted answer sentence and the correct answer sentence to the statistical model M1 in order to reduce the error with respect to the learning samples of the statistical model M1. As the loss according to this embodiment, cross entropy used in the field of language modeling is used. Specifically, first, the learning unit 113 converts the correct answer sentence into a one-hot vector having only the correct word ID with a value of "1" and the other values of "0", and performs a softmax operation on each logit constituting the predicted answer sentence to calculate a softmax operation value. The learning unit 113 calculates the cross entropy representing the difference between the predicted answer sentence and the correct answer sentence based on the softmax operation value of the predicted answer sentence and the one-hot vector of the correct answer sentence.
[0028] When step S205 is performed, the learning unit 113 updates the statistical model M1 based on the loss (step S206). Specifically, the learning unit 113 updates the learning parameters of the statistical model M1 using an arbitrary optimization method such as the error backpropagation method. The learning parameters mean the parameters that are updated by machine learning among various parameters set in the statistical model M1, such as weight parameters and biases. The calculation of the loss (step S205) and the update of the statistical model M1 (step S206) are typically performed in mini-batch units. However, it is not limited to this, and the statistical model M1 may be updated for each one or more samples constituting the batch. In FIGS. 2 and 3, the calculation of the loss (step S205) and the update of the statistical model M1 (step S206) are shown as separate steps, but in actual processing, the two steps are not clearly distinguished.
[0029] When step S206 is performed, the learning unit 113 determines whether or not the stop condition is satisfied (step S207). The stop condition can be set to any condition such as the number of repetitions of steps S201 to S207 reaching a predetermined number, the loss reaching a threshold value, the performance index value reaching a threshold value, etc. When it is determined that the stop condition is not satisfied (step S207: NO), steps S201 to S207 are repeated for other samples. And when it is determined that the stop condition is satisfied (step S207: YES), the learning unit 113 outputs the statistical model in which the learning parameters at the current update count are set as a learned model (step S208). The learned model is stored in the storage device 12 or transferred to another computer via the communication device 14 or the like.
[0030] As described above, the machine learning process by the machine learning device 1 ends.
[0031] Note that the processing procedure shown in FIG. 2 is an example and is not limited to the procedure shown in FIG. 2. For example, it is not necessary to train the statistical model using both VQA samples and non-VQA samples, and the statistical model may be trained using only non-VQA samples. In this case, the batch will be composed of only non-VQA samples.
[0032] As described above, according to the machine learning process according to the present embodiment, non-VQA samples are converted into additional samples that are samples in the VQA format, and the statistical model M1 for the VQA task is trained based on the additional samples. By converting non-VQA samples into additional samples in the VQA format, it becomes possible to increase the learning samples of the statistical model M1. Since the statistical model M1 can be trained with various learning samples, it becomes possible to improve the accuracy of the statistical model M1. In addition, since the question sentence and the answer sentence can be automatically generated from the correct labels of the non-VQA samples, the number of learning samples of the statistical model M1 can be easily increased.
[0033] Hereinafter, several examples according to this embodiment will be given to specifically describe the machine learning process according to this embodiment. Note that the overall processing procedure of the machine learning process according to each of the following examples is as shown in FIGS. 2 and 3. In the following examples, the description will focus on the parts that are different between the examples.
[0034] (Example 1) In Example 1, the machine learning of the statistical model M1 of the VQA task using VQA samples and image classification samples will be described. The image classification sample is an example of a non-VQA sample.
[0035] FIG. 5 is a diagram schematically showing the machine learning process for the statistical model M1 of the VQA task using VQA samples and image classification samples according to Example 1. As shown in FIG. 5, the VQA sample has a combination of an image 51, a question sentence 52, and a correct answer sentence 53. In the example of FIG. 5, the image 51 is an image depicting four women eating indoors, the question sentence 52 is the character string "Is there anyone wearing a hat?", and the correct answer sentence 53 is the character string "No". The image classification sample has a combination of an image 54 and a correct label 55. The correct label 55 related to the image classification task is the name (class) of the object shown in the image 54. In the example of FIG. 5, the image 54 is an image depicting a lawn mower, and the correct label 55 is the character string "lawn mower". As described above, the loss according to Example 1 is calculated based on the correct answer sentences 53, 57 which are character strings and the predicted answer sentence 58.
[0036] As shown in FIG. 5, based on the correct label 55, a question sentence 56 and a correct answer sentence 57 are generated by the conversion unit 112. The question sentence 56 and the correct answer sentence 57 constitute additional samples. The combination of the question sentence 56 and the correct answer sentence 57 related to the image classification task is classified into the following three types. Note that a character string representing the correct label is inserted into 'correct label', and a character string representing other than the correct label is inserted into 'other than correct label'.
[0037] Type 1. Question sentence "What is this?", correct answer sentence " 'correct label' " Type 2. Question sentence: "Is this the 'correct label'?", Correct answer sentence: "Yes" Type 3. Question sentence: "Is this something other than the 'correct label'?", Correct answer sentence: "No"
[0038] As can be seen from Types 1 to 3, the question sentence 56 and the correct answer sentence 57 can be defined by simple rules based on the correct label. By using the templates for each of Types 1 to 3, it is possible to automatically generate the question sentence 56 and the correct answer sentence 57. As shown in the example of Fig. 5, when the correct label 55 is "lawn mower", for Type 1, the question sentence 56 "What is this?" and the correct answer sentence 57 "Lawn mower" are generated; for Type 2, the question sentence 56 "Is this a lawn mower?" and the correct answer sentence 57 "Yes" are generated; for Type 3, the question sentence 56 "Is this a cat?" and the correct answer sentence 57 "No" are generated. For Type 3, a cat is exemplified as a class other than a lawn mower, but it is not limited to a cat and can be replaced with the name of any object other than a lawn mower.
[0039] According to the question sentence 56 and the correct answer sentence 57 of Type 1, it is possible to let the text encoder M12 learn the correct label 55 in relation to the image features. According to the question sentence 56 and the correct answer sentence 57 of Type 2, it is possible to let the image encoder M11 and the text encoder M12 learn the relationship between the correct label 55 and the image features. If there is only Type 2, there will only be positive samples and the learning will be biased. According to the question sentence 56 and the correct answer sentence 57 of Type 3, it is useful as negative samples.
[0040] FIG. 6 is a diagram schematically showing the machine learning process of a statistical model using a VQA sample and an image classification sample according to a comparative example. The comparative example is an example using the technique described in Non-Patent Document 1. The VQA sample has a combination of an image 61, a question sentence 62, and a correct answer sentence 63. The image classification sample has a combination of an image 64 and a correct label 65. The VQA sample and the image classification sample according to the comparative example are substantially the same as the VQA sample and the image classification sample according to Example 1. The statistical model M2 of the VQA task according to the comparative example has an image encoder M21, a text encoder M22, a fusion unit M23, and a decoder M24. The image encoder M21, the text encoder M22, and the fusion unit M23 are respectively the same as the image encoder M11, the text encoder M12, and the fusion unit M13 according to Example 1 shown in FIG. 5.
[0041] The decoder M24 decodes the fusion feature amount output from the fusion unit M23. Heads (output branches) M251 and M252 unique to the type of task are connected to the output end of the decoder M24. The output branches M251 and M252 are network layers including one or more fully connected layers and / or convolutional layers. When executing the VQA task, the VQA head M251 is connected, and when executing the image classification task, the image classification head M252 is connected.
[0042] The VQA head M251 outputs a predicted answer based on the decoded fused feature. More specifically, the VQA head M251 converts the decoded fused feature into a predicted answer vector. The predicted answer vector is a multi-dimensional vector having dimensions corresponding to the number of answer candidate IDs registered in the dictionary, and each element is assigned a relative value (logit) of each answer candidate ID for the predicted answer. The VQA head M251 identifies the answer candidate ID with the maximum logit, converts the identified answer candidate ID into a character string of the answer candidate using the dictionary, and outputs the character string as the predicted answer. The dictionary registers the association between the character string of the answer candidate and the answer candidate ID. For example, when there are 3000 answer candidates, there are answer candidate IDs from No. 0 to No. 2999. A loss is calculated based on the difference between the answer candidate ID that is the output of the VQA head M251 and the answer candidate ID corresponding to the correct answer sentence 63.
[0043] The image classification head M252 also outputs a predicted class based on the decoded fused feature by the same process as the VQA head M251. It outputs the class candidate with the maximum likelihood among a plurality of predetermined class candidates as the predicted class. The class candidates are also associated with class IDs, similar to the answer candidates, and the association between the class and the class ID is registered in the dictionary. A loss is calculated based on the difference between the class candidate ID that is the output of the VQA head M252 and the answer candidate ID corresponding to the correct class 66.
[0044] When training the statistical model M2 using VQA samples, since the VQA samples contain images and question texts, the relationship between both the image features and the text features is learned. When training the statistical model M2 using image classification samples, since the image classification samples do not contain text input, nothing is input to the text encoder M12, and as a result, only image features are learned. When performing learning of the VQA task, the VQA head M251 is connected to the decoder M24, and the statistical model M2 is trained. The loss is calculated based on the answer candidate ID corresponding to the correct answer and the answer candidate ID corresponding to the predicted answer. Therefore, the correct label as text included in the image classification samples is not utilized in the learning of the VQA task. For example, when training the statistical model M2 based on image classification samples related to lawn mowers, the VQA head M251 cannot learn the lawn mower as text. Therefore, at the inference stage, even if an image of a lawn mower and the question text "What is this?" are input to the learned statistical model M2, the statistical model M2 cannot output the predicted answer text "lawn mower". Furthermore, even if an image of a lawn mower and the question text "How many lawn mowers are there?" are input to the learned statistical model M2, since the text encoder M22 has not learned the text features of the lawn mower, the statistical model M2 cannot answer well.
[0045] In the comparative example, the learning of the VQA task using VQA samples and the learning of the image classification task using image classification samples are performed independently by swapping the heads M251 and M252. For this reason, the association between the ID and the candidate in the dictionary differs depending on the domain of samples such as VQA samples and image classification samples. For example, it may happen that the ID of "apple" is "159" in the statistical model M2 for the VQA task, while the ID of "apple" is "1035" in the statistical model M2 for the image classification task. Also, since the number of types of images included in various samples is different, it is also assumed that the number of answer candidates differs according to the types of the heads M251 and M252 accordingly. Due to these factors, it is difficult to share the head for different tasks.
[0046] In the comparative example, heads M251 and M252 are replaced according to the type of task, and heads M251 and M252 for the classification task of outputting a plausible answer from among a plurality of candidates are used. Therefore, it is not possible to answer a word not included in the learning sample. Also, since the input / output and head of the statistical model M2 differ for each task, only the learning sample of a single task can be included in one mini-batch. Since the tasks are switched every iteration of steps S201 to S207 in FIG. 2, the possibility that the gradient will not be optimized increases when the batch size is large.
[0047] In this regard, as shown in FIG. 5 and the like, the statistical model M1 according to the present embodiment does not use heads M251 and M252, but instead has an answer sentence converter M14 that outputs a character string. Along with this, the output of the predicted answer sentence is performed using a common tokenizer dictionary for different types of tasks. As a result, it becomes possible to unify the output format of the predicted answer sentence between VQA samples and image classification samples (additional samples). By unifying the output format, it becomes possible to mix VQA samples and image classification samples in the same mini-batch and train one statistical model M1 using the VQA samples and image classification samples without distinction. This makes it possible to reduce the possibility that the gradient will not be optimized even when the batch size is large. Along with this, since it becomes possible to train the statistical model M1 using image classification samples related to content that does not exist in the VQA samples, it becomes possible to teach the statistical model M1 a new vocabulary compared to the case of training only with VQA samples.
[0048] In the comparative example, it is necessary to use unique heads M251 and M252 for each task. That is, when performing multi-task learning, it is necessary to switch the heads M251 and M252 for each task. For this reason, in the case of multi-task learning of the VQA task and the image classification task, it is necessary to perform the prediction process of the statistical model twice, that is, the prediction process of the statistical model with the VQA head M251 by the VQA sample and the prediction process of the statistical model with the image classification bed M252 by the image classification sample. On the other hand, according to the method according to the present embodiment, the image classification sample is converted into an additional sample having a VQA format, and the prediction process of the common statistical model is performed by this additional sample and the VQA sample, so that only one prediction process is required. The correct label included in the image classification sample is used to generate a question sentence and a correct answer sentence, and this question sentence and correct answer sentence are related to the content of the correct label of the image classification task as exemplified in the above types 1 to 3. By performing the machine learning process of the statistical model by the additional sample, it becomes possible to train a statistical model capable of substantially performing the image classification task in the form of the VQA task.
[0049] FIG. 7 is a diagram showing an example of the prediction result of the statistical model trained by machine learning according to Example 1. FIG. 8 is a diagram showing an example of the prediction result of the statistical model trained by machine learning according to the comparative example. In the machine learning according to Example 1, the statistical model is learned based on both the VQA sample and the image classification data. In the machine learning according to the comparative example, the statistical model is learned only with the VQA sample. Assume that the sample of the lawn mower is not included in the VQA sample and is included only in the image classification sample.
[0050] Display screens I1 and I2 representing prediction results are displayed on the display device 15 by the display control unit 114. The display screens I1 and I2 include images I11 and I21 that are objects, and display columns I12 and I22 for question texts and prediction answer texts. As shown in FIG. 8, when learning is performed only with VQA samples, it can be seen that there are errors in the answers for items related to lawn mowers, and incorrect answers are given when asking about other unknown words. On the other hand, in the machine learning according to the present embodiment, since both VQA samples and image classification samples are learned, correct answers are obtained for items related to lawn mowers, and correct answers are also given for questions about unknown words.
[0051] (Example 2) In Example 2, machine learning of a statistical model of a VQA task using VQA samples and object detection samples will be described. The object detection sample is an example of a non-VQA sample. The same elements as those in Example 1 are denoted by the same reference numerals and the description thereof is omitted. The same operational effects as those in Example 1 are not described unless necessary.
[0052] FIG. 9 is a diagram schematically showing a machine learning process for a statistical model M1 of a VQA task using VQA samples and object detection samples according to Example 2. As shown in FIG. 9, the object detection sample has a combination of an image 91 and a correct label 92. The correct label 92 related to the object detection task is the correct parameter (Bbox) of a rectangle (bounding box) surrounding the object shown in the image 91 and the correct class of the object.
[0053] The correct parameters include the correct position and / or correct size of the bounding box. The correct position is represented by the position coordinates of the bounding box within the image 91. As an example, the position coordinates are represented by the positions of unit image regions obtained by normalizing the width and height of the image 91 to a predetermined value (e.g., 1) and then dividing each of the width and height into 10 equal parts. Note that the number of equal parts is not limited to 10 and may be any number such as 50 or 100. The position coordinates include the (left coordinate, top coordinate, right coordinate, bottom coordinate) of the bounding box as its elements. As an example, the correct size is represented by (the number of unit image regions in the width direction, the number of unit image regions in the height direction). By defining in this way, it is possible to represent the correct position and correct size as a character string, for example, like the position coordinates "3538" and size "03". Thereby, it is possible to uniformly train the statistical model. Note that the left coordinate, top coordinate, right coordinate, and bottom coordinate may each be defined by a combination of a sign indicating the direction and the position of the unit image region, such as (L3T5R3B8). The correct class is a character string representing the class of the object shown in the image 91. Thus, in the second embodiment as well, in the object detection sample, the correct label 92 is represented as a character string.
[0054] As shown in FIG. 9, an additional sample including a question sentence 93 and a correct answer sentence 94 is generated by the conversion unit 112 based on the correct label 92 of the object detection sample. Specifically, based on the correct parameters and correct class of the bounding box, which is the correct label 92, a question sentence 93 and a correct answer sentence 94 regarding the position, size, and / or class of the object shown in the image 91 are generated. The question sentence 93 and the correct answer sentence 94 regarding the object detection task are classified into the following two types as an example. Note that a character string representing the correct class is inserted into 'correct class', a character string representing a class other than the correct class is inserted into 'other than correct class', a character string representing the correct position of the bounding box, i.e., the correct position of the object, is inserted into 'position', and a character string representing the number of correct positions of the bounding box, i.e., the number of objects, is inserted into 'number of positions'.
[0055] Type 1. Question text: "Number of 'correct classes'?", Answer text: "Number of 'correct positions'" Type 2. Question text: "Number of 'other than correct classes'?", Answer text: "0" Type 3. Question text: "Where is the 'correct class' located?", Answer text: "Correct position" Type 4. Question text: "What is the name of the object at the 'correct position'?", Answer text: "Correct class"
[0056] As can be seen from Types 1 to 4, the question text 93 and the correct answer text 94 can be defined by simple rules based on the correct label 92. By using the templates for each of Types 1 to 4, it is possible to automatically generate the question text 93 and the correct answer text 94. In the example of Figure 9, assume that the correct class is "human", the number of correct positions is "4", and the correct positions are "3538", "5315", "7315", and "9315". In this case, for Type 1, the question text "Number of humans?" and the correct answer text "4" are generated; for Type 2, the question text "Number of cats?" and the correct answer text "0" are generated; for Type 3, the question text "Where are humans located?" and the correct answer text "3538,5315,7315,9315" are generated; for Type 4, the question text "What is the name of the object at position 3538?" and the correct answer text "human" are generated. Note that the combination of the question text and the correct answer text is not limited to the above 4 types. For example, question texts and correct answer texts related to the correct size may be generated.
[0057] According to the question text and answer text of Type 1, it is possible to enhance the counting ability of the object by the statistical model M1. The question text and answer text of Type 2 function as negative samples. According to the question text and answer text of Type 3, it is possible to enhance the recognition ability of the position of the object by the statistical model M1. According to the question text and answer text of Type 4, it is possible to enhance the recognition ability of the class of the object by the statistical model M1.
[0058] According to Example 2, it becomes possible to easily generate, as character strings, a question sentence and a correct answer sentence regarding the class, position, and size of an object from the correct label of an object detection sample. Thereby, it becomes possible to have the statistical model M1 learn the relationship between the question sentence and the correct answer sentence regarding the class, position, and size of the object. Also, it becomes possible to ground the image, the question sentence, and the correct answer sentence. Since the image detection sample is converted into an additional sample having a VQA format, and machine learning processing of a common statistical model is performed using this additional sample and the image detection sample, it is possible to complete multi-task learning of the VQA task and the image detection task with one machine learning process. The question sentence and the correct answer sentence are generated using the correct label included in the image detection sample, and this question sentence and correct answer sentence relate to the content of the correct label of the image detection task as exemplified in the above Types 1 to 4. By performing machine learning processing of the statistical model using the additional sample, it becomes possible to train a statistical model capable of substantially performing the image detection task in the form of the VQA task.
[0059] (Example 3) In the above Examples 1 and 2, machine learning of the statistical model of the VQA task using the VQA sample and the non-VQA sample was described. In Example 3, machine learning of the statistical model of the VQA task using two types of non-VQA samples will be described. The non-VQA task according to Example 3 may be any of an image classification task, an object detection task, an image grounding task, and an image search task, but exemplarily, it is assumed to be an image classification task and an object detection task. In this case, the two types of non-VQA samples are assumed to be an object detection sample and an image classification sample. The same elements as those in Examples 1 to 2 are denoted by the same reference numerals and the description thereof is omitted. The same operational effects as those in Examples 1 to 2 are not described unless necessary.
[0060] FIG. 10 is a diagram schematically showing a machine learning process for a statistical model M1 of a VQA task using an object detection sample and an image classification sample according to Example 3. As shown in FIG. 10, according to the method of Example 2, an additional sample including a question sentence 93 and a correct answer sentence 94 is generated based on the correct label 92 of the object detection sample, and according to the method of Example 1, an additional sample including a question sentence 56 and a correct answer sentence 57 is generated based on the correct label 55 of the image classification sample. A statistical model M1 of the VQA task is trained based on the additional sample derived from the object detection sample and the additional sample derived from the image classification sample. The statistical model M1 inputs an additional sample derived from the object detection sample or an additional sample derived from the image classification sample and outputs a predicted answer sentence 101. According to Example 3, it becomes possible to train the statistical model M1 of the VQA task based only on samples of two different non-VQA tasks.
[0061] FIG. 11 is a diagram showing an example of the prediction result of the statistical model M1 trained by machine learning according to Example 3. FIG. 12 is a diagram showing an example of the prediction result of the statistical model M2 trained by machine learning according to the comparative example. In the machine learning according to Example 3, the statistical model is learned based on both the object detection sample and the image classification data. In the machine learning according to the comparative example, the statistical model is learned only with the object detection sample. Assume that a sample of a bird of the species Ruri no Diko is not included in the object detection sample and is included only in the image classification sample.
[0062] Display screens I3 and I4 representing the prediction results are displayed on the display device 15 by the display control unit 114. The display screens I3 and I4 include images I31 and I41 as objects and display columns I32 and I42 for the question sentence and the predicted answer sentence. As shown in FIG. 12, the statistical model M2 trained only with the object detection sample cannot detect the correct class "Ruri no Diko" not included in the object detection sample. As shown in FIG. 11, the statistical model M1 trained with both the object detection sample and the image classification sample can detect the position coordinates of the correct class "Ruri no Diko" included in the image classification sample.
[0063] (Example 4) In Examples 1 to 3, the machine learning of the statistical model of the VQA task using samples having a format according to some task was described. In Example 4, the machine learning of the statistical model of the VQA task using samples having a format unrelated to the task (hereinafter referred to as non-task samples) will be described. The non-task samples also have, as their elements, an image and a label related to the image. The label may be anything as long as it is a character string related to the content of the image. In the following description, it is assumed that the label is a description text (caption) for the image. A non-task sample including an image and a caption is called an image caption sample. Note that the same elements as those in Examples 1 to 3 are denoted by the same reference numerals and the description thereof is omitted.
[0064] FIG. 13 is a diagram schematically showing the machine learning process according to Example 4. The machine learning according to Example 4 is a machine learning process for a statistical model of the VQA task using the VQA sample 31 and the image caption sample 33 which is a non-task sample. As shown in FIG. 13, the image caption sample 33 includes a combination of an image and a caption. For example, the caption of the images illustrated in FIGS. 11 and 12 is “Rurinojiko is stopping on the signboard”.
[0065] As shown in FIG. 13, a question sentence and a correct answer sentence are generated by the conversion unit 112 based on the caption. When an image and a caption are given, it is known that a task (masked language modeling) of predicting the masked word from the caption in which a part is masked is effective (Li, Liunian Harold and Yatskar, Mark and Yin, Da and Hsieh, Cho-Jui and Chang, Kai-Wei, “Visualbert: A simple and performant baseline for vision and language”, arXiv preprint arXiv:1908.03557, 2019).
[0066] The conversion unit 112 randomly selects and masks some of the words that make up the caption. The words to be masked can be of any part of speech, but it is preferable to select them from proper nouns and verbs that are easy to draw in the image. The caption after masking is set as the question sentence, and the masked word is set as the correct answer sentence. Taking the case where the original caption is "A man is walking along a white house" as a specific example. In this case, as an example, the question sentence is "[Mask] A man is walking along a white *", and the correct answer sentence is "house". The [Mask] in the question sentence is an indicator indicating that it is a masked language modeling task, and * represents the mask. By generating the question sentence and the correct answer sentence from the caption according to such rules, it becomes possible to learn the masked language modeling task simultaneously with the VQA task without switching samples or heads.
[0067] (Modification Example 1) In the above embodiment, it is assumed that the modality of the object included in various samples is an image. However, the modality of the object in this embodiment is not limited to this, and it is also applicable to videos, voices, measuring instrument outputs, and / or 3D point clouds. A video means time-series image data collected by a video camera or the like. A voice means time-series voice data collected by a microphone or the like. A measuring instrument output means time-series data of measured values output from various measuring instruments. Examples of the measuring instruments include pressure gauges, thermometers, voltmeters, ammeters, etc. attached to various devices constituting a generator. A 3D point cloud means 3D data of a plurality of sample points on an object such as LIDAR (Light Detection and Ranging).
[0068] FIG. 14 is a diagram showing an example of the network configuration of the statistical model M3 when the modality of the object is a video. The statistical model M3 includes a video encoder M31, a text encoder M32, a fusion unit M33, and an answer sentence converter M34. The video encoder M31 is an encoding network layer that outputs the feature amount of the input video (hereinafter referred to as the video feature amount). The text encoder M32 is an encoding network layer that outputs the text feature amount of the input question sentence. The fusion unit M33 outputs a fusion feature amount of the video feature amount output from the video encoder M31 and the text feature amount output from the text encoder M32. The fusion unit M33 may be a module that performs addition operations or concatenation, or may be an encoding network layer. The answer sentence converter M34 is a decoding network layer that converts the fusion feature amount output from the fusion unit M33 into a language sequence of natural language representing the answer sentence.
[0069] FIG. 15 is a diagram showing an example of the network configuration of the statistical model M4 when the modality of the object is audio. The statistical model M4 includes an audio encoder M41, a text encoder M42, a fusion unit M43, and an answer sentence converter M44. The audio encoder M41 is an encoding network layer that outputs the feature amount of the input audio (hereinafter referred to as the audio feature amount). The text encoder M42 is an encoding network layer that outputs the text feature amount of the input question sentence. The fusion unit M43 outputs a fusion feature amount of the audio feature amount output from the audio encoder M41 and the text feature amount output from the text encoder M42. The fusion unit M43 may be a module that performs addition operations or concatenation, or may be an encoding network layer. The answer sentence converter M44 is a decoding network layer that converts the fusion feature amount output from the fusion unit M43 into a language sequence of natural language representing the answer sentence.
[0070] FIG. 16 is a diagram showing an example of the network configuration of the statistical model M5 when the modality of the object is a three-dimensional point cloud. The statistical model M5 has a three-dimensional point cloud encoder M51, a text encoder M52, a fusion unit M53, and an answer sentence converter M54. The three-dimensional point cloud encoder M51 is an encoding network layer that outputs the feature amount of the input three-dimensional point cloud (hereinafter referred to as the three-dimensional point cloud feature amount). The text encoder M52 is an encoding network layer that outputs the text feature amount of the input question sentence. The fusion unit M53 outputs a fusion feature amount of the three-dimensional point cloud feature amount output from the three-dimensional point cloud encoder M51 and the text feature amount output from the text encoder M52. The fusion unit M53 may be a module that performs an addition operation or concatenation, or may be an encoding network layer. The answer sentence converter M54 is a decoding network layer that converts the fusion feature amount output from the fusion unit M53 into a language sequence of natural language representing the answer sentence.
[0071] As described above, the statistical model according to this embodiment can also handle various modalities. In the case of time-series data such as videos, voices, and outputs of measuring instruments, it is also possible to generate question sentences and correct answer sentences regarding the time axis. For example, in the case of a video from a security camera, a question sentence "During which period is there a man wearing a mask?" and a correct answer sentence "14:03 to 14:14" may be generated. For the time, it is possible to use, for example, time stamps associated with each frame.
[0072] (Modification 2) Also, in the above embodiment, it was assumed that the language used for question sentences, answer sentences, etc. in the VQA task is Japanese. However, there is no restriction on the type of language according to this embodiment, and for example, English, Chinese, Korean, German, Dutch, Portuguese, Spanish, French, etc. may also be used.
[0073] (Inference device) FIG. 17 is a diagram showing a configuration example of the inference device 2 according to the present embodiment. As shown in FIG. 17, the inference device 2 is a computer having a processing circuit 21, a storage device 22, an input device 23, a communication device 24, and a display device 25. Data communication between the processing circuit 21, the storage device 22, the input device 23, the communication device 24, and the display device 25 is performed via a bus.
[0074] The processing circuit 21 includes a processor such as a CPU and a memory such as a RAM. The processing circuit 21 includes an acquisition unit 211, a conversion unit 212, an inference unit 213, and a display control unit 214. The processing circuit 21 realizes the functions of the respective units 211 to 214 by executing an inference program. The inference program is stored in a non-transitory computer-readable recording medium such as the storage device 22. The inference program may be implemented as a single program that describes all the functions of the respective units 211 to 214, or may be implemented as a plurality of modules divided into several functional units. Further, the respective units 211 to 214 may be implemented by an integrated circuit such as an application-specific integrated circuit or an FPGA. In this case, it may be implemented by a single integrated circuit, or may be individually implemented by a plurality of integrated circuits.
[0075] The acquisition unit 211 acquires an object to be processed. The object to be processed means an object to be subjected to inference processing by a statistical model of the VQA task trained according to the various embodiments described above. Although an image or a video is typical for the object, it is not limited thereto, and data obtained by various modalities such as voice, measuring instrument output, and / or 3D point cloud may be used. A corresponding question sentence may or may not be generated for the object to be processed. When it is generated, the question sentence is associated with the object to be processed.
[0076] When a question sentence is not generated for the object to be processed, the conversion unit 212 generates a question sentence regarding the object to be processed. That is, the conversion unit 212 converts the object into a format for inference of the VQA task.
[0077] The inference unit 213 applies an object and a question sentence regarding the object to a statistical model of the VQA task to infer an answer sentence for the question sentence.
[0078] The display control unit 214 displays various information on the display device 25. For example, the display control unit 214 displays the inference result of the VQA task obtained by the inference unit 213 and the like.
[0079] The storage device 22 is composed of a ROM, an HDD, an SSD, an integrated circuit storage device, etc. The storage device 22 stores an inference program, a statistical model of the VQA task, and the like.
[0080] The input device 23 inputs various commands from the operator. As the input device 23, a keyboard, a mouse, various switches, a touch pad, a touch panel display, etc. can be used. The output signal from the input device 23 is supplied to the processing circuit 21. Note that the input device 23 may be an input device of a computer connected to the processing circuit 21 via wire or wirelessly.
[0081] The communication device 24 is an interface for performing data communication between the inference device 2 and an external device connected via a network. It is a computer that stores an object to be processed, various collection devices that collect the object, and the like.
[0082] The display device 25 displays various information under the control of the display control unit 214. As the display device 25, a CRT display, a liquid crystal display, an organic EL display, an LED display, a plasma display, or any other display known in the art can be appropriately used. Further, the display device 25 may be a projector.
[0083] Hereinafter, the inference device 2 will be described in detail. In the following description, it is assumed that the object is an image.
[0084] FIG. 18 is a diagram illustrating the processing procedure of the inference process by the inference device 2. As shown in FIG. 18, the acquisition unit 211 acquires an image to be processed (step S1801). The image to be processed may or may not have a question sentence generated.
[0085] When step S1801 is performed, the conversion unit 212 determines whether there is a question sentence for the image acquired in step S1801 (step S1802). As an example, when a question sentence is associated with the image to be processed, the conversion unit 212 determines that there is a question sentence, and when a question sentence is not associated with the image to be processed, the conversion unit 212 may determine that there is no question sentence. Note that the operator may input the presence or absence of a question sentence via the input device 23.
[0086] When it is determined in step S1802 that there is no question sentence (step S1802: NO), the conversion unit 212 generates a question sentence for the image acquired in step S1801 (step S1803). As an example, the conversion unit 212 generates a stereotyped question sentence. As the stereotyped question sentence, a general-purpose question sentence that does not depend on the content of the image to be processed is preferable. For example, "What is this?" is suitable as the stereotyped question sentence.
[0087] As another example, the conversion unit 212 may generate a question sentence based on the label associated with the image. Further, when the conversion unit 212 uses VQA sample or non-VQA sample images as the processing target, similar to the conversion unit 112, the conversion unit 212 may generate a question sentence based on the correct label included in the VQA sample or non-VQA sample.
[0088] When step S1803 is performed or when it is determined that there is a question sentence in step S1802 (step S1802: YES), the inference unit 213 applies the image and the question sentence to a statistical model of the VQA task to infer an answer sentence (step S1804). More specifically, the inference unit 21 applies the image and the question sentence to the statistical model to calculate a plurality of logit series respectively corresponding to a plurality of words constituting the inference answer sentence. Then, the inference unit 21 identifies a registered word ID having the maximum logit from each of the plurality of logit series, and applies the identified registered word ID to the tokenizer dictionary to convert it into a language sequence of the registered word. The answer sentence converter M14 predicts a character string representing the inference answer sentence by converting and combining all the logit series into the language sequence of the registered word.
[0089] When step S1804 is performed, the display control unit 214 displays the inference answer sentence obtained in step S1804 on the display device 25 (step S1805). As an example, the display control unit 214 may display not only the inference answer sentence but also the image, the question sentence, and the inference answer sentence on one screen.
[0090] As described above, the inference process by the inference device 2 ends. Since the statistical model according to the present embodiment can be trained using additional samples obtained by converting non-VQA samples into the VQA format, it can be trained based on a large number of samples, and thus the inference accuracy is improved. In addition, since it is possible to automatically generate a question sentence, it is possible to reduce the burden of generating the question sentence.
[0091] Note that the inference process shown in FIG. 18 is an example, and the inference process according to this embodiment is not limited to this only. As described in the above embodiments, the statistical model according to this embodiment can process image classification tasks, object detection tasks, image grounding tasks, image search tasks, etc. in the form of VQA tasks. Therefore, as stereotypical question sentences, those related to the content corresponding to these tasks may be prepared. At this time, the generation of question sentences and the inference of answer sentences may be performed cyclically so as to generate question sentences again from the inference answer sentences for the basic question sentence "What is this?". For example, when an image of a "lawn mower" and a question sentence "What is this?" are applied to the statistical model, an inference answer sentence "lawn mower" is output. Based on the inference answer sentence "lawn mower", an answer sentence corresponding to an object detection task, an image grounding task, or an image search task is generated. For example, as a stereotypical question sentence similar to an object detection task, "Where is the lawn mower?" etc. are appropriate. As a stereotypical question sentence similar to an image grounding task, "What is under the lawn mower?" is appropriate. By cyclically performing the generation of question sentences and the inference of answer sentences in this way, it becomes possible to answer questions corresponding to various tasks learned by the statistical model. The types of stereotypical question sentences may be arbitrarily selected by the operator via the input device 23, or may be selected comprehensively.
[0092] Thus, according to this embodiment, it becomes possible to provide a machine learning device, a machine learning method, a machine learning program, and an inference device that can efficiently learn a statistical model for VQA tasks.
[0093] Although several embodiments of the present invention have been described, these embodiments are presented as examples and are not intended to limit the scope of the invention. These novel embodiments can be implemented in various other forms, and various omissions, replacements, and changes can be made without departing from the gist of the invention. These embodiments and their modifications are included in the scope and gist of the invention, and are included in the invention described in the claims and its equivalent scope.
Explanation of Reference Numerals
[0094] 1... Machine learning device, 11... Processing circuit, 12... Memory device, 13... Input device, 14... Communication device, 15... Display device, 31... VQA sample, 32... Non-VQA sample, 33... Image caption sample, 51... Image, 52... Question text, 53... Correct answer text, 54... Image, 55... Correct label, 56... Question text, 57... Correct answer text, 58... Predicted answer text, 61... Image, 62... Question text, 63... Correct answer text, 64... Image, 65... Correct label, 66... Correct class, 91... Image, 92... Correct label, 93... Question text, 94... Correct answer text, 101... Predicted answer text, 111... Acquisition unit, 112... Conversion unit, 113... Learning unit, 114... Display control unit.
Claims
1. A conversion unit that generates a learning sample in a VQA format for a VQA task based on a sample in a non-VQA (Visual Question Answering) format, wherein the learning sample in the VQA format has, as elements, a combination of an object, a question sentence regarding the object, and an answer sentence for the question sentence, and the sample in the non-VQA format has, as elements, a combination of an object and a label related to the object, A learning unit that trains a statistical model of the VQA task based on the learning sample in the VQA format generated by the conversion unit, Comprising, The VQA task is an image detection task, The sample in the format of the image detection task among the non-VQA formats includes, as the label, the correct position and correct size of a rectangle surrounding an object shown in the image and the correct class of the object, The conversion unit generates the question sentence and the answer sentence based on the correct position and / or correct size and the correct class, A machine learning device.
2. The statistical model is, An encoder that converts the object into a first feature amount, An encoder that converts the answer sentence into a second feature amount, A fuser that generates a fused feature amount of the first feature amount and the second feature amount, A converter that converts the fused feature amount into a character string of natural language representing the answer sentence, The machine learning device according to Claim 1.
3. The converter according to Claim 2, which converts the fused feature amount into a series of relative values representing the occurrence probability of each word constituting the answer sentence.
4. A step in which a computer generates a learning sample in a VQA format for a VQA task based on a sample in a non-VQA (Visual Question Answering) format, wherein the learning sample in the VQA format has, as elements, a combination of an object, a question sentence regarding the object, and an answer sentence for the question sentence, and the sample in the non-VQA format has, as elements, a combination of an object and a label related to the object, a conversion step, A learning step in which a computer trains a statistical model of the VQA task based on the learning sample in the VQA format generated in the conversion step, Comprising, The VQA task is an image detection task, Among the non-VQA formats, the sample of the format of the image detection task includes, as the label, the correct position and correct size of a rectangle surrounding an object shown in the image and the correct class of the object, The conversion step generates the question text and the answer text based on the correct position and / or correct size and the correct class, Machine learning method.
5. A computer, A conversion function for generating a learning sample of a VQA format related to a VQA task based on a sample in a non-VQA (Visual Question Answering) format, wherein the learning sample of the VQA format has, as elements, a combination of an object, a question text related to the object, and an answer text for the question text, and the sample in the non-VQA format has, as elements, a combination of an object and a label related to the object, A learning function for training a statistical model of the VQA task based on the learning sample of the VQA format generated by the conversion function, To be realized, The VQA task is an image detection task, Among the non-VQA formats, the sample of the format of the image detection task includes, as the label, the correct position and correct size of a rectangle surrounding an object shown in the image and the correct class of the object, The conversion function generates the question text and the answer text based on the correct position and / or correct size and the correct class, Machine learning program.
6. An inference device comprising: an inference unit that applies an object and a question text related to the object to a statistical model of a VQA (Visual Question Answering) task generated by the machine learning device according to any one of Claims 1 to 3, and infers an answer text for the question text; A display control unit that displays the answer text; Inference device.
7. The inference device according to Claim 6, further comprising a conversion unit that generates the question text based on a label associated with the object.
8. The inference device according to Claim 6, further comprising a conversion unit that generates the stereotypical question text.
Citation Information
Patent Citations
Conversion processor, transliteration processor, and program
JP2018028848A
Image processing device, image processing method, and recording medium
WO2020179065A1
Information processing method, information processing device, and information processing program
WO2021182199A1