Learning device, training method, and program

JPWO2024185054A5Pending Publication Date: 2025-11-10
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
JP2025504974
Authority / Receiving Office
JP · JP
Patent Type
Applications
Filing Date
2025-08-27
Publication Date
2025-11-10

AI Technical Summary

Technical Problem

Object recognition models trained on specific datasets may experience decreased recognition accuracy when encountering images significantly different from the training data, leading to erroneous object detection.

Method used

A learning device and method that extracts and utilizes images similar to those misrecognized by the model, generating new training data by excluding correctly recognized images and incorporating additional data to retrain the model, ensuring it can recognize objects accurately across varying distributions.

Benefits of technology

Prevents the decrease in recognition accuracy by retraining the model with diverse and relevant data, enhancing its ability to identify objects in images with different distributions.

✦ Generated by Eureka AI based on patent content.
Patent Text Reader

Abstract

A learning device, wherein an extraction means extracts, from an image group including at least one image, a first image that satisfies a use condition for use in training a model that recognizes an object included in the image. A training means trains the model using training data that includes the first image.
Need to check novelty before this filing date? Find Prior Art

Description

Learning device, training method, and recording medium

[0001] The present disclosure relates to the art of object recognition.

[0002] In the field of machine learning, object recognition models that are trained to recognize objects contained in images have been proposed in recent years.

[0003] For example, Patent Literature 1 discloses a technique for generating a training dataset used for retraining an object detection model that has already been trained.

[0004] International Publication No. WO2020 / 202636

[0005] However, the technology disclosed in Patent Document 1 has a problem in that, for example, when an image that is significantly different from the group of images included in a training dataset is input to an object detection model that has been retrained using the training dataset, the object included in the image may be falsely detected.

[0006] That is, according to the technology disclosed in Patent Document 1, there is a risk that the recognition accuracy may decrease when performing object recognition, which is a problem corresponding to the above-mentioned problem.

[0007] One object of the present disclosure is to provide a learning device that can prevent a decrease in recognition accuracy that occurs when performing object recognition.

[0008] In one aspect of the present disclosure, a learning device includes an extraction means for extracting a first image from a group of images including at least one image, the first image satisfying usage conditions for use in training a model that recognizes an object included in the image, and a training means for training the model using training data including the first image.

[0009] In another aspect of the present disclosure, a training method includes extracting a first image from a group of images including at least one image, the first image satisfying usage conditions for use in training a model that recognizes an object contained in the image, and training the model using training data including the first image.

[0010] In yet another aspect of the present disclosure, a recording medium stores a program that causes a computer to execute a process of extracting a first image from a group of images including at least one image, the first image satisfying usage conditions for use in training a model that recognizes an object contained in the image, and training the model using training data including the first image.

[0011] According to the present disclosure, it is possible to prevent a decrease in recognition accuracy that occurs when performing object recognition.

[0012] 1 is a block diagram showing the hardware configuration of an information processing device according to a first embodiment. 2 is a block diagram showing the functional configuration of an information processing device according to the first embodiment. 3 is a diagram showing an overview of data changes due to processing by the information processing device according to the first embodiment. 4 is a flowchart showing an example of processing performed in the information processing device according to the first embodiment. 5 is a block diagram showing the functional configuration of an information processing device according to a second embodiment. 6 is a diagram showing an overview of data changes due to processing by the information processing device according to the second embodiment. 7 is a diagram showing an example of table data used in processing by the information processing device according to the second embodiment. 8 is a flowchart showing an example of processing related to reality determination performed in the information processing device according to the second embodiment. 9 is a flowchart showing an example of processing performed in the information processing device according to the second embodiment. 10 is a block diagram showing the functional configuration of a learning device according to a third embodiment. 11 is a flowchart for explaining processing performed in the learning device according to the third embodiment.

[0013] Hereinafter, preferred embodiments of the present disclosure will be described with reference to the drawings.

[0014] 1 is a block diagram showing the hardware configuration of an information processing device according to Embodiment 1. As shown in FIG. 1, the information processing device 100 has an interface (IF) 111, a processor 112, a memory 113, a recording medium 114, and a database (DB) 115.

[0015] The IF 111 inputs and outputs data to and from an external device. For example, an image (or video) captured by a camera or the like is input to the information processing device 100 through the IF 111.

[0016] The processor 112 is a computer such as a CPU (Central Processing Unit), and executes a program prepared in advance to control the entire information processing device 100. Specifically, the processor 112 performs, for example, processing related to training an object recognition model that has been trained to be able to recognize objects included in an image.

[0017] The memory 113 is configured by a ROM (Read Only Memory), a RAM (Random Access Memory), etc. The memory 113 is also used as a working memory while the processor 112 is executing various processes.

[0018] The recording medium 114 is a non-volatile, non-transitory recording medium such as a disk-shaped recording medium or a semiconductor memory, and is configured to be detachable from the information processing device 100. The recording medium 114 records various programs to be executed by the processor 112. When the information processing device 100 executes various processes, the programs recorded on the recording medium 114 are loaded into the memory 113 and executed by the processor 112.

[0019] The DB 115 stores, for example, images input via the IF 111 and processing results obtained by processing by the processor 112 .

[0020] [Functional Configuration] Fig. 2 is a block diagram showing the functional configuration of the information processing device according to the first embodiment. The information processing device 100 has a function as a learning device. As shown in Fig. 2, the information processing device 100 also has a data acquisition unit 11, a training processing unit 12, and an object recognition processing unit 13.

[0021] The data acquisition unit 11 acquires, as input data, a dataset DAE including a plurality of data sets in which images captured by a camera or the like are associated with labels corresponding to character strings representing objects included in the images. The data acquisition unit 11 also acquires, as input data, a dataset DAF different from the dataset DAE, which includes a plurality of data sets in which images captured by a camera or the like are associated with labels corresponding to character strings representing objects included in the images. The data acquisition unit 11 also outputs the dataset DAE and the dataset DAF to the training processing unit 12. The data acquisition unit 11 also acquires, as input data, images captured by a camera or the like during a period after retraining of the object recognition model 13A, which will be described later, is completed.

[0022] The data acquisition unit 11 may acquire, for example, data received from outside the information processing device 100 as input data. Alternatively, the data acquisition unit 11 may acquire, for example, data read from the DB 115 as input data. It is desirable that the data acquisition unit 11 acquire the dataset DAF so that the number of datasets is smaller than the number of datasets in the dataset DAE. It is also desirable that the data acquisition unit 11 acquire, as the dataset DAF, a dataset having a distribution different from the distribution of the dataset DAE, or a dataset including Out Of Distribution (OOD) data that falls outside the distribution of the dataset DAE.

[0023] The training processing unit 12 uses the data sets DAE and DAF output from the data acquisition unit 11 to perform processing related to training of the object recognition processing unit 13. Furthermore, during a period after retraining of the object recognition model 13A is completed, the training processing unit 12 outputs images included in the input data obtained by the data acquisition unit 11 to the object recognition processing unit 13. Furthermore, the training processing unit 12 has an object recognition model training unit 12A, a similar image generation unit 12B, and an additional data acquisition unit 12C.

[0024] The object recognition model training unit 12A functions as a training means. The object recognition model training unit 12A uses the dataset DAE as training data TDE to train the object recognition model 13A of the object recognition processing unit 13. After training the object recognition model 13A using the training data TDE, the object recognition model training unit 12A retrains the object recognition model 13A using new training data TDM obtained by adding additional data TDF (described later) to the training data TDE.

[0025] The similar image generation unit 12B functions as an acquisition unit. The similar image generation unit 12B uses the data set DAF output from the data acquisition unit 11 to generate a plurality of similar images that resemble an image in which an object is erroneously recognized by the object recognition model 13A.

[0026] The additional data acquisition unit 12C functions as an extraction unit, and acquires additional data TDF by excluding, from the plurality of similar images generated by the similar image generation unit 12B, images that satisfy a predetermined exclusion condition that the object recognition model 13A correctly recognizes the object.

[0027] The object recognition processing unit 13 has an object recognition model 13A that has been trained to recognize objects included in images. In other words, the object recognition processing unit 13 is configured to perform processing for recognizing objects included in images using the object recognition model 13A. Furthermore, during a period until retraining of the object recognition model 13A is completed, the object recognition processing unit 13 outputs recognition results (such as the recognition results NAF and RCF described below) obtained by processing using the object recognition model 13A to the training processing unit 12. Furthermore, during a period after retraining of the object recognition model 13A is completed, the object recognition processing unit 13 inputs images output from the training processing unit 12 into the object recognition model 13A to obtain recognition results of objects included in the images, and outputs the obtained recognition results to the outside of the information processing device 100 and / or to the DB 115. The object recognition model may also be interpreted as AI (artificial intelligence) for object detection.

[0028] [Specific Example of Training Processing] A specific example of the processing for training the object recognition model 13A performed by the training processing unit 12 will now be described with appropriate reference to Fig. 3. Fig. 3 is a diagram showing an overview of data transitions due to processing by the information processing device according to the first embodiment.

[0029] The object recognition model training unit 12A uses the data set DAE as training data TDE to train the object recognition model 13A (see FIG. 3).

[0030] The similar image generation unit 12B inputs an image PAF included in the dataset DAF to the object recognition model 13A, which has been trained using at least the training data TDE, to obtain a recognition result NAF of the object included in the image PAF. Furthermore, if a label LAF associated with an image PAF and the recognition result NAF represent the same object, the similar image generation unit 12B excludes a dataset DFX associated with the image PAF and the label LAF from the dataset DAF. Furthermore, the similar image generation unit 12B obtains each dataset, excluding the dataset DFX from the dataset DAF, as a dataset DBF (see FIG. 3 ).

[0031] The similar image generation unit 12B inputs an image PBF included in a data set DBF into an image-to-text model, thereby generating an explanatory sentence MBF that explains the image PBF in natural language. Also, the similar image generation unit 12B inputs the explanatory sentence MBF into the text-to-image model, thereby generating multiple similar images according to the explanatory sentence MBF.

[0032] The image-to-text model and the text-to-image model are disclosed, for example, in "Discovering Bugs in Vision Models using Off-the-shelf Image Generation and Captioning" by Olivia Wiles, et al.

[0033] According to the above-described processing, the similar image generation unit 12B can generate a plurality of similar images that are similar to one image PBF in which an object has been erroneously recognized by the object recognition model 13 A. Furthermore, according to the above-described processing, the similar image generation unit 12B can generate similar images based on an explanatory text MBF that explains one image PBF.

[0034] The similar image generating unit 12B acquires a dataset DCF by associating the same label LBF as the label associated with one image PBF with each of the plurality of similar images generated as described above (see FIG. 3).

[0035] The additional data acquisition unit 12C inputs one similar image SCF included in the dataset DCF to the object recognition model 13A trained using at least the training data TDE, thereby acquiring a recognition result RCF for the object included in the one similar image SCF. Furthermore, if a label LBF associated with one similar image SCF included in the dataset DCF and the recognition result RCF represent the same object, the additional data acquisition unit 12C excludes a dataset DFY associated with the one similar image SCF and the label LBF from the dataset DCF. Furthermore, the additional data acquisition unit 12C acquires each dataset, excluding dataset DFY, from the dataset DCF as additional data TDF (see FIG. 3 ).

[0036] According to the processing described above, the additional data acquisition unit 12C can acquire additional data TDF by excluding, from each similar image generated by the similar image generation unit 12B, similar images that satisfy a predetermined exclusion condition that the object recognition model 13A has correctly recognized an object. Furthermore, according to the processing described above, the additional data acquisition unit 12C can extract similar images that satisfy a usage condition for use in retraining the object recognition model 13A from an image group containing a plurality of similar images generated by the similar image generation unit 12B, and acquire additional data TDF that includes the extracted similar images. The aforementioned usage condition may be a condition that the object recognition model 13A has incorrectly recognized an object.

[0037] The object recognition model training unit 12A retrains the object recognition model 13A using new training data TDM obtained by adding additional data TDF to the training data TDE (see FIG. 3).

[0038] The process related to retraining of the object recognition model 13A may be repeatedly performed until a predetermined termination condition is met. Specifically, the process related to retraining of the object recognition model 13A may be repeatedly performed, for example, until a condition is met that the number of retrainings reaches a predetermined number. Furthermore, the process related to retraining of the object recognition model 13A may be repeatedly performed, for example, until a condition is met that the number of data sets included in the additional data TDF is equal to or less than a predetermined number. Furthermore, the process related to retraining of the object recognition model 13A may be repeatedly performed, for example, until a condition is met that the difference between the number of data sets included in the training data TDM and the number of data sets included in the training data TDE is equal to or less than a certain value.

[0039] According to this specific example, for example, the data acquisition unit 11 may reacquire the data set DAF every time the object recognition model training unit 12A retrains the object recognition model 13A a predetermined number of times.

[0040] According to this specific example, as long as the similar image generation unit 12B generates a similar image similar to one image PBF, for example, a description MCF obtained by partially modifying a description MBF may be input to the text-to-image model. Specifically, according to this specific example, for example, a description MCF obtained by modifying at least one parameter of the description MBF, such as angle, color, and size, may be input to the text-to-image model. In other words, according to this specific example, the similar image generation unit 12B may acquire a similar image based on a description MCF obtained by modifying at least one parameter of the description MBF.

[0041] According to this specific example, for example, the similar image generation unit 12B may acquire similar images similar to a given image PBF by searching the Internet and / or a database using words and phrases included in the description MBF. Also, according to this specific example, for example, the similar image generation unit 12B may acquire an image obtained by modifying at least a portion of a given image PBF as a similar image similar to the given image PBF. Also, according to this specific example, the similar image generation unit 12B may generate or acquire one similar image similar to a given image PBF. That is, according to this specific example, it is sufficient that the similar image generation unit 12B acquires a group of images including at least one similar image similar to a given image PBF.

[0042] [Processing Flow] Next, a description will be given of the flow of processing performed in the information processing apparatus according to the first embodiment. Fig. 4 is a flowchart showing an example of processing performed in the information processing apparatus according to the first embodiment.

[0043] First, the information processing device 100 acquires, as input data, a dataset DAE and a dataset DAF that is different from the dataset DAE (step S11).

[0044] Next, the information processing device 100 uses the data set DAE acquired in step S11 as training data TDE to train the object recognition model 13A (step S12).

[0045] Next, the information processing device 100 uses the data set DAF acquired in step S11 to generate a plurality of similar images that are similar to the image in which the object recognition model 13A has erroneously recognized the object (step S13).

[0046] Next, the information processing device 100 acquires additional data TDF by excluding similar images that satisfy a predetermined exclusion condition from the similar images generated in step S13 (step S14). The predetermined exclusion condition may be set as a condition that the object recognition model 13A correctly recognizes the object.

[0047] Subsequently, the information processing device 100 retrains the object recognition model 13A using new training data TDM obtained by adding the additional data TDF obtained in step S14 to the training data TDE (step S15).

[0048] Next, the information processing device 100 determines whether or not a predetermined termination condition for retraining the object recognition model 13A has been met (step S16).

[0049] If the predetermined termination condition is not satisfied (step S16: NO), the information processing device 100 performs the processes from step S13 onwards again. Furthermore, if the predetermined termination condition is satisfied (step S16: YES), the information processing device 100 ends the series of processes related to training the object recognition model 13A. Furthermore, the information processing device 100 completes the retraining of the object recognition model 13A after performing the processes from step S13 to step S16 one or more times.

[0050] As described above, according to this embodiment, a data set including images in which the object recognition model 13A, which has been trained using the training data TDE, has erroneously recognized an object, can be acquired as the additional data TDF. Furthermore, as described above, according to this embodiment, the object recognition model 13A can be retrained using new training data TDM obtained by adding the additional data TDF to the training data TDE. Therefore, according to this embodiment, it is possible to prevent a decrease in recognition accuracy that occurs when performing object recognition.

[0051] Second Embodiment Next, a second embodiment of the present disclosure will be described. Note that, for simplicity, specific descriptions of parts to which the above-described processes and the like can be applied will be omitted as appropriate.

[0052] [Functional Configuration] Fig. 5 is a block diagram showing the functional configuration of an information processing device according to the second embodiment. The information processing device 200 has a function as a learning device. As shown in Fig. 5, the information processing device 200 also has a data acquisition unit 21, a training processing unit 22, an importance estimation processing unit 23, and an object recognition processing unit 24.

[0053] The data acquisition unit 21 acquires, as input data, a dataset DAE including multiple data sets in which images captured by a camera or the like are associated with labels corresponding to character strings representing objects included in the images. The data acquisition unit 21 also acquires, as input data, character strings MDG that express in natural language a field targeted for object recognition by the object recognition processing unit 24. The data acquisition unit 21 also acquires, as input data, a dataset DDG including multiple data sets configured as combinations of sentences and images related to the character string MDG. The data acquisition unit 21 acquires a dataset DDH by assigning a label and importance to each image included in the dataset DDG. The data acquisition unit 21 also outputs the dataset DAE and the dataset DDH to the training processing unit 22. The data acquisition unit 21 also acquires, as input data, images captured by a camera or the like during a period after retraining of the object recognition model 24A, described below, is completed.

[0054] The data acquisition unit 21 may acquire data received from outside the information processing device 200 as input data. Alternatively, the data acquisition unit 21 may acquire data read from the DB 115 as input data. It is preferable that the data acquisition unit 21 acquires the dataset DDG so that the number of datasets is less than the number of datasets in the dataset DAE. It is also preferable that the data acquisition unit 21 acquires, as the dataset DDG, a dataset having a distribution different from that of the dataset DAE, or a dataset including OOD data that deviates from the distribution of the dataset DAE.

[0055] The training processing unit 22 uses the data sets DAE and DDH output from the data acquisition unit 21 to perform processing related to training of the importance estimation processing unit 23 and the object recognition processing unit 24. Furthermore, during a period after retraining of the object recognition model 24A is completed, the training processing unit 22 outputs images included in the input data obtained by the data acquisition unit 21 to the object recognition processing unit 24. Furthermore, the training processing unit 22 has an importance estimation model training unit 22A, an object recognition model training unit 22B, a similar image generation unit 22C, a reality determination processing unit 22D, and an additional data acquisition unit 22E.

[0056] The importance estimation model training unit 22A acquires a dataset DDI including a data set that does not overlap with a dataset DDK, which will be described later, from the dataset DDH output from the data acquisition unit 21. Furthermore, the importance estimation model training unit 22A uses the dataset DDI as training data TDI to train an importance estimation model 23A of the importance estimation processing unit 23.

[0057] The object recognition model training unit 22B functions as a training means. Furthermore, the object recognition model training unit 22B acquires, from the dataset DDH output from the data acquisition unit 21, a dataset DDK including a data set that does not overlap with the dataset DDI. Furthermore, the object recognition model training unit 22B uses, as training data TDL, the dataset DAE and a dataset DDL corresponding to each piece of data included in the dataset DDK with the importance level removed, to train the object recognition model 24A of the object recognition processing unit 24. Furthermore, after training the object recognition model 24A using the training data TDL, the object recognition model training unit 22B retrains the object recognition model 24A using new training data TDN that is different from the training data TDL.

[0058] The similar image generating unit 22C functions as an acquisition unit, and generates a plurality of similar images that are similar to the images included in the data set DDK.

[0059] The reality determination processing unit 22D functions as a determination unit, and determines whether each similar image generated by the similar image generating unit 22C is a realistic image.

[0060] The additional data acquisition unit 22E functions as an extraction unit. The additional data acquisition unit 22E acquires additional data by eliminating, from among the similar images generated by the similar image generation unit 22C, similar images that satisfy a first exclusion condition of being unrealistic and similar images that satisfy a second exclusion condition of the object recognition model 24A correctly recognizing the object. The additional data acquisition unit 22E acquires a dataset DDX by adding the aforementioned additional data to a dataset DDK. The additional data acquisition unit 22E also acquires, as a dataset DDM, the datasets included in the dataset DDX from which datasets that satisfy a third exclusion condition of being relatively unimportant have been eliminated. The dataset DDM can be used as a dataset that serves as the basis for new training data TDN.

[0061] The importance estimation processing unit 23 has an importance estimation model 23A that has been trained so as to be able to estimate the importance of an image. In other words, the importance estimation processing unit 23 is configured to be able to perform processing for estimating the importance of an image using the importance estimation model 23A. Note that the importance estimation model may also be read as AI for importance estimation.

[0062] The object recognition processing unit 24 has an object recognition model 24A that has been trained to be able to recognize objects included in images. In other words, the object recognition processing unit 24 is configured to be able to perform processing for recognizing objects included in images using the object recognition model 24A. Furthermore, during a period until retraining of the object recognition model 24A is completed, the object recognition processing unit 24 outputs recognition results (such as the recognition results NDL and RDX described below) obtained by processing using the object recognition model 24A to the training processing unit 22. Furthermore, during a period after retraining of the object recognition model 24A is completed, the object recognition processing unit 24 inputs images output from the training processing unit 22 into the object recognition model 24A to obtain recognition results of objects included in the images, and outputs the obtained recognition results to the outside of the information processing device 200 and / or the DB 115.

[0063] [Specific Example of Training Processing] A specific example of the processing for training the object recognition model 24A performed by the training processing unit 22 etc. will now be described with appropriate reference to Fig. 6. Fig. 6 is a diagram showing an overview of data transitions due to processing by the information processing device according to the second embodiment.

[0064] The data acquisition unit 21 acquires as input data a dataset DDG including multiple data sets configured as combinations of sentences and images related to the string MDG, for example, by searching the Internet and / or a database using words contained in the string MDG.

[0065] According to this embodiment, for example, it is desirable that a character string written in a natural language by a user on a device other than the information processing device 200 is input as a character string MDG to the data acquisition unit 21. Note that hereinafter, unless otherwise specified, a case will be described in which the character string MDG "bus accident" is input to the data acquisition unit 21.

[0066] The data acquisition unit 21 may acquire, as input data, a combination of a sentence and an image representing a real-world object and / or event related to the character string MDG. Specifically, the data acquisition unit 21 may acquire, as input data, a combination of a news image used in news related to the character string MDG, for example, "bus accident," and a news article describing the news.

[0067] The data acquisition unit 21 performs processing for associating labels with images included in each dataset of the dataset DDG based on the dataset DDG. Specifically, the data acquisition unit 21 acquires multiple clusters by classifying multiple datasets included in the dataset DDG so that similar datasets belong to the same cluster. Furthermore, the data acquisition unit 21 generates labels by summarizing sentences included in one of the datasets classified into one cluster CLM among the multiple clusters acquired as described above, and associates the generated labels with images included in each dataset. According to this processing, the data acquisition unit 21 can associate the same label with images included in each dataset classified into one cluster CLM. Furthermore, according to the above processing, when the data acquisition unit 21 acquires a dataset DDG corresponding to a character string MDG of "bus accident," for example, the data acquisition unit 21 can acquire a dataset in which one image in the dataset DDG is associated with a label of "bus overturn" and another dataset in which another image in the dataset DDG is associated with a label of "bus-car collision."

[0068] The data acquisition unit 21 performs processing for setting the importance of images included in each data set of the data set DDG based on the data set DDG. Specifically, the data acquisition unit 21 sets the importance of images included in one data set based on table data TBD as shown in Fig. 7 and sentences included in the one data set of the data set DDG, for example. Fig. 7 is a diagram showing an example of table data used in processing by the information processing device according to the second embodiment.

[0069] The table data TBD is configured as data representing, for example, a correspondence between a word or phrase related to the character string MDG and the importance of the word or phrase. Furthermore, the importance in the table data TBD is preferably set as a value indicating whether the word or phrase is likely to attract attention in the field represented by the character string MDG. Specifically, in the table data TBD, for example, the importance of a word or phrase estimated to attract attention in the field represented by the character string MDG is preferably set as a relatively high value, and the importance of a word or phrase estimated to attract little attention in the field represented by the character string MDG is preferably set as a relatively low value.

[0070] According to the processing using the table data TBD of FIG. 7 , for example, when a sentence in a data set contains the phrase "dead person," the data acquiring unit 21 can set the importance of an image included in the data set to "10." Also, according to the processing using the table data TBD of FIG. 7 , for example, when a sentence in a data set contains the phrase "injured person," the data acquiring unit 21 can set the importance of an image included in the data set to "3." Also, according to the processing using the table data TBD of FIG. 7 , for example, when a sentence in a data set contains two phrases, "dead person" and "injured person," the data acquiring unit 21 can set the importance of an image included in the data set to "13." In other words, when a sentence in a data set contains the same phrase as one included in the table data TBD, the data acquiring unit 21 can set the sum of the importance levels associated with the phrase in the table data TBD as the importance of the image included in the data set. In addition, if a sentence in a data set does not contain the same phrase as that contained in the table data TBD, the data acquisition unit 21 can set the importance of the image contained in the data set to "0".

[0071] The data acquisition unit 21 performs the processes described above to acquire a dataset DDH in which each image included in the dataset DDG is assigned a label and an importance level (see FIG. 6 ). The data acquisition unit 21 also outputs the dataset DAE and the dataset DDH to the training processing unit 22.

[0072] The importance estimation model training unit 22A uses the dataset DDI acquired from the dataset DDH as training data TDI to train the importance estimation model 23A (see FIG. 6 ). According to this processing, the importance estimation model training unit 22A can train the importance estimation model 23A so as to estimate the importance of an image based on the image and its label input to the importance estimation processing unit 23.

[0073] The object recognition model training unit 22B acquires a dataset DDK from the dataset DDH (see FIG. 6 ). The object recognition model training unit 22B also trains an object recognition model 24A of the object recognition processing unit 24 using, as training data TDL, the dataset DAE and a dataset DDL obtained by excluding importance from each data set included in the dataset DDK (see FIG. 6 ).

[0074] The similar image generation unit 22C inputs an image PDK included in a data set DDK into an image-to-3D model, thereby generating a three-dimensional image QDK representing the three-dimensional shape of an object included in the image PDK. Furthermore, the similar image generation unit 22C inputs the three-dimensional image QDK into the 3D-to-image model, thereby generating multiple similar images corresponding to the three-dimensional image QDK. The similar image generation unit 22C may set the number of similar images generated corresponding to the three-dimensional image QDK based on, for example, the importance assigned to the image PDK. Specifically, the similar image generation unit 22C may set the number of similar images generated corresponding to the three-dimensional image QDK to (P × K) based on, for example, the importance P assigned to the image PDK and a constant K.

[0075] According to the processing described above, the similar image generation unit 22C can generate a similar image that is similar to a single image PDK based on a three-dimensional image QDK that represents the three-dimensional shape of an object contained in the single image PDK.

[0076] The reality determination processing unit 22D determines whether each similar image generated by the similar image generating unit 22C is a realistic image. The reality determination processing unit 22D also associates each similar image generated by the similar image generating unit 22C with a determination result regarding whether the image is realistic.

[0077] Here, the processing related to the reality determination performed in the reality determination processing unit 22D will be described with reference to the flowchart of Fig. 8. Fig. 8 is a flowchart showing an example of the processing related to the reality determination performed in the information processing device according to the second embodiment.

[0078] First, the reality assessment processing unit 22D inputs one of the similar images generated by the similar image generation unit 22C into an Image-to-Text model, thereby generating a first explanatory text MDK that explains the one similar image SDK in natural language (step S21).

[0079] Next, the reality assessment processing unit 22D generates an image RDL corresponding to the first description MDK by inputting the sentence "a realistic image that satisfies the MDK" generated using the first description MDK into the text-to-image model (step S22). Specifically, for example, if the first description MDK is "an overturned bus," the reality assessment processing unit 22D can generate an image RDL corresponding to the first description MDK by inputting the sentence "a realistic image that satisfies the overturned bus" into the text-to-image model.

[0080] Next, the reality assessment processing unit 22D inputs the image RDL to the object recognition model 24A that has been trained using at least the training data TDL, thereby obtaining a recognition result NDL related to an object included in the image RDL. Then, based on the recognition result NDL, the reality assessment processing unit 22D determines whether the image RDL includes an object identical to the object represented by the first description MDK (step S23).

[0081] If the image RDL does not contain an object identical to the object represented by the first description MDK (step S23: NO), the reality judgment processing unit 22D associates the judgment result that the image is unrealistic with one similar image SDK (step S27), and then performs the processing of step S28 described below.

[0082] If the image RDL contains an object identical to the object represented by the first explanatory text MDK (step S23: YES), the reality assessment processing unit 22D inputs the image RDL into an Image-to-Text model to generate a second explanatory text MDL that explains the image RDL in natural language (step S24).

[0083] Next, the reality determination processing unit 22D determines whether or not the second explanatory text MDL represents an event that can occur in the real world (step S25).

[0084] Here, an example of the determination process performed in step S25 will be described.

[0085] The reality assessment processing unit 22D may determine, for example, based on a phrase included in the second explanatory text MDL, whether the second explanatory text MDL represents an event that could occur in the real world. According to this determination process, for example, when the second explanatory text MDL includes the phrase "a bus floating in the sky," the reality assessment processing unit 22D can determine that the second explanatory text MDL does not represent an event that could occur in the real world. Furthermore, according to the above-described determination process, for example, when the second explanatory text MDL includes the phrase "a bus overturned on the road," the reality assessment processing unit 22D can determine that the second explanatory text MDL represents an event that could occur in the real world.

[0086] The reality assessment processing unit 22D may determine, for example, based on a combination of words included in the second description MDL, whether the second description MDL represents an event that could occur in the real world. According to this determination process, for example, when the second description MDL contains a combination of “bus,” “sky,” and “floating,” the reality assessment processing unit 22D can determine that the second description MDL does not represent an event that could occur in the real world. Furthermore, according to the above-described determination process, for example, when the second description MDL contains a combination of “bus,” “road,” and “rollover,” the reality assessment processing unit 22D can determine that the second description MDL represents an event that could occur in the real world. Note that, according to this specific example, the reality assessment processing unit 22D may make the above determination by reading, from the DB 115, at least one of information related to a word combination that represents an event that could occur in the real world and information related to a word combination that does not represent an event that could occur in the real world. Alternatively, according to this specific example, the reality assessment processing unit 22D may make the above assessment based on, for example, the analysis results obtained by analyzing the frequency of occurrence of word combinations contained in sentences of a news article and the word combinations contained in the second explanatory text MDL.

[0087] According to this embodiment, the reality judgment processing unit 22D may perform processing in step S25 to obtain the result of the user's judgment as to whether or not the second explanatory text MDL represents an event that could occur in the real world.

[0088] If the second description MDL represents an event that could occur in the real world (step S25: YES), the reality assessment processing unit 22D associates the determination result that the image is realistic with one similar image SDK (step S26), and then performs the processing of step S28, which will be described later. On the other hand, if the second description MDL does not represent an event that could occur in the real world (step S25: NO), the reality assessment processing unit 22D associates the determination result that the image is unrealistic with one similar image SDK (step S27), and then performs the processing of step S28, which will be described later.

[0089] The reality determination processing unit 22D determines whether or not a determination result regarding whether or not the image is realistic has been associated with all the similar images generated by the similar image generating unit 22C (step S28).

[0090] If there is any similar image among all the similar images generated by the similar image generation unit 22C that is not associated with a determination result as to whether it is a realistic image (step S28: NO), the reality determination processing unit 22D returns to step S21 and processes the similar image that is not associated with the determination result. Also, if there is any similar image among all the similar images generated by the similar image generation unit 22C that is associated with a determination result as to whether it is a realistic image (step S28: YES), the reality determination processing unit 22D ends the series of processes related to reality determination.

[0091] According to the processing described above, the reality assessment processing unit 22D can determine whether a similar image SDK is a realistic image based on a first description MDK that describes the similar image SDK, an image RDL acquired using the first description MDK, and a second description MDL that describes the image RDL. Furthermore, according to the processing described above, the reality assessment processing unit 22D can determine that a similar image SDK is a realistic image when the image RDL includes an object identical to an object represented by the first description MDK and the second description MDL represents an event that could occur in the real world.

[0092] According to this specific example, the reality assessment processing unit 22D may associate a determination result that an image is realistic with a similar image that includes an object that exists in the real world and represents an event that could occur in the real world. According to this specific example, the reality assessment processing unit 22D may also associate a determination result that an image is realistic with a similar image that includes an object identical to an object included in a news image used in news related to the character string MDG and represents an event similar to an event described in a news article describing the news. Furthermore, in this specific example, a realistic image can be interpreted as an image that conforms to reality, an image that represents an event similar to a real event, or an image in which a causal relationship based on natural laws can be identified. Furthermore, in this specific example, a non-realistic image can be interpreted as an image that does not conform to reality, an image that represents an event different from a real event, or an image in which a causal relationship based on natural laws cannot be identified.

[0093] The additional data acquisition unit 22E acquires a similar image group GDL by excluding images associated with a determination result that the image is unrealistic from among the similar images generated by the similar image generation unit 22C. The additional data acquisition unit 22E inputs one similar image SDX included in the similar image group GDL to an object recognition model 24A trained using at least the training data TDL, thereby acquiring a recognition result RDX of the object included in the similar image SDX. Based on the processing performed by the similar image generation unit 22C, the additional data acquisition unit 22E identifies, from each dataset included in the dataset DDK, a dataset that includes one image PDX used to generate the one similar image SDX and a label LDX associated with the one image PDX. The additional data acquisition unit 22E also excludes one similar image SDX from the similar image group GDL if the label LDX and the recognition result RDX represent the same object.

[0094] The additional data acquisition unit 22E acquires a dataset DSK by associating a label LDK that is the same as the label LDK associated with an image PDK used to generate a similar image SDK that was not excluded from the similar image group GDL (see FIG. 6 ). The additional data acquisition unit 22E also inputs the dataset DSK to the importance estimation model 23A to acquire an importance JDK that corresponds to an estimated value of the importance of a similar image SDK included in the dataset DSK. The additional data acquisition unit 22E also acquires a dataset DSL by associating the importance JDK with a similar image SDK included in the dataset DSK (see FIG. 6 ). The additional data acquisition unit 22E also acquires a dataset DDX by adding the dataset DSL as additional data to the dataset DDK (see FIG. 6 ).

[0095] According to the processing described above, the additional data acquisition unit 22E can acquire additional data to add to the dataset DDK by eliminating, from each similar image generated by the similar image generation unit 22C, similar images that satisfy the first exclusion condition of being unrealistic and similar images that satisfy the second exclusion condition of the object recognition model 24A correctly recognizing the object.

[0096] The additional data acquisition unit 22E acquires the dataset DDM by excluding datasets that satisfy a third exclusion condition of relatively low importance from each dataset included in the dataset DDX (see FIG. 6). Specifically, the additional data acquisition unit 22E can acquire the dataset DDM by, for example, sorting each dataset included in the dataset DDX in descending order of importance and excluding datasets that are ranked below a predetermined rank from the dataset DDX. Note that the aforementioned predetermined rank is preferably set to, for example, 10,000th place.

[0097] According to the above-described processing, the additional data acquisition unit 22E can extract similar images that satisfy the usage conditions for use in retraining the object recognition model 24A from an image group containing a plurality of similar images generated by the similar image generation unit 22C, and acquire a dataset DDM containing the extracted similar images. The usage conditions may be relatively high in importance and may be realistic images in which the object recognition model 24A has erroneously recognized an object.

[0098] The object recognition model training unit 22B updates the dataset DDK to the dataset DDM, and retrains the object recognition model 24A using the dataset DAE and a dataset DDN obtained by removing importance from each data set included in the dataset DDM as new training data TDN (see FIG. 6 ).

[0099] The process related to retraining of the object recognition model 24A may be repeatedly performed until a predetermined termination condition is met. Specifically, the process related to retraining of the object recognition model 24A may be repeatedly performed, for example, until a condition is met that the number of retrainings reaches a predetermined number. Furthermore, the process related to retraining of the object recognition model 24A may be repeatedly performed, for example, until a condition is met that the importance of each data set included in the dataset DDM has reached a certain value or more.

[0100] According to this specific example, for example, the similar image generation unit 22C may acquire similar images similar to one image PDK by searching the Internet and / or a database. According to this specific example, for example, the similar image generation unit 22C may acquire an image in which at least a portion of one image PDK has been modified as a similar image similar to one image PDK. Furthermore, according to this specific example, the similar image generation unit 22C may generate or acquire one similar image similar to one image PDK. In other words, according to this specific example, it is sufficient that the similar image generation unit 22C acquires an image group including at least one similar image similar to one image PDK.

[0101] According to this specific example, the determination of whether a similar image is a realistic image is not limited to being made by the reality determination processing unit 22D, but may also be made, for example, by a user who visually confirms the similar image.

[0102] [Processing Flow] Next, a description will be given of the flow of processing performed in the information processing apparatus according to the second embodiment. Fig. 9 is a flowchart showing an example of processing performed in the information processing apparatus according to the second embodiment.

[0103] First, the information processing device 200 acquires the data set DAE and the data set DDG as input data (step S31).

[0104] Next, the information processing device 200 acquires a dataset DDH in which labels and importance levels are assigned to the dataset DDG. The information processing device 200 also trains an importance estimation model 23A using the dataset DDI as training data TDI (step S32). The information processing device 200 also trains an object recognition model 24A using the dataset DAE and the dataset DDL as training data TDL (step S32).

[0105] Next, the information processing device 200 generates a plurality of similar images that are similar to the images included in the dataset DDK that is the basis of the dataset DDL (step S33).

[0106] Next, the information processing device 200 eliminates similar images that satisfy the first exclusion condition and similar images that satisfy the second exclusion condition from the similar images generated in step S33, thereby acquiring additional data to be added to the dataset DDK (step S34). The first exclusion condition may be set as a condition that the image is unrealistic. The second exclusion condition may be set as a condition that the object recognition model 24A correctly recognizes the object.

[0107] Next, the information processing device 200 acquires a dataset DDX by adding the additional data acquired in step S34 to the dataset DDK. The information processing device 200 also acquires a dataset DDM by excluding datasets that satisfy a third exclusion condition from each dataset included in the dataset DDX (step S35). The third exclusion condition may be set as a condition of relatively low importance.

[0108] Next, the information processing device 200 retrains the object recognition model 24A using the dataset DAE and the dataset DDN obtained by excluding importance from each data set included in the dataset DDM as new training data TDN (step S36).

[0109] Next, the information processing device 200 determines whether or not a predetermined termination condition for retraining the object recognition model 24A has been met (step S37).

[0110] If the predetermined termination condition is not satisfied (step S37: NO), the information processing device 200 performs the processes from step S33 onwards again. Furthermore, if the predetermined termination condition is satisfied (step S37: YES), the information processing device 200 ends the series of processes related to training the object recognition model 24A. Furthermore, the information processing device 200 completes the retraining of the object recognition model 24A after performing the processes from step S33 to step S37 one or more times.

[0111] As described above, according to this embodiment, the object recognition model 24A can be retrained while increasing the importance of each data set included in the data set DDM. Furthermore, as described above, according to this embodiment, a data set that is relatively important and includes a realistic image in which the object recognition model 24A has erroneously recognized an object can be acquired as additional data. Furthermore, as described above, according to this embodiment, the object recognition model 24A can be retrained using the data set DAE and the data set DDN including the additional data as new training data TDN. Therefore, according to this embodiment, a decrease in recognition accuracy that occurs when performing object recognition can be prevented.

[0112] Third Embodiment FIG. 10 is a block diagram showing the functional configuration of a learning device according to a third embodiment.

[0113] The learning device 500 according to this embodiment has the same hardware configuration as the information processing device 100. The learning device 500 also has an extraction unit 511 and a training unit 512.

[0114] The extraction means 511 can be realized, for example, by using at least one of the functions of the additional data acquisition unit 12C of the first embodiment and the functions of the additional data acquisition unit 22E of the second embodiment. The training means 512 can be realized, for example, by using at least one of the functions of the object recognition model training unit 12A of the first embodiment and the functions of the object recognition model training unit 22B of the second embodiment.

[0115] FIG. 11 is a flowchart for explaining the processing performed in the learning device according to the third embodiment.

[0116] The extraction means 511 extracts a first image that satisfies the conditions for use in training the model 600 that recognizes an object contained in the image from a group of images that includes at least one image (step S51).

[0117] The training means 512 trains the model 600 using training data including the first image (step S52).

[0118] According to this embodiment, it is possible to prevent a decrease in recognition accuracy that occurs when performing object recognition.

[0119] A part or all of the above-described embodiments can be described as, but not limited to, the following supplementary notes.

[0120] (Supplementary Note 1) A learning device comprising: an extraction means for extracting a first image from a group of images including at least one image, the first image satisfying usage conditions for use in training a model that recognizes an object included in the image; and a training means for training the model using training data including the first image.

[0121] (Supplementary Note 2) The learning device according to Supplementary Note 1, further comprising an acquisition means for acquiring the image group including a similar image that is similar to the second image.

[0122] (Supplementary Note 3) The learning device according to Supplementary Note 2, wherein the extraction means extracts, as the first image, from the group of images, the similar image that satisfies the use condition that the model has misrecognized an object.

[0123] (Supplementary Note 4) The learning device of Supplementary Note 2, wherein the extraction means extracts, from the group of images, the similar image that is relatively important and satisfies the usage condition of being a realistic image in which the model has misrecognized an object, as the first image.

[0124] (Appendix 5) The learning device of Appendix 4 further includes a determination means for determining whether the similar image is a realistic image based on a first description describing the similar image, a third image acquired using the first description, and a second description describing the third image.

[0125] (Appendix 6) The learning device of Appendix 5, wherein the determination means determines that the similar image is a realistic image if the third image contains an object identical to the object represented by the first description and the second description represents an event that could occur in the real world.

[0126] (Supplementary Note 7) The learning device of Supplementary Note 2, wherein the acquisition means acquires the similar image based on either an explanatory text explaining the second image or a three-dimensional image representing a three-dimensional shape of an object included in the second image.

[0127] (Supplementary Note 8) The learning device according to Supplementary Note 7, wherein the acquisition means acquires the similar images based on another explanatory sentence in which at least one parameter included in the explanatory sentence is changed.

[0128] (Supplementary Note 9) A training method comprising: extracting a first image that satisfies usage conditions for use in training a model that recognizes an object contained in the image from a group of images that includes at least one image; and training the model using training data that includes the first image.

[0129] (Supplementary Note 10) A recording medium having recorded thereon a program that causes a computer to execute a process of extracting a first image that satisfies usage conditions for use in training a model that recognizes an object contained in the image from a group of images that includes at least one image, and training the model using training data that includes the first image.

[0130] Although the present disclosure has been described above with reference to the embodiments, the present disclosure is not limited to the above-described embodiments and examples. Various modifications that can be understood by those skilled in the art can be made to the configuration and details of the present disclosure within the scope of the present disclosure.

[0131] 11, 21 Data acquisition unit 12, 22 Training processing unit 13, 24 Object recognition processing unit 23 Importance estimation processing unit 100, 200 Information processing device

Claims

1. an extraction means for extracting a first image from a group of images including at least one image, the first image satisfying a usage condition for use in training a model that recognizes an object included in the image; training means for training the model using training data including the first image; A learning device comprising:

2. The learning device according to claim 1 , further comprising an acquisition means for acquiring the group of images including a similar image that is similar to the second image.

3. The learning device according to claim 2 , wherein the extraction means extracts, from the group of images, the similar image that satisfies the use condition that the model has erroneously recognized an object, as the first image.

4. The learning device according to claim 2, wherein the extraction means extracts, from the group of images, the similar image that is relatively important and satisfies the usage condition of being a realistic image in which the model has misrecognized an object, as the first image.

5. The learning device described in claim 4, further comprising a determination means for determining whether the similar image is a realistic image based on a first description describing the similar image, a third image acquired using the first description, and a second description describing the third image.

6. The learning device described in claim 5, wherein the determination means determines that the similar image is a realistic image if the third image contains an object identical to the object represented by the first description and the second description represents an event that could occur in the real world.

7. The learning device according to claim 2, wherein the acquisition means acquires the similar image based on either an explanatory text describing the second image or a three-dimensional image representing the three-dimensional shape of an object included in the second image.

8. The learning device according to claim 7 , wherein the acquisition means acquires the similar images based on another explanatory sentence in which at least one parameter included in the explanatory sentence is changed.

9. A computer-implemented training method comprising: extracting a first image from a group of images including at least one image, the first image satisfying a usage condition for use in training a model that recognizes an object included in the image; A training method for training the model using training data including the first image.

10. extracting a first image from a group of images including at least one image, the first image satisfying a usage condition for use in training a model that recognizes an object included in the image; A program that causes a computer to execute a process of training the model using training data that includes the first image.