Recognition device, recognition method, and program

The recognition device generates stable and uniform training data for object recognition by using a generative model to control the distribution of training data, addressing the challenges of biased data generation and improving recognition accuracy.

JP7896760B2Active Publication Date: 2026-07-29NEC CORP
View PDF 4 Cites 0 Cited by

Patent Information

Authority / Receiving Office
JP · JP
Patent Type
Patents
Current Assignee / Owner
NEC CORP
Filing Date
2023-03-02
Publication Date
2026-07-29

AI Technical Summary

Technical Problem

Existing AI data generation technologies for object recognition face challenges in generating unbiased, stable, and uniform training data, leading to unstable training processes and biased data, especially in environments like factories or construction sites where obtaining large amounts of annotation data is difficult.

Method used

A recognition device and method that uses a recognition model to generate training data by combining object and background images, employing a generative model to create additional data sets while minimizing a loss function to control the distribution, ensuring accurate recognition model training.

Benefits of technology

Enables the generation of unbiased, stable, and uniform training data for object recognition, enhancing recognition accuracy by training the model with generated data, even with limited initial data.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure 0007896760000001
    Figure 0007896760000001
  • Figure 0007896760000002
    Figure 0007896760000002
  • Figure 0007896760000003
    Figure 0007896760000003
Patent Text Reader

Abstract

Provided is a recognition device comprising: a recognition model training unit that trains a recognition model so that the recognition model receives, as input data, an image consisting of an object and a background and outputs a set of data including the type, position, and orientation of the object; a recognition unit that uses the recognition model to obtain, from the image consisting of the object and the background, a set of recognition result data, which is the result of recognizing at least the type, position, and orientation of the object; a generation unit that uses a generation model to generate, from the image consisting of the object and the background, a pair of another set of data that includes the type, position, and orientation of the object and another image that includes the object and the background as arranged on the basis of the other set of data; and a generation model training unit that trains the generation model on the basis of the value of a loss function that indicates the deviation between another set of recognition result data, which is the result of recognizing the other image using the recognition unit, and the other set of data.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to a recognition device, a recognition method, and a program.

Background Art

[0002] When performing object recognition using an AI (Artificial Intelligence) model, it is necessary to ensure a sufficient amount of a large amount of annotation data (teacher data by bounding boxes) for learning. However, when trying to create annotation data using real data, there is a problem that it causes a great deal of labor and an increase in costs due to this.

[0003] Especially in a work site or the like (such as a factory, a warehouse, a construction site, etc.), it is often difficult to obtain and create a large amount of training data, and the above problems become prominent.

[0004] Regarding this problem, there is an artificial data generation technique for generating artificial training data by combining background data and data of an object to be detected (image / 3D model, etc.). GAN (Generative Adversarial Network) is cited as an example thereof.

[0005] In Patent Document 1, an image generation method or the like capable of generating training data for learning a neural network or the like by synthesizing a background image and an image of an object to be detected is disclosed. Specifically, from an image (target image) composed of an object and a background, an image of the object is acquired as a detection target image, information indicating the position, shape, size, etc. of the object is acquired from the metadata of the image, and using a generation model, it is possible to generate another image in which an object having a different position, shape, size, etc. is arranged on the background image.

[0006] By comparing the generated image with the target image, and once the similarity between the two reaches a predetermined standard, it is possible to further fine-tune the parameters of the generation model so that the conditions for the currently generated image can be applied to the generation of other images as well. This makes it possible to generate multiple training data images from a single target image. [Prior art documents] [Patent Documents]

[0007] [Patent Document 1] Japanese Patent Publication No. 2019-159630 [Overview of the Initiative] [Problems that the invention aims to solve]

[0008] Furthermore, the disclosures in the above-mentioned prior art documents are incorporated herein by reference. The following analysis was conducted by the inventors.

[0009] In addition to the above, data generation AI technologies have also been developed that optimize the distribution of artificial training data generated by a data generation network using a desired objective function, such as the loss function of a recognition model. For example, there are data generation methods that train a generative model to maximize the loss function of the recognition model during training, and then use this to generate data.

[0010] While data generation AIs like those mentioned above can generate data useful for training, they tend to have unstable training processes and can result in biased data generation. For example, in object recognition, if the objective function is designed to maximize the loss function, images may be generated with objects placed in impossible locations, or with combinations that make recognition unnecessarily difficult.

[0011] In this regard, learning methods such as curriculum learning have been proposed, which involve controlling the distribution of the training data by initially learning the easier parts of the training data and gradually increasing the difficulty of the learning process, thereby enabling the generation of highly accurate models.

[0012] However, while curriculum-based learning can yield some results, the appropriate curriculum differs for each task (domain). Therefore, curriculum setting must be done somewhat by trial and error for each task, making it difficult to stably control the distribution of training data necessary for generating highly accurate models.

[0013] Based on the above-mentioned problems, the objective of the present invention is to provide a recognition device that can generate data useful for learning in an unbiased, stable, and uniform manner for generating recognition models for object recognition, and that can achieve outstanding recognition accuracy by training a recognition model using this data, in an AI data generation technology for object recognition. [Means for solving the problem]

[0014] According to a first aspect of the present invention or disclosure, a recognition device is provided, comprising: a recognition model learning unit that takes an image consisting of an object and a background as input data and trains a recognition model such that the output data includes the type, position, and orientation of the object; a recognition unit that uses the recognition model to acquire recognition result data which is the result of recognizing at least the type, position, and orientation of an object from an image consisting of an object and a background; a generation unit that generates a set from an image consisting of an object and a background using a generation model, the set of which includes other data including the type, position, and orientation of an object and another image which includes an object and a background arranged based on the other data; and a generation model learning unit that trains a generation model based on other recognition result data which is the result of the recognition unit recognizing the other image and the value of a loss function which indicates the deviation from the other data.

[0015] A recognition method is provided that causes a computer to perform the following steps: training a recognition model with an image consisting of an object and a background as input data such that the output data includes the type, position and orientation of the object; using the recognition model to obtain recognition result data which is the result of recognizing at least the type, position and orientation of the object from the image consisting of an object and a background; generating a pair from the image consisting of an object and a background using a generative model, which consists of another image which includes another data including the type, position and orientation of the object and an object and a background arranged based on the other data; and training the generative model based on another recognition result data which is the result of recognizing the other image and the value of a loss function which indicates the deviation from the other data.

[0016] A third aspect of the present invention and disclosure is provided, which causes a computer to perform the following steps: a process of training a recognition model with an image consisting of an object and a background as input data, such that the output data includes the type, position, and orientation of the object; a process of obtaining recognition result data, which is the result of recognizing at least the type, position, and orientation of the object from an image consisting of an object and a background, using the recognition model; a process of generating a pair from an image consisting of an object and a background using a generative model, which consists of another image containing another data including the type, position, and orientation of the object, and another image containing an object and a background arranged based on the other data; and a process of training a generative model based on another recognition result data, which is the result of recognizing the other image, and the value of a loss function that shows the deviation from the other data. [Effects of the Invention]

[0017] From each perspective of the present invention and disclosure, the present invention provides a recognition device, recognition method, and program that enable the generation of data useful for learning in an unbiased, stable, and uniform manner in an AI data generation technology for generating recognition models for object recognition, and that contribute to achieving outstanding recognition accuracy by using this data to train a recognition model. [Brief explanation of the drawing]

[0018] [Figure 1] It is a block diagram showing an example of the configuration of a recognition device according to an embodiment. [Figure 2] It is a schematic diagram for showing an outline of the processing of a recognition device according to the first embodiment. [Figure 3] It is a block diagram showing an example of the configuration of a recognition device according to the first embodiment. [Figure 4(a)] It is a schematic diagram showing a series of operations of a recognition device according to the first embodiment. [Figure 4(b)] It is a schematic diagram showing a series of operations of a recognition device according to the first embodiment. [Figure 4(c)] It is a schematic diagram showing a series of operations of a recognition device according to the first embodiment. [Figure 5] It is a flowchart for explaining an example of the operation of a recognition device according to the first embodiment. [Figure 6] It is a block diagram showing an example of the hardware configuration of a recognition device according to the first embodiment. [Figure 7] It is a schematic diagram for showing the configuration of a generation unit of a recognition device according to the second embodiment. [Figure 8] It is a schematic diagram for showing the configuration of a generation unit of a recognition device according to the second embodiment.

Embodiments for Carrying Out the Invention

[0019] [Outline of Processing of One Embodiment] First, an overview of the processing of one embodiment will be described. The reference numerals in the drawings attached to this overview are for convenience and serve as examples to aid understanding; the description in this overview is not intended to be limiting in any way. Furthermore, the connection lines between blocks in each figure include both bidirectional and unidirectional lines. Unidirectional arrows schematically indicate the flow of the main signal (data) and do not exclude bidirectional flow. In addition, although not explicitly shown, input ports and output ports exist at the input and output ends of each connection line in the circuit diagrams, block diagrams, internal configuration diagrams, and connection diagrams disclosed in this application. The same applies to input / output interfaces.

[0020] [Configuration of one embodiment] Next, the configuration of a recognition device according to one embodiment will be described with reference to the figures. Figure 1 is a block diagram showing an example of the configuration of a recognition device according to one embodiment. As shown in this figure, the recognition device 10 according to one embodiment includes a recognition model learning unit 11, a recognition unit 12, a generation unit 13, and a generation model learning unit 14.

[0021] The recognition model learning unit 11 takes an image consisting of an object and a background as input data and trains a recognition model so that the output data includes the type, position, and orientation of the object. The recognition unit 12 uses the recognition model to obtain recognition result data, which is the result of recognizing at least the type, position, and orientation of the object from the image consisting of the object and a background. The generation unit 13 generates a pair from the image consisting of the object and a background using a generative model, which consists of another data including the type, position, and orientation of the object, and another image containing the object and background arranged based on the other data. The generative model learning unit 14 learns a generative model based on another recognition result data, which is the result of the recognition unit recognizing the other image, and the value of a loss function that shows the deviation from the other data.

[0022] According to one embodiment of the recognition device 10, first, the recognition model is trained using image data consisting of an object and a background, and its annotation data. That is, an image consisting of an object and a background is taken as input data, and the recognition model is trained to output at least the type, position, and orientation of the object. Next, the above image data consisting of an object and a background is taken as input, and the generative model outputs data of another image consisting of an object and a background, and data of the type, position, and orientation of another object placed in that other image, which serves as annotation data. The recognition model recognizes the other image and obtains recognition result data. The value of the loss function is calculated using this recognition result data and the annotation data of the type, position, and orientation of another object.

[0023] Thus, in one embodiment, the recognition device generates training data using a generative model that generates training data. This makes it possible to improve recognition accuracy by training the recognition model using the generated training data, even if there is only a small amount of annotated training data. Furthermore, by calculating the value of the loss function using the recognition result data, it is possible to control the learning of the generative model based on the value of the loss function and generate annotated training data useful for learning the recognition model.

[0024] Specific embodiments will be described in more detail below with reference to the drawings. In each embodiment, the same components are denoted by the same reference numerals, and their descriptions are omitted.

[0025] [First Embodiment] Figure 2 is a schematic diagram illustrating the processing overview of the recognition device according to the first embodiment. As shown in this figure, the recognition device 10 of this embodiment consists of a database that holds background / object data, an artificial data generation AI, and a recognition AI. The artificial data generation AI generates artificial data with objects placed in the background from the background / object data. The recognition AI recognizes the artificial data and obtains an inference result. The inference result is fed back to the recognition AI and the artificial data generation AI based on the value of the loss function.

[0026] [Configuration of the first embodiment] Next, the configuration of the recognition device 10 in this embodiment will be described with reference to the figures. Figure 3 is a block diagram showing an example of the configuration of the recognition device 10 according to this embodiment. As shown in this figure, the recognition device 10 according to the first embodiment includes a recognition model learning unit 11, a recognition unit 12, a generation unit 13, a generation model learning unit 14, an initial training data holding unit 15, and a reference value holding unit 16.

[0027] The recognition model learning unit 11 takes an image consisting of an object and a background as input data and trains the recognition model so that the output data includes the type, position, and orientation of the object. An "object" is the object to be recognized in object recognition. For example, in an autonomous driving system for vehicles, this could include vehicles and people traveling on the road, signs, and traffic lights. The "background" is the space to which the object is placed, and the image with the object placed in the background becomes the object to be recognized by the recognition model. The "type" of the object is the category to which the object to be recognized belongs, and examples include "automobile," "person," and "artificial objects" such as signs and traffic lights. The "position and orientation of the object" is the coordinates or position vector that indicates the position in which the object is placed within the background, for example, the position vector relative to the center of gravity of the bounding box representing the object, and a certain range in space specified by data such as the height, width, depth, and rotation angle of the bounding box.

[0028] A "recognition model" refers to a set of operations for transforming predetermined input data to obtain output data. In this embodiment, "recognition" means detecting objects inherent in the input image data and obtaining the type, position, and orientation of the object as output data.

[0029] The "recognition model" may be implemented as a neural network model consisting of an activation function and weighting parameters that represent the strength of its connections. In this embodiment, "learning" the recognition model means adjusting the above parameters so that, for a specific input image consisting of an object and a background, it outputs data indicating the type, position, and orientation of a specific object, as shown in the annotation data.

[0030] Learning may be performed using a gradient method, where the objective function is a loss function that represents the sum of the errors between the output of the recognition model for the input image and the desired output, which is annotation data. In other words, it may be performed by calculating the combination of parameters that minimizes the loss function.

[0031] The recognition unit 12 uses a recognition model to acquire recognition result data, which is the result of recognizing at least the type, position, and orientation of an object from an image consisting of an object and a background. Specifically, it predicts the type, position, and orientation of a particular object using bounding boxes based on data acquired from devices such as optical cameras and LiDAR (Light Detection and Ranging).

[0032] The generation unit 13 generates a set of data from an image consisting of an object and a background, including the type, position, and orientation of the object, and another image containing the object and background arranged based on the aforementioned data, using a generation model. In other words, it can generate a set of background and object images that can be artificially used to train a recognition model, and annotation data for the object.

[0033] The generative model learning unit 14 learns the generative model based on the value of a loss function that shows the discrepancy between another recognition result data, which is the result of the recognition unit 12 recognizing another image generated by the generation unit 13, and another (annotation) data generated by the generation unit 13. The learning may be carried out in the direction of maximizing the value of the loss function or in the direction of minimizing it, but if it is carried out in the direction of maximizing it, the generation of images with impossible object arrangements or arrangements that are difficult to recognize, which are the problems mentioned above, will increase. Therefore, in the recognition device 10 of this embodiment, it is desirable to carry out the learning in the direction of minimizing the value of the loss function.

[0034] The initial training data holding unit 15 holds initial training data consisting of a given image consisting of an object and a background, and data including the type, position, and orientation of the object. The "given image" is not an image generated by the generation unit 13, but initial training data based on measured values, etc. Initially, the recognition AI and the generation AI are trained using this initial training data. That is, the recognition model learning unit 11 learns a recognition model using the initial training data, the recognition unit 12 recognizes the image of the initial training data using the recognition model and obtains recognition result data, and the generation model learning unit 14 learns the generation model in a direction that minimizes the value of the loss function which indicates the discrepancy between the recognition result data and the data including the type, position, and orientation of the object included in the initial training data.

[0035] The reference value holding unit 16 holds a reference value for the loss function. The "reference value" is the target value for the loss function and is a value that controls the difficulty level when the recognition unit 12 recognizes the artificial training data (other data) generated by the generation unit 13.

[0036] When this reference value is set, the recognition unit 12 of the recognition device 10 obtains a value of a loss function that indicates the discrepancy between the recognition result data, which is the result of recognizing the image generated by the generative model, and the data including the type, position, and orientation of objects in the image generated by the generative model. The generative model learning unit 14 then learns the generative model so that the value of the loss function approaches the reference value.

[0037] The reference value can be a fixed value, or it can be determined based on the value of the loss function. For example, if the value of the loss function is x, the reference value S(x) can be a constant multiple of the most recent value of the loss function, such as S(x) = 1.05 × x.

[0038] [Explanation of operation] Figures 4(a) to 4(c) are schematic diagrams showing a series of operations of the recognition device 10 of this embodiment. Referring to Figure 4(a), first, 1. The recognition AI (recognition model) is trained with a limited amount of initial training data (R0). The recognition AI is trained until the loss function L is minimized. Next, the generative AI (generative model) is trained using the trained recognition AI (R0) so that the value of the loss function of the recognition AI is minimized (G0). Referring to Figure 4(b), next, 3. The value of the loss function of R0 in G0 (L(R0)) is recorded. Next, 4. A reference value (S(L(R0))) for the loss value of the generative AI is set based on L(R0). Referring to Figure 4(c), next, 5. The generative AI is trained so that the loss approaches the reference value (G1). Finally, 6. The recognition AI is trained with the data generated by G1 (R1).

[0039] In the recognition device 10 of this embodiment, a recognition AI with higher recognition accuracy can be trained by repeating steps 3 to 6. For example, in step 3, it is possible to repeat the process until the value of the loss function reaches a predetermined error range.

[0040] Figure 5 is a flowchart illustrating an example of the operation of the recognition device 10 in this embodiment. As shown in this figure, first, initial training data is acquired (step S101). Next, the recognition model is trained using the initial training data (step S102). Then, the generative model is trained in a direction that minimizes the loss function of the recognition model (step S103).

[0041] Next, the baseline value of the retained loss function is obtained (step S104). Then, the value of the loss function is obtained (step S105). Specifically, the value of the loss function is obtained which represents the discrepancy between the recognition result data, which is the result of the recognition model recognizing an image consisting of an object and background generated by the generative model, and the data, which includes the type, position, and orientation of the object generated by the generative model.

[0042] Next, it is determined whether the acquired loss function value is within a predetermined error range relative to the reference value. If the loss function value is within the predetermined error range relative to the reference value (step S106, Y), the process is terminated. If it is not within the predetermined error range (step S106, N), the generative model is trained so that the loss function value approaches the reference value (step S107), and the recognition model is trained using the data generated by the generative model (step S108). After that, the process returns to obtaining the reference value for the loss function value (step S104).

[0043] [Hardware configuration] The recognition device 10 of this embodiment is executable by an information processing device (computer) and has the configuration illustrated in Figure 6. The recognition device 10 includes a CPU (Central Processing Unit) 61, memory 62, input / output interface 63, and a communication means such as a NIC (Network Interface Card) 64, which are interconnected by an internal bus 65.

[0044] However, the configuration shown in Figure 6 is not intended to limit the hardware configuration of the recognition device. The recognition device 10 may include hardware not shown, and it may not have an input / output interface 63 if necessary. Furthermore, the number of CPUs and other components included in these devices is not limited to the example shown in Figure 6; for example, the recognition device 10 may include multiple CPUs.

[0045] Memory 62 consists of RAM (Random Access Memory), ROM (Read Only Memory), and auxiliary storage devices (such as hard disks).

[0046] The input / output interface 63 is a means that serves as an interface for display devices and input devices (not shown). The display device is, for example, a liquid crystal display. The input device is, for example, a camera or sensor that receives an image consisting of an object and a background, and a device that receives user input such as a keyboard or mouse.

[0047] The functions of the recognition device 10 include the recognition model learning program and the generative model learning program stored in the memory 62. program The system is implemented using a group of programs (processing modules) such as recognition programs, generation programs, and reference value judgment programs, and a group of data such as initial training data, reference value data, and other parameters used by each program. The processing module is implemented, for example, by the CPU 61 executing each program stored in memory 62. Furthermore, the program can be downloaded via a network or updated using a storage medium that stores the program. Moreover, the processing module may be implemented by a semiconductor chip. In other words, there is a means to execute the functions performed by the processing module using some hardware and / or software.

[0048] [Hardware operation] In the recognition device 10, the recognition model learning program is called from memory 62 and enters execution mode on the CPU 61. This program reads the initial training data held in memory 62 and performs learning of the recognition model. Next, the recognition program is called from memory 62 and enters execution mode on the CPU 61. This program reads the initial training data and outputs recognition result data. Next, the generative model learning unit is called from memory 62 and enters execution mode on the CPU 61. This program uses a loss function that calculates the distance between a vector representing the initial training data (annotations) and a vector representing the recognition result data, and performs learning in a direction that minimizes the loss function. For example, it calculates the parameters for which the derivative of the loss function is 0 and stores them in memory 62 along with the value of the loss function at that time.

[0049] Next, the reference value determination program is called from memory 62 and put into execution mode on CPU 61. This program retrieves the reference value stored in memory 62, compares it with the calculated loss function value, and determines whether it is within the error range relative to the reference value. If it is within the error range relative to the reference value, the process terminates. If it is not within the error range relative to the reference value, the generative model learning program is called from memory 62 and put into execution mode on CPU 61. This program learns the generative model so that the value of the loss function approaches the reference value.

[0050] Next, the generation program is called from memory 62 and enters execution mode on CPU 61. This program takes an image consisting of an object and a background as input data and outputs training data consisting of another image consisting of an object and a background, and its annotation data (at least the type, position, and orientation of the object). Next, the recognition model learning program enters execution mode again on CPU 61 and learns the recognition model using the training data generated by the generation program as input.

[0051] Once learning is complete, the baseline value of the loss function is obtained again, the value of the loss function is obtained, and the baseline value judgment program determines whether the value of the loss function is within a predetermined error range relative to the baseline value. The learning and generation process is repeated until it falls within the error range.

[0052] [Explanation of effects] According to the recognition device 10 of this embodiment, by generating data for learning and learning a recognition model using this data, it is possible to proceed with learning while generating training data, even when it is difficult to obtain initial training data and the amount of data is small. In this process, it is possible to proceed with learning so that the value of the loss function approaches a predetermined reference value, and rather than simply maximizing or minimizing the value of the loss function, it is possible to generate data at a level useful for learning using a uniform method. As a result, by learning the recognition model, it is possible to provide a recognition device that delivers outstanding recognition accuracy.

[0053] [Second Embodiment] In this embodiment, the configuration of the generation unit 13 of the recognition device 10 will be described in particular. Specifically, the generation unit has a network that generates the type, position, and orientation of an object, and a network that generates an image consisting of the object and a background. The network that generates the type, position, and orientation of an object takes an image consisting of an object and a background as input and outputs the position and orientation of the object generated by random variables sampled based on a predetermined distribution. The network that generates the image consisting of the object and a background generates an image in which the object is placed in the background based on the position and orientation of the object.

[0054] [Configuration of the second embodiment] Figures 7 and 8 are schematic diagrams showing the configuration of the generation unit 13 of the recognition device 10 according to the second embodiment. Figure 7 is a network that generates the object type, position, and orientation from the data generated by the generation unit 13. Figure 8 is a network that generates an image with the object placed in the background.

[0055] In Figure 7, a bird's-eye view image is first input from a device such as a camera. The input data is encoded by an encoder to obtain the values ​​of latent variables. These values ​​are further decoded to generate a location map. The location map shows the distribution of potential object placements. Object coordinates are generated by sampling based on these location maps.

[0056] The object's arrangement is sampled by generating random numbers (noise) according to a predetermined distribution. Each coordinate is provided as a random variable, and a random number term is added to each sample to form the distribution. Similarly, the orientation (rotation) is also sampled.

[0057] In Figure 8, the object data is transformed and positioned based on the object type, position, and orientation vectors generated in Figure 7. Specifically, the transformation is performed by multiplying the position vector and orientation vector of the original object data. Finally, it is composited with the background. The generated image data may be an image with a bounding box placed in the background, as shown in the figure.

[0058] Furthermore, as shown in Figure 7, the generation network can generate vectors representing the object type, position, and orientation, as well as an occlusion mask representing the overlap between multiple objects. This vector data is multiplied with the object data in Figure 8 and reflected in the image data. In other words, the generation unit 13 can calculate the region that disappears due to the overlap of multiple objects and generate an image in which the object is placed in the background by projecting that region onto the object's image.

[0059] [Explanation of effects] In this embodiment of the recognition device, background data is used as input information to combine with object data generated from 3D models or the like, generating the position and orientation of the combined objects, and then generating specific learning data based on that.

[0060] Some or all of the embodiments described above can also be described as follows. However, these following appendices are merely illustrative examples of the present invention, and the present invention is not limited to these cases. [Note 1] The recognition device relating to the first viewpoint described above is as follows. [Note 2] Preferably, the recognition device as described in Appendix 1, further comprising: an initial training data holding unit that holds initial training data consisting of a given image consisting of an object and a background, and data including the type, position, and orientation of the object; a recognition model learning unit learns a recognition model using the initial training data; a recognition unit recognizes the image of the initial training data using the recognition model and obtains recognition result data; and a generation model learning unit learns a generation model in a direction that minimizes the value of a loss function that shows the discrepancy between the recognition result data and the data including the type, position, and orientation of the object included in the initial training data. [Note 3] The recognition device further comprises a reference value holding unit that holds a reference value for the value of the loss function, the recognition unit acquires a value of the loss function that indicates the deviation between recognition result data, which is the result of recognizing an image generated by a generative model, and data including the type, position, and orientation of objects in the image generated by the generative model, and the generative model learning unit learns the generative model so that the value of the loss function approaches the reference value, preferably as described in Appendix 1 or 2. [Note 4] The reference value is determined based on the value of the loss function, preferably by the recognition device described in Appendix 3. [Note 5] The recognition model learning unit further uses an image containing objects and a background generated by the generative model as input data, and trains the recognition model so that the output data includes the type, position, and orientation of the objects related to the image generated by the generative model, preferably as described in Appendix 3 or 4. [Note 6] The generation unit comprises a network for generating the type, position, and orientation of an object, and a network for generating an image consisting of an object and a background, wherein the network for generating the type, position, and orientation of an object takes an image consisting of an object and a background as input and outputs the position and orientation of the object generated by random variables sampled based on a predetermined distribution, and the network for generating an image consisting of an object and a background generates an image in which the object is placed in the background based on the position and orientation of the object, preferably as described in any of appendices 1 to 5. [Note 7] The generation unit calculates the area that disappears due to the overlapping of multiple objects, projects the area onto the object's image, and generates an image in which the object is placed within the background, preferably the recognition device described in Appendix 6. [Note 8] The recognition device relating to the second perspective described above is as follows. [Note 9] Preferably the recognition method as described in Appendix 8, further comprising the steps of: acquiring initial training data consisting of a set of a given image consisting of an object and a background, and data including the type, position, and orientation of the object; training a recognition model using the initial training data; recognizing the image of the initial training data with the recognition model and obtaining recognition result data; and training a generative model in a direction that minimizes the value of a loss function that shows the discrepancy between the recognition result data and the data including the type, position, and orientation of the object included in the initial training data. [Note 10] Preferably the recognition method of Appendix 9, further comprising: (a) the step of obtaining a reference value for the value of the loss function; (b) the step of obtaining a value of the loss function that shows the discrepancy between recognition result data relating to the recognition of an image consisting of an object and a background generated by a generative model and data including the type, position and orientation of the object generated by the generative model; (c) the step of training the generative model so that the value of the loss function approaches the reference value; and (d) the step of training the recognition model using an image including an object and a background generated by the generative model as input data, so that data including the type, position and orientation of the object relating to the image generated by the generative model becomes output data. [Note 11] The recognition method, preferably as described in Appendix 10, is to repeat steps (a) through (d) until the value of the loss function reaches a predetermined error range relative to the reference value. [Note 12] The program is as described above regarding the third perspective.

[0061] Furthermore, each disclosure of the above-mentioned patent documents, etc., cited herein shall be incorporated by reference. Within the framework of the full disclosure of the present invention (including the claims), further modifications and adjustments to the embodiments are possible based on the fundamental technical concept. Also, within the framework of the full disclosure of the present invention, various combinations or selections (including partial deletions) of various disclosed elements (including each element of each claim, each element of each embodiment, each element of each drawing, etc.) are possible. In other words, the present invention naturally includes various modifications and changes that a person skilled in the art could make in accordance with the full disclosure, including the claims, and the technical concept. In particular, with respect to the numerical ranges described herein, any numerical value or sub-range included within that range should be interpreted as being specifically described unless otherwise stated. [Explanation of Symbols]

[0062] 10: Recognition device 11: Recognition Model Learning Unit 12: Recognition part 13: Generation part 14: Generative Model Learning Unit 15: Initial training data storage unit 16: Reference value holding section 61: CPU 62: Memory 63: Input / Output Interface 64: NIC 65: Internal bus

Claims

1. A recognition model learning unit takes an image consisting of an object and a background as input data and trains a recognition model so that the output data includes the type, position, and orientation of the object. A recognition unit that uses the aforementioned recognition model to acquire recognition result data, which is the result of recognizing at least the type, position, and orientation of an object from an image consisting of an object and a background. A generation unit generates, using a generation model, a set of data from an image consisting of an object and a background, including the type, position, and orientation of the object, and another image containing the object and background arranged based on the aforementioned data. A generative model learning unit learns the generative model based on the value of a loss function that shows the discrepancy between the aforementioned alternative image and the recognition result data obtained by the recognition unit, A recognition device having the following features.

2. An initial training data holding unit holds initial training data consisting of a given image comprising an object and a background, and data including the type, position, and orientation of the object. It further possesses, The recognition model learning unit learns the recognition model using the initial training data. The recognition unit recognizes the image of the initial training data using the recognition model and obtains the recognition result data. The generative model learning unit learns the generative model in a direction that minimizes the value of the loss function, which represents the discrepancy between the recognition result data and the data including the type, position, and orientation of the object included in the initial training data. The recognition device according to claim 1.

3. A reference value holding unit that holds a reference value for the value of the loss function, It further possesses, The recognition unit obtains a value of a loss function that indicates the discrepancy between the recognition result data, which is the result of recognizing the image generated by the generation model, and the data including the type, position, and orientation of objects in the image generated by the generation model. The generative model learning unit learns the generative model so that the value of the loss function approaches the reference value. The recognition device according to claim 1.

4. The aforementioned reference value is determined based on the value of the loss function. The recognition device according to claim 3.

5. The recognition model learning unit further uses an image containing an object and background generated by the generation model as input data, and trains the recognition model so that the output data includes the type, position, and orientation of the object related to the image generated by the generation model. The recognition device according to claim 3.

6. The generation unit includes a network for generating the type, position, and orientation of an object, and a network for generating an image consisting of the object and a background. The network that generates the type, position, and orientation of the object takes an image consisting of the object and background as input and outputs the position and orientation of the object generated by random variables sampled based on a predetermined distribution. The network that generates an image consisting of the object and background generates an image in which the object is placed within the background based on the position and orientation of the object. The recognition device according to any one of claims 1 to 5.

7. The recognition device according to claim 6, wherein the generation unit calculates the region that disappears due to the overlapping of a plurality of objects, and generates an image in which the region is projected onto the image of the object and the object is placed in the background.

8. A recognition method that causes a computer to perform the following steps: The steps include training a recognition model using an image consisting of an object and a background as input data, such that the output data includes the type, position, and orientation of the object, The steps include: obtaining recognition result data, which is the result of recognizing at least the type, position, and orientation of an object from an image consisting of an object and a background using the aforementioned recognition model; A step of generating a set of data from an image consisting of an object and a background, including the type, position, and orientation of the object, and another image containing the object and background arranged based on the aforementioned data, using a generative model. The steps include: training the generative model based on the value of a loss function that shows the deviation between the aforementioned other image and the other image, A recognition method for [unclear].

9. A step of acquiring initial training data consisting of a set of a given image comprising an object and a background, and data including the type, position, and orientation of the object. A step of training the recognition model using the initial training data, The steps include: recognizing the image of the initial training data using the recognition model and obtaining the recognition result data; The steps include training the generative model in a direction that minimizes the value of the loss function that shows the discrepancy between the recognition result data and the data including the type, position, and orientation of the object included in the initial training data, The recognition method according to claim 8, further comprising the above.

10. A process to train a recognition model using an image consisting of an object and a background as input data, such that the output data includes the type, position, and orientation of the object. The process involves obtaining recognition result data, which is the result of recognizing at least the type, position, and orientation of an object from an image consisting of an object and a background, using the aforementioned recognition model. A process that generates, using a generative model, a set of data from an image consisting of an object and a background, including the type, position, and orientation of the object, and another image containing the object and background arranged based on the aforementioned data. A process for learning the generative model based on the value of a loss function that shows the deviation between the aforementioned other image and the aforementioned other image, A program that causes a computer to execute something.