Teacher data generation device, teacher data generation method, and learning system

The training data generation device addresses inconsistencies in object position and size by synthesizing composite images with accurate labels, improving the recognition accuracy of machine learning models for image recognition tasks.

WO2026154536A1PCT designated stage Publication Date: 2026-07-23MITSUBISHI ELECTRIC CORP
View PDF 0 Cites 0 Cited by

Patent Information

Authority / Receiving Office
WO · WO
Patent Type
Applications
Current Assignee / Owner
MITSUBISHI ELECTRIC CORP
Filing Date
2025-01-14
Publication Date
2026-07-23

AI Technical Summary

Technical Problem

Existing techniques for generating training data for image recognition tasks, such as object detection and semantic segmentation, suffer from inconsistencies in object position and size due to differing shooting conditions between background and object images, which can degrade recognition accuracy.

Method used

A training data generation device that acquires background and object images with their respective shooting conditions, generates a three-dimensional background model, determines object placement and transformation size, and synthesizes composite images with accurate labels to ensure consistency in object position and size.

Benefits of technology

Produces high-quality training images and labels that maintain object position and size consistency, enhancing the recognition accuracy of machine learning models for tasks like object detection and semantic segmentation.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure JP2025000821_23072026_PF_FP_ABST
    Figure JP2025000821_23072026_PF_FP_ABST
Patent Text Reader

Abstract

The present invention automatically generates a teacher image, in which there is no incongruity in the position and size of an object, and a label in teacher data generation in which a person has conventionally generated a label by checking an image. The present invention comprises: a background image acquisition unit (1) that acquires a background image; a background imaging condition acquisition unit (2) that acquires a background imaging condition; an object image acquisition unit (3) that acquires an object image; an object imaging condition acquisition unit (4) that acquires an object imaging condition; an image synthesis unit (5) that generates a synthesized image of the background image and the object image from the background image, the background imaging condition, the object image, and the object imaging condition; and a label generation unit (6) that generates a label. The image synthesis unit (5) generates a background three-dimensional model by three-dimensional modeling of the background, determines the two-dimensional position of an object in the background image from the three-dimensional position of the object, and pastes an object conversion image at the two-dimensional position of the object in the background image to generate a synthesized image.
Need to check novelty before this filing date? Find Prior Art

Description

Training data generation device, training data generation method, and learning system

[0001] This disclosure relates to a training data generation device, a training data generation method, and a learning system.

[0002] One type of image recognition task using deep learning is detecting specific objects within an image and outputting their range and type. This image recognition task is broadly divided into object detection tasks and semantic segmentation tasks. Object detection tasks refer to the process of detecting specific objects in an image and outputting their positions as rectangular bounding boxes. Representative machine learning models include "YOLO (You Only Look Once)" and "SSD (Single Shot MultiBox Detector)". On the other hand, semantic segmentation tasks refer to the task of classifying which object each pixel in an image belongs to. In other words, it is possible to understand the shape of objects in more detail than in object detection tasks. Representative machine learning models include "FCN (Fully Convolutional Network)" and "U-Net".

[0003] To improve the recognition accuracy of these tasks to a practical level, a large amount of training data—pairs of training images and corresponding labels—is required for each task. In particular, creating label information generally requires a person to visually and manually examine the images, set rectangular bounding boxes, or assign each pixel to a specific object, which is an enormous amount of work. Furthermore, for tasks where detection events rarely occur, such as detecting wildlife photographed in the mountains, it is difficult to capture a large number of training images in the first place.

[0004] As a method for efficiently creating teacher data, a technique for artificially generating new teacher data from existing data has been proposed. For example, an object image, which is an image of a detection target, is pasted onto a background image, which is an image of the background of the pasting destination, and a synthetic image obtained by pasting the object image onto the background image is created as a new teacher image. Further, a technique has been proposed for automatically generating label information for the teacher image from the effective pixel range of the image of the detection target (see, for example, Non-Patent Document 1 and Non-Patent Document 2).

[0005] G. Georgakis, A. Mousavian, A. C. Berg, and J. Kosecka. Synthesizing Training Data for Object Detection in Indoor Scenes. arXiv:1702.07836, 2017.D. Dwibedi, I. Misra, and M. Hebert. Cut, Paste and Learn: Surprisingly Easy Synthesis for Instance Detection. arXiv:1708.01642, 2017.

[0006] However, in the techniques disclosed in Non-Patent Document 1 and Non-Patent Document 2, there is a problem in that a sense of incongruity may occur in the position and size of the object in the teacher image, which is a synthetic image, because the shooting conditions of the background image and the shooting conditions of the object image are different. When a machine learning model is learned using a teacher image with a sense of incongruity, it may not contribute to improving the recognition accuracy and may even deteriorate the recognition accuracy.

[0007] The present disclosure has been made to solve the above-described problems, and an object of the present disclosure is to provide a teacher data generation device, a teacher data generation method, and a learning system that generate high-quality teacher images without a sense of incongruity in the position and size of an object and labels corresponding to the teacher images.

[0008] The training data generation device of this disclosure comprises: a background image acquisition unit that acquires a background image; a background shooting condition acquisition unit that acquires background shooting conditions which are the shooting conditions for the background image; an object image acquisition unit that acquires an object image which is an image of an object to be detected; an object shooting condition acquisition unit that acquires object shooting conditions which are the shooting conditions for the object image; an image synthesis unit that generates a training image which is a composite image of the background image and the object image from the background image, background shooting conditions, object image and object shooting conditions; and a label generation unit that generates labels corresponding to objects in the composite image. The background image is captured by a background shooting camera, the background shooting conditions include the camera height which indicates the height from the reference plane of the background shooting camera when the background image was captured, the camera optical axis vertical angle which indicates the angle in the height direction of the optical axis from the background shooting camera toward the subject when the background image was captured, and the camera field of view which indicates the field of view of the background shooting camera when the background image was captured, the object shooting conditions include the effective pixel range which is information indicating the range of an object in the object image, and further, The image synthesis unit generates a three-dimensional background model based on the background shooting conditions, sets the area in the background model where objects can be placed as the object placement area, determines the three-dimensional position where objects are placed in the object placement area of ​​the background model, determines the two-dimensional position where objects are placed in the background image from the three-dimensional position, calculates the object transformation size, which is the apparent size of the object at the two-dimensional position, generates an object transformation image by transforming the object image to the size of the object transformation size, pastes the object transformation image at the two-dimensional position of objects in the background image to generate a composite image, and the label generation unit generates a label from the effective pixel range, the object transformation size, and the two-dimensional position of objects in the composite image.It is characterized by generating object pixel range information that indicates the pixel range of an object in a composite image.

[0009] The training data generation device of this disclosure includes an image synthesis unit that generates a three-dimensional background model based on background shooting conditions, sets the area in the three-dimensional background model where objects can be placed as an object placement area, determines the three-dimensional object position, which is the three-dimensional position where an object is placed in the object placement area of ​​the three-dimensional background model, determines the two-dimensional object position, which is the two-dimensional position where an object is placed in the background image, from the three-dimensional object position, calculates the object transformation size, which is the apparent size of the object at the two-dimensional object position, generates an object transformation image by transforming the object image to the size of the object transformation size, pastes the object transformation image onto the two-dimensional object position in the background image to generate a composite image, and a label generation unit that generates information on the object pixel range, which indicates the pixel range of the object in the composite image, as a label from the effective pixel range, the object transformation size, and the two-dimensional object position in the composite image, thereby enabling the generation of high-quality training images that do not feel unnatural in terms of the position and size of the objects, and labels corresponding to the training images.

[0010] This figure shows an example of a background image in Embodiment 1. This figure shows an example of a composite image in Embodiment 1. This figure is for explaining labels in Embodiment 1. This block diagram shows the configuration of the training data generation device according to Embodiment 1. This figure is for explaining background shooting conditions in Embodiment 1. This is a flowchart for explaining the operation of the training data generation device according to Embodiment 1. This figure shows an example of a three-dimensional background model in Embodiment 1. This figure shows an example of a two-dimensional map in Embodiment 1. This figure shows an example of area type and area object type information in Embodiment 1. This figure shows another example of a three-dimensional background model in Embodiment 1. This is a flowchart for explaining the process of determining the three-dimensional background model and object placement area in Embodiment 1. This figure shows yet another example of a three-dimensional background model in Embodiment 1. This is a flowchart for explaining another process of determining the three-dimensional background model and object placement area in Embodiment 1. This block diagram shows the configuration of the training data generation device according to Embodiment 2. This is a flowchart for explaining the operation of the training data generation device according to Embodiment 2. This figure shows an example of a composite image in Embodiment 2. This block diagram shows the configuration of the training data generation device according to Embodiment 3. This is a flowchart for explaining the operation of the training data generation device according to Embodiment 3. This figure shows an example of a composite image in Embodiment 3. This figure is for explaining labels in Embodiment 3. This is a block diagram showing the configuration of the training data generation device according to Embodiment 4. This is a flowchart illustrating the operation of the training data generation device according to Embodiment 4. This is a diagram showing an example of a composite image in Embodiment 4. This is a diagram illustrating the labels in Embodiment 4. This is a block diagram showing the configuration of the training data generation device according to Embodiment 5. This is a flowchart illustrating the operation of the training data generation device according to Embodiment 5. This is a block diagram showing the configuration of the training data generation device according to Embodiment 6. This is a flowchart illustrating the operation of the training data generation device according to Embodiment 6. This is a block diagram showing the configuration of the learning system according to Embodiment 7.This is a schematic diagram showing an example of the hardware configuration of the training data generation device according to Embodiments 1 to 6. This is a schematic diagram showing another example of the hardware configuration of the training data generation device according to Embodiments 1 to 6. This is a schematic diagram showing an example of the hardware configuration of the learning system according to Embodiment 7. This is a schematic diagram showing another example of the hardware configuration of the learning system according to Embodiment 7.

[0011] Hereinafter, a teacher data generation device, a teacher data generation method, and a learning system relating to embodiments for implementing this disclosure will be described in detail with reference to the drawings. In each figure, the same reference numerals indicate the same or corresponding parts.

[0012] Embodiment 1. Figure 1 is a diagram showing an example of a background image 10 in Embodiment 1. Figure 2 is a diagram showing a composite image 11 in Embodiment 1. The composite image 11 is obtained by combining an object image, which is an image of the object to be detected, with the background image 10 by converting it to an object-converted image 12. In the example shown in Figure 2, the object is a "cat". Figure 3 is a diagram showing the object pixel range 13, which indicates the pixel range of the object in the composite image 11 in Embodiment 1. In the following description, the object pixel range 13 will be shown as a bounding box, which is a rectangular area surrounding the range of the object. Figure 3 is a diagram showing the object pixel range 13 superimposed on the composite image 11. The object pixel range 13 is information of a label corresponding to the object. Another piece of information in the label corresponding to the object is "cat," which indicates the type of object. The training data generation device in Embodiment 1 generates a composite image 11 as shown in Figure 2 as a training image, and generates information of the object pixel range 13 as shown in Figure 3 as a label corresponding to the object.

[0013] For the sake of brevity, Figures 1 to 3 are shown as examples, but the generated training data is not limited to what is shown in Figure 2 or Figure 3. There may be multiple types of background images and object images, and since a large amount of training data is actually required, multiple background images and multiple object images may be used in combination. Alternatively, multiple object images may be composited onto a single background image.

[0014] Furthermore, in the following description, the object pixel range 13, which is the label information indicating the object's position, will be described as a bounding box. However, the object pixel range 13 may also be a semantic segmentation label that indicates the object's region in pixel units as the label information indicating the object's position.

[0015] Figure 4 is a block diagram showing the configuration of the training data generation device 100 according to Embodiment 1. The training data generation device 100 according to Embodiment 1 includes a background image acquisition unit 1, a background shooting condition acquisition unit 2, an object image acquisition unit 3, an object shooting condition acquisition unit 4, an image synthesis unit 5, and a label generation unit 6.

[0016] The background image 10 captured by the background camera 20 and the background shooting conditions, which are the shooting conditions for the background image 10 when it was captured by the background camera 20, are stored in the background data storage device 21. The background image acquisition unit 1 acquires the background image 10 from the background data storage device 21. The background shooting condition acquisition unit 2 acquires the background shooting conditions, which are the shooting conditions for the background image 10 when it was captured by the background camera 20, from the background data storage device 21. Figure 5 is a diagram illustrating the background shooting conditions. The background shooting conditions include the camera height 32, which indicates the height of the background camera 20 from the reference plane 31 when the background image 10 was captured, the camera optical axis vertical angle 34, which indicates the vertical angle of the optical axis 33 from the background camera 20 toward the subject when the background image 10 was captured, and the camera field of view 35, which indicates the field of view of the background camera 20 when the background image 10 was captured. The reference plane 31 is, for example, the ground or floor. The reference plane 31 can be any plane on which an object is placed. If the reference plane 31 is the ground, the background shooting conditions may include a two-dimensional map of the location where the background image 10 was taken, a two-dimensional camera position indicating the two-dimensional position of the background shooting camera 20 on the two-dimensional map when the background image 10 was taken, and a camera optical axis azimuth angle indicating the azimuth angle of the optical axis from the background shooting camera 20 toward the subject when the background image 10 was taken, with north as the reference direction. Alternatively, the conditions may include a three-dimensional map of the location where the background image 10 was taken, a two-dimensional camera position indicating the two-dimensional horizontal position of the background shooting camera 20 on the three-dimensional map when the background image 10 was taken, and a camera optical axis azimuth angle indicating the azimuth angle of the optical axis from the background shooting camera 20 toward the subject when the background image 10 was taken, with north as the reference direction.

[0017] Although the background image acquisition unit 1 acquires the background image 10 from the background data storage device 21 and the background shooting condition acquisition unit 2 acquires the background shooting conditions from the background data storage device 21, the background image acquisition unit 1 may acquire the background image 10 directly from the background shooting camera 20, and the background shooting condition acquisition unit 2 may acquire the background shooting conditions, which are the shooting conditions for the background image 10, directly from the background shooting camera 20.

[0018] The object image, which is an image of the object to be detected captured by the object-capture camera 22, and the object-capture conditions, which are the conditions under which the object image was captured when the object was captured by the object-capture camera 22, are stored in the object data storage device 23. The object image acquisition unit 3 acquires the object image, which is an image of the object to be detected, from the object data storage device 23. The object-capture conditions acquisition unit 4 acquires the object-capture conditions, which are the conditions under which the object image was captured when the object was captured by the object-capture camera 22, from the object data storage device 23. The object-capture conditions include the effective pixel range, which is information indicating the extent of the object in the object image. The object-capture conditions further include the object height, which is the actual height of the object when the object image was captured, or the object width, which is the actual width of the object when the object image was captured. The object height is, for example, the length from the bottom edge to the top edge of the object as viewed from the object-capture camera when the object image was captured, and the object width is, for example, the length from the leftmost edge to the rightmost edge of the object as viewed from the object-capture camera when the object image was captured. The object-capture conditions further include object type information indicating the type of object.

[0019] Although the object image acquisition unit 3 acquires object images from the object data storage device 23 and the object shooting condition acquisition unit 4 acquires object shooting conditions from the object data storage device 23, the object image acquisition unit 3 may acquire object images directly from the object shooting camera 22, and the object shooting condition acquisition unit 4 may acquire object shooting conditions for object images directly from the object shooting camera 22.

[0020] The image synthesis unit 5 generates a training image, which is a composite image 11, by combining the background image 10 and the object image from the background image 10, background shooting conditions, object image, and object shooting conditions. The label generation unit 6 generates labels corresponding to the objects in the composite image 11. Details of the image synthesis unit 5 and the label generation unit 6 will be described later.

[0021] The combination of the composite image 11 and the labels corresponding to the objects in the composite image 11 is output to the learning device 24 as training data. The learning device 24 uses the acquired training data to generate a trained model for inferring the range of objects from the image. If the labels in the training data include information about the type of object, the learning device 24 uses the acquired training data to generate a trained model for inferring the range and type of object from the image. The generated trained model is stored in the trained model storage device 25. The object detection device 110 detects the object to be detected from the detection image, which is an image acquired by the detection image camera 26, and outputs information about the object's pixel range, which indicates the pixel range of the object. If the labels in the training data include information about the type of object, the object detection device 110 detects the object to be detected from the detection image, which is an image acquired by the detection image camera 26, and outputs the object's type and information about the object's pixel range, which indicates the pixel range of the object. The object detection device 110 comprises a detection image acquisition unit 111 and an inference unit 112. The object detection device 110 acquires a detection image from the detection image capture camera 26 and outputs it to the inference unit 112. The inference unit 112 acquires the detection image from the detection image acquisition unit 111. The inference unit 112 uses the trained model acquired from the trained model storage device 25 to detect the object to be detected from the detection image, infers information on the object pixel range indicating the pixel range of the object, and outputs the inference result, which is information on the object pixel range indicating the pixel range of the object. If the label of the training data includes information on the type of object, the inference unit 112 uses the trained model acquired from the trained model storage device 25 to detect the object to be detected from the detection image, infers the type of object and information on the object pixel range indicating the pixel range of the object, and outputs the inference result, which is information on the object type and the object pixel range indicating the pixel range of the object.The object detection device 110 is, for example, a monitoring device that detects suspicious persons, vehicles, animals, etc., a medical image analysis device that detects fractures, tumors, abnormal tissue, etc., and a parts detection device that detects parts flowing down a line in a factory, etc.

[0022] Figure 6 is a flowchart illustrating the operation of the training data generation device 100 according to Embodiment 1. Step S01 is a background image acquisition step, step S02 is a background shooting condition acquisition step, step S03 is an object image acquisition step, step S04 is an object shooting condition acquisition step, steps S05 to S08 are image synthesis steps, and step S09 is a label generation step.

[0023] In step S01, the background image acquisition unit 1 acquires the background image 10, outputs the acquired background image 10 to the image synthesis unit 5, and proceeds to step S02. In step S02, the background shooting condition acquisition unit 2 acquires the background shooting conditions, which are the shooting conditions for the background image 10 when the background image 10 was captured by the background shooting camera 20, outputs the acquired background shooting conditions to the image synthesis unit 5, and proceeds to step S03. In step S03, the object image acquisition unit 3 acquires the object image, which is the image of the object to be detected, outputs the acquired object image to the image synthesis unit 5, and proceeds to step S04. In step S04, the object shooting condition acquisition unit 4 acquires the object shooting conditions, which are the shooting conditions for the object image when the object was captured by the object shooting camera 22, outputs the acquired object shooting conditions to the image synthesis unit 5 and the label generation unit 6, and proceeds to step S05.

[0024] In step S05, the image synthesis unit 5 generates a three-dimensional background model of the location shown in the background image 10 based on the background shooting conditions acquired from the background shooting conditions acquisition unit 2 and the object shooting conditions acquired from the object shooting conditions acquisition unit 4, and proceeds to step S06. By generating a three-dimensional background model, in later steps it is possible to determine the position of each pixel in the background image 10 in the three-dimensional space that models the background, and the object image can be synthesized into the background image in a position and size that does not look out of place. Three methods for three-dimensional modeling will be described below.

[0025] 1. When Map Information is Not Available The following describes the process for when map information for the shooting location of the background image 10 is not available as a background shooting condition. In this process, it is assumed that the reference plane of the shooting location is entirely planar. The reference plane is, for example, the ground or floor. The reference plane can be any plane on which objects are placed. If information about the reference plane of the shooting location is available, the obtained information may be reflected in the reference plane of the background three-dimensional model. Next, the positional relationship between the reference plane 31 and the optical axis 33 of the background shooting camera 20 can be determined from the camera height 32 and the camera optical axis vertical angle 34 of the background shooting condition. Furthermore, the projection position of the reference plane in the background image 10 can be determined from the camera field of view 35 of the background shooting condition, and for example, the background three-dimensional model 40 shown in Figure 7 can be generated as a background three-dimensional model corresponding to the background image 10. In the background three-dimensional model 40 shown in Figure 7, the part indicated by the shaded area represents the reference plane. In the example shown in Figure 7, the entire reference plane is set as the object placement area 41, which is the area on which objects can be placed.

[0026] 2. When two-dimensional map information is available, the reference plane is the ground, and the processing when the following are obtained as background shooting conditions: a two-dimensional map of the shooting location of the background image 10, the camera two-dimensional position indicating the two-dimensional position of the background shooting camera 20 on the two-dimensional map when the background image 10 was taken, and the camera optical axis azimuth angle indicating the azimuth angle of the optical axis from the background shooting camera 20 toward the subject when the background image 10 was taken, with north as the reference direction, will be explained. Figure 8 shows an example of a two-dimensional map 50 of the shooting location of the background image 10. In the two-dimensional map 50 shown in Figure 8, the position of the background shooting camera 20 when the background image 10 was taken is the center of the lower end of the two-dimensional map 50, and the position of the background shooting camera 20 on the two-dimensional map 50 can be determined from the camera two-dimensional position. In the two-dimensional map 50 shown in Figure 8, the depth direction in the background image 10 is the upward direction in the two-dimensional map 50, and the correspondence between the depth direction in the background image 10 and the direction in the two-dimensional map 50 can be determined from the camera optical axis azimuth angle. The two-dimensional map 50 includes information on the area type for each region of the shooting location of the background image 10. In Embodiment 1, a grassland area 52 with the area type "grassland", a tree area 51 with the area type "tree", and a tree stump area 53 with the area type "tree stump" are shown. Figure 9 is a diagram showing an example of area type and area object type information in Embodiment 1. The area object type information indicates the type of object that can be placed for each area type, and the image synthesis unit 5 is provided with area object type information in advance. In the example shown in Figure 9, it is shown that an object with object type information "cat" can be placed in the area with the area type "grassland", and that there are no objects that can be placed in the areas with the area type "tree" and "tree stump". The two-dimensional map 50 includes area object type information. Next, the background three-dimensional model 40 shown in Figure 7 is generated using the method described in "1." above. Furthermore, the background three-dimensional model 40a shown in Figure 10 is generated by reflecting the information of the two-dimensional map 50 shown in Figure 8 and the area object type information shown in Figure 9 onto the reference plane of the background three-dimensional model 40 using the information of the two-dimensional position of the camera and the azimuth angle of the camera optical axis.In the background 3D model 40a shown in Figure 10, based on the area object type information shown in Figure 9, the areas where objects of a certain type can be placed are indicated by diagonal lines as object-placement-possible areas 41a, and the areas where objects cannot be placed are indicated as object-unplacement-prohibited areas 42a. In Figure 10, the object is a "cat," and the area type is "grassland," and "cat" is indicated in the area object type information for grassland area 52. Therefore, in the background 3D model 40a of Figure 10, the area corresponding to grassland area 52 in Figure 8 is indicated as object-placement-possible areas 41a.

[0027] Figure 11 shows a flowchart illustrating the process of determining the above-mentioned three-dimensional background model 40a and object placement area 41a. The process shown in Figure 11 is included in the process of step S05 shown in Figure 6. In step S51, the image synthesis unit 5 obtains a two-dimensional map of the shooting location of the background image 10 from the background shooting condition acquisition unit 2 and proceeds to step S52. In step S52, the image synthesis unit 5 obtains object type information by extracting object type information indicating the type of object from the object shooting conditions and proceeds to step S53. In step S53, by comparing the area object type information as shown in Figure 9 with the object type information, the image synthesis unit 5 determines the area on the two-dimensional map where the type of object to be placed can be placed and proceeds to step S54. In step S54, the image synthesis unit 5 generates a three-dimensional background model, for example, by the method shown in "1." above and proceeds to step S55. In step S55, the image synthesis unit 5 obtains information on the camera's two-dimensional position and camera optical axis azimuth angle by extracting this information from the background shooting conditions, and proceeds to step S56. In step S56, based on the placement area determined in step S53, the image synthesis unit 5 sets an object placement area, which is the area in the space of the three-dimensional background model where an object can be placed, and terminates the process.

[0028] 3. When three-dimensional map information is available, the reference plane is the ground, and the processing when the following are obtained as background shooting conditions: a three-dimensional map of the shooting location of the background image 10, the camera's two-dimensional position indicating the horizontal two-dimensional position of the background shooting camera 20 on the three-dimensional map when the background image 10 was taken, and the camera's optical axis azimuth angle indicating the azimuth angle of the optical axis from the background shooting camera 20 toward the subject when the background image 10 was taken, with north as the reference direction. The three-dimensional map shall include information on the topography of the shooting location of the background image 10 and the three-dimensional shape of objects present at the shooting location of the background image 10. The information on the three-dimensional shape may be measured in advance, or it may be based on depth information obtained by depth estimation when the background image 10 was taken by the background shooting camera 20. The three-dimensional map may further be based on three-dimensional shape information obtained using LiDAR (Light Detection and Ranging), etc. The image synthesis unit 5 generates a background three-dimensional model by extracting a portion of the three-dimensional map within the camera's shooting range from the information of the camera's two-dimensional position and the camera's optical axis azimuth angle. The generated three-dimensional model has normal vectors on all surfaces of the terrain of the shooting location and the three-dimensional shape surfaces of objects present at the shooting location. Generally, objects can be placed on a roughly horizontal plane, so the position is on a plane where the normal vector is roughly pointing upwards. Therefore, the object placement area is defined as the region where the height component of the normal vector of the three-dimensional shape surface of the terrain is pointing upwards, and the region where the height component of the normal vector of the three-dimensional shape surface of the object is pointing upwards. Specifically, the object placement area may also be defined as a position on the three-dimensional shape surface of the terrain or the three-dimensional shape surface of an object where the angle between the normal vector and the horizontal plane is greater than or equal to a certain value. For example, the object placement area may be defined as a position where the angle between the normal vector and the horizontal plane is between 60 and 90 degrees. Figure 12 shows a background 3D model 40b generated based on information from a 3D map, and three examples of normal vectors 43 are shown.In the background three-dimensional model 40b shown in Figure 12, by treating the object-unplaceable area 42b as a solid in three-dimensional space, it is possible to set an object-placeable area 41b that takes occlusion into consideration.

[0029] Figure 13 shows a flowchart illustrating the process of determining the above-mentioned background three-dimensional model 40b and object placement area 41b. The process shown in Figure 13 is included in the process of step S05 shown in Figure 6. In step S51a, the image synthesis unit 5 obtains a three-dimensional map of the shooting location of the background image 10 from the background shooting condition acquisition unit 2 and proceeds to step S52a. In step S52a, the image synthesis unit 5 obtains information on the camera two-dimensional position and camera optical axis azimuth angle by extracting information on the camera two-dimensional position and camera optical axis azimuth angle from the background shooting conditions and proceeds to step S53a. In step S53a, the image synthesis unit 5 generates a background three-dimensional model by cutting out a part of the three-dimensional map and proceeds to step S54a. In step S54a, the image synthesis unit 5 determines the normal vector for all shape surfaces of the background three-dimensional model, that is, for all points on the surfaces of terrain and objects in the background three-dimensional model, and proceeds to step S55a. In step S55a, the image synthesis unit 5 sets all shape surfaces whose normal vectors point upward as object placement areas, that is, sets all terrain and object surface points whose normal vectors point upward as object placement areas, and then terminates the process.

[0030] In this example, the object placement area was determined using only the normal vector of the background 3D model. However, it is also possible to set the object placement area by combining the method using the area type and area object type information of the 2D map, as shown in "2.", with the method using the normal vector of the background 3D model. In that case, the area should only be set as an object placement area if it is determined to be object placement possible using both of the above methods.

[0031] In step S06, the image synthesis unit 5 sets the object's three-dimensional position, which is the three-dimensional position in the three-dimensional space of the background three-dimensional model generated in step S05, as the object-placeable region, and proceeds to step S07. The object's three-dimensional position is, for example, the position where the object touches the object-placeable region, which is the reference plane. By setting the object's three-dimensional position as the object-placeable region of the background three-dimensional model, the object can be placed in a three-dimensional position that is visible in the background image and where the object may exist. The object's three-dimensional position may be determined randomly within the range of the object-placeable region.

[0032] In step S07, the image synthesis unit 5 calculates which position in the background image 10 corresponds to the three-dimensional object position in the three-dimensional space of the background three-dimensional model determined in step S06, and determines the two-dimensional object position, which is the placement position of the object in the background image 10. Furthermore, the image synthesis unit 5 uses information on the object height or object width to calculate the object transformation size, which is the apparent size of the object at the two-dimensional object position in the background image 10, and proceeds to step S08.

[0033] In step S08, the image synthesis unit 5 generates an object transformation image 12 by transforming the object image to the size indicated by the object transformation size, pastes the object transformation image 12 onto the two-dimensional object position on the background image 10 to generate a composite image 11, outputs the generated composite image 11, and proceeds to step S09. If there are multiple object images, the order in which each object transformation image 12 is pasted is determined based on the front-to-back relationship of the placement positions of each object in the background three-dimensional model.

[0034] In step S09, the label generation unit 6 generates information for the object pixel range 13, which indicates the pixel range of the object in the composite image 11, as a label from the effective pixel range information, which is information indicating the range of the object in the object image, the apparent size information of the object at the two-dimensional position of the object in the background image 10, and the information of the two-dimensional position of the object in the composite image 11. The generated object pixel range 13 information is then output, and the process ends. The label generation unit 6 may also output the object type information obtained from the object shooting condition acquisition unit 4 as the label information for the corresponding object.

[0035] Furthermore, as background shooting conditions and object shooting conditions, when the background image 10 is captured by the background shooting camera 20 and when the object image is captured by the object shooting camera 22, at least one of the camera parameters of the background shooting camera 20, namely the angle of view, focal length, position, angle, exposure time, shutter speed, ISO sensitivity, aperture value, and F-number, may be acquired as background camera parameters, and the same camera parameters acquired as the background camera parameters may be acquired as object camera parameters, namely the camera parameters of the object shooting camera 22. The object image may then be transformed based on the difference between the same camera parameters included in the background camera parameters and the object camera parameters.

[0036] As described above, the training data generation device 100 according to Embodiment 1 comprises: a background image acquisition unit 1 that acquires a background image 10; a background shooting condition acquisition unit 2 that acquires background shooting conditions which are the shooting conditions for the background image 10; an object image acquisition unit 3 that acquires an object image which is an image of an object to be detected; an object shooting condition acquisition unit 4 that acquires object shooting conditions which are the shooting conditions for the object image; an image synthesis unit 5 that generates a training image which is a composite image 11 of the background image 10 and the object image from the background image 10, background shooting conditions, object image and object shooting conditions; and a label generation unit 6 that generates labels corresponding to the objects in the composite image 11. The background image 10 is captured by a background shooting camera 20, and the background shooting conditions include: a camera height 32 indicating the height from the reference plane 31 of the background shooting camera 20 when the background image 10 was captured; a camera optical axis vertical angle 34 indicating the vertical angle of the optical axis 33 from the background shooting camera 20 toward the subject when the background image 10 was captured; and a camera field of view 35 indicating the field of view of the background shooting camera 20 when the background image 10 was captured. The data includes the effective pixel range, which is information indicating the extent of the object in the object image; further includes the object height, which is the actual height of the object; or the object width, which is the actual width of the object; and further includes object type information indicating the type of object. The image synthesis unit 5 generates a three-dimensional background model based on the background shooting conditions, which is a three-dimensional model of the location shown in the background image 10; sets the area in the background three-dimensional model where an object can be placed as the object placement area; determines the three-dimensional object position, which is the three-dimensional position where the object is placed in the object placement area of ​​the background three-dimensional model; determines the two-dimensional object position, which is the two-dimensional position where the object is placed in the background image 10, which is the two-dimensional position where the object is placed from the three-dimensional object position; finds the object conversion size, which is the apparent size of the object at the two-dimensional object position; generates an object conversion image 12 by converting the object image to the size of the object conversion size; pastes the object conversion image 12 onto the two-dimensional object position in the background image 10 to generate a composite image 11; and the label generation unit 6,By generating information on the object pixel range 13, which indicates the pixel range of the object in the composite image 11, as a label from the effective pixel range, object transformation size, and the two-dimensional position of the object in the composite image 11, it is possible to generate high-quality training images that do not show any inconsistencies in the position and size of the object, and labels corresponding to the training images.

[0037] Embodiment 2. Figure 14 is a block diagram showing the configuration of the training data generation device 100a according to Embodiment 2. Comparing the training data generation device 100a according to Embodiment 2 shown in Figure 14 with the training data generation device 100 according to Embodiment 1 shown in Figure 4, the background shooting condition acquisition unit 2 has become the background shooting condition acquisition unit 2a, and the image synthesis unit 5 has become the image synthesis unit 5a. The other configurations of the training data generation device 100a according to Embodiment 2 are the same as those of the training data generation device 100 according to Embodiment 1.

[0038] Figure 15 is a flowchart illustrating the operation of the training data generation device 100a according to Embodiment 2. In Figure 15, the processes of steps S01, S03 to S07, and S09 are the same as the processes of steps S01, S03 to S07, and S09 of the training data generation device 100 according to Embodiment 1 shown in Figure 6. Step S02a is the background shooting condition acquisition step, and steps S05 to S07 and S08a are image synthesis steps. Below, the differences in the operation of the training data generation device 100a according to Embodiment 2 compared to the operation of the training data generation device 100 according to Embodiment 1 will be explained.

[0039] In step S01, the background image acquisition unit 1 acquires the background image 10, outputs the acquired background image 10 to the image synthesis unit 5a, and proceeds to step S02a. In step S02a, the background shooting condition acquisition unit 2a acquires the background shooting conditions, which are the shooting conditions for the background image 10 when the background image 10 was captured by the background shooting camera 20, outputs the acquired background shooting conditions to the image synthesis unit 5a, and proceeds to step S03. Here, the background shooting conditions include, in addition to the background shooting conditions shown in Embodiment 1, information about when the background image was captured, such as the position of the light source, the luminous intensity of the light source, and the color of the light source. For example, the background shooting conditions include information about when the background image 10 shown in Figure 1 was captured, such as the position of the light source 14, the luminous intensity of the light source 14, and the color of the light source 14.

[0040] The image synthesis unit 5a performs the processing from step S05 to step S07 and proceeds to step S08a. In step S08a, the image synthesis unit 5a generates an object conversion image 12 by converting the object image to the size indicated by the object conversion size, pastes the object conversion image 12 onto the two-dimensional object position on the background image 10, and further draws the shadow of the object on the background image 10 using information on the light source position and light source intensity to generate the composite image 11a shown in Figure 16, outputs the generated composite image 11a and proceeds to step S09. In the composite image 11a of Figure 16, the shadow 15 of the object is shown. In step S08a, for example, in the three-dimensional space of the background three-dimensional model, the three-dimensional shadow, which is the shadow in the three-dimensional space of the background three-dimensional model, may be determined from the placed object, light source position and light source intensity, and the shadow 15 of the object may be determined from the three-dimensional shadow using information on the conversion from the three-dimensional object position to the two-dimensional object position and the conversion from the actual size of the object to the object conversion size. In step S08a, the brightness of the object conversion image 12 may be changed based on the light source intensity information, or the color of the object conversion image 12 may be changed based on the light source color information.

[0041] Note that, as background shooting conditions and object shooting conditions, for each of the light sources when the background image 10 is shot by the background shooting camera 20 and the light source when the object image is shot by the object shooting camera 22, at least one of the light source color which is the color of the light source as the background light source parameter which is the parameter of the light source when the background image 10 is shot, the light source luminous intensity which is the intensity of the light of the light source, and the light source angle which is the angle of the light source with respect to the optical axis of each camera of the light source is acquired, and the same parameter as that acquired as the background light source parameter is acquired as the object light source parameter which is the parameter of the light source when the object image is shot, and the object image may be converted based on the difference of the same parameter included in the background light source parameter and the object light source parameter.

[0042] As described above, in the teacher data generation device 100a according to the second embodiment, the background shooting conditions include the light source position which is the position of the light source 14 when the background image 10 is shot and the light source luminous intensity which is the intensity of the light of the light source 14 when the background image 10 is shot. Since the image composition unit 5a generates the composite image 11a in which the shadow of the object is drawn using the information of the light source position and the light source luminous intensity, the composite image 11a including the shadow 15 of the object can be generated, and high-quality teacher data without a sense of incongruity in the grounding feeling of the object can be obtained.

[0043] Furthermore, in the teacher data generation device 100a according to the second embodiment, since the image composition unit 5a changes the luminance of the object conversion image 12 based on the information of the light source luminous intensity, high-quality teacher data without a sense of incongruity in the luminance of the object can be obtained.

[0044] Furthermore, in the teacher data generation device 100a according to the second embodiment, the background shooting conditions include the light source color which is the color of the light of the light source 14 when the background image 10 is shot. Since the image composition unit 5a changes the color of the object conversion image 12 based on the information of the light source color, high-quality teacher data without a sense of incongruity in the color of the object can be obtained.

[0045] Embodiment 3. Figure 17 is a block diagram showing the configuration of the training data generation device 100b according to Embodiment 3. Comparing the training data generation device 100b according to Embodiment 3 shown in Figure 17 with the training data generation device 100 according to Embodiment 1 shown in Figure 4, the object image acquisition unit 3 has become the object image acquisition unit 3b, the object shooting condition acquisition unit 4 has become the object shooting condition acquisition unit 4b, the image synthesis unit 5 has become the image synthesis unit 5b, and the label generation unit 6 has become the label generation unit 6b. The other configurations of the training data generation device 100b according to Embodiment 3 are the same as those of the training data generation device 100 according to Embodiment 1.

[0046] Figure 18 is a flowchart illustrating the operation of the training data generation device 100b according to Embodiment 3. In Figure 18, the processes of steps S01, S02, and S05 are the same as the processes of steps S01, S02, and S05 of the training data generation device 100 according to Embodiment 1 shown in Figure 6. Step S03b is the object image acquisition step, step S04b is the object shooting condition acquisition step, steps S05, S06b, S07b, and S08b are image synthesis steps, and step S09b is the label generation step. Below, the differences between the operation of the training data generation device 100b according to Embodiment 3 and the operation of the training data generation device 100 according to Embodiment 1 will be explained.

[0047] In step S01, the background image acquisition unit 1 acquires the background image 10 and outputs the acquired background image 10 to the image synthesis unit 5b. In step S02, the background shooting condition acquisition unit 2 acquires the background shooting conditions and outputs the acquired background shooting conditions to the image synthesis unit 5b, and the process proceeds to step S03b. In step S03b, the object image acquisition unit 3b acquires an object image, which is an image of the object to be detected, and also acquires an undetected object image, which is an image of an undetected object that is not to be detected. The acquired object image and undetected object image are output to the image synthesis unit 5b, and the process proceeds to step S04b. In step S04b, the object shooting condition acquisition unit 4b acquires the object shooting conditions, which are the shooting conditions for the object image when an object is photographed by the object shooting camera 22, and also acquires the undetected object shooting conditions, which are the shooting conditions for the undetected object image when an undetected object is photographed by the object shooting camera 22. The acquired object shooting conditions and undetected object shooting conditions are output to the image synthesis unit 5b and the label generation unit 6b, and the process proceeds to step S05. The non-detection object capture conditions include the non-detection object effective pixel range, which is information indicating the extent of the non-detection object in the non-detection object image. The non-detection object capture conditions further include the non-detection object height, which is the actual height of the non-detection object when the non-detection object image was captured, or the non-detection object width, which is the actual width of the non-detection object when the non-detection object image was captured.

[0048] In step S05, the image synthesis unit 5b generates a background three-dimensional model by three-dimensionally modeling the location shown in the background image 10 based on the background shooting conditions acquired from the background shooting condition acquisition unit 2, and proceeds to step S06b. In step S06b, the image synthesis unit 5b sets the object three-dimensional position, which is the three-dimensional position for placing the object in the three-dimensional space of the background three-dimensional model generated in step S05, in the object placement possible region, and also sets the non-detected object three-dimensional position, which is the three-dimensional position for placing the non-detected object in the three-dimensional space of the background three-dimensional model, in the object placement possible region, and proceeds to step S07b.

[0049] In step S07b, the image synthesis unit 5b calculates which position in the background image 10 the object three-dimensional position in the three-dimensional space of the background three-dimensional model determined in step S06b corresponds to, obtains the object two-dimensional position, which is the placement position of the object in the background image 10, and further calculates the object conversion size, which is the apparent size of the object at the object two-dimensional position of the background image 10, using the information on the object height or object width. Next, the image synthesis unit 5b calculates which position in the background image 10 the non-detected object three-dimensional position in the three-dimensional space of the background three-dimensional model determined in step S06b corresponds to, obtains the non-detected object two-dimensional position, which is the placement position of the non-detected object in the background image 10, and further calculates the non-detected object conversion size, which is the apparent size of the non-detected object at the non-detected object two-dimensional position of the background image 10, using the information on the non-detected object height or non-detected object width, and proceeds to step S08b.

[0050] In step S08b, first, the image synthesis unit 5b generates an object transformation image 12 by transforming an object image to the size indicated by the object transformation size, and generates a non-detected object transformation image 16 by transforming a non-detected object image to the size indicated by the non-detected object transformation size. Next, the image synthesis unit 5b determines the pasting order, which is the order in which the object transformation image 12 and the non-detected object transformation image 16 are pasted onto the background image 10, based on information about the three-dimensional positions of objects and the three-dimensional positions of non-detected objects in the background three-dimensional model. For example, the object transformation image 12 or the non-detected object transformation image 16 are pasted in order from those that are further back in the three-dimensional space. The image synthesis unit 5b pastes the object transformation image 12 and the non-detected object transformation image 16 onto the background image 10 according to the determined pasting order to generate a composite image 11b. At this time, the object transformation image 12 is pasted onto the corresponding two-dimensional position of the object, and the non-detected object transformation image 16 is pasted onto the corresponding two-dimensional position of the non-detected object. The image synthesis unit 5b outputs the generated composite image 11b and proceeds to step S09b.

[0051] Figure 19 shows an example of a composite image 11b in Embodiment 3. In the composite image 11b shown in Figure 19, since the box, which is a non-detectable object in the background 3D model, is in front of the cat, which is the object to be detected, the object transformation image 16 is pasted on the background image 10 after the object transformation image 12 is pasted on, so the cat is hidden by the box, and in the composite image 11b, an object transformation image 12b is shown in which only a part of the cat object is visible.

[0052] In step S09b, the label generation unit 6b uses the effective pixel range information, which indicates the extent of the object in the object image, the apparent size information of the object at the object's two-dimensional position in the background image 10, the two-dimensional position information of the object in the composite image 11b, and the display area information of the object transformation image 12b in the composite image 11b to determine the area in the composite image 11b where the object transformation image 12b is shown as a label. It then generates object pixel range 13b information, which indicates the pixel range of the object in the composite image 11b, outputs the generated object pixel range 13b information, and terminates the process. Figure 20 is a diagram illustrating the bounding box, which is a label in Embodiment 3. In the example shown in Figure 20, part of the cat, which is an object, is hidden by the box, which is a non-detected object. Therefore, the object pixel range 13b in Figure 20 does not represent the entire cat, but rather the area of ​​the cat's image that is not hidden by the box.

[0053] As described above, the training data generation device 100b according to Embodiment 3 includes: an object image acquisition unit 3b that acquires an undetected object image, which is an image of an undetected object to be undetected; an object shooting condition acquisition unit 4b that acquires an undetected object shooting condition, which is the shooting condition for the undetected object image; the undetected object shooting condition includes an effective pixel range of the undetected object, which is information indicating the range of the undetected object in the undetected object image; and further includes the undetected object height, which is the actual height of the undetected object when the undetected object image was taken, or the undetected object width, which is the actual width of the undetected object when the undetected object image was taken; an image synthesis unit 5b that generates a training image, which is a composite image 11b of the background image 10, the object image and the undetected object image, from the background image 10, the background shooting condition, the object image, the object shooting condition, the undetected object image and the undetected object shooting condition; and the image synthesis unit 5b arranges the undetected object in the object-placeable region of the background three-dimensional model. The process involves determining the three-dimensional position of the undetected object, determining the two-dimensional position of the undetected object in the background image 10 based on the three-dimensional position of the undetected object, determining the undetected object conversion size, which is the apparent size of the undetected object at the two-dimensional position of the undetected object, generating an undetected object conversion image by converting the undetected object image to the size of the undetected object conversion size, determining the pasting order, which is the order in which the object conversion image and the undetected object conversion image are pasted, pasting the object conversion image 12 at the two-dimensional object position in the background image 10 according to the pasting order, pasting the undetected object conversion image 16 at the two-dimensional position of the undetected object in the background image to generate a composite image 11b, and the label generation unit 6b generates information for the object pixel range 13b in accordance with the display area of ​​the object conversion image 12b in the composite image 11b, so that the undetected object andBy generating a composite image 11b that includes an image of an object partially hidden by an undetected object, and generating information on the object pixel range 13b corresponding to the region where the object image is shown in the composite image 11b, it is possible to generate highly robust training data.

[0054] Embodiment 4. Figure 21 is a block diagram showing the configuration of the training data generation device 100c according to Embodiment 4. Comparing the training data generation device 100c according to Embodiment 4 shown in Figure 21 with the training data generation device 100 according to Embodiment 1 shown in Figure 4, a background object detection unit 7 has been added, the image synthesis unit 5 has become the image synthesis unit 5c, and the label generation unit 6 has become the label generation unit 6c. The other configurations of the training data generation device 100c according to Embodiment 4 are the same as those of the training data generation device 100 according to Embodiment 1.

[0055] Figure 22 is a flowchart illustrating the operation of the training data generation device 100c according to Embodiment 4. In Figure 22, the processes of steps S01 to S04 and step S05 are the same as the processes of steps S01 to S04 and step S05 of the training data generation device 100 according to Embodiment 1 shown in Figure 6. Step S10 is a background object detection step, steps S05, S06c, S07c and S08c are image synthesis steps, and step S09c is a label generation step. Below, the differences between the operation of the training data generation device 100c according to Embodiment 4 and the operation of the training data generation device 100 according to Embodiment 1 will be explained.

[0056] In step S01, the background image acquisition unit 1 acquires the background image 10, outputs the acquired background image 10 to the background object detection unit 7 and the image synthesis unit 5c, and proceeds to step S02. In step S02, the background shooting condition acquisition unit 2 acquires the background shooting conditions, outputs the acquired background shooting conditions to the image synthesis unit 5c, and proceeds to step S03. In step S03, the object image acquisition unit 3 acquires the object image, which is an image of the object to be detected, outputs the acquired object image to the image synthesis unit 5c, and proceeds to step S04. In step S04, the object shooting condition acquisition unit 4 acquires the object shooting conditions, which are the shooting conditions for the object image when the object is captured by the object shooting camera 22, outputs the acquired object shooting conditions to the image synthesis unit 5c and the label generation unit 6c, and proceeds to step S10.

[0057] In step S10, the background object detection unit 7 acquires the background image 10 from the background image acquisition unit 1, detects objects in the background from the background image 10 as background objects, and determines the effective pixel range of the background objects, which is information indicating the extent of the background objects in the background image 10. Furthermore, the background object detection unit 7 extracts the image of the effective pixel range of the background objects from the background image 10 as the background object image, outputs the background object image and the effective pixel range of the background objects to the image synthesis unit 5c, and proceeds to step S05. The means by which the background object detection unit 7 detects background objects from the background image 10 may be a depth estimation method or an existing machine learning model that performs a semantic segmentation task.

[0058] In step S05, the image synthesis unit 5c generates a three-dimensional background model based on the background shooting conditions acquired from the background shooting condition acquisition unit 2, and proceeds to step S06c. In step S06c, the image synthesis unit 5c sets the object placement area as the three-dimensional position where an object is placed in the three-dimensional space of the three-dimensional background model generated in step S05, and also determines the background object three-dimensional position, which is the three-dimensional position of the background object in the three-dimensional space of the background model, and proceeds to step S07c. The background object three-dimensional position may be determined, for example, by setting the position of the lower end of the effective pixel range of the background object in the background image 10 as the position where the background object is in contact with the reference plane. Alternatively, if the depth of the background object is estimated by the background object detection unit 7, the background object three-dimensional position may be determined using the position of the background object in the background image 10 and the depth information of the background object.

[0059] In step S07c, the image synthesis unit 5c calculates which position in the background image 10 corresponds to the three-dimensional object position in the three-dimensional space of the background three-dimensional model determined in step S06c, and obtains the two-dimensional object position, which is the placement position of the object in the background image 10. Furthermore, using information on the object height or object width, it calculates the object transformation size, which is the apparent size of the object at the two-dimensional object position in the background image 10. The image synthesis unit 5c determines the position in the background image 10 where the background object is displayed as the two-dimensional background object position, and proceeds to step S08c.

[0060] In step S08c, the image synthesis unit 5c first generates an object conversion image 12 by converting the object image to the size indicated by the object conversion size. Next, the image synthesis unit 5c determines the pasting order, which is the order in which the object conversion image 12 and the background object images are pasted onto the background image 10, based on information about the three-dimensional positions of the objects in the background three-dimensional model and the three-dimensional positions of the objects in the background. For example, the object conversion image 12 or the background object images are pasted in order from those that are furthest in the three-dimensional space, either the object three-dimensional position or the background object three-dimensional position. The image synthesis unit 5c pastes the object conversion image 12 and the background object images onto the background image 10 according to the determined pasting order to generate a composite image 11c. At this time, the object conversion image 12 is pasted onto the corresponding two-dimensional position of the object, and the background object images are pasted onto the corresponding two-dimensional position of the background object. The image synthesis unit 5c outputs the generated composite image 11c and proceeds to step S09c.

[0061] Figure 23 shows an example of a composite image 11c in Embodiment 4. In the composite image 11c shown in Figure 23, the tree stump, which is an object in the background in the background three-dimensional model, is in front of the cat, which is the object to be detected. Therefore, the object transformation image 17 is pasted onto the background image 10 after the object transformation image 12 is pasted onto the background image 10, so the cat is hidden by the tree stump, and the composite image 11c shows an object transformation image 12c in which only a part of the cat object is shown.

[0062] In step S09c, the label generation unit 6c obtains the area in the composite image 11c where the object transformation image 12c is shown as a label from the effective pixel range information, which is information indicating the extent of the object in the object image; the apparent size information of the object at the two-dimensional position of the object in the background image 10; the two-dimensional position information of the object in the composite image 11c; and the display area information of the object transformation image 12c in the composite image 11c. It then generates information for the object pixel range 13c, which indicates the pixel range of the object in the composite image 11c, outputs the generated object pixel range 13c information, and terminates the process. Figure 24 is a diagram illustrating the object pixel range which is a label in Embodiment 4. In the example shown in Figure 24, part of the cat object is hidden by the tree stump, which is a background object. Therefore, the object pixel range 13c in Figure 24 does not represent the entire cat object, but rather the area of ​​the cat image that is not hidden by the tree stump.

[0063] As described above, the training data generation device 100c according to Embodiment 4 includes a background object detection unit 7 that detects objects in the background from the background image 10 as background objects, determines the effective pixel range of background objects that indicates the range of background objects in the background image 10, and extracts the image of the effective pixel range of background objects in the background image 10 as a background object image. The image synthesis unit 5c determines the three-dimensional position of background objects, which is the three-dimensional position of background objects in the three-dimensional space of the background three-dimensional model, sets the position where background objects are displayed in the background image 10 as the two-dimensional position of background objects, and the object in the background three-dimensional model. Based on information on the three-dimensional position of the object and the three-dimensional position of the object within the background, the paste order is determined, which is the order in which the object transformation image and the object within the background image are pasted onto the background image 10. According to the paste order, the object transformation image is pasted onto the two-dimensional position of the object in the background image 10, and the object within the background image is pasted onto the two-dimensional position of the object within the background image 10 to generate a composite image 11c. The label generation unit 6c generates information on the object pixel range 13c to match the display area of ​​the object transformation image 12c in the composite image 11c, so that highly robust training data can be created that does not feel unnatural in terms of the front-to-back relationship of the object placement.

[0064] Embodiment 5. Figure 25 is a block diagram showing the configuration of the training data generation device 100d according to Embodiment 5. Comparing the training data generation device 100d according to Embodiment 5 shown in Figure 25 with the training data generation device 100 according to Embodiment 1 shown in Figure 4, the object image acquisition unit 3 has become the object image acquisition unit 3d, and the object shooting condition acquisition unit 4 has become the object shooting condition acquisition unit 4d. The other configurations of the training data generation device 100d according to Embodiment 5 are the same as those of the training data generation device 100 according to Embodiment 1. The object data storage device 23 stores the data generated by the fictional object generation device 27. The fictional object generation device 27 generates a fictional object image generated from a three-dimensional model of the object to be detected, and object shooting conditions, which are the virtual shooting conditions when the object image is generated from the three-dimensional model, and stores the generated object image and object shooting conditions in the object data storage device 23.

[0065] Figure 26 is a flowchart illustrating the operation of the training data generation device 100d according to Embodiment 5. In Figure 26, steps S01, S02, and steps S05 to S09 are the same as steps S01, S02, and steps S05 to S09 of the training data generation device 100 according to Embodiment 1 shown in Figure 6. Step S03d is the object image acquisition step, and step S04d is the object shooting condition acquisition step. Below, the differences in the operation of the training data generation device 100d according to Embodiment 5 compared to the operation of the training data generation device 100 according to Embodiment 1 will be explained.

[0066] In step S03d, the object image acquisition unit 3d acquires a hypothetical object image generated from a three-dimensional model of the object to be detected, outputs the acquired object image to the image synthesis unit 5, and proceeds to step S04d. By generating an object image using a three-dimensional model, a variety of object images consistent with the background shooting environment can be easily generated. In step S04d, the object shooting condition acquisition unit 4d acquires the object shooting conditions, which are the virtual shooting conditions when the object image is generated from the three-dimensional model, outputs the acquired object shooting conditions to the image synthesis unit 5 and the label generation unit 6, and proceeds to step S05. In the training data generation device 100d according to Embodiment 5, the effective pixel range indicates the range of the object when the object image is generated using the three-dimensional model, the object height is the height of the hypothetical object assumed when the object image is generated using the three-dimensional model, and the object width is the width of the hypothetical object assumed when the object image is generated using the three-dimensional model.

[0067] In steps S05 to S09, the same processing as in the training data generation device 100 according to Embodiment 1 is performed. The training data generation device 100d according to Embodiment 5 can generate a training image, which is a composite image of the background image 10 and a fictional object image generated from a three-dimensional model of the object to be detected, and can generate labels corresponding to the objects in the generated composite image.

[0068] As described above, in the training data generation device 100d according to Embodiment 5, the object image acquisition unit 3d acquires a fictitious object image generated from a three-dimensional model of an object, and the object shooting condition acquisition unit 4d acquires the object shooting conditions, which are the virtual shooting conditions at the time the object image was generated. As a result, the synthesized object is more consistent with the background shooting environment, and diverse and high-quality training data can be generated.

[0069] Embodiment 6. Figure 27 is a block diagram showing the configuration of the training data generation device 100e according to Embodiment 6. Comparing the training data generation device 100e according to Embodiment 6 shown in Figure 27 with the training data generation device 100 according to Embodiment 1 shown in Figure 4, the object image acquisition unit 3 has become the object image acquisition unit 3e, and the object shooting condition acquisition unit 4 has become the object shooting condition acquisition unit 4e. The other configurations of the training data generation device 100e according to Embodiment 6 are the same as those of the training data generation device 100 according to Embodiment 1. The object data storage device 23 stores the data generated by the fictional object generation device 27. The fictional object generation device 27 generates a fictional object image generated by an image generation AI (Artificial Intelligence) with respect to the object to be detected, and object shooting conditions, which are the virtual shooting conditions when the object image is generated by the image generation AI, and stores the generated object image and object shooting conditions in the object data storage device 23.

[0070] Figure 28 is a flowchart illustrating the operation of the training data generation device 100e according to Embodiment 6. In Figure 28, steps S01, S02, and steps S05 to S09 are the same as steps S01, S02, and steps S05 to S09 of the training data generation device 100 according to Embodiment 1 shown in Figure 6. Step S03e is the object image acquisition step, and step S04e is the object shooting condition acquisition step. Below, the differences in the operation of the training data generation device 100e according to Embodiment 6 compared to the operation of the training data generation device 100 according to Embodiment 1 will be explained.

[0071] In step S03e, the object image acquisition unit 3e acquires a hypothetical object image generated by the image generation AI for the object to be detected, outputs the acquired object image to the image synthesis unit 5, and proceeds to step S04e. The image generation AI uses a so-called t2i (Text-to-Image) model, for example, which takes natural language text as a prompt as input and outputs a digital image that conforms to that prompt. By using the image generation AI, a variety of object images with specific features indicated by the prompt can be easily generated. In step S04e, the object shooting condition acquisition unit 4e acquires the object shooting conditions, which are the hypothetical shooting conditions when the object image is generated by the image generation AI, outputs the acquired object shooting conditions to the image synthesis unit 5 and the label generation unit 6, and proceeds to step S05. In the training data generation device 100e according to Embodiment 6, the effective pixel range indicates the range of the object when the object image is generated by the image generation AI, the object height is the height of a hypothetical object assumed when the object image is generated by the image generation AI, and the object width is the width of a hypothetical object assumed when the object image is generated by the image generation AI. For example, when generating an object image, the effective pixel range, object height, and object width information may be input as prompts, and the information input at this time may be used as the object shooting conditions.

[0072] In steps S05 to S09, the same processing as in the training data generation device 100 according to Embodiment 1 is performed. The training data generation device 100e according to Embodiment 6 can generate a training image, which is a composite image of the background image 10 and a fictional object image generated by the image generation AI, and can generate labels corresponding to the objects in the generated composite image.

[0073] As described above, the training data generation device 100e according to Embodiment 6 has an object image acquisition unit 3e that acquires object images of objects generated by the image generation AI, and an object shooting condition acquisition unit 4e that acquires object shooting conditions, which are the virtual shooting conditions when the object image was generated, so that diverse and high-quality training data can be created.

[0074] Embodiment 7. Figure 29 is a block diagram showing the configuration of the learning system according to Embodiment 7. The learning system 101 according to Embodiment 7 shown in Figure 29 comprises a teacher data generation device 100f, a teacher data receiver 28, and a learning device 24. Compared with the teacher data generation device 100 according to Embodiment 1 shown in Figure 4, the teacher data transmission unit 8 is added to the teacher data generation device 100f in Embodiment 7 shown in Figure 29. The other configurations of the teacher data generation device 100f in Embodiment 7 are the same as those of the teacher data generation device 100 according to Embodiment 1.

[0075] First, the operation of the training data generation device 100f in Embodiment 7 will be explained in terms of how it differs from the operation of the training data generation device 100 in Embodiment 1. The training data, which is a combination of a training image (a composite image 11) generated in the image synthesis unit 5 and a label corresponding to an object in the composite image 11 (a label generated in the label generation unit 6), is output to the training data transmission unit 8. The training data transmission unit 8 transmits the training data, which is a combination of the training image (a composite image 11) and the label, to the training data receiver 28 via communication. The communication means may be wired or wireless.

[0076] The training data receiver 28 receives training data, which is a combination of a training image (a composite image 11) and a label transmitted from the training data transmission unit 8, and outputs the received training data, which is a combination of a training image and a label, to the learning device 24. The learning device 24 uses the acquired training data to generate a trained model for inferring the range and type of an object from an image.

[0077] As described above, the learning system according to Embodiment 7 is a learning system 101 comprising a teacher data generation device 100f, a teacher data receiver 28, and a learning device 24, wherein the teacher data generation device 100f comprises a background image acquisition unit 1 for acquiring a background image 10, a background shooting condition acquisition unit 2 for acquiring background shooting conditions which are the shooting conditions for the background image 10, an object image acquisition unit 3 for acquiring an object image which is an image of an object to be detected, an object shooting condition acquisition unit 4 for acquiring object shooting conditions which are the shooting conditions for the object image, an image synthesis unit 5 for generating a teacher image which is a composite image 11 of the background image 10 and the object image from the background image 10, background shooting conditions, object image and object shooting conditions, a label generation unit 6 for generating labels corresponding to objects in the composite image 11, and a teacher data transmission unit 8 for transmitting the composite image 11 and labels to the teacher data receiver 28, wherein the background image 10 is captured by a background shooting camera 20, and the background shooting conditions include a camera height 32 which indicates the height from the reference plane of the background shooting camera 20 when the background image 10 was captured, and the background image 10 was captured The object shooting conditions include the camera optical axis vertical angle 34, which indicates the vertical angle of the optical axis from the background shooting camera 20 toward the subject, and the camera field of view 35, which indicates the field of view of the background shooting camera 20 when the background image 10 was captured. The object shooting conditions include the effective pixel range, which is information indicating the range of the object in the object image, and further include the object height, which is the actual height of the object, or the object width, which is the actual width of the object, and further include object type information indicating the type of object. The image synthesis unit 5 generates a three-dimensional background model by three-dimensionally modeling the locations shown in the background image 10 based on the background shooting conditions, sets the area in the three-dimensional background model where an object can be placed as the object placement area, determines the object three-dimensional position, which is the three-dimensional position where the object is placed in the object placement area of ​​the three-dimensional background model, determines the object two-dimensional position, which is the two-dimensional position where the object is placed in the background image 10 from the object three-dimensional position, and finds the object transformation size, which is the apparent size of the object at the object two-dimensional position.An object-converted image 12 is generated by converting the object image to the object-conversion size, and a composite image 11 is generated by pasting the object-converted image 12 onto the two-dimensional object position on the background image 10. The label generation unit 6 generates information of the object pixel range 13, which indicates the pixel range of the object in the composite image 11, as a label from the effective pixel range, the object-converted size, and the two-dimensional object position in the composite image 11. The training data receiver 28 receives the composite image 11 and the label, and outputs the received composite image 11 and label to the learning device 24. The learning device 24 generates a trained model using the acquired composite image 11 and label. Therefore, even if the training data generation device 100f is located far from the learning device 24, the learning device 24 can acquire the training data and generate a trained model.

[0078] Figure 30 is a schematic diagram showing an example of the hardware configuration of the training data generation devices 100, 100a, 100b, 100c, 100d, and 100e according to Embodiments 1 to 6. The background image acquisition unit 1, background shooting condition acquisition units 2, 2a, object image acquisition units 3, 3b, 3d, 3e, object shooting condition acquisition units 4, 4b, 4d, 4e, image synthesis units 5, 5a, 5b, 5c, label generation units 6, 6b, 6c, and background object detection unit 7 are implemented by a processor 201 such as a CPU (Central Processing Unit) that executes a program stored in memory 202. Memory 202 is also used as a temporary storage device in each process executed by the processor 201. Memory 202 is, for example, a non-volatile or volatile semiconductor memory such as RAM, ROM, flash memory, or EPROM, a magnetic disk, an optical disk, or a combination thereof. The processor 201 and memory 202 are connected to the bus 203. The background data storage device 21, object data storage device 23, and learning device 24 may also be connected to the bus 203, and the background camera 20, object camera 22, and learning device 24 may also be connected to the bus 203.

[0079] Figure 31 is a schematic diagram showing another example of the hardware configuration of the training data generation devices 100, 100a, 100b, 100c, 100d, and 100e according to Embodiments 1 to 6. In Figure 31, the processing circuit 204 is connected to the bus 203. When the processing circuit 204 is dedicated hardware, examples include a single circuit, a composite circuit, a programmed processor, an ASIC (Application Specific Integrated Circuit), an FPGA (Field Programmable Gate Array), or a combination thereof. Each of the background image acquisition unit 1, background shooting condition acquisition unit 2, 2a, object image acquisition unit 3, 3b, 3d, 3e, object shooting condition acquisition unit 4, 4b, 4d, 4e, image synthesis unit 5, 5a, 5b, 5c, label generation unit 6, 6b, 6c, and background object detection unit 7 may be implemented by the processing circuit 204, or the background image acquisition unit 1, background shooting condition acquisition unit 2, 2a, object image acquisition unit 3, 3b, 3d, 3e, object shooting condition acquisition unit 4, 4b, 4d, 4e, image synthesis unit 5, 5a, 5b, 5c, label generation unit 6, 6b, 6c, and background object detection unit 7 may be implemented together by the processing circuit 204. Furthermore, the background image acquisition unit 1, background shooting condition acquisition units 2, 2a, object image acquisition units 3, 3b, 3d, 3e, object shooting condition acquisition units 4, 4b, 4d, 4e, image synthesis units 5, 5a, 5b, 5c, label generation units 6, 6b, 6c, and background object detection unit 7 may be partially implemented by dedicated hardware and partially implemented by software or firmware. The background data storage device 21, object data storage device 23, and learning device 24 may be connected to the bus 203, and the background shooting camera 20, object shooting camera 22, and learning device 24 may be connected to the bus 203.

[0080] Figure 32 is a schematic diagram showing an example of the hardware configuration of the learning system 101 in Embodiment 7. The teacher data transmission unit 8 is implemented by a teacher data transmitter 205. The background image acquisition unit 1, background shooting condition acquisition unit 2, object image acquisition unit 3, object shooting condition acquisition unit 4, image synthesis unit 5, and label generation unit 6 are implemented by a processor 201a such as a CPU that executes a program stored in memory 202a. Memory 202a is also used as a temporary storage device in each process executed by the processor 201a. Memory 202a is, for example, a non-volatile or volatile semiconductor memory such as RAM, ROM, flash memory, or EPROM, a magnetic disk, an optical disk, or a combination thereof. The processor 201a, memory 202a, and teacher data transmitter 205 are connected to a bus 203a. The teacher data transmitter 205 is connected to a teacher data receiver 28 by wired or wireless connection, and the teacher data receiver 28 is connected to a learning device 24. The background data storage device 21 and the object data storage device 23 may be connected to the bus 203, and the background camera 20 and the object camera 22 may also be connected to the bus 203.

[0081] Figure 33 is a schematic diagram showing another example of the hardware configuration of the learning system 101 in Embodiment 7. In Figure 33, the processing circuit 204a and the teacher data transmitter 205 are connected to the bus 203a. When the processing circuit 204a is dedicated hardware, examples include a single circuit, a composite circuit, a programmed processor, an ASIC, an FPGA, or a combination thereof. The background image acquisition unit 1, the background shooting condition acquisition unit 2, the object image acquisition unit 3, the object shooting condition acquisition unit 4, the image synthesis unit 5, and the label generation unit 6 may each be implemented by the processing circuit 204a, or the background image acquisition unit 1, the background shooting condition acquisition unit 2, the object image acquisition unit 3, the object shooting condition acquisition unit 4, the image synthesis unit 5, and the label generation unit 6 may be implemented together by the processing circuit 204a. Furthermore, the background image acquisition unit 1, the background shooting condition acquisition unit 2, the object image acquisition unit 3, the object shooting condition acquisition unit 4, the image synthesis unit 5, and the label generation unit 6 may be partially implemented by dedicated hardware and partially implemented by software or firmware. The background data storage device 21 and the object data storage device 23 may be connected to the bus 203, and the background camera 20 and the object camera 22 may also be connected to the bus 203.

[0082] While this disclosure describes various exemplary embodiments, the various features, aspects, and functions described in one or more embodiments are not limited to the application of a particular embodiment, but are applicable individually or in various combinations to the embodiments. Therefore, countless variations not illustrated are conceivable within the scope of the art disclosed in this specification. These include, for example, modifying, adding, or omitting at least one component, or even extracting at least one component and combining it with components from other embodiments.

[0083] 1 Background image acquisition unit, 2, 2a Background shooting condition acquisition unit, 3, 3b, 3d, 3e Object image acquisition unit, 4, 4b, 4d, 4e Object shooting condition acquisition unit, 5, 5a, 5b, 5c Image synthesis unit, 6, 6b, 6c Label generation unit, 7 Background object detection unit, 8 Training data transmission unit, 10 Background image, 11, 11a, 11b, 11c Synthesis image, 12, 12b, 12c Object transformation image, 13, 13b, 13c Object pixel range, 14 Light source, 15 Object shadow, 16 Non-detected object transformation image, 17 Background object image, 20 Background shooting camera, 21 Background data storage device, 22 Object shooting camera, 23 Object data storage device, 24 Learning device, 25 Trained model storage device, 26 Detected image shooting camera, 27 Fictional object generation device, 28 Training data receiver, 31 Reference plane, 32 Camera height, 33 Optical axis, 34 Camera optical axis vertical angle, 35 Camera field of view, 40, 40a, 40b Background 3D model, 41, 41a, 41b Object placement area, 42a, 42b Object placement non-area, 43 Normal vector, 50 2D map, 51 Tree area, 52 Grassland area, 53 Stump area, 100, 100a, 100b, 100c, 100d, 100e, 100f Training data generator, 101 Learning system, 201, 201a Processor, 202, 202a Memory, 203, 203a Bus, 204, 204a Processing circuit, 205 Training data transmitter.

Claims

1. The system comprises: a background image acquisition unit for acquiring a background image; a background shooting condition acquisition unit for acquiring background shooting conditions which are the shooting conditions for the background image; an object image acquisition unit for acquiring an object image which is an image of an object to be detected; an object shooting condition acquisition unit for acquiring object shooting conditions which are the shooting conditions for the object image; an image synthesis unit for generating a training image which is a composite image of the background image and the object image from the background image, the background shooting conditions, the object image and the object shooting conditions; and a label generation unit for generating a label corresponding to the object in the composite image, wherein the background image is captured by a background shooting camera, and the background shooting conditions include: a camera height indicating the height of the background shooting camera from a reference plane when the background image was captured; a camera optical axis vertical angle indicating the vertical angle of the optical axis from the background shooting camera toward the subject when the background image was captured; and a camera field of view indicating the field of view of the background shooting camera when the background image was captured. The object shooting conditions include an effective pixel range which is information indicating the extent of the object in the object image, and further include an object height which is the actual height of the object, or an object width which is the actual width of the object, and further include object type information which indicates the type of the object. The image synthesis unit generates a three-dimensional background model which is a three-dimensional model of the location shown in the background image based on the background shooting conditions, sets the area in the background three-dimensional model where the object can be placed as an object-placeable area, determines the three-dimensional object position which is the three-dimensional position where the object is placed in the object-placeable area of ​​the background three-dimensional model, determines the two-dimensional object position which is the two-dimensional position where the object is placed in the background image from the three-dimensional object position, finds the object conversion size which is the apparent size of the object at the two-dimensional object position, and generates an object conversion image which is the object image converted to the size of the object conversion size.A training data generation device characterized in that it generates a composite image by pasting the object conversion image onto the two-dimensional object position in the background image, and the label generation unit generates information of the object pixel range indicating the pixel range of the object in the composite image as the label, based on the effective pixel range, the object conversion size, and the two-dimensional object position in the composite image.

2. The training data generation apparatus according to claim 1, wherein the reference plane is the ground or floor surface, and the image synthesis unit sets the area of ​​the reference plane in the background three-dimensional model to the object placement area.

3. The training data generation device according to claim 1, wherein the background shooting conditions include a two-dimensional map of the location where the background image is shot, a two-dimensional camera position indicating the two-dimensional position of the background shooting camera on the two-dimensional map when the background image is shot, and a camera optical axis azimuth angle indicating the azimuth angle of the optical axis from the background shooting camera toward the subject when the background image is shot, the two-dimensional map includes information on the type of area for each area of ​​the location where the background image is shot, the image synthesis unit is provided in advance with area object type information, which is information on the type of object that can be placed for each area type, and the object placement area is set by comparing the object type information of the object with the area object type information, thereby setting the object placement area, which is the area in the background three-dimensional model where the object can be placed.

4. The training data generation apparatus according to claim 1, wherein the background shooting conditions include a three-dimensional map of the location where the background image is taken, a two-dimensional camera position indicating the two-dimensional horizontal position of the background shooting camera on the three-dimensional map when the background image is taken, and a camera optical axis azimuth angle indicating the azimuth angle of the optical axis from the background shooting camera toward the subject when the background image is taken, the three-dimensional map includes topographic information of the location where the background image is taken and information on the three-dimensional shape of objects present at the location where the background image is taken, and the object placement area is characterized in that, in the background three-dimensional model, at least the height component of the normal vector of the three-dimensional shape surface of the terrain points upward and the height component of the normal vector of the three-dimensional shape surface of the object points upward.

5. The training data generation apparatus according to any one of claims 1 to 4, wherein the background shooting conditions include a light source position, which is the position of the light source when the background image is captured, and a light source intensity, which is the intensity of the light from the light source when the background image is captured, and the image synthesis unit generates the composite image in which the shadow of the object is drawn using the information of the light source position and the light source intensity.

6. The training data generation device according to claim 5, characterized in that the image synthesis unit changes the brightness of the object conversion image based on the light source luminosity information.

7. The training data generation device according to claim 5 or 6, wherein the background shooting conditions include the light source color, which is the color of the light from the light source when the background image was captured, and the image synthesis unit changes the color of the object conversion image based on the information of the light source color.

8. The object image acquisition unit acquires an undetected object image, which is an image of an undetected object that is not to be detected; the object shooting condition acquisition unit acquires an undetected object shooting condition, which is the shooting condition for the undetected object image; the undetected object shooting condition includes an effective pixel range of the undetected object, which is information indicating the range of the undetected object in the undetected object image; and further includes an undetected object height, which is the actual height of the undetected object when the undetected object image was taken, or an undetected object width, which is the actual width of the undetected object when the undetected object image was taken; the image synthesis unit generates a training image, which is a composite image of the background image, the object image, and the undetected object image, from the background image, the background shooting condition, the object image, the object shooting condition, the undetected object image, and the undetected object shooting condition; the image synthesis unit determines the three-dimensional position of the undetected object, which is the three-dimensional position in which the undetected object is placed in the object-placeable region of the background three-dimensional model. The following steps are performed: determine the two-dimensional position of the undetected object in the background image from the three-dimensional position of the undetected object; determine the undetected object conversion size, which is the apparent size of the undetected object at the two-dimensional position of the undetected object; generate an undetected object conversion image by converting the undetected object image to the size of the undetected object conversion size; determine the pasting order, which is the order in which the object conversion image and the undetected object conversion image are pasted onto the background image, from the information of the three-dimensional position of the object in the background three-dimensional model and the three-dimensional position of the undetected object; paste the object conversion image onto the two-dimensional position of the object in the background image according to the pasting order; and generate the composite image by pasting the undetected object conversion image onto the two-dimensional position of the undetected object in the background image.The training data generation device according to any one of claims 1 to 4, characterized in that the label generation unit generates information on the object pixel range in accordance with the display area of ​​the object conversion image in the composite image.

9. A training data generation device according to any one of claims 1 to 4, comprising: a background object detection unit that detects objects in the background from the background image as background objects, determines an effective pixel range for background objects that indicates the range of the background objects in the background image, and extracts an image of the effective pixel range for background objects in the background image as a background object image; the image synthesis unit determines the three-dimensional position of the background object, which is the three-dimensional position of the background object in the three-dimensional space of the three-dimensional background model; the position in which the background object is displayed in the background image is defined as the two-dimensional position of the background object; the paste order, which is the order in which the object conversion image and the background object image are pasted onto the background image, is determined from the information of the three-dimensional position of the object in the three-dimensional background model and the three-dimensional position of the background object; the object conversion image is pasted onto the two-dimensional position of the object in the background image according to the paste order, and the background object image is pasted onto the two-dimensional position of the background object in the background image to generate the composite image; and the label generation unit generates information of the object pixel range in accordance with the display area of ​​the object conversion image in the composite image.

10. The training data generation apparatus according to any one of claims 1 to 4, characterized in that the object image acquisition unit acquires the object image generated from the three-dimensional model of the object, and the object shooting condition acquisition unit acquires the object shooting conditions, which are the virtual shooting conditions when the object image was generated.

11. The training data generation device according to any one of claims 1 to 4, characterized in that the object image acquisition unit acquires the object image generated by the image generation AI, and the object shooting condition acquisition unit acquires the object shooting conditions, which are the virtual shooting conditions when the object image was generated.

12. The process includes: a background image acquisition step for acquiring a background image; a background shooting condition acquisition step for acquiring background shooting conditions which are the shooting conditions for the background image; an object image acquisition step for acquiring an object image which is an image of an object to be detected; an object shooting condition acquisition step for acquiring object shooting conditions which are the shooting conditions for the object image; an image synthesis step for generating a training image which is a composite image of the background image and the object image from the background image, the background shooting conditions, the object image and the object shooting conditions; and a label generation step for generating a label corresponding to the object in the composite image, wherein the background image is captured by a background shooting camera, and the background shooting conditions include: a camera height indicating the height of the background shooting camera from the reference plane when the background image was captured; a camera optical axis vertical angle indicating the vertical angle of the optical axis from the background shooting camera toward the subject when the background image was captured; and a camera field of view indicating the field of view of the background shooting camera when the background image was captured. The object shooting conditions include an effective pixel range which is information indicating the extent of the object in the object image, and further include an object height which is the actual height of the object when the object image was taken, or an object width which is the actual width of the object when the object image was taken, and further include object type information which indicates the type of the object, and the image synthesis step includes generating a background 3D model which is a 3D model of the location shown in the background image based on the background shooting conditions, setting the area in the background 3D model where the object can be placed as an object-placeable area, determining the object 3D position which is the 3D position in which the object is placed in the object-placeable area of ​​the background 3D model, determining the object 2D position which is the 2D position in which the object is placed in the background image from the object 3D position, and determining the object transformation size which is the apparent size of the object at the object 2D position.A method for generating training data, characterized in that: an object conversion image is generated by converting the object image to the size of the object conversion size; the object conversion image is pasted onto the two-dimensional position of the object in the background image to generate the composite image; and the label generation step generates information on the object pixel range, which indicates the pixel range of the object in the composite image, as the label, from the effective pixel range, the object conversion size, and the two-dimensional position of the object in the composite image.

13. A learning system comprising a training data generation device, a training data receiver, and a learning device, wherein the training data generation device comprises: a background image acquisition unit for acquiring a background image; a background shooting condition acquisition unit for acquiring background shooting conditions which are the shooting conditions for the background image; an object image acquisition unit for acquiring an object image which is an image of an object to be detected; an object shooting condition acquisition unit for acquiring object shooting conditions which are the shooting conditions for the object image; an image synthesis unit for generating a training image which is a composite image of the background image and the object image from the background image, the background shooting conditions, the object image and the object shooting conditions; a label generation unit for generating a label corresponding to the object in the composite image; and a training data transmission unit for transmitting the composite image and the label to the training data receiver, wherein the background image is captured by a background shooting camera, and the background shooting conditions include: a camera height indicating the height of the background shooting camera from a reference plane when the background image was captured; a camera optical axis vertical angle indicating the angle in the vertical direction of the optical axis from the background shooting camera toward the subject when the background image was captured; and a camera field of view indicating the field of view of the background shooting camera when the background image was captured. The object shooting conditions include an effective pixel range which is information indicating the extent of the object in the object image, and further include an object height which is the actual height of the object, or an object width which is the actual width of the object, and further include object type information which indicates the type of the object. The image synthesis unit generates a three-dimensional background model based on the background shooting conditions, which is a three-dimensional model of the location shown in the background image, sets the area in the background three-dimensional model where the object can be placed as an object placement area, determines the three-dimensional object position which is the three-dimensional position where the object is placed in the object placement area of ​​the background three-dimensional model, and determines the two-dimensional object position which is the two-dimensional position where the object is placed in the background image from the three-dimensional object position.A learning system characterized by: determining the object transformation size, which is the apparent size of the object at the object's two-dimensional position; generating an object transformation image by transforming the object image to the size of the object transformation size; pasting the object transformation image onto the object's two-dimensional position in the background image to generate a composite image; the label generation unit generating object pixel range information indicating the pixel range of the object in the composite image as the label, based on the effective pixel range, the object transformation size, and the object's two-dimensional position in the composite image; the training data receiver receiving the composite image and the label, and outputting the received composite image and the label to the learning device; and the learning device generating a trained model using the acquired composite image and the label.