Training data generation device, training data generation method, and program
The learning data generation device generates images with controlled occlusion rates to enhance the discriminability of hidden object parts, addressing the limitations of conventional learning data in machine learning-based image recognition systems.
Patent Information
- Application Number
- PCT/JP2024/045369
- Authority / Receiving Office
- WO · WO
- Patent Type
- Applications
- Current Assignee / Owner
- Priority Date
- 2023-12-25
- Filing Date
- 2024-12-23
- Publication Date
- 2025-07-03
AI Technical Summary
Conventional image recognition technologies using machine learning face challenges in improving recognition accuracy due to the limited availability of diverse and effective learning data, particularly when parts of objects are hidden, as existing methods often utilize images with occlusion rates below a threshold, leading to suboptimal training data.
A learning data generation device and method that generates images where a part of a target model is masked by a hidden model, ensuring the occlusion rate of specified feature regions remains below a threshold, thereby preserving the discriminability of the target model.
This approach allows for the creation of accurate learning data that enhances the recognition model's ability to identify objects even when parts are hidden, improving recognition accuracy and applicability in real-world scenarios.
Smart Images

Figure JP2024045369_03072025_PF_FP_ABST
Abstract
Description
Training data generation device, training data generation method, and program
[0001] The present disclosure relates to a training data generation device, a training data generation method, and a program.
[0002] Image recognition techniques using machine learning such as deep learning have been known for some time. In deep learning, it is known that the recognition accuracy can be improved as the number and types of training data (e.g., training data) used to train a recognition model increase.
[0003] Therefore, even when recognizing a recognition target that is partially occluded in an image, highly accurate recognition can be achieved by training a recognition model in advance using a large number of different images in which the target is partially occluded as training data. However, it is not realistic to prepare such a large number of different training data using only real images.
[0004] For this reason, for example, Patent Document 1 discloses a technology that uses CG (Computer Graphics) to generate, as training data, an image in which a part of a target 3D model is hidden by a background 3D model.
[0005] Japanese Patent Application Laid-Open No. 2018-163554
[0006] However, the conventional techniques described above only use images whose occlusion rate of a target 3D model is equal to or less than a threshold as training data, which often results in images that are not useful for training a recognition model being used as training data.
[0007] The present disclosure has been made in consideration of the above circumstances, and aims to provide a training data generation device, a training data generation method, and a program that can accurately generate images useful for training a recognition model as training data.
[0008] One aspect of the training data generation device of the present disclosure includes a memory that stores a target model and a occlusion model, and a processor that generates, as training data, an image in which a part of the target model is occluded by the occlusion model so that the occlusion rate of a feature region specified on a screen displaying the target model is less than a first threshold.
[0009] One aspect of the training data generation method of the present disclosure includes the steps of: a processor reading a target model and an occlusion model from a memory; and the processor generating, as training data, an image in which a part of the target model is occluded by the occlusion model so that the occlusion rate of a feature region specified on a screen displaying the target model is less than a first threshold.
[0010] One aspect of the program of the present disclosure is for causing a computer to execute the steps of reading a target model and an occlusion model from memory, and generating, as training data, an image in which a portion of the target model is occluded by the occlusion model so that the occlusion rate of a feature region specified on a screen displaying the target model is less than a first threshold value.
[0011] According to the present disclosure, it is possible to provide a training data generation device, a training data generation method, and a program that can accurately generate images useful for training a recognition model as training data.
[0012] FIG. 1 is a block diagram showing an example of the hardware configuration of a training data generation device according to this embodiment. FIG. 2 is a block diagram showing an example of the functional configuration of a training data generation device according to this embodiment. FIG. 3 is a diagram showing an example of a target model according to this embodiment. FIG. 4 is a diagram showing an example of a target model according to this embodiment. FIG. 5 is a diagram showing an example of an occlusion model according to this embodiment. FIG. 6 is a diagram showing an example of an occlusion model according to this embodiment. FIG. 7 is a diagram showing an example of an occlusion model list according to this embodiment. FIG. 8 is a diagram showing an example of a background image according to this embodiment. FIG. 9 is a diagram showing an example of a background image list according to this embodiment. FIG. 10 is a diagram showing an example of a setting screen for generating training data according to this embodiment. FIG. 11 is a diagram showing an example of an updated setting screen according to this embodiment. FIG. 12 is a diagram showing an example of a state in which feature regions are set for a target model on the display area of the setting screen according to this embodiment. FIG. 13 is a diagram showing an example of an occlusion region specified for a target model according to this embodiment. FIG. 14 is a plan view for explaining an example of a method for generating an image by an image generation unit according to this embodiment. FIG. 15 is a diagram showing an example of an image generated by the image generation unit according to this embodiment. FIG. 16 is a diagram showing an example of an image generated by the image generation unit according to this embodiment. FIG. 17 is a diagram showing an example of an image generated by the image generation unit according to this embodiment. FIG. 18 is a diagram showing an example of an image generated by the image generation unit of this embodiment. FIG. 19 is a diagram showing an example of an image generated by the image generation unit of this embodiment. FIG. 20 is a flowchart showing an example of a learning data generation process performed in the learning data generation device of this embodiment. FIG. 21 is a diagram showing an example of a display area of Modification 2. FIG. 22 is a plan view showing an example of a moving image generation method by the image generation unit of Modification 3. FIG. 23 is a diagram showing an example of a moving image generated by the image generation unit of Modification 3. FIG. 24 is a diagram showing an example of a moving image generated by the image generation unit of Modification 3.
[0013] Hereinafter, an embodiment of the present disclosure (hereinafter simply referred to as "the present embodiment") will be described in detail with reference to the drawings. Note that the present disclosure is not limited to the following embodiment. Furthermore, the following embodiment and modified examples can be combined as appropriate.
[0014] The training data generation device of this embodiment generates images in which a portion of a target model is occluded by an occlusion model, and outputs the images as training data. Specifically, the training data generation device of this embodiment sets a characteristic portion of the target model as a feature region, and outputs images in which the occlusion rate of the feature region is less than a threshold value as training data.
[0015] In this manner, in this embodiment, images in which the characteristic portions of the target model remain unoccluded to a certain extent are used as training data. That is, in this embodiment, images that are likely to retain their distinctiveness as a target model even if some portions are occluded are used as training data. Therefore, according to this embodiment, images that are useful for training a recognition model that is also used to recognize a recognition target that is partially occluded can be generated with high accuracy as training data.
[0016] In the following, an example will be described in which the training data generation device of this embodiment generates training data used to train a recognition model for recognizing official vehicles using CG (Computer Graphics), but the present invention is not limited to this. Note that official vehicles refer to official vehicles in a broad sense, such as emergency vehicles such as police cars and ambulances, and route buses that are used for public transportation.
[0017] 1 is a block diagram showing an example of the hardware configuration of a training data generation device 10 of this embodiment. As shown in FIG. 1, the training data generation device 10 includes a control device 11, a main memory device 13, an auxiliary memory device 15, a display device 17, an input device 19, a communication device 21, and various buses 23. The control device 11, the main memory device 13, the auxiliary memory device 15, the display device 17, the input device 19, and the communication device 21 are connected via the various buses 23. As such, the training data generation device 10 of this embodiment has a general hardware configuration using a typical computer.
[0018] The control device 11 controls the overall operation of the training data generation device 10. The control device 11 may be, for example, at least one of a central processing unit (CPU) and a graphics processing unit (GPU), but is not limited to these. There may be any number of CPUs or GPUs as long as they are one or more, and they may be single-core or multi-core.
[0019] Examples of the main storage device 13 include, but are not limited to, a read-only memory (ROM) and a random access memory (RAM). The ROM stores various programs, such as a program for controlling the training data generation device 10 and a program for generating training data according to this embodiment. The RAM is used as a working area when the control device 11 performs various controls based on the programs stored in the ROM.
[0020] The auxiliary storage device 15 stores the various programs described above and data for generating training data in this embodiment. The various programs described above may be stored in at least one of the main storage device 13 and the auxiliary storage device 15. Examples of the auxiliary storage device 15 include, but are not limited to, existing storage devices capable of magnetic, electrical, or optical storage, such as a hard disk drive (HDD), a solid state drive (SSD), and a digital versatile disc (DVD). The auxiliary storage device 15 may be built into the training data generation device 10 or may be externally attached to the training data generation device 10 via an interface such as a universal serial bus (USB). The auxiliary storage device 15 may also be a network-attached storage (NAS) connected via a network such as a local area network (LAN) or a wide area network (WAN).
[0021] The display device 17 displays various screens used when the training data generation device 10 generates training data, and serves as a user interface with the user (operator). Examples of the display device 17 include, but are not limited to, various displays such as a liquid crystal display, an organic electroluminescence (EL) display, and a touch panel display. The display device 17 may be a built-in display built into the training data generation device 10, or an external display connected to the training data generation device 10 via a display interface such as HDMI (registered trademark).
[0022] The input device 19 is used for various inputs, selections, and specifications used when the training data generation device 10 generates training data, and serves as a user interface with the user (operator). Examples of the input device 19 include, but are not limited to, a keyboard, a mouse, and a touch panel. The input device 19 may be built into the training data generation device 10 or may be externally attached to the training data generation device 10 via an interface such as a USB.
[0023] Examples of the communication device 21 include, but are not limited to, a communication device for a wired LAN and a wireless communication device for a wireless LAN. The communication device 21 may be used to externally acquire the program or data for generating the training data of this embodiment, or may be used to externally output the training data generated by the training data generation device 10.
[0024] In addition to the above configuration, the training data generation device 10 may further include hardwired circuits specific to the training data generation device 10, such as an IC (Integrated Circuit), an ASIC (Application Specific Integrated Circuit), and an FPGA (Field-Programmable Gate Array), in order to realize the training data generation function.
[0025] 2 is a block diagram showing an example of the functional configuration of the training data generation device 10 according to this embodiment. As shown in FIG. 2 , the training data generation device 10 includes an operation unit 101, an operation reception unit 103, a display control unit 105, a display unit 107, an image acquisition unit 109, a target model storage unit 111, an occlusion model storage unit 115, a background image storage unit 119, a screen generation unit 123, a condition setting unit 125, an image generation unit 131, an image (training data) output unit 133, and a training data storage unit 135. The image acquisition unit 109 includes the target model acquisition unit 113, the occlusion model acquisition unit 117, and the background image acquisition unit 121. The condition setting unit 125 includes a feature region setting unit 127 and an obscuration ratio setting unit 129.
[0026] The operation unit 101 can be realized by, for example, the input device 19 described with reference to FIG.
[0027] The operation reception unit 103, the display control unit 105, the image acquisition unit 109, the screen generation unit 123, the condition setting unit 125, the image generation unit 131, and the image output unit 133 can be realized, for example, by the control device 11 and the main memory device 13 described in FIG. 1.
[0028] For example, the control device 11 reads out a program for generating learning data according to this embodiment that is stored in the main storage device 13 (ROM) or the auxiliary storage device 15, or that is externally acquired from the communication device 21 via a network, and loads the program into the main storage device 13 (RAM). The control device 11 executes various processes in accordance with the loaded program, thereby realizing each of the above-described functional units. Here, the description has been given taking an example in which each of the above-described functional units is realized as software, but at least a portion of each of the above-described functional units may also be realized as hardware. In this case, the functional units realized as hardware may be realized, for example, by the above-described hardwired circuit. Furthermore, any of the above-described functional units may also be realized by a combination of software and hardware.
[0029] The display unit 107 can be realized, for example, by the display device 17 described in Fig. 1. The target model storage unit 111, the hidden model storage unit 115, the background image storage unit 119, and the training data storage unit 135 can be realized, for example, by at least one of the main storage unit 13 and the auxiliary storage unit 15 described in Fig. 1.
[0030] Based on user operations, the operation unit 101 performs various operation inputs required for generating training data on the training data generation device 10. Examples of the various operation inputs required for generating training data include, but are not limited to, an operation input for displaying a setting screen for generating training data, an operation input for specifying a target model to be used for generating training data, an operation input for specifying various conditions to be used for generating training data, and an operation input for instructing the start of training data generation.
[0031] The operation reception unit 103 receives various operation inputs from the operation unit 101 .
[0032] The display control unit 105 controls the display of a setting screen generated or updated by a screen generation unit 123 (to be described later) on the display unit 107 .
[0033] The display unit 107 displays a setting screen under the control of the display control unit 105 .
[0034] The image acquisition unit 109 acquires various image data stored in the target model storage unit 111, the occlusion model storage unit 115, and the background image storage unit 119. As described above, the image acquisition unit 109 includes the target model acquisition unit 113, the occlusion model acquisition unit 117, and the background image acquisition unit 121.
[0035] The target model storage unit 111 stores a target model. As described above, the training data generation device 10 of this embodiment generates training data used to train a recognition model for recognizing official vehicles in a broad sense, such as police cars, ambulances, and route buses, using CG. Therefore, in this embodiment, the target model will be described as a model of an official vehicle in a broad sense, such as a police car, an ambulance, or a route bus (in a broad sense, a model of a moving object), but is not limited thereto.
[0036] As described above, in this embodiment, images that serve as training data are generated using CG. Therefore, in this embodiment, a case where the target model is a 3D model used in CG, i.e., a three-dimensional image, will be described as an example, but the present invention is not limited to this. For example, if the images that serve as training data are generated by image synthesis, the target model may be an object, i.e., a two-dimensional image. Furthermore, an image cropped from a live image (e.g., an image of a police car, ambulance, or route bus cropped from a photograph) may be used as the target model.
[0037] Fig. 3 is a diagram showing an example of an object model 211 according to this embodiment. Fig. 4 is a diagram showing an example of an object model 221 according to this embodiment. The object model 211 is a 3D model of a police car. The object model 221 is a 3D model of an ambulance.
[0038] The target model acquisition unit 113 acquires a target model from the target model storage unit 111. For example, the target model acquisition unit 113 acquires a list of file names of the target models stored in the target model storage unit 111 in response to an instruction from the screen generation unit 123 (described later). Furthermore, for example, the target model acquisition unit 113 acquires a target model in response to an instruction from the image generation unit 131 (described later).
[0039] The occlusion model storage unit 115 stores occlusion models for occluding target models. As described above, in this embodiment, the target models are models of official vehicles in a broad sense. For this reason, in this embodiment, the occlusion models are described taking as examples models of road signs, traffic lights, trucks, etc. that may occlude models of these official vehicles on the road, but are not limited to these. In this way, in this embodiment, the occlusion models are models of objects that appear frequently in locations where the target models appear.
[0040] As described above, in this embodiment, the target model is a 3D model used in CG. Therefore, in this embodiment, the occlusion model is also a 3D model used in CG, i.e., a three-dimensional image, but the present invention is not limited to this. For example, if the images to be used as learning data are generated by image synthesis, the occlusion model may be an object, i.e., a two-dimensional image. Furthermore, an image cropped from a real image (e.g., an image of a road sign or traffic light cropped from a photograph) may be used as the occlusion model.
[0041] In this embodiment, the hidden model storage unit 115 stores each hidden model and a hidden model list that lists the file names of each hidden model. FIG. 5 is a diagram showing an example of a hidden model 311 in this embodiment. FIG. 6 is a diagram showing an example of a hidden model 321 in this embodiment. FIG. 7 is a diagram showing an example of a hidden model list 301 in this embodiment. The hidden model 311 is a 3D model of a road sign indicating a national highway. The hidden model 321 is a 3D model of a traffic light. The hidden model list 301 is a file in CSV (Comma Separated Values) format with the file name "ModelList.csv" that lists the file names of each hidden model. Note that in the example shown in FIG. 7, the file name extensions of each hidden model are omitted. For example, image data with the file name "Traffic_sign_000" in the occlusion model list 301 indicates the occlusion model 311, and image data with the file name "Traffic_light_000" in the occlusion model list 301 indicates the occlusion model 321. In this embodiment, it is assumed that the occlusion model list 301 is a list of occlusion models for the target model 211 and the target model 221, but the present invention is not limited to this.
[0042] The occlusion model acquisition unit 117 acquires occlusion models and occlusion model lists from the occlusion model storage unit 115. For example, the occlusion model acquisition unit 117 acquires occlusion model lists in response to instructions from a screen generation unit 123 (described later), or acquires occlusion models in response to instructions from an image generation unit 131 (described later).
[0043] The background image storage unit 119 stores a background image that serves as the background for the target model and the occlusion model. As described above, in this embodiment, the target model is a model of a public vehicle in a broad sense, and the occlusion model is a model of a road sign or the like that occludes the model of the public vehicle on the road. For this reason, in this embodiment, a case in which the background image is an image showing a road will be described as an example, but this is not limited to this. Note that the background image may be a two-dimensional image or a three-dimensional image. The background image may also be a real-life image such as a photograph. When the background image is a real-life image, it is preferable to use a field of view image from a camera that captures images used for recognition in a recognition model generated by learning training data as the background image. Training data generated using such a background image has a context similar to the real-life image used for recognition, thereby improving the recognition accuracy of the recognition model.
[0044] In this embodiment, the background image storage unit 119 stores each background image and a background image list that lists the file names of each background image. FIG. 8 is a diagram showing an example of a background image 411 in this embodiment. FIG. 9 is a diagram showing an example of a background image list 401 in this embodiment. The background image 411 is an image of a road. The background image list 401 is a CSV file with the file name "BackgroundList.csv" that lists the file names of each background image. Note that in the example shown in FIG. 9, the file name extensions of the background images are omitted. For example, image data with the file name "Background_000" in the background image list 401 represents the background image 411.
[0045] Background image acquisition unit 121 acquires a background image or a background image list from background image storage unit 119. For example, background image acquisition unit 121 acquires a background image list in response to an instruction from screen generation unit 123 (described later), or acquires a background image in response to an instruction from image generation unit 131 (described later).
[0046] The screen generation unit 123 generates and updates a setting screen for generating learning data in accordance with various operation inputs received by the operation reception unit 103 .
[0047] When operation input instructing display of a setting screen for generating learning data is accepted by operation accepting unit 103, screen generating unit 123 newly generates the setting screen. Specifically, screen generating unit 123 instructs target model acquiring unit 113 to acquire a list of file names of target models stored in target model storage unit 111. Screen generating unit 123 uses this acquired list of file names of target models to newly generate a setting screen for generating learning data.
[0048] Fig. 10 is a diagram showing an example of a setting screen 501 for generating training data according to this embodiment. The setting screen 501 shown in Fig. 10 is a setting screen newly generated by the screen generating unit 123. As shown in Fig. 10, the setting screen 501 is roughly divided into a display of items 503 related to models and images used in generating training data, and a display of items 505 related to the concealment of the target model.
[0049] Here, we will first explain item 503. Item 503 includes a pull-down list 511 that allows the user to specify a target model, a display area 513 that displays the target model specified in pull-down list 511, a display area 515 that displays a list of hidden models, and a display area 517 that displays a list of background images.
[0050] 10 , since the setting screen 501 is in an initial state, the pull-down list 511, the display area 513, the display area 515, and the display area 517 are all blank. However, if the user selects the pull-down list 511 using the operation unit 101, a list of the file names of the target models acquired by the screen generation unit 123 is displayed, allowing the user to select the target model to be used to generate training data. In this way, the pull-down list 511 is a first area that accepts the specification of a target model to be used to generate training data from among multiple target models.
[0051] For this reason, the screen generation unit 123 updates the setting screen when an operation input for specifying a target model to be used in generating training data is accepted by the operation acceptance unit 103. Specifically, the screen generation unit 123 instructs the target model acquisition unit 113 to acquire the specified target model, instructs the occlusion model acquisition unit 117 to acquire a occlusion model list, and instructs the background image acquisition unit 121 to acquire a background image list. The screen generation unit 123 updates the setting screen using the acquired target model, occlusion model list, and background image list.
[0052] 11 is a diagram showing an example of an updated setting screen 501 according to this embodiment. The setting screen 501 shown in FIG. 11 is a setting screen updated by the screen generation unit 123 when an operation input for specifying a target model is accepted by the operation acceptance unit 103.
[0053] In the example shown in FIG. 11 , a target model with the file name "Police_car_000" is specified in the pull-down list 511. Here, the target model with the file name "Police_car_000" is the target model 211, and therefore the target model 211 is displayed in the display area 513. The display area 513 is a third area that displays the specified target model. As described above, in this embodiment, when the screen generation unit 123 receives an instruction via the operation reception unit 103 to select a target model to be used for generating learning data from among multiple target models, the screen generation unit 123 displays a screen of the target model on the display unit 107 via the display control 105. Furthermore, the display area 515 displays "ModelList.csv", which is the file name of the hidden model list 301, and the display area 517 displays "BackgroundList.csv", which is the file name of the background image list 401.
[0054] As a result, the target model, occlusion model, and background image to be used for generating training data are set on the setting screen 501. The occlusion model is automatically set from the occlusion model list 301 by the image generation unit 131 (described later), and the background image is automatically set from the background image list 401 by the image generation unit 131 (described later). However, the method for setting the occlusion model and background image to be used for generating training data is not limited to this. For example, the screen generation unit 123 may display a list of occlusion models, similar to the target model, and allow the user to select a occlusion model to be used for generating training data from the list. That is, the display area 515 may be a pull-down list for selecting a occlusion model. In this case, the display area 515 serves as a second area that accepts the specification of a occlusion model to be used for generating training data from among multiple occlusion models. Furthermore, for example, the screen generation unit 123 may display a list of background images, similar to the target model, and allow the user to select a background image to be used for generating training data from the list. The screen generation unit 123 may also be configured to allow a plurality of occlusion models to be used for generating training data to be specified.
[0055] Next, item 505 will be described. Note that the contents of item 505 have not been updated since the state in FIG. 10 , and therefore will continue to be described with reference to FIG. 11 . Item 505 includes a feature region designation button 521 for starting designation of a feature region for the target model designated in pull-down list 511, an input box 531 for inputting a threshold value for the concealment rate of the feature region, a check box 541 for indicating whether or not to designate the concealment rate of the target model, an input box 543 for inputting a threshold value for the concealment rate of the target model, a check box 551 for indicating whether or not to designate an obscured region of the target model, an input box 553 for inputting the position of the obscured region of the target model, and an input box 555 for inputting a threshold value for the obscured region of the target model.
[0056] The condition setting unit 125 will now be described. The condition setting unit 125 sets various conditions specified in item 505. As described above, the condition setting unit 125 includes a feature region setting unit 127 and a concealment rate setting unit 129. When the feature region specification button 521 is pressed, the feature region setting unit 127 becomes able to set a feature region for the target model, and sets the feature region specified by the user for the target model. The concealment rate setting unit 129 sets the concealment rate threshold for the feature region input in input box 531, the concealment rate threshold for the target model input in input box 543, the position of the concealment region of the target model input in input box 553, and the maximum range of the concealment region of the target model input in input box 555. Item 505 will be described in detail below, with appropriate reference to the feature region setting unit 127 and the concealment rate setting unit 129.
[0057] When the operation receiving unit 103 receives an operation input of pressing the feature region designation button 521, the screen generation unit 123 transitions to a state in which a feature region can be set for the target model 211 on the display area 513. As a result, when the user performs an operation input to set a feature region for the target model 211 using the operation unit 101, the operation receiving unit 103 accepts the designation of the feature region of the target model 211 displayed in the display area 513. Furthermore, the feature region setting unit 127 sets the input feature region for the target model 211, and the screen generation unit 123 visualizes and displays the set feature region on the target model 211.
[0058] FIG. 12 is a diagram illustrating an example of a state in which a feature region 213 has been set for a target model 211 on the display area 513 of the setting screen 501 according to this embodiment. In the example illustrated in FIG. 12 , the target model 211 is a 3D model of a police car, and the 3D model is configured in parts, such as a red light, tires, headlights, and frame. This allows the user to specify feature regions for the target model 211 on a part-by-part basis. When the user specifies a part on the target model 211 to be set as a feature region using the operation unit 101, the feature region setting unit 127 sets the specified part as a feature region, and the screen generation unit 123 visualizes and displays the set feature region on the target model 211. In this way, in this embodiment, when the screen generation unit 123 receives an instruction to set a feature region for the target model 211 displayed in the display area 513 via the operation reception unit 103, the feature region set by the feature region setting unit 127 is visualized and displayed on the target model 211 displayed in the display area 513.
[0059] 12 shows the case where the red light of the target model 211 is designated by the user as the feature region 213. Therefore, the feature region 213 is set at the red light, and a hatched region 525 is displayed on the red light to visualize that the feature region 213 has been set.
[0060] In this embodiment, an example of specifying a feature region by selecting a part has been described. However, the method of specifying a feature region is not limited to this. For example, a user may specify a feature region by using the operation unit 101 to enclose an arbitrary region on the target model 211 with an arbitrary polygon. Alternatively, a user may specify a feature region by using the operation unit 101 to enclose an arbitrary region on the target model 211 with a circle or the like. Alternatively, a feature region may be automatically specified for the target model 211. In this case, the shape of the target model 211 may be analyzed to extract a portion that best represents the characteristic of the target model 211, and the extracted portion may be specified as a feature region. Note that the number of feature regions that can be specified is not limited to one, and two or more may be specified. In this case, when the screen generation unit 123 receives an instruction to set multiple feature regions for the target model displayed in the display area 513 via the operation reception unit 103, the screen generation unit 123 visualizes and displays each of the feature regions set by the feature region setting unit 127 on the target model displayed in the display area 513.
[0061] When the operation receiving unit 103 receives an operation input specifying a threshold value for the concealment rate of a characteristic region in the input box 531, the screen generating unit 123 displays the specified threshold value in the input box 531. In this way, the screen generating unit 123 receives the specification of a first threshold value, which is a threshold value for the concealment rate of a characteristic region, via the operation receiving unit 103. Furthermore, the concealment rate setting unit 129 sets the specified threshold value for the concealment rate for the characteristic region set by the feature region setting unit 127.
[0062] As a result, when generating, as training data, an image in which a part of a target model is concealed by an occlusion model, the image generation unit 131 (described later) can generate, as training data, an image in which the concealment rate of a feature region is less than a set threshold. Note that the concealment rate of a feature region means the proportion of the area of the feature region that is concealed by the occlusion model to the entire feature region.
[0063] When the operation accepting unit 103 accepts an operation input for checking the check box 541 and further accepts an operation input for specifying the threshold value of the concealment rate of the target model in the input box 543, the screen generating unit 123 displays the specified threshold value in the input box 543. In this way, the screen generating unit 123 accepts the specification of the second threshold value, which is the threshold value of the concealment rate of the target model, via the operation accepting unit 103. Furthermore, the concealment rate setting unit 129 sets the specified threshold value of the concealment rate for the target model specified in the pull-down list 511.
[0064] As a result, when generating an image in which a portion of a target model is concealed by an occlusion model as training data, the image generation unit 131 (described later) can also use the concealment rate of the target model to generate the training data. In this embodiment, the concealment rate of the target model refers to the proportion of the area of the target model that is concealed by the occlusion model to the entire area of the target model. However, the concealment rate of the target model is not limited to this, and may also be the proportion of the area of the non-feature area that is concealed by the occlusion model to the entire non-feature area. A non-feature area is an area obtained by excluding the feature area from the entire area of the target model. Note that in this embodiment, whether or not to use the concealment rate of the target model can be specified using a check box 541. Therefore, if the check box 541 is not checked, the concealment rate of the target model is not used to generate the training data.
[0065] Assume that the operation accepting unit 103 accepts an operation input to check the check box 551, and further accepts an operation input to specify the position of the hidden area of the target model in the input box 553 and an operation input to specify the maximum range of the hidden area of the target model in the input box 555. In this case, the screen generating unit 123 displays information about the specified position in the input box 553 and the specified maximum range in the input box 555. Furthermore, the concealment rate setting unit 129 sets the specified position and the specified maximum range of the hidden area for the target model specified in the pull-down list 511.
[0066] FIG. 13 is a diagram showing an example of a concealment area 559 specified for the target model 211 in this embodiment. In the example shown in FIG. 13 , it is assumed that a specified position of "left" is specified in input box 553 and a maximum range of "50%" is specified in input box 555 for the target model 211. In this case, the concealment rate setting unit 129 sets a rectangle 557 surrounding the target model 211. Furthermore, since "left" is specified in input box 553, the concealment rate setting unit 129 sets the start position of the concealment area to "left." Furthermore, since "50%" is specified in input box 555, the concealment rate setting unit 129 sets the size of the concealment area to 50% of the size of rectangle 557. As a result, the concealment rate setting unit 129 sets the left half of rectangle 557 as concealment area 559, as shown in FIG. 13 .
[0067] Therefore, when generating an image in which a portion of the target model is concealed by an occlusion model as training data, the image generation unit 131 (described later) can specify the region of the target model that is to be concealed by the occlusion model. When the occlusion region 559 described in FIG. 13 is set for the target model 211, occlusion by the occlusion model will only occur within the occlusion region 559. In this embodiment, whether or not to use the occlusion region of the target model can be specified using a check box 551. Therefore, when the check box 551 is not checked, the occlusion region of the target model is not used to generate training data.
[0068] In this embodiment, the threshold value of the concealment rate of the characteristic region is not optional, but it may be optional, like the concealment rate of the target model and the concealment region of the target model.
[0069] Next, an input box 561 for inputting the number of images to be generated as learning data, and a learning data generation start button 563 will be described.
[0070] When the operation receiving unit 103 receives an operation input specifying the number of images to be generated as learning data in the input box 561, the screen generating unit 123 sets the specified number and displays the specified number in the input box 561.
[0071] When the operation receiving unit 103 receives an operation input of pressing the start generation button 563, the screen generating unit 123 instructs the image generating unit 131 to start generating learning data under the various conditions set on the setting screen 501.
[0072] When instructed to start generating training data, the image generation unit 131 generates, as training data, an image in which a portion of the target model is concealed with an occlusion model so that the concealment rate of the feature region specified on the setting screen displaying the target model is less than a first threshold. Specifically, when instructed to start generating training data by the screen generation unit 123, the image generation unit 131 acquires, from the screen generation unit 123, information on the model and image used to generate the training data set in item 503 and information on the number of images to be generated as training data. Furthermore, the image generation unit 131 acquires, from the image acquisition unit 109, the target model, the occlusion model, and a background image used to generate the training data, based on the information on the model and image used to generate the training data acquired from the screen generation unit 123. Furthermore, the image generation unit 131 acquires, from the condition setting unit 125, the various conditions set in item 505. Based on the acquired information, the image generation unit 131 generates an image including the target model in which the feature region is set and the occlusion model, and if the feature region is concealed by the occlusion model so as to satisfy the occlusion condition, the generated image is used as training data.
[0073] The method for generating training data by the image generation unit 131 of this embodiment will be specifically described below with reference to the drawings. In practice, background images are also used to generate training data, but for the sake of convenience, the background image may be omitted from the description. Generating training data using a background image allows the context to be closer to a real-life image (such as a photograph) even for training data generated by CG, thereby improving the recognition accuracy of the recognition model.
[0074] FIG. 14 is a plan view showing an example of an image generation method by the image generation unit 131 of this embodiment. As described above, in this embodiment, training data is generated using CG. To this end, as shown in FIG. 14 , the image generation unit 131 defines a three-dimensional CG space 601 and places 3D models, such as the target model 211 and the occlusion model 311, and a virtual camera 611, within the CG space 601. The virtual camera 611 projects (renders) various models and backgrounds included in a field of view 603 onto a projection surface (not shown) to generate an image. A known rendering method, such as the Z-buffer method, may be used. The image generation unit 131 can arbitrarily position the target model 211 and the occlusion model 311 within the field of view 603. However, the training data generation device 10 of this embodiment generates, as training data, an image in which a portion of the feature region of the target model 211 is occluded by the occlusion model 311. For this reason, it is desirable to place the occlusion model 311 closer to the target model 211 in the depth direction (z direction) (toward the virtual camera 611).
[0075] FIG. 15 is a diagram showing an example of an image 651A generated by the image generation unit 131 of this embodiment. The image 651A is an image generated by the virtual camera 611 in the example shown in FIG. 14. As shown in FIG. 14, the image 651A is an image including the target model 211 and the occlusion model 311. Here, the image generation unit 131 determines whether or not to use the generated image 651A as image data using various conditions acquired from the condition setting unit 125. Here, it is assumed that the various conditions include a feature region 213 set in the target model 211, and a first threshold value (e.g., 20%) set as the threshold for the concealment rate of the feature region. It is assumed that the concealment rate and the occlusion region of the target model are not set.
[0076] In this case, the image generation unit 131 uses image 651A as training data if the concealment rate of the feature region 213 in image 651A is less than a first threshold, and does not use image 651A as training data if the concealment rate of the feature region 213 is equal to or greater than the first threshold. In this way, it is possible to use as training data images in which the feature portions of the target model remain unconcealed to a certain extent, that is, images that are likely to retain their distinctiveness as a target model even if they are partially occluded. Therefore, it is possible to accurately generate images as training data that are useful for training a recognition model that is also used to recognize a recognition target that is partially occluded.
[0077] 15 , a portion of the target model 211 is occluded by the occlusion model 311, but the feature region 213 is not occluded by the occlusion model 311. Therefore, the occlusion rate of the feature region 213 is 0%, which is less than the first threshold, and the image generation unit 131 uses the image 651A as training data.
[0078] The concealment rate of the feature region 213 can be determined, for example, by comparing an image in which only the feature region 213 of the target model 211 is projected onto the projection surface with an image in which the feature region 213 of the target model 211 and the occlusion model 311 are projected onto the projection surface in the state shown in FIG. 14 . The latter image is projected with the occlusion model 311 masked with the same color as the background. The image in which only the feature region 213 of the target model 211 is projected onto the projection surface is an image showing the silhouette of the feature region 213, and is an image in which the entire feature region 213 is represented without any omissions. Here, the area of the image in which only the feature region 213 of the target model 211 is projected onto the projection surface is defined as A. On the other hand, in the image in which the feature region 213 of the target model 211 and the occlusion model 311 are projected onto the projection surface, if at least a portion of the feature region 213 is occluded by the occlusion model 311, the image will have the occlusion portion missing from the silhouette of the feature region 213. In other words, it is an image that represents the area of the feature region 213 that is not obscured by the occlusion portion. Here, the area of the image obtained by projecting the feature region 213 of the target model 211 and the occlusion model 311 onto the projection plane is assumed to be B. Under the above conditions, the obscuration rate of the feature region 213 is ((A-B) / A))*100%.
[0079] 16 is a diagram showing an example of an image 651B generated by the image generation unit 131 of this embodiment. Image 651B is an image generated by the virtual camera 611 in which the target model 211 and the occlusion model 311 are arranged in a different manner from the arrangement shown in FIG. 14 . As shown in FIG. 16 , image 651B is an image including the target model 211 and the occlusion model 311. Note that the various conditions acquired from the condition setting unit 125 are the same as those in FIG. 15 . That is, as the various conditions, a feature region 213 is set in the target model 211, a first threshold (e.g., 20%) is set as the threshold for the concealment rate of the feature region, and the concealment rate of the target model and the occlusion region of the target model are not set.
[0080] 16 , a portion of the target model 211 is occluded by the occlusion model 311, but the feature region 213 is entirely occluded by the occlusion model 311. Therefore, the occlusion rate of the feature region 213 is 100%, which is equal to or greater than the first threshold, and therefore the image generation unit 131 does not use the image 651B as learning data.
[0081] FIG. 17 is a diagram showing an example of an image 651C generated by the image generation unit 131 of this embodiment. The image 651C is an image generated by the virtual camera 611 in a state in which the target model 211 and the occlusion model 321 are placed in the CG space 601. As shown in FIG. 17, the image 651C is an image including the target model 211 and the occlusion model 321. Regarding the various conditions acquired from the condition setting unit 125, it is assumed that a feature region 213 is set in the target model 211, and a first threshold value (e.g., 20%) is set as the threshold for the concealment rate of the feature region. Furthermore, it is assumed that a second threshold value (e.g., 30%) is set as the threshold for the concealment rate of the target model. It is assumed that no concealment rate is set for the target model.
[0082] In this case, if the concealment rate of the feature region 213 in image 651C is less than the first threshold and the concealment rate of the target model 211 is less than the second threshold, the image generation unit 131 selects image 651C as training data. However, if the concealment rate of the feature region 213 is equal to or greater than the first threshold or if the concealment rate of the target model 211 is equal to or greater than the second threshold, the image generation unit 131 does not select image 651C as training data. In this way, it is possible to more accurately generate images as training data that are useful for training a recognition model that is also used to recognize a recognition target whose parts are concealed.
[0083] In the example shown in FIG. 17 , a portion of the target model 211 is occluded by the occlusion model 321, but nearly half of the feature region 213 is occluded by the occlusion model 321. Therefore, the occlusion rate of the feature region 213 is equal to or greater than a first threshold (e.g., 20%). On the other hand, while a portion of the target model 211 is occluded by the occlusion model 321, the majority of the target model 211 is not occluded. Therefore, the occlusion rate of the target model 211 is less than a second threshold (e.g., 30%). As described above, in this embodiment, the occlusion rate of the target model 211 is less than the second threshold (e.g., 30%), but the occlusion rate of the feature region 213 is equal to or greater than the first threshold. Therefore, the image generating unit 131 does not use image 651C as training data. Note that the occlusion rate of the target model 211 can be calculated using the same method as that for the feature region 213.
[0084] FIG. 18 is a diagram showing an example of an image 651D generated by the image generation unit 131 of this embodiment. The image 651D is an image generated by the virtual camera 611 in a state in which the target model 221 and the occlusion model 321 are placed in the CG space 601. As shown in FIG. 18, the image 651D is an image including the target model 221 and the occlusion model 321. Regarding the various conditions acquired from the condition setting unit 125, it is assumed that feature regions 223A and 223B are set in the target model 211, and that first and third thresholds (e.g., 20%) are set as thresholds for the concealment rate of the feature regions. It is assumed that the threshold for the concealment rate of the target model and the occlusion region of the target model are not set. Thus, in the example shown in FIG. 18, it is assumed that multiple feature regions are set in the target model 221.
[0085] In this case, if the concealment rate of the feature region 223A in image 651D is equal to or greater than the first threshold, but the concealment rate of the feature region 223B is less than the third threshold, the image generation unit 131 selects image 651D as training data. However, if the concealment rate of the feature region 223A is equal to or greater than the first threshold and the concealment rate of the feature region 223B is equal to or greater than the third threshold, the image generation unit 131 does not select image 651D as training data. When a target model has multiple features, even if a first feature is concealed to a certain extent or more, if a second feature remains unconcealed to a certain extent or more, there is a good chance that the target model will not lose its distinctiveness. The same is true in the reverse case. Therefore, by including such images as training data, it is possible to accurately generate images useful for training a recognition model that is also used to recognize partially concealed recognition targets. The first and third thresholds may be the same or different values.
[0086] 18 , a portion of the target model 221 is occluded by the occlusion model 321, a large portion of the feature region 223A is occluded by the occlusion model 321, and region 223B is not occluded by the occlusion model 321. Therefore, the occlusion rate of the feature region 223A is equal to or greater than the first threshold (e.g., 20%), but the occlusion rate of the feature region 223B is less than the third threshold (e.g., 20%). Thus, in this embodiment, even if the occlusion rate of the feature region 223A is equal to or greater than the first threshold, if the occlusion rate of the feature region 223B is less than the third threshold, the image generation unit 131 uses image 651D as training data.
[0087] FIG. 19 is a diagram showing an example of an image 651E generated by the image generation unit 131 of this embodiment. Image 651E is an image generated by the virtual camera 611 in a state in which a background image 401 is placed in addition to the target model 211 and the occlusion model 311 in the CG space 601. As shown in FIG. 19 , image 651E is an image including the target model 211, the occlusion model 311, and the background image 411. Regarding the various conditions acquired from the condition setting unit 125, it is assumed that a feature region 213 is set in the target model 211, and a first threshold value (e.g., 20%) is set as the threshold value for the concealment rate of the feature region. Furthermore, it is assumed that "50%" from the "left" of the target model 211, i.e., the left half of the target model 211, is set as the concealment region of the target model. It is assumed that a threshold value for the concealment rate of the target model is not set.
[0088] 19 , a portion of the left half of the target model 221 is occluded by the occlusion model 321, and the feature region 213 is not occluded by the occlusion model 311. Therefore, the occlusion rate of the feature region 213 is 0%, which is less than the first threshold, and the image generation unit 131 uses the image 651E as training data.
[0089] The image output unit 133 outputs the learning data generated by the image generation unit 131 to the learning data storage unit 135. When an instruction to generate learning data is accepted by the screen generation unit 123, the image output unit 133 outputs, as learning data, to the learning data storage unit 135 an image in which the feature region of the target model is occluded by the occlusion model so that the occlusion condition is satisfied.
[0090] The learning data storage unit 135 stores, as learning data, the images generated by the image generation unit 131. Specifically, the learning data storage unit 135 stores the learning data output from the image output unit 133.
[0091] FIG. 20 is a flowchart showing an example of the training data generation process performed by the training data generation device 10 of this embodiment.
[0092] First, when an operation input for specifying a target model is received, the screen generating unit 123 sets the specified target model as a target model to be used for generating training data (step S11).
[0093] Next, the screen generation unit 123 acquires the occlusion model list from the occlusion model acquisition unit 117, and acquires the background image list from the background image acquisition unit 121 (step S13).
[0094] Next, the screen generation unit 123 sets a hijacking model to be used for generating training data from the hijacking model list, and sets a background image to be used for generating training data from the background image list (step S15). Note that the screen generation unit 123 may select a hijacking model to be set from the hijacking model list in any way, but in this embodiment, the screen generation unit 123 selects and sets the hijacking model in the order of the list. The same applies to the background image.
[0095] Next, when the characteristic region setting unit 127 receives an operation input of pressing the characteristic region designation button 521, it further receives an operation input of designating a characteristic region in the target model on the display area 513, and sets the characteristic region in the target model (step S17).
[0096] Next, when an operation input specifying the threshold value of the concealment rate of the characteristic region is received in the input box 531, the concealment rate setting unit 129 sets the specified threshold value of the concealment rate for the characteristic region (step S19).
[0097] Next, when the screen generation unit 123 receives an operation input to check the check box 541, it determines to set a threshold value for the obscuration rate of the target model (YES in step S21). In this case, when the operation input to specify a threshold value for the obscuration rate of the target model in the input box 543 is received, the obscuration rate setting unit 129 sets the specified threshold value for the obscuration rate for the target model (step S23).
[0098] On the other hand, if the screen generation unit 123 determines that the threshold value of the concealment rate of the target model is not to be set (NO in step S21), the process of step S23 is not performed.
[0099] Next, when the operation input of checking check box 551 is received, screen generation unit 123 determines that a concealment area is to be set on the target model (YES in step S25). In this case, when the operation input of specifying the position of the concealment area of the target model in input box 553 is received, concealment rate setting unit 129 sets the position of the specified concealment area on the target model (step S27).
[0100] Next, when an operation input specifying the maximum range of the target model is received in the input box 553, the concealment rate setting unit 129 sets the maximum range of the specified concealment region for the target model (step S29).
[0101] Through steps S27 and S29, the concealment rate setting unit 129 sets, for the target model, a concealment area identified from the position and maximum range of the designated concealment area.
[0102] On the other hand, if the screen generation unit 123 determines that a hidden area is not to be set in the target model (NO in step S25), the processes of steps S27 and S29 are not performed.
[0103] Next, when an operation input specifying the number of images to be generated as learning data is received in the input box 561, the screen generating unit 123 sets the specified number (step S31).
[0104] Next, when the image generation unit 131 receives an operation input of pressing the generation start button 563, it generates an image including the target model and the occlusion model in which the feature region is set, using the target model, the occlusion model, and the background image set in steps S11 to S15 (step S33).
[0105] Next, the image generation unit 131 determines whether the feature region in the generated image is occluded by the occlusion model so as to satisfy the occlusion condition, according to the conditions set in steps S17 to S29. The image output unit 133 outputs the image in which the feature region is occluded by the occlusion model so as to satisfy the occlusion condition as training data to the training data storage unit 135 (step S35).
[0106] Next, if the number of images generated as training data satisfies the number specified in step S31 (YES in step S37), the image generation unit 131 ends the process. On the other hand, if the number of images generated as training data does not satisfy the number specified in step S31 (NO in step S37), the process returns to step S15. In this case, the screen generation unit 123 sets a new occlusion model to be used for generating training data from the occlusion model list, and sets a new background image to be used for generating training data from the background image list (step S15).
[0107] As described above, in this embodiment, images in which the characteristic portions of the target model remain unoccluded to a certain extent are used as training data. That is, in this embodiment, images that are likely to retain their distinctiveness as a target model even if some parts are occluded are used as training data. Therefore, according to this embodiment, it is possible to accurately generate images that are useful for training a recognition model that is also used to recognize a recognition target that is partially occluded as training data.
[0108] Furthermore, in this embodiment, images are used as training data in which not only are the characteristic portions of the target model left unoccluded to a certain extent, but also the target model itself remains unoccluded to a certain extent. In other words, in this embodiment, images are used as training data that are more likely to retain their distinctiveness as a target model even if some parts are occluded. Therefore, according to this embodiment, it is possible to generate, with high accuracy, images that are useful for training a recognition model that is also used to recognize a recognition target that is partially occluded.
[0109] Furthermore, in this embodiment, when a target model has multiple feature parts, even if a first feature part is occluded to a certain extent or more, if a second feature part remains unoccluded to a certain extent or more, there is a good possibility that the target model has not lost its distinctiveness, and therefore such images are also used as training data. Therefore, according to this embodiment, it is possible to accurately generate images as training data that are useful for training a recognition model that is also used to recognize a recognition target that is partially occluded.
[0110] Furthermore, by generating a recognition model by learning the training data generated by the method of this embodiment, highly accurate recognition becomes possible even when recognizing a partially obscured recognition target. In particular, as described in this embodiment, by generating a recognition model for recognizing official vehicles as a recognition model and using this recognition model for traffic light control, it becomes possible to perform traffic light control appropriate for official vehicles and private vehicles. For example, a surveillance camera is installed at a traffic light, and the surveillance camera captures images of vehicles passing in front of the traffic light and inputs the images into the recognition model. If the recognition model recognizes a vehicle passing in front of the traffic light as an official vehicle, the signal light can be set to remain green even when it is time to change from green to red. This reduces the opportunities for official vehicles (especially emergency vehicles) to slow down due to a red light, contributing to faster emergency response by emergency vehicles.
[0111] When learning data generated by the method of this embodiment, it is preferable to add a separate annotation to the learning data and treat it as teacher data.
[0112] (Variation 1) In the above embodiment, an example has been described in which a user manually inputs various conditions on a setting screen for generating training data. In contrast, in Variation 1, input of various conditions on the setting screen may be automated. For example, the recognition accuracy of a recognition model generated by learning generated training data may be tested, and parameters of various conditions may be automatically set according to the test results.
[0113] (Variation 2) In the above embodiment, an example has been described in which the positions on the background image to place the target model and the occlusion model when generating training data are arbitrarily set by the image generation unit 131. In contrast to this, in Variation 1, the positions on the background image to place the target model and the occlusion model may be specified by the user.
[0114] In the second modification, first, a list of background images is displayed on the setting screen 501, similar to the target model, and the user is allowed to select a background image to be used for generating training data from the list. That is, a pull-down list for allowing the user to select a background image is provided on the setting screen 501, replacing the display area 517 displaying the background image list. Then, a separate display area 519 is provided on the setting screen 501, in which the background image selected by the user using the pull-down list is displayed. Note that a similar modification may also be performed on the occlusion model.
[0115] FIG. 21 is a diagram showing an example of a display area 519 in Modification Example 2. In the example shown in FIG. 21 , a background image 411 is displayed in the display area 519. Also in the example shown in FIG. 21 , a placement area in which a target model is to be placed can be set for the background image 411. Note that a placement area designation button like the feature area designation button 521 described in the above embodiment may be separately provided on the setting screen 501, and pressing this placement area designation button may transition to a state in which a placement area can be set for the background image 411. When the user performs an operation input to set a placement area for the background image 411 using the operation unit 101, the condition setting unit 125 sets the input placement area on the background image 411, and the screen generation unit 123 visualizes and displays the set placement area on the background image 411.
[0116] For example, the user can specify the placement area by specifying an arbitrary area on the background image 411 by enclosing it with an arbitrary polygon using the operation unit 101. In the example shown in Fig. 21 , the placement area 413 of the target model specified in this way is set in the background image 411 and visualized and displayed. In the example shown in Fig. 21 , the placement area 413 is the area of a road.
[0117] Although the example in which the user specifies the placement area has been described in Modification 2, the placement area may also be automatically detected by the image generation unit 131. For example, the image generation unit 131 may use a method such as plane detection to detect the placement area 413 from the topography of the background image 411 as a movable area in which the target model can move.
[0118] When arranging the target model and the occlusion model in the CG space, if an arrangement area has been set, the image generation unit 131 arranges the target model and the occlusion model in this arrangement area. For example, when using the background image 411 described in Fig. 21 , the image generation unit 131 arranges the target model at any position in the arrangement area 413 and arranges the occlusion model at any position, thereby generating an image. For example, the image 651E described in Fig. 19 is generated.
[0119] According to variant example 2, the target model can be placed in an area of the background image that is likely to actually be located there, so that images that are more useful for training a recognition model that is also used to recognize a recognition target that is partially occluded can be generated with high accuracy as training data.
[0120] (Variation 3) In the above embodiment, an example was described in which still images were generated as training data. However, in Variation 3, an example is described in which video images are generated as training data. In the above embodiment, an example was described in which the training data generation device uses CG to generate training data used to train a recognition model for recognizing official vehicles. In contrast, in Variation 3, an example is described in which CG is used to generate training data used to train a recognition model for recognizing the action of replenishing goods in a situation where goods are being replenished.
[0121] In the following, an example will be described in which, in a scene of stock replenishment, the combination of a store clerk model and a product model is the target model, and a passerby model crossing in front of the store clerk is the occlusion model, but this is not limiting. Therefore, in Variation 3, the product model becomes the feature region of the target model.
[0122] FIG. 22 is a plan view showing an example of a moving image generation method by the image generation unit 131 in Modification 3. Modification 3 also generates learning data using CG. As shown in FIG. 22 , the image generation unit 131 defines a three-dimensional CG space 801 and places, within the CG space 801, a store clerk model 701, a product model serving as a feature region 703, a shelf model 705 to which the product models are replenished, a passerby model serving as an occlusion model 711, and a virtual camera 811. As described above, the combination of the store clerk model 701 and the product model is the target model 709. The store clerk model 701, the product model serving as the feature region 703, the shelf model 705, and the occlusion model 711 (passerby model) are all 3D models. The virtual camera 811 projects (renders) various models and backgrounds included in its field of view 803 onto a projection surface (not shown) to generate an image. In addition, in the third modification, in order to generate a moving image, the virtual camera 811 performs the above-described image generation in units of frames (for example, 1 / 60 seconds).
[0123] Furthermore, the occlusion model 711 moves within the CG space 801 along the paths indicated by arrows a1 and a2, and is assumed to be at the position of the occlusion model 711 indicated by the solid line. In other words, the occlusion model 711 indicated by the dashed line indicates the past position of the occlusion model 711. In Modification 3, the virtual camera 811 generates a video of a scene in which the occlusion model 711 moves along the paths indicated by arrows a1 and a2 in front of the store clerk model 701 who is replenishing the product models that become the feature region 703 in the product shelf model 705.
[0124] FIG. 23 is a diagram showing an example of a moving image 851 generated by the image generation unit 131 of Modification Example 3. In the example shown in FIG. 23, some of the images constituting the moving image 851 are illustrated in chronological order. An image 851A constituting the moving image 851 is an image generated by the image generation unit 131 before the occlusion model 711 starts moving along the arrow a1 (see FIG. 22). An image 851B constituting the moving image 851 is an image generated by the image generation unit 131 after the occlusion model 711 has finished moving along the arrow a1 (see FIG. 22) and before it starts moving along the arrow a2 (see FIG. 22). An image 851C constituting the moving image 851 is an image generated by the image generation unit 131 after the occlusion model 711 has finished moving along the arrow a2 (see FIG. 22).
[0125] The image generation unit 131 determines whether each image should be used as learning data, as described in the above embodiment. However, since the target in Modification Example 3 is a moving image, the time element is also taken into account when determining whether the moving image constituting each image should be used as learning data. This is because, without taking the time element into account, it is impossible to determine whether the behavior of the store clerk model 701 is replenishing products on shelves or removing products from shelves. Specifically, the image generation unit 131 determines whether a state determined not to be used as learning data in the determination of whether to use the moving image as learning data as described in the above embodiment continues for a certain period of time (a certain number of frames). If this state continues for a certain period of time, the image generation unit 131 determines not to use the moving image as learning data. On the other hand, if there is no scene in the moving image where a state determined not to be used as learning data in the determination of whether to use the moving image as learning data as described in the above embodiment continues for a certain period of time, the image generation unit 131 determines that the moving image should be used as learning data. Note that the threshold for the certain period of time may be set on the setting screen 501. For example, an input box for inputting a threshold value for a certain period of time may be added to the setting screen 501, like the input box 531 for inputting a threshold value for the concealment rate of the characteristic region described in the above embodiment.
[0126] 23 , in image 851B, feature region 703 is occluded by occlusion model 711. However, in image 851C, occlusion of feature region 703 by occlusion model 711 is eliminated, and the occlusion of feature region 703 is eliminated in a short time. As such, moving image 851 does not contain any scenes in which a state in which it is determined not to be used as learning data in the determination of whether or not to use learning data as described in the above embodiment continues for a certain period of time, and therefore image generation unit 131 uses moving image 851 as learning data.
[0127] FIG. 24 is a diagram showing an example of a moving image 853 generated by the image generation unit 131 of Modification Example 3. The moving image 853 is a video of a scene in which the occlusion model 711 stands motionless in front of the store clerk model 701. In the example shown in FIG. 24 , a portion of each image constituting the moving image 853 is illustrated in chronological order. In all of images 853A, 853B, and 853C constituting the moving image 853, the occlusion model 711 stands motionless in front of the store clerk model 701, and the feature region 703 is occluded by the occlusion model 711. As described above, in the moving image 853, a state in which it was determined not to be used as learning data in the determination of whether to use the moving image as learning data described in the above embodiment continues for a certain period of time, and therefore the image generation unit 131 does not use the moving image 853 as learning data.
[0128] As described above, in Modification 3, the image generation unit 131 generates, as training data, video in which a part of the target model is concealed with an occlusion model so that the state in which the concealment rate of the feature region is equal to or greater than the first threshold is maintained for less than a certain period of time. As described above, according to Modification 3, video that is likely to retain its distinctiveness as a target model even if a part is continuously occluded is used as training data. Therefore, according to Modification 3, video that is useful for training a recognition model that is also used to recognize a recognition target whose part is occluded can be accurately generated as training data.
[0129] (Program) The program executed by the training data generation device 10 in the above embodiment and each of the above modified examples is provided by being stored in a computer-readable storage medium such as a CD-ROM, CD-R, memory card, DVD, or flexible disk (FD) in the form of a file in an installable or executable format.
[0130] The programs executed by the training data generation device 10 of the above embodiment and each of the above modifications may be stored on a computer connected to a network such as the Internet and provided by being downloaded via the network. The programs executed by the training data generation device 10 of the above embodiment and each of the above modifications may be provided or distributed via a network such as the Internet. The programs executed by the training data generation device 10 of the above embodiment and each of the above modifications may be provided by being pre-installed in a ROM or the like.
[0131] The programs executed by the training data generation device 10 of the above embodiment and each of the above modifications have a modular configuration for implementing the above-mentioned units on a computer. In terms of actual hardware, for example, the CPU reads the training program from the HDD onto the RAM and executes it, thereby implementing the above-mentioned units on the computer.
[0132] As described above, according to the above embodiment and each of the above modifications, images useful for training a recognition model can be generated with high accuracy as training data.
[0133] The above-described embodiment and each of the modifications merely illustrate examples of implementations of the present disclosure, and the technical scope of the present disclosure should not be construed as being limited by these. Therefore, the present disclosure can be implemented in various forms without departing from the spirit or main features thereof. For example, the above-described embodiment and each of the modifications may be appropriately combined in their respective constituent units. Furthermore, for example, some components may be deleted from all components in the above-described embodiment and each of the modifications.
[0134] The present disclosure includes the following aspects.
[0135] (1) A training data generation device comprising: a memory that stores a target model and an occlusion model; and a processor that generates, as training data, an image in which a part of the target model is occluded by the occlusion model so that the occlusion rate of a feature region specified on a screen displaying the target model is less than a first threshold.
[0136] In the above configuration (1), images in which the characteristic portions of the target model remain unoccluded to a certain extent are used as training data. In other words, in the above configuration (1), images that are likely to retain their distinctiveness as a target model even if some parts are occluded are used as training data. Therefore, with the above configuration (1), it is possible to accurately generate images that are useful for training a recognition model that is also used to recognize a recognition target that is partially occluded as training data.
[0137] (2) The training data generation device according to (1), wherein the processor receives a designation of the first threshold and a designation of a second threshold that is a threshold for the concealment rate of the target model, and generates, as the training data, an image in which a part of the target model is concealed by the concealment model so that the concealment rate of the feature region is less than the first threshold and the concealment rate of the target model is less than the second threshold.
[0138] In the above configuration (2), images in which not only the characteristic parts of the target model remain unoccluded to a certain extent, but also the target model itself remains unoccluded to a certain extent are used as training data. In other words, in the above configuration (2), images that are more likely to retain their distinctiveness as a target model even if they are partially occluded are used as training data. Therefore, with the above configuration (2), images that are useful for training a recognition model that is also used to recognize a recognition target that is partially occluded can be generated with high accuracy as training data.
[0139] (3) The training data generation device according to (2), wherein the second threshold is greater than the first threshold.
[0140] According to the above configuration (3), images that are useful for training a recognition model that is also used to recognize a recognition target that is partially occluded can be generated more accurately as training data.
[0141] (4) The training data generation device according to (1), wherein the processor accepts designation of a plurality of locations of the target model as the feature regions.
[0142] In the configuration (4) above, when a target model has multiple feature parts, even if one of the feature parts is occluded to a certain extent or more, as long as other feature parts remain unoccluded to a certain extent or more, there is a good possibility that the target model will not lose its distinctiveness, and therefore such images are also used as training data. Therefore, with the configuration (4) above, it is possible to accurately generate images as training data that are useful for training a recognition model that is also used to recognize a recognition target that is partially occluded.
[0143] (5) The training data generation device described in (1) above, wherein the memory further stores, as a background image, a field of view image from an image capture device that captures an image used for recognition by a recognition model generated by learning the training data, and the processor generates, as the training data, an image in which a part of the target model is concealed by the occlusion model on the background image so that the concealment rate of the feature region is less than the first threshold.
[0144] The configuration (5) above makes it possible to generate training data having a background similar to that of a real-life image (such as a photograph) used for recognition in a recognition model. Since the context of such training data is similar to that of the real-life image used for recognition, the recognition accuracy of the recognition model can be improved. Therefore, the configuration (5) above makes it possible to generate, with high accuracy, images that are more useful for training a recognition model that is also used to recognize a recognition target whose parts are occluded.
[0145] (6) The training data generation device according to (1), wherein the memory further stores a background image; the target model is a model of a moving body; and the processor detects a movable area in which the target model can move from a topography in the background image; and generates, as the training data, an image in which a part of the target model is concealed by the occlusion model in the movable area of the background image so that the concealment rate of the feature area is less than the first threshold.
[0146] In the configuration (6) above, the target model can be placed in an area of the background image that is likely to actually be located there, so that images that are more useful for training a recognition model that is also used to recognize a recognition target that is partially occluded can be generated with high accuracy as training data.
[0147] (7) The training data generation device according to (1), wherein the occlusion model is a model of an object that appears frequently in a location where the target model appears, and the memory stores a plurality of the occlusion models for the target model.
[0148] In the configuration (7) above, the target model can be concealed with an occlusion model that is likely to actually conceal the target model, so that images that are more useful for training a recognition model that is also used to recognize a recognition target that is partially occluded can be generated with high accuracy as training data.
[0149] (8) The training data generation device according to (1), wherein the processor generates, as training data, a video in which a part of the target model is concealed by the occlusion model so that a state in which an occlusion rate of the feature region is equal to or higher than a first threshold is maintained for less than a certain period of time.
[0150] In the above configuration (8), video images that are likely to retain their recognizability as an object model even when a portion of the image is continuously occluded are used as training data. Therefore, the above configuration (8) makes it possible to accurately generate training data that are useful for training a recognition model that can also be used to recognize a recognition object whose portion is occluded. In particular, the above configuration (8) is suitable for generating training data for a recognition model used in behavioral analysis, which cannot be analyzed without taking into account the time factor. For example, when analyzing the product replenishing behavior of a store clerk, it is impossible to determine whether the clerk is replenishing products on a shelf or removing products from a shelf without taking into account the time factor.
[0151] (9) A learning data generation method including: a step of a processor reading out a target model and an occlusion model from a memory; and a step of the processor generating, as learning data, an image in which a part of the target model is occluded by the occlusion model so that an occlusion rate of a feature region specified on a screen displaying the target model is less than a first threshold.
[0152] In the above configuration (9), images in which the characteristic portions of the target model remain unoccluded to a certain extent are used as training data. That is, in the above configuration (9), images that are likely to retain their distinctiveness as a target model even if they are partially occluded are used as training data. Therefore, according to the above configuration (9), it is possible to accurately generate images that are useful for training a recognition model that is also used to recognize a recognition target that is partially occluded as training data.
[0153] (10) A program for causing a computer to execute the steps of: reading a target model and an occlusion model from a memory; and generating, as training data, an image in which a part of the target model is occluded by the occlusion model so that the occlusion rate of a feature region specified on a screen displaying the target model is less than a first threshold.
[0154] In the above configuration (10), images in which the characteristic portions of the target model remain unoccluded to a certain extent are used as training data. That is, in the above configuration (10), images that are likely to retain their distinctiveness as a target model even if they are partially occluded are used as training data. Therefore, according to the above configuration (10), images that are useful for training a recognition model that is also used to recognize a recognition target that is partially occluded can be generated with high accuracy as training data.
[0155] The present disclosure also includes the following aspects.
[0156] (11) The training data generation device according to (1), wherein the processor generates images with the obscuration ratio set to a different obscuration ratio for each direction.
[0157] According to the above configuration (11), it is possible to generate, with high accuracy, images that are useful for training a recognition model that is also used to recognize a recognition target that is partially occluded, as training data.
[0158] (12) A training data generation device comprising: a processor that executes the following processes: a process of displaying on a display unit a screen including a first area that accepts designation of a target model to be used for generating training data from among a plurality of target models; a second area that accepts designation of a occlusion model to be used for generating the training data from among a plurality of occlusion models; and a third area that displays the designated target model; a process of accepting designation of a feature region of the target model displayed in the third area; a process of generating an image in which a part of the designated target model is occluded with the designated occlusion model so that the occlusion rate of the feature region is less than a first threshold; and a process of outputting the generated image as the training data.
[0159] In the configuration (12) above, the selection of the target model and the occlusion model and the specification of the feature area of the target model are accepted on the screen to generate the occlusion image, so that the learning data as intended by the operator can be generated easily and accurately.
[0160] (13) A training data generation device comprising: a processor that, when receiving an instruction to select a target model to be used for generating training data from among a plurality of target models, displays a screen of the target model on a display unit; when receiving an instruction to set a feature region for the displayed target model, visualizes and displays the set feature region on the target model; and when receiving an instruction to generate the training data, outputs, as the training data, an image that is concealed by an occlusion model so that the concealment rate of the feature region of the target model is less than a first threshold.
[0161] In the configuration (13) above, images in which the characteristic portions of the target model remain unoccluded to a certain extent are used as training data. That is, in the configuration (13) above, images that are likely to retain their distinctiveness as a target model even if they are partially occluded are used as training data. Therefore, with the configuration (13) above, it is possible to generate, with high accuracy, images that are useful for training a recognition model that is also used to recognize a recognition target that is partially occluded, as training data.
[0162] This application is based on Japanese Patent Application No. 2023-217932 filed on December 25, 2023, the contents of which are incorporated herein by reference in their entirety.
[0163] REFERENCE SIGNS LIST 10 Learning data generation device 11 Control device 13 Main memory device 15 Auxiliary memory device 17 Display device 19 Input device 21 Communication device 23 Various buses 101 Operation unit 103 Operation reception unit 105 Display control unit 107 Display unit 109 Image acquisition unit 111 Object model storage unit 113 Object model acquisition unit 115 Hiding model storage unit 117 Hiding model acquisition unit 119 Background image storage unit 121 Background image acquisition unit 123 Screen generation unit 125 Condition setting unit 127 Feature region setting unit 129 Hiding rate setting unit 131 Image generation unit 133 Image output unit 135 Learning data storage unit
Claims
1. A learning data generation device comprising: a memory that stores a target model and a concealment model; and a processor that generates, as learning data, an image in which a part of the target model is concealed by the concealment model such that a concealment rate of a specified feature region on a screen for displaying the target model is less than a first threshold value.
2. The learning data generation device according to claim 1, wherein the processor accepts specification of the first threshold value and specification of a second threshold value that is a threshold value of a concealment rate of the target model, and generates, as the learning data, an image in which a part of the target model is concealed by the concealment model such that the concealment rate of the feature region is less than the first threshold value and the concealment rate of the target model is less than the second threshold value.
3. The learning data generation device according to claim 2, wherein the second threshold value is greater than the first threshold value.
4. The learning data generation device according to claim 1, wherein the processor accepts specification of a plurality of locations of the target model as the feature region.
5. The learning data generation device according to claim 1, wherein the memory further stores, as a background image, a field-of-view image from an imaging device that captures an image used for recognition in a recognition model generated by learning the learning data, and the processor generates, as the learning data, an image in which a part of the target model is concealed by the concealment model on the background image such that the concealment rate of the feature region is less than the first threshold value.
6. The learning data generation device according to claim 1, wherein the memory further stores a background image, the target model is a model of a moving body, the processor detects a movable region in which the target model can move from the terrain in the background image, and generates, as the learning data, an image in which a part of the target model is concealed by the concealment model on the movable region of the background image such that the concealment rate of the feature region is less than the first threshold value.
7. The learning data generation device according to claim 1, wherein the concealment model is a model of an object having a high appearance frequency at a location where the target model appears, and the memory stores a plurality of the concealment models for the target model.
8. The learning data generation device according to claim 1, wherein the processor generates, as learning data, a moving image obtained by hiding a part of the target model with the hiding model such that a state in which a hiding rate of the feature region is equal to or greater than a first threshold is less than a certain period of time.
9. A learning data generation method including: a step in which a processor reads a target model and a hiding model from a memory; and a step in which the processor generates, as learning data, an image obtained by hiding a part of the target model with the hiding model such that a hiding rate of a feature region specified on a screen for displaying the target model is less than a first threshold.
10. A program for causing a computer to execute: a step of reading a target model and a hiding model from a memory; and a step of generating, as learning data, an image obtained by hiding a part of the target model with the hiding model such that a hiding rate of a feature region specified on a screen for displaying the target model is less than a first threshold.
Citation Information
Patent Citations
Information processing apparatus
JP2021033707A
Learning data generation system, learning method of machine learning model and learning data generation method
JP2023047195A
Image processing device and image processing method
JP2023161432A