Training data generation device, training data generation method, and program

The learning data generation device efficiently generates images of human-object interactions using affordance information, addressing inefficiencies in manual scene setup and improving recognition accuracy by automating the positioning and posture of models.

WO2025142822A1PCT designated stage expired Publication Date: 2025-07-03PANASONIC INTELLECTUAL PROPERTY MANAGEMENT CO LTD
View PDF 3 Cites 0 Cited by

Patent Information

Application Number
PCT/JP2024/045418
Authority / Receiving Office
WO · WO
Patent Type
Applications
Current Assignee / Owner
Priority Date
2023-12-25
Filing Date
2024-12-23
Publication Date
2025-07-03

AI Technical Summary

Technical Problem

Existing image recognition techniques using machine learning face challenges in efficiently generating large quantities of learning data for scenes involving interactions between human and object models, as manually setting the interaction scenarios is inefficient.

Method used

A learning data generation device and method that utilizes affordance information to automatically arrange human and object models in interactive scenes, allowing for efficient generation of images where the human model performs defined actions on the object model, using a processor to handle affordance information and display units for user interaction.

Benefits of technology

Enables efficient generation of learning data for interaction scenes between human and object models, improving recognition accuracy by automating the positioning and posture settings and preventing unrealistic interactions, thus enhancing the quality of recognition models.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure JP2024045418_03072025_PF_FP_ABST
    Figure JP2024045418_03072025_PF_FP_ABST
Patent Text Reader

Abstract

A training data generation device according to one embodiment of the present disclosure comprises: a memory that stores a person model, an object model, and affordance information which defines an action that can be taken with respect to the object model; and a processor that generates, as training data, an image in which the person model and the object model are arranged so that the person model performs the action defined by the affordance information with respect to the object model.
Need to check novelty before this filing date? Find Prior Art

Description

Training data generation device, training data generation method, and program

[0001] The present disclosure relates to a training data generation device, a training data generation method, and a program.

[0002] Image recognition techniques using machine learning such as deep learning have been known for some time. In deep learning, it is known that the greater the number and types of training data (e.g., training data) used to train a recognition model, the higher the recognition accuracy. However, it is not realistic to prepare such a large amount of training data of various types using only real images.

[0003] For this reason, for example, Patent Document 1 discloses a technique in which images for training data are generated using CG (Computer Graphics), and annotation information is added to the images to create the training data.

[0004] Japanese Patent Application Laid-Open No. 2019-23858

[0005] However, when generating images of scenes in which interactions between a person model and an object model are occurring as learning data, even using the conventional technology described above, the user needs to set the interaction situation in detail, which makes it inefficient.

[0006] The present disclosure has been made in consideration of the above circumstances, and aims to provide a training data generation device, a training data generation method, and a program that can efficiently generate images of scenes in which interactions between human models and object models are occurring as training data.

[0007] One aspect of the training data generation device of the present disclosure includes a memory that stores a human model, an object model, and affordance information that defines actions that can be taken on the object model, and a processor that generates, as training data, images in which the human model and the object model are arranged so that the human model performs the action defined in the affordance information on the object model.

[0008] One aspect of the training data generation device of the present disclosure includes a processor that displays a screen of affordance information of an object model on a display unit, accepts instructions defining actions that can be taken with respect to the object model, sets the affordance information, accepts instructions to generate training data, and outputs, as training data, an image in which the human model and the object model are arranged so that the human model performs the action defined in the affordance information with respect to the object model.

[0009] One aspect of the training data generation method of the present disclosure includes the steps of: a processor acquiring, from a memory, a human model, an object model, and affordance information that defines an action that can be taken with respect to the object model; and a processor generating, as training data, an image in which the human model and the object model are arranged so that the human model performs the action defined in the affordance information with respect to the object model.

[0010] One aspect of the training data generation method of the present disclosure includes the steps of: a processor causing a display unit to display an editing screen for affordance information of an object model; the processor receiving an instruction defining an action that can be taken with respect to the object model and registering the action as the affordance information; and the processor receiving an instruction to generate training data and outputting, as training data, an image in which the human model and the object model are arranged so that the human model performs the action defined in the affordance information with respect to the object model.

[0011] One aspect of the program of the present disclosure is for causing a computer to execute the steps of: acquiring from a memory a person model, an object model, and affordance information that defines actions that can be taken on the object model; and generating, as training data, images in which the person model and the object model are arranged so that the person model performs the action defined in the affordance information on the object model.

[0012] One aspect of the program of the present disclosure is for causing a computer to execute the steps of: displaying an editing screen for affordance information of an object model on a display unit; accepting instructions defining actions that can be taken on the object model and registering the actions as the affordance information; and accepting instructions to generate training data and outputting, as training data, an image in which the human model and the object model are arranged so that the human model performs the action defined in the affordance information on the object model.

[0013] According to the present disclosure, it is possible to provide a training data generation device, a training data generation method, and a program that can efficiently generate images of scenes in which interactions between a person model and an object model are occurring as training data.

[0014] FIG. 1 is a block diagram showing an example of the hardware configuration of a training data generation device according to this embodiment. FIG. 2 is a block diagram showing an example of the functional configuration of the training data generation device according to this embodiment. FIG. 3 is a diagram showing an example of an object model according to this embodiment. FIG. 4 is a diagram showing an example of an object model according to this embodiment. FIG. 5 is a diagram showing an example of a person model according to this embodiment. FIG. 6 is a diagram showing an example of a person model list according to this embodiment. FIG. 7 is a diagram showing an example of a background image according to this embodiment. FIG. 8 is a diagram showing an example of a background image list according to this embodiment. FIG. 9 is a diagram showing an example of a setting screen for generating training data according to this embodiment. FIG. 10 is a diagram showing an example of an updated setting screen according to this embodiment. FIG. 11 is a diagram showing an example of a setting screen after addition of affordance information according to this embodiment. FIG. 12 is an explanatory diagram of examples of X-axis rotation, Z-axis rotation, and Y-axis rotation of an object model according to this embodiment. FIG. 13 is a diagram showing another example of affordance information according to this embodiment. FIG. 14 is an explanatory diagram of examples of horizontal rotation, vertical rotation, and free rotation of an object model according to this embodiment. FIG. 15 is a plan view showing an example of a method for generating an image by the image generation unit and the position and orientation determination unit according to this embodiment. FIG. 16 is a diagram showing an example of an image generated by the image generation unit according to this embodiment. Fig. 17 is a plan view showing another example of an image generation method by the image generation unit and the position and orientation determination unit of this embodiment. Fig. 18 is a diagram showing an example of an image generated by the image generation unit of this embodiment. Fig. 19 is a diagram showing an example of an image generated by the image generation unit of this embodiment. Fig. 20 is a diagram showing an example of an image generated by the image generation unit of this embodiment. Fig. 21 is a flowchart showing an example of training data generation processing performed in the training data generation device of this embodiment.

[0015] Hereinafter, an embodiment of the present disclosure (hereinafter simply referred to as "the present embodiment") will be described in detail with reference to the drawings. Note that the present disclosure is not limited to the following embodiment. Furthermore, the following embodiment and modified examples can be combined as appropriate.

[0016] The training data generation device of this embodiment uses affordance information to generate, as training data, images in which a human model and an object model are arranged so that the human model performs an action defined in the affordance information with respect to an object model. The affordance information defines actions that can be taken with respect to an object model. In this way, this embodiment uses affordance information, so the generation of the above-described images can be automated. In other words, this embodiment does not require the user to set in detail the position and posture of the human model and the position and posture of the object model so that the human model performs an action defined in the affordance information with respect to the object model. Therefore, this embodiment makes it possible to efficiently generate, as training data, images of scenes in which interactions between a human model and an object model are occurring.

[0017] In the following, a case where the training data generation device of this embodiment generates training data used for training a recognition model using CG (Computer Graphics) will be described as an example, but the present invention is not limited to this. Examples of recognition models include, but are not limited to, a recognition model used to search for a person interacting with an object and a recognition model used for behavior analysis of a person interacting with an object. The recognition model may be any recognition model that is used to recognize a scene in which an interaction between a person and an object is occurring.

[0018] 1 is a block diagram showing an example of the hardware configuration of a training data generation device 10 according to this embodiment. As shown in FIG. 1, the training data generation device 10 includes a control device 11 (an example of a processor), a main storage device 13 (an example of a memory), an auxiliary storage device 15 (an example of a memory), a display device 17, an input device 19, a communication device 21, and various buses 23. The control device 11, the main storage device 13, the auxiliary storage device 15, the display device 17, the input device 19, and the communication device 21 are connected via the various buses 23. As such, the training data generation device 10 according to this embodiment has a general hardware configuration using a typical computer.

[0019] The control device 11 controls the overall operation of the training data generation device 10. The control device 11 may be, for example, at least one of a central processing unit (CPU) and a graphics processing unit (GPU), but is not limited to these. There may be any number of CPUs or GPUs as long as they are one or more, and they may be single-core or multi-core.

[0020] Examples of the main storage device 13 include, but are not limited to, a read-only memory (ROM) and a random access memory (RAM). The ROM stores various programs, such as a program for controlling the training data generation device 10 and a program for generating training data according to this embodiment. The RAM is used as a working area when the control device 11 performs various controls based on the programs stored in the ROM.

[0021] The auxiliary storage device 15 stores the various programs described above and data for generating training data in this embodiment. The various programs described above may be stored in at least one of the main storage device 13 and the auxiliary storage device 15. Examples of the auxiliary storage device 15 include, but are not limited to, existing storage devices capable of magnetic, electrical, or optical storage, such as a hard disk drive (HDD), a solid state drive (SSD), and a digital versatile disc (DVD). The auxiliary storage device 15 may be built into the training data generation device 10 or may be externally attached to the training data generation device 10 via an interface such as a universal serial bus (USB). The auxiliary storage device 15 may also be a network-attached storage (NAS) connected via a network such as a local area network (LAN) or a wide area network (WAN).

[0022] The display device 17 displays various screens used when the training data generation device 10 generates training data, and serves as a user interface with the user (operator). Examples of the display device 17 include, but are not limited to, various displays such as a liquid crystal display, an organic electroluminescence (EL) display, and a touch panel display. The display device 17 may be a built-in display built into the training data generation device 10, or an external display connected to the training data generation device 10 via a display interface such as HDMI (registered trademark).

[0023] The input device 19 is used for various inputs, selections, and specifications used when the training data generation device 10 generates training data, and serves as a user interface with the user (operator). Examples of the input device 19 include, but are not limited to, a keyboard, a mouse, and a touch panel. The input device 19 may be built into the training data generation device 10 or may be externally attached to the training data generation device 10 via an interface such as a USB.

[0024] Examples of the communication device 21 include, but are not limited to, a communication device for a wired LAN and a wireless communication device for a wireless LAN. The communication device 21 may be used to externally acquire the program or data for generating the training data of this embodiment, or may be used to externally output the training data generated by the training data generation device 10.

[0025] In addition to the above configuration, the training data generation device 10 may further include hardwired circuits specific to the training data generation device 10, such as an IC (Integrated Circuit), an ASIC (Application Specific Integrated Circuit), and an FPGA (Field-Programmable Gate Array), in order to realize the training data generation function.

[0026] 2 is a block diagram showing an example of the functional configuration of the training data generation device 10 according to this embodiment. As shown in FIG. 2, the training data generation device 10 includes an operation unit 101, an operation reception unit 103, a display control unit 105, a display unit 107, an image acquisition unit 109, an object model storage unit 111, a person model storage unit 115, a background image storage unit 119, a screen generation unit 123, an affordance information setting unit 125, an image generation unit 127, a training data output unit 131, and a training data storage unit 133. The image acquisition unit 109 includes the object model acquisition unit 113, the person model acquisition unit 117, and the background image acquisition unit 121. The image generation unit 127 includes a position and orientation determination unit 129.

[0027] The operation unit 101 can be realized by, for example, the input device 19 described with reference to FIG.

[0028] The operation reception unit 103, the display control unit 105, the image acquisition unit 109, the screen generation unit 123, the affordance information setting unit 125, the image generation unit 127, and the learning data output unit 131 can be realized, for example, by the control unit 11 and the main memory unit 13 described in Figure 1.

[0029] For example, the control device 11 reads out a program for generating learning data according to this embodiment that is stored in the main storage device 13 (ROM) or the auxiliary storage device 15, or that is externally acquired from the communication device 21 via a network, and loads the program into the main storage device 13 (RAM). The control device 11 executes various processes in accordance with the loaded program, thereby realizing each of the above-described functional units. Here, the description has been given taking an example in which each of the above-described functional units is realized as software, but at least a portion of each of the above-described functional units may also be realized as hardware. In this case, the functional units realized as hardware may be realized, for example, by the above-described hardwired circuit. Furthermore, any of the above-described functional units may also be realized by a combination of software and hardware.

[0030] The display unit 107 can be realized, for example, by the display device 17 described in Fig. 1. The object model storage unit 111, the person model storage unit 115, the background image storage unit 119, and the learning data storage unit 133 can be realized, for example, by at least one of the main storage unit 13 and the auxiliary storage unit 15 described in Fig. 1.

[0031] Based on user operations, the operation unit 101 performs various operation inputs required for training data generation on the training data generation device 10. Examples of the various operation inputs required for training data generation include, but are not limited to, an operation input for displaying a setting screen for generating training data, an operation input for specifying an object model to be used for training data generation, an operation input for setting affordance information to be used for training data generation, and an operation input for instructing the start of training data generation.

[0032] The operation reception unit 103 receives various operation inputs from the operation unit 101 .

[0033] The display control unit 105 controls the display of a setting screen generated or updated by a screen generation unit 123 (to be described later) on the display unit 107 .

[0034] The display unit 107 displays a setting screen under the control of the display control unit 105 .

[0035] The image acquisition unit 109 acquires various image data stored in the object model storage unit 111, the person model storage unit 115, and the background image storage unit 119. As described above, the image acquisition unit 109 includes the object model acquisition unit 113, the person model acquisition unit 117, and the background image acquisition unit 121.

[0036] The object model storage unit 111 stores an object model. As described above, in this embodiment, images serving as learning data are generated using CG. For this reason, in this embodiment, the object model is described as a 3D model used in CG, i.e., a three-dimensional image, but the present invention is not limited to this. In addition, for the object model in this embodiment, a local coordinate system with coordinate axes XYZ is set so that the front direction of the object model is the +X-axis direction.

[0037] Fig. 3 is a diagram showing an example of an object model 211 of this embodiment. Fig. 4 is a diagram showing an example of an object model 221 of this embodiment. The object model 211 is a 3D model of a dolly. A local coordinate system with coordinate axes XYZ is set for the object model 211 so that the +X axis direction is the front direction. The front direction of the object model 211 is the traveling direction of the dolly. The object model 221 is a 3D model of an umbrella. A local coordinate system with coordinate axes XYZ is set for the object model 221 so that the +X axis direction is the front direction.

[0038] The object model acquisition unit 113 acquires an object model from the object model storage unit 111. For example, the object model acquisition unit 113 acquires a list of file names of object models stored in the object model storage unit 111 in response to an instruction from the screen generation unit 123 (described later). Furthermore, for example, the object model acquisition unit 113 acquires an object model in response to an instruction from the image generation unit 127 (described later).

[0039] The character model storage unit 115 stores character models. The character models perform actions defined by affordance information, which will be described later, on object models. As described above, in this embodiment, the object models are 3D models used in CG. Therefore, in this embodiment, the character models are also 3D models used in CG, i.e., three-dimensional images, but the present invention is not limited to this. In addition, for the character models in this embodiment, a local coordinate system with coordinate axes XYZ is set so that the front direction of the character model is the +X-axis direction.

[0040] In this embodiment, the character model storage unit 115 stores each character model and a character model list that lists the file names of each character model. FIG. 5 is a diagram showing an example of a character model 311 in this embodiment. FIG. 6 is a diagram showing an example of a character model list 301 in this embodiment. The character model 311 is a 3D model of a young man. A local coordinate system with coordinate axes XYZ is set for the character model 311 so that the +X axis direction is the front direction. The character model list 301 is a file in CSV (Comma Separated Values) format named "ModelList.csv" that lists the file names of each character model. Note that in the example shown in FIG. 6, the file name extensions of each character model are omitted. For example, image data named "Human_model_000" in the character model list 301 indicates the character model 311.

[0041] The character model acquisition unit 117 acquires a character model or a character model list from the character model storage unit 115. For example, the character model acquisition unit 117 acquires a character model list in response to an instruction from a screen generation unit 123 (described later), or acquires a character model in response to an instruction from an image generation unit 127 (described later).

[0042] The background image storage unit 119 stores background images that serve as backgrounds for object models and character models. The background images may be two-dimensional or three-dimensional images. In this embodiment, the background image storage unit 119 stores each background image and a background image list that lists the file names of each background image. FIG. 7 is a diagram showing an example of a background image 411 in this embodiment. FIG. 8 is a diagram showing an example of a background image list 401 in this embodiment. The background image 411 is an image of a cloudy sky that looks like it might soon rain. The background image list 401 is a CSV file named "BackgroundList.csv" that lists the file names of each background image. Note that in the example shown in FIG. 8, the file name extensions of the background images are omitted. For example, image data with the file name "Background_000" in the background image list 401 represents the background image 411.

[0043] Background image acquisition unit 121 acquires a background image or a background image list from background image storage unit 119. For example, background image acquisition unit 121 acquires a background image list in response to an instruction from screen generation unit 123 (described later), or acquires a background image in response to an instruction from image generation unit 127 (described later).

[0044] The screen generation unit 123 generates and updates a setting screen for generating learning data in accordance with various operation inputs received by the operation reception unit 103 .

[0045] When operation input instructing display of a setting screen for generating learning data is accepted by operation accepting unit 103, screen generating unit 123 newly generates the setting screen. Specifically, screen generating unit 123 instructs object model acquiring unit 113 to acquire a list of file names of object models stored in object model storage unit 111. Screen generating unit 123 newly generates a setting screen for generating learning data using this acquired list of file names of object models.

[0046] Fig. 9 is a diagram showing an example of a setting screen 501 for generating training data according to this embodiment. The setting screen 501 shown in Fig. 9 is a setting screen newly generated by the screen generating unit 123. As shown in Fig. 9, the setting screen 501 is roughly divided into a display of items 503 related to models and images used in generating training data, and a display of items 505 related to affordance information.

[0047] First, a description will be given of item 503. Item 503 includes a pull-down list 511 for allowing the user to specify an object model, a display area 513 in which the object model specified in pull-down list 511 is displayed, a display area 515 in which a person model list is displayed, and a display area 517 in which a background image list is displayed.

[0048] 9, since the setting screen 501 is in an initial state, the pull-down list 511, the display area 513, the display area 515, and the display area 517 are all blank. However, if the user selects the pull-down list 511 using the operation unit 101, a list of file names of object models acquired by the screen generation unit 123 is displayed, and the user can select an object model to be used to generate training data.

[0049] For this reason, when the operation receiving unit 103 receives an operation input for specifying an object model to be used in generating learning data, the screen generating unit 123 updates the setting screen. Specifically, the screen generating unit 123 instructs the object model acquiring unit 113 to acquire the specified object model, instructs the person model acquiring unit 117 to acquire a person model list, and instructs the background image acquiring unit 121 to acquire a background image list. The screen generating unit 123 updates the setting screen using these acquired object models, person model lists, and background image lists. Note that a plurality of person model lists and a plurality of background image lists may be prepared, and the screen generating unit 123 may acquire a person model list and a background image list linked to the acquired object model.

[0050] 10 is a diagram showing an example of an updated setting screen 501 according to this embodiment. The setting screen 501 shown in Fig. 10 is a setting screen updated by the screen generating unit 123 when an operation input for specifying an object model is accepted by the operation accepting unit 103.

[0051] In the example shown in FIG. 10 , an object model with the file name "Trolley_000" is specified in the pull-down list 511. Since the object model with the file name "Trolley_000" is the object model 211, the object model 211 is displayed in the display area 513. In this manner, in this embodiment, when the screen generation unit 123 receives an instruction via the operation reception unit 103 to select an object model to be used for generating learning data from among a plurality of object models, the screen generation unit 123 displays a screen of the object model on the display unit 107 via the display control unit 105. The display area 515 displays "ModelList.csv", which is the file name of the person model list 301, and the display area 517 displays "BackgroundList.csv", which is the file name of the background image list 401.

[0052] As a result, the object model, character model, and background image to be used for generating the training data are set on the setting screen 501. Note that the character model is automatically set from a character model list 301 by the image generation unit 127 (described later), and the background image is automatically set from a background image list 401 by the image generation unit 127 (described later). However, the method for setting the character model and background image to be used for generating the training data is not limited to this. For example, the screen generation unit 123 may display a list of character models, similar to the object models, and allow the user to select a character model to be used for generating the training data from the list. For another example, the screen generation unit 123 may display a list of background images, similar to the object models, and allow the user to select a background image to be used for generating the training data from the list.

[0053] Next, item 505 will be described. Note that the contents of item 505 have not been updated since the state in Fig. 9 , so the description will continue with reference to Fig. 10. Item 505 includes an add button 521, an edit button 523, a delete button 525, affordance information 531, a save button 541, and a load button 543.

[0054] The add button 521 is a button for adding affordance information, the edit button 523 is a button for editing affordance information, and the delete button 525 is a button for deleting affordance information.

[0055] The affordance information 531 is information that defines an action that can be taken with respect to an object model. In this embodiment, the affordance information 531 is information that includes object parts, person parts / actions, X-axis rotation, Z-axis rotation, and Y-axis rotation. The object parts are information that define parts of the object model at which actions that can be taken with respect to the object model are performed. The person parts are information that define parts of the person model at which actions that can be taken with respect to the object model are performed. The actions are information that define actions that can be taken with respect to the object model, and examples of actions include, but are not limited to, grabbing (with hands), grasping (with hands), pinching (with hands), pushing (with hands), rotating (with hands), and kicking (with feet).

[0056] As a premise for the X-axis rotation, Z-axis rotation, and Y-axis rotation, when an object model is placed in a CG space (described later), the default orientation (home position) is such that the forward direction (+X-axis direction) of the object model coincides with the forward direction (+X-axis direction) of the person model. The X-axis rotation is information indicating the rotatable angle relative to the X-axis direction of a local coordinate system set for the object model. The Z-axis rotation is information indicating the rotatable angle relative to the Z-axis direction of a local coordinate system set for the object model. The Y-axis rotation is information indicating the rotatable angle relative to the Y-axis direction of a local coordinate system set for the object model. In this manner, the X-axis rotation, Z-axis rotation, and Y-axis rotation define the allowable range of the relative orientation that the object model can take with respect to the person model. Note that, although the rotation angles of the object model are set in the local coordinate system here, this is not limiting and they may also be set in the world coordinate system.

[0057] The Save button 541 is a button for saving the affordance information 531 as a file. The Load button 543 is a button for reading out the saved file as the affordance information 531.

[0058] Here, the affordance information setting unit 125 will be described. The affordance information setting unit 125 adds, edits, and deletes the affordance information 531. When the user uses the operation unit 101 to specify new information (record) in the affordance information 531 and presses the Add button 521, the affordance information setting unit 125 adds the specified information (record) to the affordance information 531. This sets (registers) the new affordance information. Furthermore, when the user uses the operation unit 101 to specify editing content for information (record) already set in the affordance information 531 and presses the Edit button 523, the affordance information setting unit 125 updates the information (record) with the specified editing content. This edits the affordance information. Furthermore, when the user uses the operation unit 101 to specify a record already set in the affordance information 531 and presses the Delete button 525, the affordance information setting unit 125 deletes the record. This deletes the affordance information.

[0059] The following specifically describes the addition of affordance information 531, appropriately referring to the screen generation unit 123. When the operation reception unit 103 receives an operation input (instruction) for registering new information in an empty record of the affordance information 531, the screen generation unit 123 displays the information input in the empty record. Furthermore, when the screen generation unit 123 receives an operation input (instruction) for pressing the add button 521 via the operation reception unit 103, the affordance information setting unit 125 adds the record into which the new information has been input to the affordance information 531.

[0060] Fig. 11 is a diagram showing an example of the setting screen 501 after adding affordance information 531 of this embodiment. In the example shown in Fig. 11, a new record 532 has been added to the affordance information 531. In the record 532, a handle is set as the object part, grabbing with both hands / front hand is set as the person part / action, and none is set as the X-axis rotation, Z-axis rotation, and Y-axis rotation.

[0061] The user can specify an object part on the object model 211 displayed in the display area 513. Specifically, when the screen generation unit 123 receives, via the operation reception unit 103, an operation input (instruction) specifying a part to be set as an object part on the object model 211 displayed in the display area 513, the affordance information setting unit 125 sets the specified part as the object part. In the example shown in Fig. 11 , the user specifies part 213, which is the handle part of the object model 211, and the affordance information setting unit 125 sets part 213 (handle) as the object part.

[0062] Furthermore, when the screen generating unit 123 receives an operation input (instruction) specifying information for the person's body part / action of record 532 via the operation receiving unit 103, the affordance information setting unit 125 sets the specified information for the person's body part / action. In the example shown in Fig. 11 , the user specifies both hands as the person's body part and "grabbing with normal hands" as the action, so that the affordance information setting unit 125 sets "grabbing with both hands / normal hands" as the person's body part / action.

[0063] Furthermore, when the screen generating unit 123 receives, via the operation receiving unit 103, an operation input (instruction) specifying a rotatable angle for each of the X-axis rotation, Z-axis rotation, and Y-axis rotation of the record 532, the affordance information setting unit 125 sets the specified angle to the X-axis rotation, Z-axis rotation, and Y-axis rotation. In the example shown in Fig. 11 , the user specifies "none" for the X-axis rotation, Z-axis rotation, and Y-axis rotation, causing the affordance information setting unit 125 to set "none" for the X-axis rotation, Z-axis rotation, and Y-axis rotation.

[0064] As a result, as shown in FIG. 11 , a new record 532 is added to the affordance information 531. The record 532 is affordance information indicating an action in which the person model grasps the handle of the object model 211, which is a cart, with both hands and the palm of the hand. FIG. 12 is an explanatory diagram of an example of X-axis rotation, Z-axis rotation, and Y-axis rotation of the object model 211 in this embodiment. In the record 532, since X-axis rotation, Z-axis rotation, and Y-axis rotation are all set to "no," the object model 211 does not perform any of the X-axis rotation, Z-axis rotation, and Y-axis rotation shown in FIG. 12 . Therefore, the object model 211 does not rotate relative to the person model, and the orientation that the object model 211 can take relative to the person model is unique. In this way, by setting all of X-axis rotation, Z-axis rotation, and Y-axis rotation to "no," it is possible to define an orientation (unique orientation) that the object model 211 can take. As described above, the record 532 is affordance information that further indicates that the orientation (front direction) of the object model 211 is a default orientation that matches the front direction of the human model.

[0065] When editing record 532, the user uses operation unit 101 to specify (instruct) the editing content of the information set in record 532 and presses (instructs) edit button 523, causing affordance information setting unit 125 to update record 532 with the specified editing content. When deleting record 532, the user uses operation unit 101 to specify (instruct) record 532 and presses (instructs) delete button 525, causing affordance information setting unit 125 to delete record 532.

[0066] Here, another example of the affordance information 531 will also be described. Fig. 13 is a diagram showing another example of the affordance information 531 of this embodiment. In Fig. 13, the affordance information 531 of the object model 221 will be described, not the object model 211. In the example shown in Fig. 13, the file name "Umbrella_000" of the object model 221 is specified in the pull-down list 511, and the object model 221 is displayed in the display area 513.

[0067] 13 , new records 535 and 536 are added to the affordance information 531. In record 535, the umbrella handle is set as the object part, "grabbing with one hand / forward hand" is set as the person part / action, and none, 0 to 360 degrees, and 0 to 360 degrees are set as the X-axis rotation, Z-axis rotation, and Y-axis rotation. In record 536, the umbrella body is set as the object part, "grabbing with one hand / forward hand" is set as the person part / action, and none, 0 to 360 degrees, and 0 to 360 degrees are set as the X-axis rotation, Z-axis rotation, and Y-axis rotation.

[0068] The method for setting the object part, person part / action, X-axis rotation, Z-axis rotation, and Y-axis rotation of records 535 and 536 is the same as that of record 532, and therefore will not be described here. In the example shown in Fig. 13, the user specifies part 223, which is the grip part of object model 221, and the affordance information setting unit 125 sets part 223 (grip part) as the object part of record 535. Similarly, the user specifies part 225, which is the torso part of object model 221, and the affordance information setting unit 125 sets part 225 (torso part) as the object part of record 535.

[0069] Record 535 is affordance information indicating an action in which the human model grasps the handle of the object model 221, which is an umbrella, with one hand and with the right hand. FIG. 14 is an explanatory diagram of an example of X-axis rotation, Z-axis rotation, and Y-axis rotation of the object model 221 in this embodiment. In record 535, X-axis rotation is set to none, and Z-axis rotation and Y-axis rotation are set to 0 to 360 degrees. Therefore, the object model 221 does not rotate around the X-axis shown in FIG. 14, but can rotate around the Z-axis and Y-axis. Therefore, the object model 221 does not rotate relative to the human model around the X-axis, but can rotate relative to the Z-axis and Y-axis. Therefore, the orientation that the object model 221 can take relative to the human model does not change around the X-axis, but can change around the Z-axis and Y-axis. In this way, by setting the rotation angle around any of the X-axis, Z-axis, and Y-axis, the allowable range of postures that the object model 221 can take can be defined. As described above, the record 535 is affordance information that further indicates that the orientation (front direction) of the object model 221 can be changed around the Z axis and Y axis from the default orientation that matches the front direction of the human model.

[0070] Record 536 is affordance information indicating an action in which a human model grasps the body of the object model 221, which is an umbrella, with one hand and with the right hand. In record 536, X-axis rotation is set to none, and Z-axis rotation and Y-axis rotation are set to 0 to 360 degrees. Therefore, the object model 221 does not rotate around the X-axis shown in FIG. 14 , but can rotate around the Z-axis and Y-axis. Therefore, the object model 221 does not rotate around the X-axis relative to the human model, but can rotate around the Z-axis and Y-axis relative to the human model. Therefore, the orientation of the object model 221 relative to the human model does not change around the X-axis, but can change around the Z-axis and Y-axis. In other words, record 536 is affordance information further indicating that the orientation (front direction) of the object model 221 can change around the Z-axis and Y-axis from a default orientation that matches the front direction of the human model.

[0071] Next, an input box 551 for inputting the number of images to be generated as learning data, and a learning data generation start button 553 will be described.

[0072] When the operation receiving unit 103 receives an operation input specifying the number of images to be generated as learning data in the input box 551, the screen generating unit 123 sets the specified number and displays the specified number in the input box 551.

[0073] When the operation receiving unit 103 receives an operation input of pressing the start generation button 553, the screen generating unit 123 instructs the image generating unit 127 to start generating learning data under the various conditions set on the setting screen 501.

[0074] When an instruction to start generating learning data is received, the image generation unit 127 generates, as learning data, an image in which a person model and an object model are arranged so that the person model performs an action defined in the affordance information on an object model. Specifically, the image generation unit 127 generates, as learning data, an image in which a person model and an object model are arranged so that the person model performs an action defined in the affordance information on an object part of the object model. For example, the image generation unit 127 generates, as learning data, an image in which a person model and an object model are arranged so that the person model grasps an object part of the object model. Furthermore, for example, the image generation unit 127 generates, as learning data, an image in which a person model and an object model are arranged so that the object model assumes a posture defined in the affordance information when the person model performs an action defined in the affordance information on the object model. Furthermore, for example, the image generation unit 127 generates, as learning data, an image in which a person model and an object model are arranged so that the object model assumes a posture within a permissible range defined in the affordance information when the person model performs an action defined in the affordance information on the object model.

[0075] In detail, when the image generation unit 127 is instructed by the screen generation unit 123 to start generating training data, the image generation unit 127 acquires, from the screen generation unit 123, information on the models and images used to generate the training data, set in item 503, and information on the number of images to be generated as training data. Furthermore, the image generation unit 127 acquires, from the image acquisition unit 109, the object models, person models, and background images used to generate the training data, based on the information on the models and images used to generate the training data acquired from the screen generation unit 123. Furthermore, the image generation unit 127 acquires the affordance information 531 set in item 505 from the affordance information setting unit 125. As described above, the image generation unit 127 includes the position and orientation determination unit 129. Based on the acquired information, the position and orientation determination unit 129 determines the arrangement of the person models and object models so that the person models perform actions defined by the affordance information with respect to the object models. The image generation unit 127 generates, as training data, images in the arrangement determined by the position and orientation determination unit 129.

[0076] The following describes in detail a method for generating training data by the image generation unit 127 and the position and orientation determination unit 129 of this embodiment. In practice, a background image is also used to generate training data, but for the sake of convenience, the background image may be omitted from the description here. When training data is generated using a background image, the context can be made closer to a real image (such as a photograph) even for training data generated by CG, thereby improving the recognition accuracy of the recognition model.

[0077] Fig. 15 is a plan view showing an example of an image generation method by the image generation unit 127 and the position and orientation determination unit 129 of this embodiment. In the example shown in Fig. 15, an image is generated using information set on the setting screen 501 shown in Fig. 11. Specifically, the object model 211, the person model 311, and the affordance information of the record 532 are used. A description of the background image will be omitted.

[0078] As described above, in this embodiment, training data is generated using CG. To this end, the image generation unit 127 defines a three-dimensional CG space 601 as shown in Fig. 15 , and the position and orientation determination unit 129 places an object model 211 and a person model 311, which are 3D models, and a virtual camera 611 within the CG space 601. The virtual camera 611 projects (renders) various models and backgrounds included in a field of view 603 onto a projection surface (not shown) to generate an image. As a rendering method, for example, a known method such as the Z-buffer method may be used.

[0079] The position and orientation determination unit 129 determines the placement of the human model 311 and the object model 211 within the field of view 603 so that the human model 311 performs the action defined by the affordance information of the record 532 on the object model 211 .

[0080] First, the position and orientation determination unit 129 determines the position and orientation of the human model 311. The position and orientation determination unit 129 determines the position of the human model 311 to be any position within the field of view 603. The position and orientation determination unit 129 also determines the orientation of the human model 311 by referring to the human body part / action in the affordance information of the record 532. In this case, since the human body part / action is grabbing with both hands / forward hand, the position and orientation determination unit 129 determines the orientation of the human model 311 to be a grabbing posture with both hands and forward hand. Note that the position and orientation determination unit 129 may also determine the orientation of the human model 311 by referring to the object body part and the object model 211 in the affordance information of the record 532, the position of the ground in the field of view 603, and the like. In this way, the grip strength of the hands of the human model 311 when grabbing the handle (body part 213) of the object model 211 and the position of the outstretched arms can be accurately determined. The position and orientation determination unit 129 positions the human model 311 so that it assumes the determined orientation at the determined position.

[0081] Next, the position and orientation determination unit 129 determines the position and orientation of the object model 211. The position and orientation determination unit 129 determines the position of the object model 211 by referring to the object part and person part / action in the affordance information of the record 532. Here, the object part is a handle, and the person part / action is grasping with both hands / forward hand, so the position and orientation determination unit 129 determines the position of the object model 211 to be a position where the person model 311 can grasp the handle of the object model 211 with both hands and forward hand. The position and orientation determination unit 129 also determines the orientation of the object model 211 by referring to the X-axis rotation, Z-axis rotation, and Y-axis rotation in the affordance information of the record 532. Here, there is no X-axis rotation, Z-axis rotation, or Y-axis rotation, so the position and orientation determination unit 129 determines the orientation of the object model 211 to be the default orientation. The default orientation is an orientation in which the orientation (front direction) of the object model 211 matches the front direction of the person model 311. The position and orientation determination unit 129 places the object model 211 so that it takes the determined orientation at the determined position.

[0082] Therefore, in the example shown in Figure 15, in a default orientation in which the orientation (front direction) of the object model 211 matches the front direction of the person model 311, the person model 311 is positioned within the field of view 603 so as to grasp the handle of the object model 211 with both hands and the right hand.

[0083] Fig. 16 is a diagram showing an example of an image 651A generated by the image generation unit 127 of this embodiment. The image 651A is an image generated by the virtual camera 611 in the example shown in Fig. 15 . The image generation unit 127 generates the image 651A and uses it as learning data. As shown in Fig. 16 , the image 651A is an image in which, in a state in which the front direction of the object model 211 coincides with the front direction of the person model 311, the person model 311 is performing the action of grasping a part 213 serving as a handle of the object model 211 with both hands 313 and the normal hand.

[0084] Here, other examples of images (learning data) generated by the image generation unit 127 will also be described. Fig. 17 is a plan view showing another example of an image generation method by the image generation unit 127 and the position and orientation determination unit 129 of this embodiment. In the example shown in Fig. 17, an image is generated using information set on the setting screen 501 shown in Fig. 13. Specifically, the object model 221, the person model 311, and the affordance information of the record 535 are used. Note that a description of the background image will be omitted.

[0085] The image generation unit 127 defines a three-dimensional CG space 701 as shown in FIG. 17, and the position and orientation determination unit 129 places the object model 221 and the person model 311, which are 3D models, and a virtual camera 711 in the CG space 701.

[0086] The position and orientation determination unit 129 determines the position of the human model 311 to be any position within the field of view 703. The position and orientation determination unit 129 also determines the orientation of the human model 311 by referring to the human body part / action in the affordance information of the record 535. In this case, the human body part / action is grabbing with one hand / forward hand, so the position and orientation determination unit 129 determines the orientation of the human model 311 to be a grabbing posture with one hand and forward hand.

[0087] The position and orientation determination unit 129 determines the position of the object model 221 by referring to the object part and person part / action in the affordance information of the record 535. Here, the object part is the gripping part, and the person part / action is grasping with one hand / forward hand, so the position and orientation determination unit 129 determines the position of the object model 221 to be a position where the person model 311 can grasp the gripping part of the object model 221 with one hand and forward hand. The position and orientation determination unit 129 also determines the orientation of the object model 221 by referring to the X-axis rotation, Z-axis rotation, and Y-axis rotation in the affordance information of the record 535. Here, the X-axis rotation, Z-axis rotation, and Y-axis rotation are none, 0 to 360 degrees, and 0 to 360 degrees, respectively, so the position and orientation determination unit 129 determines the orientation of the object model 221 to be the default orientation, and adjusts the orientation around the Z-axis and Y-axis. For example, the position and orientation determination unit 129 adjusts the orientation of the object model 221 around the Y axis so that the holding portion of the object model 221 and the portion of the hand of the human model 311 that grasps the object model 221 are parallel to each other. The position and orientation determination unit 129 places the object model 221 so that it takes the determined orientation at the determined position.

[0088] 17 , the orientation (front direction) of the object model 221 is a default orientation that matches the front direction of the person model 311. The posture of the object model 221 is also adjusted around the Y axis so that the holding portion of the object model 221 and the part of the hand of the person model 311 that grasps the object model 221 are parallel. In this state, the person model 311 is positioned within the field of view 703 so as to grasp the holding portion (part 223) of the object model 221 with one hand and with the right hand.

[0089] FIG. 18 is a diagram showing an example of an image 751A generated by the image generation unit 127 of this embodiment. The image 751A is an image generated by the virtual camera 711 in the example shown in FIG. 17 . The image generation unit 127 generates the image 751A and uses it as learning data. As shown in FIG. 18 , the image 751A is in a state in which the front direction of the object model 221 coincides with the front direction of the person model 311. The posture of the object model 221 is adjusted around the Y axis so that the part 223, which is the holding portion of the object model 221, and the part of the hand 313 of the person model 311 that grasps the object model 221 are parallel. In this state, the image 751A is an image in which the person model 311 grasps the part 223, which is the holding portion of the object model 221, with one hand 313 and the normal hand.

[0090] FIG. 19 is a diagram showing an example of an image 751B generated by the image generation unit 127 of this embodiment. Image 751B is an image generated by the virtual camera 711 in the example shown in FIG. 17 , in which the person model 311 and the object model 221 are arranged using the affordance information of record 536 instead of the affordance information of record 535. The arrangement method of the person model 311 and the object model 221 is the same as in the case of FIG. 18 , and therefore description thereof will be omitted. The image generation unit 127 generates image 751B and uses it as learning data. As shown in FIG. 19 , image 751B is in a state in which the front direction of the object model 221 coincides with the front direction of the person model 311. Furthermore, the posture of the object model 221 is adjusted around the Y axis so that the body part 225 of the object model 221 and the part of the hand 313 of the person model 311 that grasps the object model 221 are parallel to each other. Therefore, the angle of the object model 221 around the Y axis is different compared to the example shown in Fig. 18. Image 751B is an image in which, in this state, the human model 311 is grasping body part 225, which is the torso part of the object model 221, with one hand 313 and the normal hand.

[0091] 20 is a diagram showing an example of an image 751C generated by the image generation unit 127 of this embodiment. The image 751C is an image generated by the virtual camera 711 in a state where a background image 411 is placed in addition to the object model 211 and the person model 311 in the CG space 701. The image 751C is an image obtained by adding the background image 411 to the image 751B shown in FIG.

[0092] The learning data output unit 131 outputs the learning data generated by the image generation unit 127 to the learning data storage unit 133. When an instruction to generate learning data is accepted by the screen generation unit 123, the learning data output unit 131 outputs, as learning data, to the learning data storage unit 133 an image in which a person model and an object model are arranged so that the person model performs an action defined by the affordance information on the object model.

[0093] The learning data storage unit 133 stores, as learning data, the images generated by the image generation unit 127. Specifically, the learning data storage unit 133 stores the learning data output from the learning data output unit 131.

[0094] FIG. 21 is a flowchart showing an example of the training data generation process performed by the training data generation device 10 of this embodiment.

[0095] First, when an operation input for specifying an object model is received, the screen generating unit 123 sets the specified object model as an object model to be used for generating learning data (step S11).

[0096] Next, if the affordance information is to be changed (YES in step S13), the process proceeds to step S15, and if the affordance information is not to be changed (NO in step S13), the process proceeds to step S37.

[0097] Next, if affordance information is to be added (YES in step S15), the process proceeds to step S17, and if affordance information is not to be added (NO in step S15), the process proceeds to step S27.

[0098] Next, when the screen generation unit 123 receives an operation input specifying the part to be set as the object part of the affordance information on the object model displayed in the display area 513, the affordance information setting unit 125 sets the specified part as the object part (step S17).

[0099] Next, when the screen generating unit 123 receives an operation input of information specifying a human body part, the affordance information setting unit 125 sets the specified information as the human body part (step S19).

[0100] Next, when the screen generating unit 123 receives an operation input of information specifying the action of the human model, the affordance information setting unit 125 sets the specified information as the action of the human model (step S21).

[0101] Next, when the screen generation unit 123 receives an operation input specifying the rotation angles of the object model around the X-axis, Z-axis, and Y-axis, the affordance information setting unit 125 sets the specified rotation angles for the X-axis, Z-axis, and Y-axis rotation (step S23).

[0102] Next, when the screen generation unit 123 receives an operation input of pressing the add button 521, the affordance information setting unit 125 adds the information (record) specified in steps S17 to S23 to the affordance information and saves it (step S25).

[0103] Next, if the affordance information is to be edited (YES in step S27), the process proceeds to step S29, and if the affordance information is not to be edited (NO in step S27), the process proceeds to step S31.

[0104] Next, when the screen generation unit 123 receives an operation input specifying the record of the affordance information to be edited (step S29), the process proceeds to step S17, and the same process as adding affordance information is performed on the affordance information to be edited. However, in step S25, the operation input to press the edit button 523 instead of the add button 521 is received.

[0105] Next, if the affordance information is to be deleted (YES in step S31), the process proceeds to step S33, and if the affordance information is not to be deleted (NO in step S31), the process returns to step S13.

[0106] Next, the screen generation unit 123 accepts an operation input specifying the record of affordance information to be deleted (step S33), and when it further accepts an operation input to press the delete button 525, the affordance information setting unit 125 deletes the specified record (step S35).

[0107] Next, the screen generation unit 123 acquires the character model list and the background image list, and sets the character model and background image to be used for generating the learning data (step S37). Note that the screen generation unit 123 may select a character model from the character model list in any way, but in this embodiment, the character models are selected and set in the order of the character model list. The same applies to the background image.

[0108] Next, when the image generation unit 127 receives an operation input of pressing the generation start button 553, it places a person model in the CG space and also places an object model to adapt to the person model using the affordance information (step S39).

[0109] Next, the image generation unit 127 generates images of the arranged human model and object model, and the learning data output unit 131 stores the generated images as learning data in the learning data storage unit 133 (step S41).

[0110] Next, if learning data is to be generated using different affordance information (YES in step S43), return to step S39; if learning data is not to be generated using different affordance information (NO in step S43), proceed to step S45.

[0111] Next, if the training data is to be generated using a different character model / background image (YES in step S45), the process returns to step S37; if the training data is not to be generated using a different character model / background image (NO in step S45), the process proceeds to step S47.

[0112] Next, if training data is to be generated using another object model (YES in step S47), the process returns to step S11, and if training data is not to be generated using another object model (NO in step S47), the process ends.

[0113] As described above, in this embodiment, by using affordance information, images in which a person model and an object model are arranged so that the person model performs an action defined in the affordance information on an object model are automatically generated as training data. In other words, in this embodiment, there is no need for the user to set in detail the position and posture of the person model and the position and posture of the object model so that the person model performs an action defined in the affordance information on the object model. Therefore, according to this embodiment, images of scenes in which interactions between a person model and an object model are occurring can be efficiently generated as training data.

[0114] Furthermore, in this embodiment, the affordance information includes information indicating actions that can be taken with respect to an object model, so it is possible to prevent the generation of images as training data in which a human model performs an action with respect to an object model that cannot be predicted from the object model. Therefore, according to this embodiment, it is possible to efficiently generate images as training data in which interactions between a human model and an object model are occurring.

[0115] Furthermore, in this embodiment, the affordance information includes information indicating the parts of the object model where possible actions can be performed on the object model, so it is possible to prevent the generation of images as training data in which a human model performs an action on a part of the object model that cannot be predicted from the object model. Therefore, according to this embodiment, it is possible to efficiently generate images as training data of scenes in which interactions between a human model and an object model are occurring.

[0116] Furthermore, in this embodiment, the affordance information includes information indicating the parts of the human model where possible actions can be performed on the object model, so it is possible to prevent the generation of images as training data in which a human model performs an action on an object model at a part of the human model that cannot be predicted from the object model. Therefore, according to this embodiment, it is possible to efficiently generate images as training data of scenes in which interactions between a human model and an object model are occurring.

[0117] Furthermore, in this embodiment, the affordance information includes information indicating the rotatable angles of the object model with respect to each axis direction, which makes it possible to prevent the generation of images as training data in which a human model performs an action on an object model in an unnatural orientation. Therefore, according to this embodiment, images of scenes in which an interaction between a human model and an object model is occurring can be efficiently generated as training data.

[0118] Furthermore, if a recognition model is generated by learning the training data generated by the method of this embodiment, highly accurate recognition can be performed in scenes where a person is interacting with an object. Examples of recognition models include, but are not limited to, recognition models used to search for people interacting with an object and recognition models used to analyze the behavior of people interacting with an object. For example, if images of typical poses when a person model holds an object model are generated as training data and the generated training data is learned to generate a recognition model, highly accurate recognition can be performed in scenes where a person corresponding to the person model holds an object corresponding to the object model.

[0119] When generating a recognition model by learning the training data generated by the method of this embodiment, it is preferable to add affordance information as annotations and treat the training data as teacher data, which makes it possible to generate a training model capable of performing highly accurate recognition.

[0120] By using affordance information, the positional relationship between a person model and an object model can be learned for each part, which allows for detailed learning of the relationship between the object model's position and the person model, improving recognition accuracy in situations where interactions between people and objects occur. For example, it can be used to learn in advance where an object model is placed on a person model.

[0121] Furthermore, by using affordance information, it is possible to grasp the features of an object model on a part-by-part basis, making it possible to learn the features such as the shape of each part, thereby improving recognition accuracy in situations where a person interacts with an object. For example, it is possible to extract the unique features of an object model and learn how they relate to the hands and body of a human model.

[0122] Furthermore, by using affordance information and context, the person model can learn how and in what situations the object model is used, thereby improving recognition accuracy in situations where there is interaction between a person and an object. Context can be understood using background images, etc. This makes it possible to understand and recognize behaviors that are appropriate to the environment, such as using an umbrella in the rain or using a suitcase at the airport.

[0123] Furthermore, by using affordance information, it is possible to determine which part of the human model is interacting with the object model. Therefore, by using the joint points of the human model's skeleton to learn which joint points the object model is close to, it is possible to accurately recognize patterns such as how the object model is being held, thereby improving recognition accuracy in situations where interaction between a human and an object is occurring.

[0124] Furthermore, by using affordance information and time information, it is possible to learn a series of actions of a person model relative to an object model, which can further improve recognition accuracy, especially in scenes in videos where interactions between people and objects are occurring.

[0125] (Modification) In the above embodiment, an example has been described in which the user manually specifies affordance information, but the specification of affordance information may be automated by estimating the affordance information.

[0126] An example of a technique for estimating affordance information is the technique described in the document at https: / / ieeexplore.ieee.org / stamp / stamp.jsp?tp=&arnumber=9459755. Figure 8 of this document proposes a method for estimating actions that can be assumed to be taken on an object captured in a real-life photo, and the parts of the object where those actions will be performed.

[0127] For example, suppose that an object model 211 is selected in a pull-down list 511 on the setting screen 501 shown in FIG. 10 , and the object model 211 is displayed in a display area 513. In this case, the affordance information setting unit 125 may use the above-described estimation technology to estimate that the action that can be taken on the object model 211 is "grab" and that the handle portion of the object model 211 is the object part. In this case, the affordance information setting unit 125 may automatically set the object part as a handle and the person part / action as both hands / grab, etc., as affordance information. Note that, as in the above embodiment, the X-axis rotation, Z-axis rotation, and Y-axis rotation are manually performed by the user. Furthermore, instead of completely automating the setting of affordance information, the user may be prompted to confirm whether or not the setting has been made. Furthermore, when object parts on which actions are to be performed on multiple objects are estimated using the technology for estimating affordance information, the multiple object parts may be presented as candidate parts, and the user may select one or more object parts from the candidate parts. Furthermore, when a part of an object on which an action is to be performed is estimated using a technology for estimating affordance information, area information relating to the part of the object can be included in the affordance information. For example, when a handle is estimated, the area information may include a user operation to change the entire area of ​​the handle to a partial area such as the center part.

[0128] In a modified example, the screen generation unit 123 may visualize and display the estimated object part in the display area 513 or highlight the estimated object part to inform the user of the estimated object part.

[0129] (Program) The program executed by the training data generation device 10 in the above embodiment and each of the above modified examples is provided by being stored in a computer-readable storage medium such as a CD-ROM, CD-R, memory card, DVD, or flexible disk (FD) in the form of a file in an installable or executable format.

[0130] The programs executed by the training data generation device 10 of the above embodiment and each of the above modifications may be stored on a computer connected to a network such as the Internet and provided by being downloaded via the network. The programs executed by the training data generation device 10 of the above embodiment and each of the above modifications may be provided or distributed via a network such as the Internet. The programs executed by the training data generation device 10 of the above embodiment and each of the above modifications may be provided by being pre-installed in a ROM or the like.

[0131] The programs executed by the training data generation device 10 of the above embodiment and each of the above modifications have a modular configuration for implementing the above-mentioned units on a computer. In terms of actual hardware, for example, the CPU reads the training program from the HDD onto the RAM and executes it, thereby implementing the above-mentioned units on the computer.

[0132] As described above, according to the above embodiment and each of the above modifications, images useful for training a recognition model can be generated with high accuracy as training data.

[0133] The above-described embodiment and each of the modifications merely illustrate examples of implementations of the present disclosure, and the technical scope of the present disclosure should not be construed as being limited by these. Therefore, the present disclosure can be implemented in various forms without departing from the spirit or main features thereof. For example, the above-described embodiment and each of the modifications may be appropriately combined in their respective constituent units. Furthermore, for example, some components may be deleted from all components in the above-described embodiment and each of the modifications.

[0134] The present disclosure includes the following aspects.

[0135] (1) A training data generation device comprising: a memory that stores a human model, an object model, and affordance information that defines actions that can be taken on the object model; and a processor that generates, as training data, images in which the human model and the object model are arranged so that the human model performs the action defined in the affordance information on the object model.

[0136] In the above configuration (1), by using affordance information, images in which a person model and an object model are arranged so that the person model performs an action defined in the affordance information on an object model are automatically generated as training data. In other words, in the above configuration (1), there is no need for a user to set the position and posture of the person model and the position and posture of the object model in detail so that the person model performs an action defined in the affordance information on the object model. Furthermore, in the above configuration (1), because the affordance information defines actions that can be taken on the object model, it is possible to prevent the generation of images as training data in which a person model performs an action on an object model that cannot be predicted from the object model. Therefore, according to the above configuration (1), images of scenes in which interactions between a person model and an object model are occurring can be efficiently generated as training data.

[0137] (2) The training data generation device according to (1), wherein the affordance information further defines a part of the object model on which the action is performed, and the processor generates, as the training data, an image in which the human model and the object model are arranged so that the human model performs the action on the part of the object model.

[0138] In the above configuration (2), the affordance information includes parts of the object model where possible actions can be performed on the object model, so it is possible to prevent the generation of images as training data in which a human model performs an action on a part of the object model that cannot be predicted from the object model. Therefore, according to the above configuration (2), it is possible to efficiently generate images as training data of scenes in which interactions between a human model and an object model are occurring.

[0139] (3) The training data generation device according to (2), wherein the action is grabbing with the hand of the human model, and the processor generates, as the training data, an image in which the human model and the object model are arranged so that the human model is grabbing the part of the object model.

[0140] In the above configuration (3), the affordance information includes parts of the human model where possible actions can be performed on the object model, so it is possible to prevent the generation of images as training data in which the human model performs an action on the object model using parts of the human model that cannot be predicted from the object model. Therefore, the above configuration (3) makes it possible to efficiently generate images of scenes in which interactions between the human model and the object model are occurring as training data.

[0141] (4) The training data generation device according to (1), wherein the affordance information further defines a posture that the object model can take, and the processor generates, as the training data, an image in which the human model and the object model are arranged so that the object model takes the posture when the human model performs the action on the object model.

[0142] In the configuration (4) above, the affordance information includes the posture of the object model, so it is possible to prevent the generation of images as training data in which a human model performs an action on an object model in an unnatural orientation. Therefore, the configuration (4) above makes it possible to efficiently generate images as training data in which interactions between a human model and an object model are occurring.

[0143] (5) The training data generation device according to (1), wherein the affordance information further defines an allowable range of poses that the object model can take, and the processor generates, as the training data, an image in which the human model and the object model are arranged so that the object model takes a pose that falls within the allowable range when the human model performs the action on the object model.

[0144] In the configuration (5) above, the affordance information includes an allowable range of poses that the object model can assume, so it is possible to prevent the generation of images as training data in which a human model performs an action on an object model in an unnatural orientation. Therefore, the configuration (5) above makes it possible to efficiently generate images as training data in which interactions between a human model and an object model are occurring.

[0145] (6) The training data generation device according to (1), wherein the processor performs at least one of adding, editing, and deleting the affordance information.

[0146] (7) A training data generation device comprising: a processor that displays an editing screen for affordance information of an object model on a display unit; accepts instructions defining actions that can be taken on the object model and sets the actions as the affordance information; accepts instructions to generate training data; and outputs, as training data, an image in which the human model and the object model are arranged so that the human model performs the action defined in the affordance information on the object model.

[0147] In the above configuration (7), by using affordance information, images in which a person model and an object model are arranged so that the person model performs an action defined in the affordance information on an object model are automatically generated as training data. In other words, in the above configuration (7), there is no need for a user to set the position and posture of the person model and the position and posture of the object model in detail so that the person model performs an action defined in the affordance information on an object model. Furthermore, in the above configuration (7), because the affordance information defines actions that can be taken on an object model, it is possible to prevent the generation of images as training data in which a person model performs an action on an object model that cannot be predicted from the object model. Therefore, according to the above configuration (7), images of scenes in which interactions between a person model and an object model are occurring can be efficiently generated as training data.

[0148] (8) The training data generation device according to (7), wherein the processor receives an instruction defining a part of the object model on which the action is to be performed, further sets the instruction as the affordance information, and outputs, as the training data, an image in which the human model and the object model are arranged so that the human model performs the action on the part of the object model.

[0149] In the above configuration (8), the affordance information includes parts of the object model where possible actions can be performed on the object model, so it is possible to prevent the generation of images as training data in which a human model performs an action on a part of the object model that cannot be predicted from the object model. Therefore, according to the above configuration (8), it is possible to efficiently generate images as training data of scenes in which interactions between a human model and an object model are occurring.

[0150] (9) The training data generation device according to (8), wherein the processor further causes a screen of the object model to be displayed on the display unit, and the processor receives an instruction to specify the part on the screen of the object model as an instruction to define the part.

[0151] In the above configuration (9), the part of the object model on which the action that can be taken on the object model is to be performed can be specified on the screen of the object model, so that the part can be specified easily.

[0152] (10) The training data generation device according to (8), wherein the action is grabbing with a hand of the human model, and the processor outputs, as the training data, an image in which the human model and the object model are arranged so that the human model is grabbing the part of the object model.

[0153] In the above configuration (10), the affordance information includes parts of the human model where possible actions can be performed on the object model, so it is possible to prevent the generation of images as training data in which the human model performs actions on the object model using parts of the human model that cannot be predicted from the object model. Therefore, according to the above configuration (10), it is possible to efficiently generate images of scenes in which interactions between the human model and the object model are occurring as training data.

[0154] (11) The training data generation device according to (7), wherein the processor receives an instruction defining a posture that the object model can take, further sets the instruction as the affordance information, and outputs, as the training data, an image in which the human model and the object model are arranged so that the object model takes the posture when the human model performs the action on the object model.

[0155] In the configuration (11) above, the affordance information includes the posture of the object model, so it is possible to prevent the generation of images as training data in which a human model performs an action on an object model in an unnatural orientation. Therefore, the configuration (11) above makes it possible to efficiently generate images as training data in which an interaction between a human model and an object model is occurring.

[0156] (12) The training data generation device according to (7), wherein the processor receives an instruction defining an allowable range of poses that the object model can take, further sets the instruction as the affordance information, and outputs, as the training data, an image in which the human model and the object model are arranged so that the object model takes a pose that falls within the allowable range when the human model performs the action on the object model.

[0157] In the configuration of (12) above, the affordance information includes an allowable range of poses that the object model can assume, so it is possible to prevent the generation of images as training data in which a human model performs an action on an object model in an unnatural orientation. Therefore, the configuration of (12) above makes it possible to efficiently generate images as training data in which interactions between a human model and an object model are occurring.

[0158] (13) The training data generation device according to (7), wherein the processor receives an instruction to add, edit, or delete the affordance information, and performs processing on the affordance information in accordance with the received instruction.

[0159] (14) A training data generation method including: a step of a processor acquiring, from a memory, a human model, an object model, and affordance information that defines an action that can be taken with respect to the object model; and a step of the processor generating, as training data, an image in which the human model and the object model are arranged so that the human model performs the action defined in the affordance information with respect to the object model.

[0160] In the above configuration (14), by using affordance information, images in which a person model and an object model are arranged so that the person model performs an action defined in the affordance information on an object model are automatically generated as training data. In other words, in the above configuration (14), there is no need for a user to set the position and posture of the person model and the position and posture of the object model in detail so that the person model performs an action defined in the affordance information on an object model. Furthermore, in the above configuration (14), because the affordance information defines actions that can be taken on an object model, it is possible to prevent the generation of images as training data in which a person model performs an action on an object model that cannot be predicted from the object model. Therefore, according to the above configuration (14), images of scenes in which an interaction between a person model and an object model is occurring can be efficiently generated as training data.

[0161] (15) A learning data generation method including: a step in which a processor causes a display unit to display an editing screen for affordance information of an object model; a step in which the processor receives an instruction defining an action that can be taken with respect to the object model and registers the action as the affordance information; and a step in which the processor receives an instruction to generate learning data and outputs, as learning data, an image in which the human model and the object model are arranged so that the human model performs the action defined in the affordance information with respect to the object model.

[0162] In the above configuration (15), by using affordance information, images in which a person model and an object model are arranged so that the person model performs an action defined in the affordance information on an object model are automatically generated as training data. In other words, in the above configuration (15), there is no need for a user to set the position and posture of the person model and the position and posture of the object model in detail so that the person model performs an action defined in the affordance information on the object model. Furthermore, in the above configuration (15), because the affordance information defines actions that can be taken on the object model, it is possible to prevent the generation of images as training data in which a person model performs an action on an object model that cannot be predicted from the object model. Therefore, according to the above configuration (15), images of scenes in which an interaction between a person model and an object model is occurring can be efficiently generated as training data.

[0163] (16) A program causing a computer to execute the steps of: acquiring from a memory a human model, an object model, and affordance information that defines an action that can be taken on the object model; and generating, as learning data, an image in which the human model and the object model are arranged so that the human model performs the action defined in the affordance information on the object model.

[0164] In the above configuration (16), by using affordance information, images in which a person model and an object model are arranged so that the person model performs an action defined in the affordance information on an object model are automatically generated as training data. In other words, in the above configuration (16), there is no need for a user to set the position and posture of the person model and the position and posture of the object model in detail so that the person model performs an action defined in the affordance information on an object model. Furthermore, in the above configuration (16), because the affordance information defines actions that can be taken on an object model, it is possible to prevent the generation of images as training data in which a person model performs an action on an object model that cannot be predicted from the object model. Therefore, according to the above configuration (16), images of scenes in which an interaction between a person model and an object model is occurring can be efficiently generated as training data.

[0165] (17) A program for causing a computer to execute the steps of: displaying an editing screen for affordance information of an object model on a display unit; accepting an instruction to define an action that can be taken with respect to the object model and registering it as the affordance information; and accepting an instruction to generate learning data and outputting, as learning data, an image in which the human model and the object model are arranged so that the human model performs the action defined in the affordance information with respect to the object model.

[0166] In the above configuration (17), by using affordance information, images in which a person model and an object model are arranged so that the person model performs an action defined in the affordance information on an object model are automatically generated as training data. In other words, in the above configuration (17), there is no need for a user to set the position and posture of the person model and the position and posture of the object model in detail so that the person model performs an action defined in the affordance information on an object model. Furthermore, in the above configuration (17), because the affordance information defines actions that can be taken on an object model, it is possible to prevent the generation of images as training data in which a person model performs an action on an object model that cannot be predicted from the object model. Therefore, according to the above configuration (17), images of scenes in which an interaction between a person model and an object model is occurring can be efficiently generated as training data.

[0167] The present disclosure also includes the following aspects.

[0168] (18) The learning data generation device according to (1) or (7), wherein the processor generates the learning data by executing the following processes: a process of storing, as basic learning data, the posture of the object model relative to the person model in a normal use state of the object model; and a process of storing, as additional learning data, images including images in which the object model is rotated vertically or horizontally within a specified allowable angle range.

[0169] This application is based on Japanese Patent Application No. 2023-217934 filed on December 25, 2023, and Japanese Patent Application No. 2023-217947 filed on December 25, 2023, the contents of which are all incorporated by reference into this application.

[0170] REFERENCE SIGNS LIST 10 Learning data generation device 11 Control device 13 Main memory device 15 Auxiliary memory device 17 Display device 19 Input device 21 Communication device 23 Various buses 101 Operation unit 103 Operation reception unit 105 Display control unit 107 Display unit 109 Image acquisition unit 111 Object model storage unit 113 Object model acquisition unit 115 Person model storage unit 117 Person model acquisition unit 119 Background image storage unit 121 Background image acquisition unit 123 Screen generation unit 125 Affordance information setting unit 127 Image generation unit 129 Position and orientation determination unit 131 Learning data output unit 133 Learning data storage unit

Claims

1. A learning data generation device comprising: a memory that stores a person model, an object model, and affordance information that defines possible actions with respect to the object model; and a processor that generates, as learning data, an image in which the person model and the object model are arranged such that the person model performs the action defined by the affordance information with respect to the object model.

2. The affordance information further defines a part of the object model on which the action is to be performed, and the processor generates, as the learning data, an image in which the person model and the object model are arranged such that the person model performs the action on the part of the object model. The learning data generation device according to claim 1.

3. The action is to grasp with the hand of the person model, and the processor generates, as the learning data, an image in which the person model and the object model are arranged such that the person model grasps the part of the object model. The learning data generation device according to claim 2.

4. The affordance information further defines a posture that the object model can take, and the processor generates, as the learning data, an image in which the person model and the object model are arranged such that the object model takes the posture when the person model performs the action with respect to the object model. The learning data generation device according to claim 1.

5. The affordance information further defines an allowable range of postures that the object model can take, and the processor generates, as the learning data, an image in which the person model and the object model are arranged such that the object model takes a posture within the allowable range when the person model performs the action with respect to the object model. The learning data generation device according to claim 1.

6. The processor performs at least one of addition, editing, and deletion of the affordance information. The learning data generation device according to claim 1.

7. A learning data generation device comprising a processor that causes a display unit to display a screen of affordance information of an object model, receives an instruction to define an action that can be taken with respect to the object model, sets the instruction as the affordance information, receives an instruction to generate learning data, and outputs, as learning data, an image in which a person model and the object model are arranged such that the person model performs the action defined by the affordance information with respect to the object model.

8. The learning data generation device according to claim 7, wherein the processor receives an instruction to define a part of the object model on which the action is to be performed, further sets the instruction as the affordance information, and outputs, as the learning data, an image in which the person model and the object model are arranged such that the person model performs the action with respect to the part of the object model.

9. The learning data generation device according to claim 8, wherein the processor further causes the display unit to display a screen of the object model, and the processor receives, as the instruction to define the part, an instruction to specify the part on the screen of the object model.

10. The action is to grasp with a hand of the person model, and the learning data generation device according to claim 8, wherein the processor outputs, as the learning data, an image in which the person model and the object model are arranged such that the person model grasps the part of the object model.

11. The learning data generation device according to claim 7, wherein the processor receives an instruction to define a posture that the object model can take, further sets the instruction as the affordance information, and outputs, as the learning data, an image in which the person model and the object model are arranged such that the object model takes the posture when the person model performs the action with respect to the object model.

12. The learning data generation device according to claim 7, wherein the processor receives an instruction to define a tolerance range of a posture that the object model can take, further sets the instruction as the affordance information, and outputs, as the learning data, an image in which the person model and the object model are arranged such that the object model takes a posture within the tolerance range when the person model performs the action with respect to the object model.

13. The processor according to claim 7, which receives an instruction to add, edit, or delete the affordance information and performs processing according to the received instruction on the affordance information.

14. A learning data generation method including: a step in which a processor acquires from a memory a person model, an object model, and affordance information defining an action that can be taken with respect to the object model; and a step in which the processor generates, as learning data, an image in which the person model and the object model are arranged such that the person model performs the action defined by the affordance information with respect to the object model.

15. A learning data generation method including: a step in which a processor causes a display unit to display an editing screen for affordance information of an object model; a step in which the processor receives an instruction defining an action that can be taken with respect to the object model and registers the instruction as the affordance information; and a step in which the processor receives an instruction to generate learning data and outputs, as learning data, an image in which a person model and the object model are arranged such that the person model performs the action defined by the affordance information with respect to the object model.

16. A program for causing a computer to execute: a step in which a person model, an object model, and affordance information defining an action that can be taken with respect to the object model are acquired from a memory; and a step in which an image in which the person model and the object model are arranged such that the person model performs the action defined by the affordance information with respect to the object model is generated as learning data.

17. A program for causing a computer to execute: a step in which an editing screen for affordance information of an object model is displayed on a display unit; a step in which an instruction defining an action that can be taken with respect to the object model is received and registered as the affordance information; and a step in which an instruction to generate learning data is received and an image in which a person model and the object model are arranged such that the person model performs the action defined by the affordance information with respect to the object model is output as learning data.

Citation Information

Patent Citations

  • Scene recognizing device

    JP2000293685A

  • Learning device, learning method and learning program

    JP2023028298A

  • Object affordance detection method and apparatus

    WO2022188493A1