Learning data generation apparatus, learning data generation method, and program

The training data generation device and method create videos with annotated interactions between human and object models, addressing the limitation of existing recognition models by enabling the development of models that can recognize such interactions effectively and efficiently.

JP2025152578APending Publication Date: 2025-10-10PANASONIC INTELLECTUAL PROPERTY MANAGEMENT CO LTD
View PDF 1 Cites 0 Cited by

Patent Information

Application Number
JP2024054531
Authority / Receiving Office
JP · JP
Patent Type
Applications
Current Assignee / Owner
Filing Date
2024-03-28
Publication Date
2025-10-10

AI Technical Summary

Technical Problem

Existing recognition models are unable to recognize interactions between people and objects, as they are primarily designed to identify individuals such as persons or vehicles and lack the capability to understand interactions between them.

Method used

A training data generation device and method that generates videos of human and object models in a virtual space, adds annotation information to these interactions, and produces training data for recognition models to identify such interactions.

Benefits of technology

Enables the generation of training data that facilitates the development of recognition models capable of understanding and recognizing interactions between people and objects, reducing annotation costs through automation.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure 2025152578000001_ABST
    Figure 2025152578000001_ABST
Patent Text Reader

Abstract

To provide a learning data generation apparatus, a learning data generation method, and a program configured to generate learning data useful for learning of a recognition model for recognizing interaction between a person and an object.SOLUTION: A learning data generation apparatus includes: a video generation unit which generates a video including a person model and a first object model; an annotation information generation unit which generates annotation information in accordance with interaction between the person model and the first object model; and a learning data generation unit which generates learning data by adding the annotation information to the video.SELECTED DRAWING: Figure 2
Need to check novelty before this filing date? Find Prior Art

Description

[Technical Field]

[0001] The present disclosure relates to a training data generation device, a training data generation method, and a program. [Background technology]

[0002] A known method of data augmentation for expanding (increasing) learning data for machine learning is to generate learning data using computer graphics (CG).

[0003] For example, the technology disclosed in Patent Document 1 generates an image that narrows down to only specific CG models, such as a person model or a vehicle model, from a scene data image that includes multiple CG models. The technology disclosed in Patent Document 1 generates training data by adding information about the specific CG model to the generated image as annotation information. The annotation information may include a rectangle surrounding the specific CG model, or the type and movement of the specific CG model. The technology disclosed in Patent Document 1 learns from the generated training data to generate a recognition model that recognizes whether the recognition target is a person or vehicle corresponding to the specific CG model. [Prior art documents] [Patent documents]

[0004] [Patent Document 1] Japanese Patent Application Publication No. 2019-23858 Summary of the Invention [Problem to be solved by the invention]

[0005] However, the technology disclosed in the above-mentioned Patent Document 1 is only intended to recognize whether the recognition target is a person or a vehicle itself, and is not intended to recognize interactions that occur between a person and an object, such as a person handling the object.

[0006] For this reason, the recognition model disclosed in Patent Document 1 cannot recognize interactions between people and objects, and even if the training data disclosed in Patent Document 1 is learned, a recognition model capable of recognizing such interactions cannot be generated.

[0007] The present disclosure has been made in consideration of the above circumstances, and aims to provide a training data generation device, a training data generation method, and a program capable of generating training data useful for training a recognition model for recognizing interactions between a person and an object. [Means for solving the problem]

[0008] One aspect of the training data generation device of the present disclosure includes a video generation unit that generates a video including a human model and a first object model, an annotation information generation unit that generates annotation information according to an interaction between the human model and the first object model, and a training data generation unit that generates training data by adding the annotation information to the video.

[0009] One aspect of the training data generation method of the present disclosure includes a video generation step in which a processor generates a video including a human model and a first object model; an annotation information generation step in which a processor generates annotation information according to an interaction between the human model and the first object model; and a training data generation step in which a processor adds the annotation information to the video to generate training data.

[0010] One aspect of the program of the present disclosure is for causing a computer to execute a video generation step of generating a video including a human model and a first object model, an annotation information generation step of generating annotation information according to an interaction between the human model and the first object model, and a training data generation step of generating training data by adding the annotation information to the video. [Effects of the Invention]

[0011] According to the present disclosure, it is possible to provide a training data generation device, a training data generation method, and a program capable of generating training data useful for training a recognition model for recognizing interactions between a person and an object. [Brief explanation of the drawings]

[0012] [Figure 1] FIG. 1 is a block diagram showing an example of the hardware configuration of a training data generating device according to this embodiment. [Figure 2] FIG. 2 is a block diagram showing an example of the functional configuration of the training data generating device according to this embodiment. [Figure 3] FIG. 3 is a diagram showing an example of a person model registration screen according to this embodiment. [Figure 4] FIG. 4 is a diagram showing an example of an object model registration screen according to this embodiment. [Figure 5] FIG. 5 is a diagram showing an example of an object model registration screen according to this embodiment. [Figure 6] FIG. 6 is a diagram showing an example of a motion registration screen according to this embodiment. [Figure 7] FIG. 7 is a diagram showing an example of an operation class registration screen according to this embodiment. [Figure 8] FIG. 8 is a diagram showing an example of an action class registration screen according to this embodiment. [Figure 9] FIG. 9 is a diagram showing an example of an operation class registration screen according to this embodiment. [Figure 10] FIG. 10 is a diagram showing an example of a video generation screen according to this embodiment. [Figure 11] FIG. 11 is a diagram illustrating an example of the interaction detection method according to this embodiment. [Figure 12] FIG. 12 is a diagram showing an example of annotation information according to this embodiment. [Figure 13] FIG. 13 is a diagram showing an example of a moving image generated by the moving image generating unit of this embodiment. [Figure 14] FIG. 14 is a diagram showing an example of a confirmation screen displayed by the display control unit of this embodiment. [Figure 15] FIG. 15 is a diagram showing an example of a display setting screen according to this embodiment. [Figure 16] FIG. 16 is a diagram showing an example of a confirmation screen after changing the display settings of the timeline according to this embodiment. [Figure 17] FIG. 17 is a flowchart showing an example of the training data generation process performed by the training data generation device of this embodiment. [Figure 18] FIG. 18 is a flowchart showing an example of the process of checking and correcting annotation information shown in step S125 of the flowchart in FIG. DETAILED DESCRIPTION OF THE INVENTION

[0013] Hereinafter, an embodiment of the present disclosure (hereinafter simply referred to as "the present embodiment") will be described in detail with reference to the drawings. Note that the present disclosure is not limited to the following embodiment. Furthermore, the following embodiment and modified examples can be combined as appropriate.

[0014] The training data generation device of this embodiment reproduces interactions between people and objects that may occur in a real environment in a virtual space using person models corresponding to the people and object models corresponding to the objects. In this case, the training data generation device of this embodiment generates videos capturing the person models and object models from a predetermined viewpoint in the virtual space, and detects interactions between the person models and object models in the virtual space. The training data generation device of this embodiment generates training data by adding interaction detection results as annotation information to the generated videos.

[0015] In this manner, in this embodiment, training data is generated by adding annotation information indicating the content of the interaction to a video containing an interaction between a person model and an object model. Therefore, according to this embodiment, training data useful for training a recognition model for recognizing interactions between a person and an object can be generated. Furthermore, according to this embodiment, the generation and addition of annotation information can be automated, allowing training data to be generated efficiently.

[0016] Furthermore, the training data generation device of this embodiment defines an action class in association with an object model, and detects the action content indicated by the action class as an interaction when a human model comes into contact with an object model. In this way, this embodiment can automate the detection of interactions, thereby reducing annotation costs even in situations where the annotation costs would be high if annotation information indicating the presence or absence of an interaction on a frame-by-frame basis were to be generated.

[0017] The following description will be given taking as an example a case where the training data generation device of this embodiment uses CG (Computer Graphics) to generate training data used to train a recognition model that recognizes worker tasks (interactions between the worker and the objects used in the tasks) that occur during logistics. However, the training data of this embodiment and the recognition model generated by training the training data are not limited to this. The recognition model of this embodiment may be any trained model that recognizes interactions between a person and an object. Similarly, the training data of this embodiment may be any data that can be used to train a recognition model that recognizes interactions between a person and an object.

[0018] Fig. 1 is a block diagram showing an example of the hardware configuration of a training data generation device 10 of this embodiment. As shown in Fig. 1, the training data generation device 10 includes a control device 11, a main memory device 13, an auxiliary memory device 15, a display device 17, an input device 19, a communication device 21, and various buses 23. The control device 11, the main memory device 13, the auxiliary memory device 15, the display device 17, the input device 19, and the communication device 21 are connected via the various buses 23. As such, the training data generation device 10 of this embodiment has a general hardware configuration using a normal computer.

[0019] The control device 11 controls the overall operation of the training data generation device 10. The control device 11 may be, for example, at least one of a CPU (Central Processing Unit) and a GPU (Graphics Processing Unit), but is not limited to these. There may be any number of CPUs or GPUs as long as they are one or more, and they may be single-core or multi-core.

[0020] Examples of the main storage device 13 include, but are not limited to, a ROM (Read Only Memory) and a RAM (Random Access Memory). The ROM stores various programs, such as a program for controlling the training data generation device 10 and a program for generating training data according to this embodiment. The RAM is used as a working area when the control device 11 performs various controls based on the programs stored in the ROM.

[0021] The auxiliary storage device 15 stores the various programs described above and data for generating training data according to this embodiment. The various programs described above may be stored in at least one of the main storage device 13 and the auxiliary storage device 15. Examples of the auxiliary storage device 15 include, but are not limited to, existing storage devices capable of magnetic, electrical, or optical storage, such as a hard disk drive (HDD), a solid state drive (SSD), and a digital versatile disc (DVD). The auxiliary storage device 15 may be built into the training data generation device 10 or may be externally connected to the training data generation device 10 via an interface such as a universal serial bus (USB). The auxiliary storage device 15 may also be a network-attached storage (NAS) connected via a network such as a local area network (LAN) or a wide area network (WAN).

[0022] The display device 17 displays various screens used when the training data generation device 10 generates training data, and serves as a user interface with the user (operator). Examples of the display device 17 include, but are not limited to, various displays such as a liquid crystal display, an organic electroluminescence (EL) display, and a touch panel display. The display device 17 may be a built-in display built into the training data generation device 10, or an external display connected to the training data generation device 10 via a display interface such as HDMI (registered trademark).

[0023] The input device 19 is used for various inputs used when the training data generation device 10 generates training data, and serves as a user interface with a user (operator). Examples of the input device 19 include, but are not limited to, a keyboard, a mouse, and a touch panel. The input device 19 may be built into the training data generation device 10 or may be externally attached to the training data generation device 10 via an interface such as a USB.

[0024] Examples of the communication device 21 include, but are not limited to, a communication device for a wired LAN and a wireless communication device for a wireless LAN. The communication device 21 may be used to externally acquire the program or data for generating the training data of this embodiment, or may be used to externally output the training data generated by the training data generation device 10.

[0025] In addition to the above configuration, the training data generation device 10 may further include hardwired circuits such as an IC (Integrated Circuit), an ASIC (Application Specific Integrated Circuit), and an FPGA (Field-Programmable Gate Array) specific to the training data generation device 10 in order to realize the training data generation function.

[0026] 2 is a block diagram showing an example of the functional configuration of the training data generation device 10 according to this embodiment. As shown in FIG. 2, the training data generation device 10 includes a model storage unit 101, a motion storage unit 103, a model registration unit 111, a motion registration unit 113, an action class registration unit 115, an environment setting unit 117, a virtual space control unit 121, a confirmation unit 131, and a training data generation unit 141. The virtual space control unit 121 includes an interaction detection unit 123, a video generation unit 125, a position detection unit 127, and an annotation information generation unit 129. The confirmation unit 131 includes a display control unit 133 and an annotation information correction unit 135.

[0027] The model storage unit 101 and the motion storage unit 103 can be realized by, for example, at least one of the main storage unit 13 and the auxiliary storage unit 15 described with reference to FIG.

[0028] The model registration unit 111, motion registration unit 113, action class registration unit 115, environment setting unit 117, virtual space control unit 121, interaction detection unit 123, video generation unit 125, position detection unit 127, annotation information generation unit 129, confirmation unit 131, display control unit 133, annotation information correction unit 135, and learning data generation unit 141 can be realized, for example, by the control device 11 and main memory device 13 described in Figure 1.

[0029] For example, the control device 11 reads out a program for generating learning data according to this embodiment, which is stored in the main storage device 13 (ROM) or the auxiliary storage device 15, or which is externally acquired from the communication device 21 via a network, and loads the program into the main storage device 13 (RAM). The control device 11 executes various processes in accordance with the loaded program, thereby realizing each of the above-described functional units. Here, the description has been given taking an example in which each of the above-described functional units is realized as software, but at least a part of each of the above-described functional units may also be realized as hardware. In this case, the functional units realized as hardware may be realized, for example, by the above-described hardwired circuit. Furthermore, any of the above-described functional units may be realized by a combination of software and hardware.

[0030] The model storage unit 101 stores character models and object models to be made to appear in the virtual space. In this embodiment, the virtual space is described as a three-dimensional CG space, but the present invention is not limited to this. An example of a character model is three-dimensional model data in which a character is reproduced using CG, but the present invention is not limited to this and may also be two-dimensional model data. An example of an object model is three-dimensional model data in which an object handled by a character is reproduced using CG, but the present invention is not limited to this and may also be two-dimensional model data.

[0031] Both the person model and the object model may be model data created based on a real person or object, or may be model data created based on a fictional person or object. Furthermore, both the person model and the object model may be model data created in a CG environment by a developer or the like of the training data generation device 10 (hereinafter simply referred to as "developer"), or may be model data prepared in advance on a platform that builds a CG environment.

[0032] The motion storage unit 103 stores motion data for operating the human model or object model stored in the model storage unit 101 in a virtual space. The motion data may be data created in a CG environment by the developer of the training data generation device 10, or may be data prepared in advance on a platform that builds the CG environment. The motion data may also be data generated by capturing the movements of actual people or objects using motion capture technology.

[0033] The model registration unit 111 registers person models and object models to be made to appear in the virtual space. Specifically, the model registration unit 111 registers person models and object models selected by a developer from among the person models and object models stored in the model storage unit 101 in the virtual space control unit 121.

[0034] For example, when a developer performs an operation input requesting the display of a human model registration screen using the input device 19, the model registration unit 111 displays the human model registration screen on the display device 17. The human model registration screen is a UI (User Interface) screen for registering a human model.

[0035] Fig. 3 is a diagram showing an example of a character model registration screen 201 according to this embodiment. As shown in Fig. 3, the character model registration screen 201 includes a pull-down list 211, a display area 213, and a register button 215. The pull-down list 211 is an operation button for selecting a character model. The display area 213 is a display area in which the character model selected in the pull-down list 211 is displayed. The register button 215 is an operation button for registering the character model selected in the pull-down list 211.

[0036] 3, in order to select person model 221, which is a CG model of a worker, the developer uses input device 19 to perform an operational input to select a file with the file name "PersonA.fbx" on pull-down list 211. As a result, model registration unit 111 retrieves person model 221 with the file name "PersonA.fbx" from model storage unit 101 and displays it in display area 213. In this situation, when the developer uses input device 19 to perform an operational input to press registration button 215, model registration unit 111 registers person model 221 in virtual space control unit 121 as a person model to appear in the virtual space.

[0037] In this embodiment, the case where the number of character models appearing in the virtual space (to be registered) is one will be described as an example, but the present invention is not limited to this and there may be more than one. When multiple character models appear in the virtual space, for example, the character model registration process described in the character model registration screen 201 shown in Fig. 3 may be performed for each character model appearing in the virtual space (to be registered).

[0038] Furthermore, for example, when the developer performs an operation input requesting the display of an object model registration screen using the input device 19, the model registration unit 111 displays the object model registration screen on the display device 17. The object model registration screen is a UI screen for registering an object model.

[0039] 4 and 5 are diagrams showing an example of an object model registration screen 251 of this embodiment. As shown in FIGS. 4 and 5, the object model registration screen 251 includes a pull-down list 261, a display area 263, and a register button 265. The pull-down list 261 is an operation button for selecting an object model. The display area 263 is a display area in which the object model selected in the pull-down list 261 is displayed. The register button 265 is an operation button for registering the object model selected in the pull-down list 261.

[0040] 4, in order to select object model 271, which is a CG model of scissors, the developer uses the input device 19 to perform an operational input to select a file with the file name "scissors.fbx" on the pull-down list 261. As a result, the model registration unit 111 retrieves the object model 271 with the file name "scissors.fbx" from the model storage unit 101 and displays it in the display area 263. In this situation, when the developer uses the input device 19 to perform an operational input to press the registration button 265, the model registration unit 111 registers the object model 271 in the virtual space control unit 121 as an object model to appear in the virtual space.

[0041] 5, in order to select object model 273, which is a CG model of paper, the developer selects a file with the file name "paper.fbx" on the pull-down list 261. As a result, the model registration unit 111 retrieves the object model 273 with the file name "paper.fbx" from the model storage unit 101 and displays it in the display area 263. In this situation, when the developer presses the registration button 265, the model registration unit 111 registers the object model 273 in the virtual space control unit 121 as an object model to appear in the virtual space.

[0042] In this embodiment, an example will be described in which there are a plurality of object models (to be registered) appearing in the virtual space, but the present invention is not limited to this and there may be only one object model.

[0043] The motion registration unit 113 registers motion data for moving each of the human model and object model registered by the model registration unit 111. Specifically, the motion registration unit 113 registers, in the virtual space control unit 121, motion data selected by the developer from the motion data stored in the motion storage unit 103 for moving each of the human model and object model registered by the model registration unit 111.

[0044] For example, when a developer performs an operation input using the input device 19 to request the display of a motion registration screen, the motion registration unit 113 displays the motion registration screen on the display device 17. The motion registration screen is a UI screen for registering motion data for each of the human model and object model registered by the model registration unit 111.

[0045] Fig. 6 is a diagram showing an example of a motion registration screen 301 of this embodiment. Fig. 6 shows the motion registration screen 301 in the case where a person model 221, an object model 271, and an object model 273 have been registered by the model registration unit 111 as person models and object models (to be registered) to appear in the virtual space. As shown in Fig. 6, the motion registration screen 301 includes a pull-down list 311, a register button 313, a pull-down list 321, a register button 323, a pull-down list 331, and a register button 333.

[0046] The pull-down list 311 is an operation button for selecting motion data for the person model 221 (PersonA). The register button 313 is an operation button for registering the motion data selected in the pull-down list 311. The pull-down list 321 is an operation button for selecting motion data for the object model 271 (scissors). The register button 323 is an operation button for registering the motion data selected in the pull-down list 321. The pull-down list 331 is an operation button for selecting motion data for the object model 273 (paper). The register button 333 is an operation button for registering the motion data selected in the pull-down list 331.

[0047] 6, the developer uses the input device 19 to perform an operation input to select a file with the file name "PersonA-cut.anim" on the pull-down list 311. In this situation, when the developer uses the input device 19 to perform an operation input to press the registration button 313, the motion registration unit 113 registers the motion data with the file name "PersonA-cut.anim" in the virtual space control unit 121 as motion data for moving the person model 221.

[0048] Similarly, a file with the file name "scissors-cut.anim" is selected by the developer on pull-down list 321. In this situation, when the developer presses register button 323, the motion registration unit 113 registers the motion data with the file name "scissors-cut.anim" in the virtual space control unit 121 as motion data for moving the object model 271. Similarly, a file with the file name "paper-cut.anim" is selected by the developer on pull-down list 331. In this situation, when the developer presses register button 333, the motion registration unit 113 registers the motion data with the file name "paper-cut.anim" in the virtual space control unit 121 as motion data for moving the object model 273.

[0049] The motion data with the file name "PersonA-cut.anim" is data for causing the person model 221 to perform an action of cutting the object model 273 (paper) using the object model 271 (scissors). The file name "scissors-cut.anim" is data for causing the object model 271 (scissors) to perform an action of cutting the object model 273 (paper) when handled by the person model 221. The file name "paper-cut.anim" is data for causing the object model 273 (paper) to perform an action of being cut by the object model 271 (scissors).

[0050] The action class registration unit 115 registers action classes defined in association with the object models registered by the model registration unit 111. An action class indicates the content of an interaction that occurs between an object model for which the action class is defined and a person model. An action class is an example of content information that defines the content that occurs in an object model when the person model comes into contact with the object model. Registering an action class and defining the action class in association with an object model is an example of associating content information with an object model. Note that an action class may indicate the content of an interaction that occurs between a person model and a single object model, or may indicate the content of an interaction that occurs between a person model and multiple object models.

[0051] For example, when a developer uses the input device 19 to perform an operation input requesting the display of an action class registration screen, the action class registration unit 115 displays the action class registration screen on the display device 17. The action class registration screen is a UI screen for registering an action class.

[0052] 7 to 9 are diagrams showing an example of an action class registration screen 401 of this embodiment. The action class registration screen 401 shown in Fig. 7 shows a case where an action class is defined and registered in association with a single object model, while the action class registration screen 301 shown in Fig. 8 and Fig. 9 shows a case where an action class is defined and registered in association with a plurality of object models.

[0053] As shown in FIG. 7, the action class registration screen 301 includes a pull-down list 411 for selecting the number of object models, a pull-down list 421 for selecting an object model, a display area 423, a text box 441, and a registration button 443.

[0054] The pull-down list 411 is an operation button for selecting the number (type) of object models for which an action class is defined. The pull-down list 421 is an operation button for selecting an object model for which an action class is defined. The display area 423 is a display area in which the object model selected in the pull-down list 421 is displayed. The text box 441 is an input field for inputting an action class. The register button 343 is an operation button for defining and registering the action class input in the text box 341 in association with the object model selected in the pull-down list 321.

[0055] In the example shown in Fig. 7, the developer uses the input device 19 to input an operation to set the number of object models for which an action class is defined to "1" on the pull-down list 411. Therefore, on the action class registration screen 401 shown in Fig. 7, the pull-down list and display area for the object models are a single set of a pull-down list 421 and a display area 423. Note that when there is one object model for which an action class is defined, it is assumed that this object model will be possessed and handled (directly handled) by the human model, and therefore on the action class registration screen 401 shown in Fig. 7, this object model is referred to as a "possessive item."

[0056] 7, it is assumed that an object model 271 registered by the model registration unit 111 is selected as an object model (possessive item) for which an action class is defined. For this purpose, the developer uses the input device 19 to perform an operation input to select a file with the file name "scissors.fbx" on the pull-down list 421. As a result, the action class registration unit 115 obtains the object model 271 with the file name "scissors.fbx" from the model registration unit 111 and displays it in the display area 423.

[0057] 7, the developer uses input device 19 to perform an operation input to input "scissors grip" into text box 441 as an action class defined in association with object model 271. In this situation, when the developer uses input device 19 to perform an operation input to press registration button 443, action class registration unit 115 defines the action class "scissors grip" in association with object model 271 and registers it in virtual space control unit 121.

[0058] Next, an example of defining and registering an action class in association with a plurality of object models will be described with reference to Fig. 8. As shown in Fig. 8, an action class registration screen 401 includes a pull-down list 411 for selecting the number of object models, pull-down lists 421 and 425 for selecting an object model, display areas 423 and 427, a text box 441, and a registration button 443.

[0059] In the example shown in FIG. 8, the developer uses the input device 19 to input an operation to set the number of object models for which an action class is defined to "2" on the pull-down list 411. Therefore, on the action class registration screen 401 shown in FIG. 8, the pull-down list and display area for the object models are divided into two sets: a pull-down list 421 and a display area 423, and a pull-down list 425 and a display area 427. Note that when there are multiple object models for which an action class is defined, the multiple object models are assumed to include the above-mentioned possessed items and object models that are handled (indirectly handled) by the person model via the possessed items and are affected by the possessed items. Therefore, on the action class registration screen 401 shown in FIG. 8, the latter object models are referred to as "acting items."

[0060] 8, as in FIG. 7, a file with the file name "scissors.fbx" is selected as a possessed item in pull-down list 421, and an object model 271 is displayed in display area 423. Furthermore, in action class registration screen 401 shown in FIG. 8, "2" is selected in pull-down list 411, and therefore, unlike action class registration screen 401 shown in FIG. 7, a pull-down list 425 and a display area 427 for action items are included. In the example shown in FIG. 8, the developer uses input device 19 to perform an operation input on pull-down list 425 to select a file with the file name "paper.fbx" as an object model (action item) for which an action class is defined. As a result, action class registration unit 115 obtains object model 273 with the file name "paper.fbx" from model registration unit 111 and displays it in display area 427.

[0061] 8, the developer uses the input device 19 to perform an operation input to input "cut" into the text box 341 as an action class defined in association with the object model 271 and the object model 273. In this situation, when the developer uses the input device 19 to perform an operation input to press the register button 443, the action class registration unit 115 defines the action class "cut" in association with the object model 271 and the object model 273, and registers it in the virtual space control unit 121.

[0062] Next, another example of registering an action class defined in association with a plurality of object models will be described with reference to Fig. 9. As shown in Fig. 9, an action class registration screen 401 includes a pull-down list 411 for selecting the number of object models, pull-down lists 421, 425, and 429 for selecting an object model, display areas 423, 427, and 431, a text box 441, and a registration button 443.

[0063] 9, an operation input to set the number of object models for which an action class is defined to "3" is performed on pull-down list 411. Therefore, on action class registration screen 401 shown in Fig. 9, the pull-down list and display area for object models are three sets: pull-down list 421 and display area 423, pull-down list 425 and display area 427, and pull-down list 429 and display area 431.

[0064] 9, a file with the file name "screwdriver.fbx" is selected on pull-down list 421 as a possessed item, and an object model 281 with the file name "screwdriver.fbx" is displayed in display area 423. A file with the file name "bolt.fbx" is selected on pull-down list 425 as an acting item, and an object model 283 with the file name "bolt.fbx" is displayed in display area 427. A file with the file name "PC.fbx" is selected on pull-down list 429 as another acting item that is affected by the action of object model 283, and an object model 285 with the file name "PC.fbx" is displayed in display area 431. It is assumed that object models 281, 283, and 285 are object models that have already been registered by model registration unit 111.

[0065] 9, an operation input is performed to input "assembly" into text box 341 as an action class defined in association with object models 281, 283, and 285. In this situation, when the developer performs an operation input to press registration button 443 using input device 19, action class registration unit 115 defines the action class "assembly" in association with object model 281, object model 283, and object model 285, and registers it in virtual space control unit 121.

[0066] The environment setting unit 117 performs various settings related to the environment of the virtual space. For example, the environment setting unit 117 sets a background image and a background model to be placed in the virtual space to the virtual space control unit 121. In this embodiment, the background image and background model may be, for example, a workbench on which a worker works, but are not limited to this. Furthermore, for example, the environment setting unit 117 sets parameters related to the angle of view of a virtual camera placed in the virtual space and parameters related to lighting placed in the virtual space to the virtual space control unit 121.

[0067] The virtual space control unit 121 performs various controls related to the virtual space. Specifically, the virtual space control unit 121 generates the virtual space, controls the movements of human models and object models in the virtual space, detects interactions between human models and object models, photographs the human models and object models in the virtual space, and generates annotation information.

[0068] For example, the virtual space control unit 121 generates a virtual space and places in the generated virtual space the person model and object model registered by the model registration unit 111, the background image and background model set by the environment setting unit 117, a virtual camera that captures images of the virtual space, and a light source. The virtual space control unit 121 also adjusts the angle of view of the virtual camera using parameters related to the angle of view of the virtual camera set by the environment setting unit 117, and adjusts the brightness of the light source and the like using parameters related to lighting set by the environment setting unit 117.

[0069] For example, in the virtual space where the above-described environmental settings have been applied, the virtual space control unit 121 causes each of the human model and the object model registered by the model registration unit 111 to operate using the motion data registered by the motion registration unit 113. Furthermore, when the virtual space control unit 121 detects contact between a human model and an object model in motion in the virtual space, it detects, as an interaction, an action indicated by the action class registered in association with the object model by the action class registration unit 115. Furthermore, the virtual space control unit 121 generates a video capturing the human model and the object model from a virtual camera, and detects the positions of the human model and the object model in the generated video. The virtual space control unit 121 generates annotation information by integrating the detected interaction and the positions of the human model and the object model in the video.

[0070] For example, when a developer uses the input device 19 to perform an operation input requesting the display of a video generation screen, the virtual space control unit 121 displays the video generation screen on the display device 17. The video generation screen is a UI screen for starting the generation of the above-mentioned video and annotation information.

[0071] Fig. 10 is a diagram showing an example of a moving image generation screen 501 of this embodiment. As shown in Fig. 10, the moving image generation screen 501 includes a text box 511 for inputting the moving image name of the moving image to be generated, and a rendering start button 521 for starting generation of the moving image and annotation information. Note that on the moving image generation screen 501 shown in Fig. 10, "00:30" (30 seconds) is fixedly displayed as the moving image length indicating the length of the moving image, but this is automatically set according to the playback time of the motion data registered by the motion registration unit 113.

[0072] 10, the developer uses the input device 19 to input the name of the video, "cut.mp4," into the text box 511. In this situation, when the developer uses the input device 19 to press the rendering start button 521, the virtual space control unit 121 generates a virtual space, performs the above-mentioned environmental settings, and starts generating the video and annotation information.

[0073] The following describes how video and annotation information are generated in a virtual space using the interaction detection unit 123, video generation unit 125, position detection unit 127, and annotation information generation unit 129 included in the virtual space control unit 121.

[0074] As described above, the virtual space control unit 121 operates each of the human model and object model registered by the model registration unit 111 in the virtual space using the motion data registered by the motion registration unit 113. The interaction detection unit 123 detects interactions that may occur between a human model and an object model in motion in the virtual space. Specifically, the interaction detection unit 123 detects interactions between a human model and an object model for each frame. Note that, in this embodiment, an example will be described in which the frame rate is 10 fps (frames per second), but the present invention is not limited to this.

[0075] In this embodiment, the interaction detection unit 123 detects the presence or absence of contact between a human model and an object model on a frame-by-frame basis. When the interaction detection unit 123 detects contact between a human model and an object model, the interaction detection unit 123 detects, as an interaction, an action indicated by an action class associated with the object model and registered by the action class registration unit 115. On the other hand, when the interaction detection unit 123 does not detect contact between a human model and an object model, it detects "no interaction."

[0076] The presence or absence of contact between a person model and an object model can be achieved by using a collision function (collision detection function) that is pre-installed on the platform that creates the CG environment. The collision function sets a collider, which is a transparent, three-dimensional object such as a rectangular parallelepiped or sphere that surrounds each person model and object model, and detects whether the person model and object model are in contact with each other based on whether the colliders are in contact with each other.

[0077] A collider set on a person model is an example of a person model area for interaction detection. A collider set on an object model is an example of an object model area for interaction detection. The interaction detection unit 123 detects an interaction based on whether or not the colliders of the person model and the object model are in contact with each other. Specifically, when the colliders of the person model and the object model are in contact with each other, the interaction detection unit 123 determines that the person model and the object model are in contact with each other, and detects, as an interaction, an action indicated by an action class registered in association with the object model. Note that detecting, as an interaction, an action indicated by an action class registered in association with the object model is an example of detecting, as an interaction, the content indicated by content information associated with the object model. On the other hand, when the colliders of the person model and the object model are not in contact with each other, the interaction detection unit 123 determines that the person model and the object model are not in contact with each other, and detects "no interaction."

[0078] FIG. 11 is an explanatory diagram of an example of the interaction detection method of this embodiment. In the example shown in FIG. 11, the motion transitions and presence or absence of interactions in the virtual space of the person model 221, the object model 271, the object model 273, and the object model 275 are shown in chronological order. Note that, for convenience of explanation, in the example shown in FIG. 11, the person model 221 is displayed with only the hand model 222, not the entire body. Furthermore, the object model 275 is a CG model of a pen. In the example shown in FIG. 11, it is assumed that an action class "scissor grip" is defined in association with the object model 271 (see FIG. 7), an action class "cut" is defined in association with the object model 271 and the object model 273 (see FIG. 8), and an action class "pen hold" is defined in association with the object model 275.

[0079] 11, a collider 522 is set for hand model 222, a collider 571 for object model 271, a collider 573 for object model 273, and a collider 575 for object model 275. Note that in the example shown in Fig. 11, for convenience of explanation, hand model 222, object model 271, object model 273, and object model 275 are shown in two dimensions, but in reality they are three-dimensional models. Similarly, collider 522, collider 571, collider 573, and collider 575 are shown as visible rectangles, but in reality they are transparent (invisible) rectangular parallelepiped objects.

[0080] First, at time 0s, the collider 522 of the hand model 222 is not in contact with the colliders of any of the object models. Therefore, the interaction detection unit 123 does not detect contact between the person model 221 and any of the object models, resulting in "no interaction." As a result, the interaction detection unit 123 associates "no interaction" with each object model as the interaction detection result at time 0s (frame).

[0081] Next, the virtual space control unit 121 moves the hand model 222 so as to approach the object model 271. As a result, at the time point of 10 seconds, the collider 522 of the hand model 222 comes into contact with the collider 571 of the object model 271, and the interaction detection unit 123 detects contact between the person model 221 and the object model 271, i.e., the presence of an interaction. The interaction detection unit 123 detects "scissors grip" as the interaction between the person model 221 and the object model 271 because the action class defined in association with the object model 271 for which contact was detected is "scissors grip."

[0082] As a result, the interaction detection unit 123 associates "scissors grip" with the object model 271 and associates "no interaction" with the object models 273 and 275 as the interaction detection result at the 10 s point (frame). Note that the collider 522 of the hand model 222 is not in contact with the colliders of any of the object models up until 10 s. Therefore, in the interaction detection results for each frame from 0 s to 9.9 s, "no interaction" is associated with each object model.

[0083] Next, the virtual space control unit 121 moves the hand model 222 and the object model 271 while causing the hand model 222 to hold the object model 271 so that the object model 271 approaches the object model 273. As a result, at the time point of 20 s, the collider 522 of the hand model 222 comes into contact with the collider 571 of the object model 271, and the collider 571 of the object model 271 comes into contact with the collider 573 of the object model 273. Therefore, the interaction detection unit 123 detects contact between the person model 221 and the object model 271, and contact between the object model 271 and the object model 273 (indirect contact between the person model 221 and the object model 273 via the object model 271). The interaction detection unit 123 detects "disconnection" as the interaction between the person model 221, the object model 271, and the object model 273, because the action class defined in association with the object model 271 and the object model 273 where contact was detected is "disconnection."

[0084] As a result, the interaction detection unit 123 associates "cut" with object models 271 and 273, and associates "no" interaction with object model 275, as the interaction detection result at the 20 s point (frame). Note that collider 522 of hand model 222 is in contact only with collider 571 of object model 271 from 10 s to 19.9 s. Therefore, in the interaction detection results for each frame from 10 s to 19.9 s, "scissor grip" is associated with object model 271, and "no" interaction is associated with object models 273 and 275.

[0085] Next, the virtual space control unit 121 causes the hand model 222 to release its grip on the object model 271 and moves the hand model 222 so as to approach the object model 275. As a result, at the time point of 30 s, the collider 522 of the hand model 222 comes into contact with the collider 575 of the object model 275, and the interaction detection unit 123 detects contact between the person model 221 and the object model 275. The action class defined in association with the object model 275 for which contact was detected is "pen grip," and therefore the interaction detection unit 123 detects "pen grip" as the interaction between the person model 221 and the object model 275.

[0086] As a result, the interaction detection unit 123 associates "pen grip" with the object model 275 and associates "no interaction" with the object models 271 and 273 as the interaction detection result at the time point (frame) of 30 s. Note that the collider 522 of the hand model 222 does not contact the colliders of any of the object models from the time the hand model 222 releases its grip on the object model 271 until 30 s. Therefore, in the interaction detection result for each frame from 20 s until the hand model 222 releases its grip on the object model 271, "disconnection" is associated with the object models 271 and 273, and "no interaction" is associated with the object model 275. In addition, in the interaction detection result for each frame from the time the hand model 222 releases its grip on the object model 271 until 29.9 s, "no interaction" is associated with each object model.

[0087] The video generation unit 125 generates a video including a person model and one or more object models. Specifically, the video generation unit 125 controls a virtual camera arranged in a virtual space and generates a video capturing the person model and the object model from the virtual camera. For example, the video generation unit 125 generates an image by having the virtual camera project (render) various models, such as person models and object models, and backgrounds, included in the angle of view (field of view) of the virtual camera onto a projection surface (not shown). In rendering, a world coordinate system, which is a coordinate system in the virtual space, is converted into a camera coordinate system whose origin is the viewpoint position, which is the position where the virtual camera is arranged in the virtual space, and the various models, backgrounds, etc. expressed in the camera coordinate system are converted into a coordinate system on a two-dimensional projection surface by perspective projection transformation. Note that a well-known rendering method, such as the Z-buffer method, may be used. The video generation unit 125 generates a video captured in the virtual space by having the virtual camera perform the above-mentioned image generation for each frame.

[0088] In this embodiment, the virtual camera is fixedly positioned in the virtual space, but is not limited to this. For example, the virtual camera may be positioned above the virtual space, with an angle of view set to overlook the virtual space. In this embodiment, the virtual camera is positioned in the virtual space so as to correspond to the position of a camera that captures images of the worker's work status in the real environment to recognize the worker's work using a recognition model, but is not limited to this. For example, the virtual camera may be positioned in the virtual space so that the relative position between the worker and the camera in the real space corresponds to the relative position between the human model and the camera in the virtual space.

[0089] The position detection unit 127 detects the positions of the person model and the object model on the moving image generated by the moving image generation unit 125. Specifically, the position detection unit 127 detects the positions of the person model and the object model on the image each time an image is generated frame by frame by the moving image generation unit 125. Note that since the person model and the object model on the moving image are converted into two dimensions, the position detection unit 127 detects the positions of the person model and the object model in two-dimensional coordinates. In this embodiment, the position detection unit 127 detects the position of each model to be detected, such as a person model and an object model, by a rectangle surrounding the model, but is not limited to this.

[0090] For example, the position detection unit 127 may detect the position of each model of the detection target by calculating the two-dimensional coordinates on the video from the three-dimensional coordinates in the camera coordinate system of each model of the detection target based on coordinate transformation information performed during rendering. Also, for example, the position detection unit 127 may detect the position of each model of the detection target by performing person recognition processing or object recognition processing on the video generated by the video generation unit 125.

[0091] The annotation information generation unit 129 generates annotation information according to interactions between human models and object models. Note that the annotation information generation unit 129 may further generate annotation information according to interactions between object models. In this embodiment, the annotation information generation unit 129 generates annotation information indicating the results of interaction detection by the interaction detection unit 123. Therefore, the annotation information in this embodiment indicates the content of interactions between human models and object models for each frame. The content of the interaction may, for example, indicate whether or not there is an interaction, and if there is an interaction, the action content indicated by the action class detected as the interaction.

[0092] Specifically, the annotation information generation unit 129 generates annotation information by integrating the interaction detection result by the interaction detection unit 123 and the position detection result of the human model and the object model by the position detection unit 127. Note that both the interaction detection unit 123 and the position detection unit 127 perform detection on a frame-by-frame basis. Therefore, the annotation information generation unit 129 can integrate the interaction detection result by the interaction detection unit 123 and the position detection result of the human model and the object model by the position detection unit 127 using frames as a key.

[0093] Fig. 12 is a diagram showing an example of annotation information according to this embodiment. The example shown in Fig. 12 shows, as annotation information, information that combines the detection result of interactions in the situation in the virtual space described in Fig. 11 and the position detection result of each model on a video that captures the situation in the virtual space described in Fig. 11.

[0094] In the example shown in FIG. 12, the annotation information is information that associates a frame, a time, rectangular coordinates of a human model, and rectangular coordinates and interactions of each object model. The interaction of each object model corresponds to the detection result of the interaction, and the rectangular coordinates of each model correspond to the position detection result of each model on the video. In the example shown in FIG. 12, the human model (hand) indicates hand model 222 of human model 221, the object model (scissors) indicates object model 271, the object model (paper) indicates object model 273, and the object model (pen) indicates object model 275. In the example shown in FIG. 12, the rectangular coordinates indicate the x and y coordinates of the upper left and lower right of a rectangle surrounding the model, but the present invention is not limited to this.

[0095] 11, the interaction detection result is associated with each object model. Therefore, when the person model 221 is not in contact with any object model, as at time 0 s, the interaction of each object model is "absent." When the person model 221 is in contact with the object model 271, as at time 10 s, the interaction of the object model 271 is "present" and "scissors grip," and the interactions of the object models 273 and 275 are "absent." When the person model 221 is in contact with the object model 271 and indirectly in contact with the object model 273, as at time 20 s, the interaction of the object models 271 and 273 is "present" and "disconnected," and the interaction of the object model 275 is "absent." Also, when the person model 221 is in contact with the object model 275, as at 30 seconds, the interaction between the object models 271 and 273 is "absent," and the interaction between the object model 275 is "present" and "pen gripped."

[0096] When annotation information is generated by the annotation information generation unit 129, the confirmation unit 131 displays a confirmation screen on the display device 17 to prompt the developer to confirm and correct the annotation information. This allows the developer to confirm whether the content of the interaction in the generated annotation information is appropriate or not, and to correct any inappropriate parts, even when annotation information is automatically generated as in this embodiment.

[0097] In particular, in this embodiment, the confirmation unit 131 displays the annotation information generated by the annotation information generation unit 129 and the moving image generated by the moving image generation unit 125 in synchronization with each other on a frame-by-frame basis. This allows the developer to check whether the content of the interaction in the annotation information is appropriate while checking the scene in which the interaction actually occurred on the moving image, thereby improving the efficiency of the confirmation work.

[0098] Hereinafter, a method for displaying annotation information and moving images, and modification of annotation information will be described using the above-described display control unit 133 and annotation information modification unit 135 included in the confirmation unit 131. Note that the method for displaying annotation information and moving images, and modification of annotation information will be described using an example different from the examples described in Fig. 11 and Fig. 12. First, the different example will be described with reference to Fig. 13.

[0099] Fig. 13 is a diagram showing an example of a video 601 generated by the video generation unit 125 of this embodiment. The video 601 shown in Fig. 13 is a video of a packing operation performed by a worker in a real environment, reproduced in a virtual space using a human model 221. Note that the video data captured to recognize the packing operation of the worker in the real environment is expected to be captured, for example, from above the work site. For this reason, the video 601 shown in Fig. 13 is captured from a virtual camera placed above the work site in the virtual space so as to correspond to the shooting position of the video in the real environment.

[0100] 13 is a video of a scene in which a person model 221 is packing a packing box model 621 in a virtual space using a tape model 631, a tape model 633, and a tape model 635. A workbench model 611 is placed in the virtual space, and a packing box model 621, a tape model 631, a tape model 633, a tape model 635, and a replacement tape model 641 are placed on the workbench model 611.

[0101] The tape model 631, the tape model 633, and the tape model 635 are reproductions of different types of tape, and are attached to the tape cutter. The replacement tape model 641 is a reproduction of the same type of tape as the tape model 631, and is used as a replacement for the tape model 631. For example, when the tape of the tape model 631 runs out, the replacement tape model 641 is attached to the tape cutter of the tape model 631 in place of the tape model 631. The person model 221 packs the packing box model 621 on the workbench model 611 by applying tape to the packing box model 621 while switching between tape models.

[0102] 13, tape model 631, tape model 633, tape model 635, and replacement tape model 641 correspond to the above-mentioned object models, but are not limited to these. Also, in the example shown in Fig. 13, an action class "apply tape A" is defined in association with tape model 631, an action class "apply tape B" is defined in association with tape model 633, an action class "apply tape C" is defined in association with tape model 635, and an action class "refill tape A" is defined in association with tape model 631 and replacement tape model 641.

[0103] Based on the above assumptions, the video shown in Fig. 13 is assumed to include a scene in which the person model 221 applies tape to the packaging box model 621 in the order of tape model 633 (tape B), tape model 631 (tape A), and tape model 635 (tape C), and refills the tape model 631 (tape A). Note that while a detailed description of the annotation information of the video shown in Fig. 13 will be omitted, as described in Fig. 12, the information associates frames, time, rectangular coordinates of the person model, and rectangular coordinates and interactions of each object model.

[0104] 14 is a diagram showing an example of a confirmation screen 701 displayed on the display device 17 by the display control unit 133 of this embodiment. As shown in FIG. 14, the confirmation screen 701 includes a video display area 702, a timeline display area 703, an annotation information display area 704, and a display setting button 771.

[0105] The display control unit 133 displays the video 601 and a view angle change button 715 in the video display area 702, displays a play button 751 and a seek bar 753 in the timeline display area 703, and displays interaction information 761 as annotation information for the video 601 in the annotation information display area 704.

[0106] Furthermore, the display control unit 133 superimposes and displays annotation information of the moving image 601 on the moving image 601 displayed in the moving image display area 702. Specifically, the display control unit 133 superimposes and displays, on the moving image 601, as annotation information, a rectangle (an example of a frame) surrounding each model and a line segment indicating the presence or absence of an interaction.

[0107] 14, the display control unit 133 displays a rectangle 721 superimposed on the person model 221, a rectangle 731 on the tape model 631, a rectangle 733 on the tape model 633, a rectangle 735 on the tape model 635, and a rectangle 741 on the replacement tape model 641 on the video 601. Also, in the example shown in Fig. 14, the display control unit 133 displays line segments 732, 734, 736, and 742 connecting the rectangle 721 of the person model 221 to the rectangle 731 of the tape model 631, the rectangle 733 of the tape model 633, the rectangle 735 of the tape model 635, and the rectangle 741 of the replacement tape model 641, superimposed on the video 601. Also, in the example shown in Fig. 14, the display control unit 133 displays the line segment 732 as a dotted line because there is an interaction between the person model 221 and the tape model 631. On the other hand, the display control unit 133 displays the line segments 734, 736, and 742 as solid lines because there is no interaction between the person model 221 and the tape model 633, tape model 635, and replacement tape model 641.

[0108] Specifically, when playing back any frame of the video 601, the display control unit 133 refers to the annotation information and acquires the rectangular coordinates of each model and the interaction of the object model in that frame. Based on the acquired rectangular coordinates of each model, the display control unit 133 displays the above-mentioned rectangle superimposed on the video 601. Furthermore, based on the acquired rectangular coordinates of the human model and the rectangular coordinates of each object model, the display control unit 133 displays the above-mentioned line segment superimposed on the video 601. For example, the display control unit 133 generates a line segment connecting the center coordinate of the rectangular coordinates of the human model and the center coordinate of the rectangular coordinates of each object model, and displays the line segment superimposed on the video 601. Furthermore, the display control unit 133 determines the display mode of the above-mentioned line segment based on the acquired interaction. For example, the display control unit 133 displays the line segment as a dotted line when there is an interaction, and displays the line segment as a solid line when there is no interaction.

[0109] In this way, in this embodiment, annotation information is displayed superimposed on the video, so that developers can check the annotation information on the video, improving the efficiency of the checking work. Furthermore, in this embodiment, the display mode of the annotation information superimposed on the video is changed depending on whether or not there is interaction, making it easier for developers to check the annotation information, and further improving the efficiency of the checking work.

[0110] In the present embodiment, an example has been described in which the display mode of annotation information is changed depending on whether an interaction is present or absent, by switching between displaying a line segment as a solid line and a dotted line. However, the present invention is not limited to this. For example, any display mode may be changed depending on whether an interaction is present or absent, such as by changing the color or thickness of the line segment. Alternatively, a predetermined line segment may be displayed only when an interaction is present.

[0111] A play button 751 displayed in the timeline display area 703 is a button for playing the video 601. A seek bar 753 displayed in the timeline display area 703 is a UI component that displays the playback position (playback frame) of the video 601 using a slider.

[0112] The interaction information 761 displayed in the annotation information display area 704 includes items indicating article models and interactions. The items include tape A corresponding to tape model 631, tape B corresponding to tape model 633, tape C corresponding to tape model 635, and replacement tape A corresponding to replacement tape model 641. The interactions include whether or not there is an interaction, and if there is an interaction, the action (the content of the interaction resulting from contact with the human model).

[0113] In this embodiment, the display control unit 133 displays the interaction information 761 in synchronization with the video 601 displayed in the video display area 702 on a frame-by-frame basis. That is, each time a frame being played back in the video 601 is updated, the display control unit 133 refers to the annotation information, acquires the interaction of the object model in that frame, and updates the display content of the interaction information 761 to the acquired interaction content. Therefore, according to this embodiment, the developer can check whether the content of the interaction in the annotation information is appropriate while checking the scene in which the interaction actually occurred on the video, thereby improving the efficiency of the checking work.

[0114] Furthermore, a correction button is associated with the presence or absence of interaction and the action. When the developer uses the input device 19 to press the correction button and input the correction content, the annotation information correction unit 135 corrects the presence or absence of interaction or the action associated with the correction button to the input content and corrects the corresponding part of the annotation information. For example, the developer uses the input device 19 to press the correction button 763 associated with the presence or absence of interaction of tape A and inputs the correction content. In this case, the annotation information correction unit 135 corrects the presence or absence of interaction to the input content and corrects the corresponding part of the annotation information. Also, for example, the developer uses the input device 19 to press the correction button 765 associated with the action of tape A and inputs the correction content. In this case, the annotation information correction unit 135 corrects the action to the input content and corrects the corresponding part of the annotation information.

[0115] This allows developers to check whether the interaction content of the generated annotation information is appropriate, even when annotation information is automatically generated as in this embodiment, and to correct any parts that are not appropriate.

[0116] The display setting button 771 is a button for displaying a display setting screen for performing display settings of the confirmation screen 701. For example, when the developer performs an operation input of pressing the display setting button 771 using the input device 19, the display control unit 133 displays the display setting screen on the display device 17.

[0117] Fig. 15 is a diagram showing an example of a display setting screen 801 of this embodiment. As shown in Fig. 15, the display setting screen 801 includes check boxes 811, 813, and 815 for making display settings for interactions, check boxes 817 and 819 for making display settings for the timeline, and an OK button 821.

[0118] In the example shown in Fig. 15, three patterns can be set for the interaction display settings: "Show all," "Show only existing," and "Do not display." Note that "Show all" is set by selecting check box 811, "Show only existing" by selecting check box 813, and "Do not display" by selecting check box 815. In the example shown in Fig. 14, all interactions were displayed, so in the example shown in Fig. 15, check box 811 is selected.

[0119] 15, two patterns, "simple" and "detailed," can be set as the timeline display settings. "Simple" is set by selecting check box 817, and "detailed" is set by selecting check box 819. In the example shown in FIG. 14, the timeline was displayed in a simple format with only a seek bar, so in the example shown in FIG. 15, check box 817 is selected.

[0120] The OK button 821 is a button for confirming the display settings of the interaction and timeline to the contents selected in the check boxes. For example, on the display setting screen 801 shown in Fig. 15, suppose that the developer uses the input device 19 to deselect the check box 817, select the check box 819, and then press the OK button 821. In this case, the display control unit 133 switches the display content of the timeline display area 703 of the confirmation screen 701 from the simple version to the detailed version.

[0121] Fig. 16 is a diagram showing an example of a confirmation screen 701 after changing the timeline display settings of this embodiment. The confirmation screen 701 shown in Fig. 16 has the same display content as the confirmation screen 701 shown in Fig. 14 except for the display content of the timeline display area 703, although the display size of the video display area 702 and the like are different. For this reason, Fig. 16 will explain the timeline display area 703.

[0122] 16, the timeline display area 703 includes a play button 751, a play bar 754, and a timeline 781. In this way, by changing the timeline display setting from "simple" to "detailed," the seek bar 753 is switched to the timeline 781.

[0123] The timeline 781 is configured for each item (object model). Specifically, the timeline 781 is configured with timelines for tape A (tape model 631), tape B (tape model 633), tape C (tape model 635), and replacement tape A (replacement tape model 641). Each timeline displays chronological changes in interactions with the person model. In the example shown in FIG. 16, each timeline is shown without shading during periods without interaction and with shading corresponding to the interaction during periods with interaction, but the display manner is not limited to this. Therefore, by checking the timeline 781, it can be seen that in the video 601, the person model 221 applies tape B, tape A, and tape C to the packaging box model 621 in this order, and refills tape A. The playback bar 754 is a bar that displays the playback position (playback frame) of the video 601.

[0124] The learning data generation unit 141 generates learning data by adding annotation information generated by the annotation information generation unit 129 to the moving image generated by the moving image generation unit 125. Specifically, the learning data generation unit 141 generates learning data by adding annotation information confirmed by the confirmation unit 131 to the moving image generated by the moving image generation unit 125. The learning data generation unit 141 outputs the generated learning data to the auxiliary storage device 15 or to the outside via the communication device 21.

[0125] FIG. 17 is a flowchart showing an example of the training data generation process performed by the training data generation device 10 of this embodiment.

[0126] First, the model registration unit 111 registers a character model selected by the developer from among the character models stored in the model storage unit 101 in the virtual space control unit 121 (step S101).

[0127] Next, the model registration unit 111 registers one or more object models selected by the developer from the object models stored in the model storage unit 101 in the virtual space control unit 121 (step S103).

[0128] Next, the motion registration unit 113 registers in the virtual space control unit 121 the motion data selected by the developer from the motion data stored in the motion memory unit 103 in order to operate the person model and object model registered by the model registration unit 111 (step S105).

[0129] Next, the action class registration unit 115 registers an action class defined in association with the object model registered by the model registration unit 111 (step S107).

[0130] Next, the environment setting unit 117 performs various settings relating to the environment of the virtual space on the virtual space control unit 121 (step S109).

[0131] Next, the virtual space control unit 121 starts taking (generating) a video capturing a person model and an object model in the virtual space and generating annotation information (step S111).

[0132] Next, the interaction detection unit 123 detects whether or not there is contact between the person model and the object model in the virtual space (step S113).

[0133] When the interaction detection unit 123 detects contact between a person model and an object model (Yes in step S113), it detects, as an interaction, an action indicated by an action class registered in association with the object model by the action class registration unit 115 (step S115).

[0134] On the other hand, if the interaction detection unit 123 does not detect contact between the human model and the object model (No in step S113), it detects "no" interaction.

[0135] Next, the video generating unit 125 controls a virtual camera placed in the virtual space, and generates an image capturing the human model and the object model from the virtual camera (step S117).

[0136] Next, the position detection unit 127 detects the positions of the human model and the object model on the image generated by the video generation unit 125 (step S119).

[0137] Next, the virtual space control unit 121 checks whether the playback of the motion that moves the human model and object model in the virtual space has finished (step S121). If the playback of the motion has not finished (No in step S121), the process returns to step S113. On the other hand, if the playback of the motion has finished (Yes in step S121), the process proceeds to step S123. The processes in steps S113 to S121 are performed on a frame-by-frame basis.

[0138] Next, the annotation information generating unit 129 integrates the interaction detection result by the interaction detecting unit 123 and the position detection result of the human model and object model by the position detecting unit 127 to generate annotation information (step S123).

[0139] Next, the confirmation unit 131 displays a confirmation screen for the annotation information generated by the annotation information generation unit 129 on the display device 17, and allows the developer to confirm and correct the annotation information (step S125).

[0140] Next, the learning data generation unit 141 generates learning data by adding the annotation information confirmed by the confirmation unit 131 to the moving image generated by the moving image generation unit 125 (step S127).

[0141] FIG. 18 is a flowchart showing an example of the annotation information checking and correction process shown in step S125 of the flowchart in FIG.

[0142] First, the display control unit 133 displays a confirmation screen 701 including a moving image display area 702, a timeline display area 703, an annotation information display area 704, and a display setting button 771 on the display device 17 (step S201). The display control unit 133 displays the moving image 601 in the moving image display area 702 and the annotation information (interaction information 761) in the annotation information display area 704, and further displays the annotation information superimposed on the moving image 601.

[0143] Next, the display control unit 133 checks whether there is a request to change the display settings of the confirmation screen 701 based on whether there is an operation input to press the display setting button 771 (step S203).

[0144] When the display setting button 771 is pressed to confirm a request to change the display settings of the confirmation screen 701 (Yes in step S203), the display control unit 133 displays the display setting screen 801 on the display device 17 and accepts changes to the display settings of the interactions and the timeline (step S205). As a result, the display settings of the interactions and the timeline on the confirmation screen 701 are changed and displayed. Note that if the display setting button 771 is not pressed and the request to change the display settings of the confirmation screen 701 is not confirmed (No in step S203), the processing of step S205 is not performed.

[0145] Next, when an operation input is made to press the play button 751 on the confirmation screen 701, the display control unit 133 starts playing the video 601, updates the playback frame, and updates the display content of the interaction information 761 to that of the updated playback frame (step S207).

[0146] Next, the annotation information correction unit 135 checks whether there is a request to correct the interaction based on whether there is an interaction in the interaction information 761 and whether there is an operation input to press a correction button associated with the action (step S209). In practice, the annotation information correction unit 135 checks whether there is a request to correct the interaction when the playback of the video 601 is paused by pressing the play button 751 (pause button) during playback of the video 601.

[0147] If the request to modify the interaction is confirmed by pressing the Modify button (Yes in step S209), the annotation information modifying unit 135 accepts the input of the modification content of the interaction and modifies the interaction 761 (step S211). In this way, the annotation information is modified. Note that if the Modify button is not pressed and the request to modify the interaction is not confirmed (No in step S209), the processing of step S211 is not performed.

[0148] Next, the display control unit 133 checks whether the video 601 has been played to the end (step S213). If the video 601 has not been played to the end (No in step S213), the process returns to step S207. On the other hand, if the video 601 has been played to the end (Yes in step S213), the process ends.

[0149] In this manner, in this embodiment, training data is generated by adding annotation information indicating the content of the interaction to a video containing an interaction between a person model and an object model. Therefore, according to this embodiment, training data useful for training a recognition model for recognizing interactions between a person and an object can be generated. Furthermore, according to this embodiment, the generation and addition of annotation information can be automated, allowing training data to be generated efficiently.

[0150] In this embodiment, an action class is defined in association with an object model, and when a human model comes into contact with an object model, the action content indicated by the action class is detected as an interaction. In this way, in this embodiment, since the detection of interactions can be automated, annotation costs can be reduced even in situations where the annotation costs of generating annotation information indicating the presence or absence of interactions on a frame-by-frame basis would be high.

[0151] Furthermore, in this embodiment, once annotation information is generated, a confirmation screen is displayed to allow the developer to confirm and modify the annotation information. This allows the developer to confirm whether the content of the interaction in the generated annotation information is appropriate, even when annotation information is automatically generated as in this embodiment, and to modify any inappropriate parts.

[0152] In particular, in this embodiment, in addition to annotation information, video is displayed in synchronization with the video on a frame-by-frame basis, allowing developers to check whether the content of the interaction in the annotation information is appropriate while checking the scene in which the interaction actually occurred on the video, thereby improving the efficiency of the checking work.

[0153] In this embodiment, annotation information is also superimposed on the video, allowing developers to check the annotation information on the video, improving the efficiency of the checking work. In addition, in this embodiment, the display mode of the annotation information superimposed on the video changes depending on whether or not there is interaction, making it easier for developers to check the annotation information, further improving the efficiency of the checking work.

[0154] In addition, in this embodiment, in addition to video and annotation information, the chronological changes in interactions with human models can be displayed on a timeline for each object model, making it easier for developers to check annotation information and further improving the efficiency of the checking process.

[0155] (Variation 1) In the above embodiment, an example has been described in which a collider is used to detect contact between a person model and an object model, but the method for detecting contact between a person model and an object model is not limited to this. The interaction detection unit 123 may detect actual contact between a person model and an object model, or may detect an interaction using a positional relationship, such as a distance, between the person model and the object model. In the latter case, the interaction detection unit 123 may, for example, set reference points on the person model and the object model and detect an interaction based on the distance and positional relationship between the reference points.

[0156] (program) The programs executed by the training data generation device 10 of the above embodiment and the above modified example are provided as files in an installable or executable format stored on a computer-readable storage medium such as a CD-ROM, CD-R, memory card, DVD, or flexible disk (FD).

[0157] The programs executed by the training data generation device 10 of the above embodiment and the above modification may be stored on a computer connected to a network such as the Internet and provided by being downloaded via the network. The programs executed by the training data generation device 10 of the above embodiment and the above modification may be provided or distributed via a network such as the Internet. The programs executed by the training data generation device 10 of the above embodiment and the above modification may be provided by being pre-installed in a ROM or the like.

[0158] The programs executed by the training data generation device 10 of the above embodiment and the above modification have a modular configuration for implementing the above-mentioned units on a computer. In terms of actual hardware, for example, the CPU reads the training program from a HDD onto a RAM and executes it, thereby implementing the above-mentioned units on a computer.

[0159] As described above, according to the above embodiment and the above modification, it is possible to generate learning data that is useful for training a recognition model for recognizing interactions between a person and an object.

[0160] The above-described embodiment and modifications merely illustrate examples of specific embodiments of the present disclosure, and the technical scope of the present disclosure should not be construed as being limited by these. Therefore, the present disclosure can be implemented in various forms without departing from the spirit or main features thereof. For example, the above-described embodiment and modifications may be appropriately combined in their respective constituent units. Furthermore, for example, some components may be deleted from all components in the above-described embodiment and modifications.

[0161] The present disclosure includes the following aspects.

[0162] (1) a moving image generating unit that generates a moving image including a person model and a first object model; an annotation information generation unit that generates annotation information in accordance with an interaction between the human model and the first object model; a learning data generation unit that generates learning data by adding the annotation information to the video; A training data generation device comprising:

[0163] In the above configuration (1), training data is generated by adding annotation information generated in response to an interaction to a video containing an interaction between a person model and an object model. Therefore, the above configuration (1) makes it possible to generate training data useful for training a recognition model for recognizing interactions between a person and an object. Furthermore, the above configuration (1) makes it possible to automate the generation and addition of annotation information, thereby enabling efficient generation of training data.

[0164] (2) the video further includes a second object model; the annotation information generation unit further generates the annotation information in response to an interaction between the first object model and the second object model. The training data generation device according to (1) above.

[0165] In the above configuration (2), training data is generated by adding annotation information generated in response to interactions to videos containing interactions between object models. Therefore, the above configuration (2) makes it possible to generate training data useful for training a recognition model for recognizing interactions between objects. Furthermore, the above configuration (2) makes it possible to automate the generation and addition of annotation information, thereby enabling efficient generation of training data.

[0166] (3) further comprising an interaction detection unit that detects the interaction for each frame; the annotation information indicates, for each frame, the content of an interaction between the human model and the first object model or the second object model; The training data generation device according to (1) or (2) above.

[0167] The configuration (3) above automates the detection of interactions, thereby reducing annotation costs even in situations where the annotation costs of generating annotation information indicating the presence or absence of interactions on a frame-by-frame basis would be high.

[0168] (4) The first object model or the second object model is associated with content information that defines a content that occurs when the first object model or the second object model comes into contact with the person model; the interaction detection unit detects, when it is determined that the human model and the first object model or the second object model have come into contact with each other, the content indicated by the content information as the interaction. The training data generation device according to (3) above.

[0169] In the above configuration (4), an action class is defined in association with an object model, and when a human model and an object model come into contact, the action content indicated by the action class is detected as an interaction. In this way, the above configuration (4) can automate the detection of interactions, so annotation costs can be reduced even in situations where the annotation costs of generating annotation information indicating the presence or absence of interactions on a frame-by-frame basis would be high.

[0170] (5) a person model area for interaction detection is set in the person model; an object model area for interaction detection is set in the first object model or the second object model; the interaction detection unit detects the interaction based on whether or not there is contact between the person model area and the object model area. The training data generation device according to (3) above.

[0171] According to the above configuration (5), contact between the human model and the first object model can be detected by simple processing, thereby reducing the processing load.

[0172] (6) The interaction detection unit detects the interaction by utilizing a positional relationship between the person model and the first object model or the second object model. The training data generation device according to (3) above.

[0173] (7) The processor: a moving image generating step of generating a moving image including the person model and the first object model; an annotation information generating step of generating annotation information in accordance with an interaction between the human model and the first object model; a learning data generation step of generating learning data by adding the annotation information to the video; A training data generation method including:

[0174] In the above configuration (7), training data is generated by adding annotation information generated in response to an interaction to a video containing an interaction between a person model and an object model. Therefore, the above configuration (7) makes it possible to generate training data useful for training a recognition model for recognizing interactions between a person and an object. Furthermore, the above configuration (7) makes it possible to automate the generation and addition of annotation information, thereby enabling efficient generation of training data.

[0175] (8) a moving image generating step of generating a moving image including the person model and the first object model; an annotation information generating step of generating annotation information in accordance with an interaction between the human model and the first object model; a learning data generation step of generating learning data by adding the annotation information to the video; A program that causes a computer to execute the following.

[0176] In the above configuration (8), training data is generated by adding annotation information generated in response to an interaction to a video containing an interaction between a person model and an object model. Therefore, the above configuration (8) makes it possible to generate training data useful for training a recognition model for recognizing interactions between a person and an object. Furthermore, the above configuration (8) makes it possible to automate the generation and addition of annotation information, thereby enabling efficient generation of training data. [Explanation of symbols]

[0177] 10. Training data generation device 11 Control device 13 Main memory 15 Auxiliary storage 17 Display device 19 Input Devices 21 Communication equipment 23 Various buses 101 Model memory section 103 Motion Memory Unit 111 Model Registration Department 113 Motion Registration Unit 115 Operation class registration section 117 Environment Settings 121 Virtual Space Control Unit 123 Interaction detection unit 125 Video Generation Unit 127 Position detection unit 129 Annotation Information Generation Unit 131 Confirmation Department 133 Display control unit 135 Annotation information correction section 141 Learning data generation unit

Claims

1. a moving image generating unit that generates a moving image including the person model and the first object model; an annotation information generation unit that generates annotation information in accordance with an interaction between the human model and the first object model; a learning data generation unit that generates learning data by adding the annotation information to the video; A training data generation device comprising:

2. the animation further includes a second object model; the annotation information generation unit further generates the annotation information in response to an interaction between the first object model and the second object model. The training data generating device according to claim 1 .

3. further comprising an interaction detection unit that detects the interaction for each frame; the annotation information indicates, for each frame, the content of an interaction between the human model and the first object model or the second object model; The training data generating device according to claim 1 or 2.

4. content information that defines a content that occurs when the first object model or the second object model comes into contact with the person model is associated with the first object model or the second object model; the interaction detection unit detects, when it is determined that the human model and the first object model or the second object model have come into contact with each other, content indicated by the content information as the interaction. The training data generating device according to claim 3 .

5. a person model area for interaction detection is set in the person model; an object model area for interaction detection is set in the first object model or the second object model; the interaction detection unit detects the interaction based on whether or not there is contact between the person model area and the object model area. The training data generating device according to claim 3 .

6. the interaction detection unit detects the interaction by utilizing a positional relationship between the person model and the first object model or the second object model. The training data generating device according to claim 3 .

7. The processor: a moving image generating step of generating a moving image including the person model and the first object model; an annotation information generating step of generating annotation information in accordance with an interaction between the human model and the first object model; a learning data generation step of generating learning data by adding the annotation information to the video; A training data generation method including:

8. a moving image generating step of generating a moving image including the person model and the first object model; an annotation information generating step of generating annotation information in accordance with an interaction between the human model and the first object model; a learning data generation step of generating learning data by adding the annotation information to the video; A program that causes a computer to execute the following.

Citation Information

Patent Citations

  • Learning data generation device, learning data generation method, machine learning method, and program

    JP2019023858A