Learning data generation device, annotation information display device, learning data generation method, annotation information display method, and program

The training data generation device generates videos of human and object models in a virtual space, detecting and annotating interactions to train recognition models for identifying interactions between people and objects, addressing the limitation of existing models and reducing annotation costs.

WO2025206193A1PCT designated stage Publication Date: 2025-10-02PANASONIC INTELLECTUAL PROPERTY MANAGEMENT CO LTD
View PDF 8 Cites 0 Cited by

Patent Information

Application Number
PCT/JP2025/012496
Authority / Receiving Office
WO · WO
Patent Type
Applications
Current Assignee / Owner
Priority Date
2024-03-28
Filing Date
2025-03-27
Publication Date
2025-10-02

AI Technical Summary

Technical Problem

Existing recognition models are unable to recognize interactions between people and objects, as they are primarily designed to identify individuals such as persons or vehicles and lack the capability to understand interactions between them.

Method used

A training data generation device and method that generates videos of human and object models in a virtual space, detects interactions, and adds annotation information to these videos to create training data for recognition models, allowing for the automation of interaction detection and reduced annotation costs.

Benefits of technology

Enables the generation of training data that effectively trains recognition models to identify interactions between people and objects, improving the model's ability to recognize such interactions and reducing the need for manual, frame-by-frame annotation.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure JP2025012496_02102025_PF_FP_ABST
    Figure JP2025012496_02102025_PF_FP_ABST
Patent Text Reader

Abstract

One aspect of a learning data generation device according to the present disclosure comprises one or more processors. The one or more processors generate a moving image including a person model and a first object model, generate annotation information in accordance with the interaction between the person model and the first object model, and add the annotation information to the moving image and generate learning data.
Need to check novelty before this filing date? Find Prior Art

Description

Learning data generation device, annotation information display device, learning data generation method, annotation information display method, and program

[0001] The present disclosure relates to a training data generation device, an annotation information display device, a training data generation method, an annotation information display method, and a program.

[0002] A known method of data augmentation for expanding (increasing) learning data for machine learning is to generate learning data using computer graphics (CG).

[0003] For example, the technology disclosed in Patent Document 1 generates an image that narrows down a scene data image containing multiple CG models to include only specific CG models, such as a person model or a vehicle model. The technology disclosed in Patent Document 1 generates training data by adding information about the specific CG model to the generated image as annotation information. Examples of annotation information include a rectangle surrounding the specific CG model and the type and movement of the specific CG model. The technology disclosed in Patent Document 1 learns the generated training data to generate a recognition model that recognizes whether a recognition target is a person or a vehicle corresponding to the specific CG model.

[0004] Japanese Patent Application Laid-Open No. 2019-23858

[0005] However, the technology disclosed in the above-mentioned Patent Document 1 is only intended to recognize whether the recognition target is a person or a vehicle itself, and is not intended to recognize interactions that occur between a person and an object, such as a person handling the object.

[0006] For this reason, the recognition model disclosed in Patent Document 1 cannot recognize interactions between people and objects, and even if the training data disclosed in Patent Document 1 is learned, a recognition model capable of recognizing such interactions cannot be generated.

[0007] The present disclosure has been made in consideration of the above circumstances, and aims to provide a training data generation device, an annotation information display device, a training data generation method, an annotation information display method, and a program that are capable of generating training data useful for training a recognition model for recognizing interactions between people and objects.

[0008] One aspect of the training data generation device of the present disclosure is a training data generation device including one or more processors, wherein the one or more processors generate a video including a human model and a first object model, generate annotation information according to an interaction between the human model and the first object model, and add the annotation information to the video to generate training data.

[0009] One aspect of the annotation information display device of the present disclosure is an annotation information display device that includes one or more processors, wherein the one or more processors display a video including a human model and an object model in a video display area, display annotation information indicating an interaction between the human model and the object model in the annotation information display area, and further display the annotation information superimposed on the video.

[0010] In one aspect of the training data generation method of the present disclosure, one or more processors generate a video including a person model and a first object model, generate annotation information according to an interaction between the person model and the first object model, and add the annotation information to the video to generate training data.

[0011] One aspect of the annotation information display method of the present disclosure is that one or more processors display a video including a human model and an object model in a video display area, display annotation information indicating an interaction between the human model and the object model in an annotation information display area, and further display the annotation information superimposed on the video.

[0012] One aspect of the program of the present disclosure is for causing a computer to execute a moving image generation step of generating a moving image including a human model and a first object model, an annotation information generation step of generating annotation information according to an interaction between the human model and the first object model, and a training data generation step of generating training data by adding the annotation information to the moving image.

[0013] One aspect of the program of the present disclosure causes a computer to execute a video display step of displaying a video including a human model and an object model in a video display area, and an annotation information display step of displaying annotation information indicating an interaction between the human model and the object model in an annotation information display area, wherein in the video display step, the annotation information is further displayed superimposed on the video.

[0014] According to the present disclosure, it is possible to provide a training data generation device, an annotation information display device, a training data generation method, an annotation information display method, and a program that are capable of generating training data useful for training a recognition model for recognizing interactions between people and objects.

[0015] FIG. 1 is a block diagram showing an example of the hardware configuration of a training data generation device of this embodiment. FIG. 2 is a block diagram showing an example of the functional configuration of the training data generation device of this embodiment. FIG. 3 is a diagram showing an example of a human model registration screen of this embodiment. FIG. 4 is a diagram showing an example of an object model registration screen of this embodiment. FIG. 5 is a diagram showing an example of an object model registration screen of this embodiment. FIG. 6 is a diagram showing an example of a motion registration screen of this embodiment. FIG. 7 is a diagram showing an example of an action class registration screen of this embodiment. FIG. 8 is a diagram showing an example of an action class registration screen of this embodiment. FIG. 9 is a diagram showing an example of an action class registration screen of this embodiment. FIG. 10 is a diagram showing an example of a video generation screen of this embodiment. FIG. 11 is an explanatory diagram of an example of an interaction detection method of this embodiment. FIG. 12 is a diagram showing an example of annotation information of this embodiment. FIG. 13 is a diagram showing an example of a video generated by the video generation unit of this embodiment. FIG. 14 is a diagram showing an example of a confirmation screen displayed by the display control unit of this embodiment. FIG. 15 is a diagram showing an example of a display setting screen of this embodiment. FIG. 16 is a diagram showing an example of a confirmation screen after changing the display settings of the timeline of this embodiment. FIG. 17 is a flowchart showing an example of a training data generation process performed by the training data generation device of this embodiment. FIG. 18 is a flowchart showing an example of the process of checking and correcting annotation information shown in step S125 of the flowchart in FIG.

[0016] Hereinafter, an embodiment of the present disclosure (hereinafter simply referred to as "the present embodiment") will be described in detail with reference to the drawings. Note that the present disclosure is not limited to the following embodiment. Furthermore, the following embodiment and modified examples can be combined as appropriate.

[0017] The training data generation device of this embodiment reproduces interactions between people and objects that may occur in a real environment in a virtual space using person models corresponding to the people and object models corresponding to the objects. In this case, the training data generation device of this embodiment generates videos capturing the person models and object models from a predetermined viewpoint in the virtual space, and detects interactions between the person models and object models in the virtual space. The training data generation device of this embodiment generates training data by adding interaction detection results as annotation information to the generated videos.

[0018] In this manner, in this embodiment, training data is generated by adding annotation information indicating the content of the interaction to a video containing an interaction between a person model and an object model. Therefore, according to this embodiment, training data useful for training a recognition model for recognizing interactions between a person and an object can be generated. Furthermore, according to this embodiment, the generation and addition of annotation information can be automated, allowing training data to be generated efficiently.

[0019] Furthermore, the training data generation device of this embodiment defines an action class in association with an object model, and detects the action content indicated by the action class as an interaction when a human model comes into contact with an object model. In this way, this embodiment can automate the detection of interactions, thereby reducing annotation costs even in situations where the annotation costs would be high if annotation information indicating the presence or absence of an interaction on a frame-by-frame basis were to be generated.

[0020] The following description will be given taking as an example a case where a training data generation device (an example of an annotation information display device) of this embodiment uses CG (Computer Graphics) to generate training data used to train a recognition model that recognizes worker tasks (interactions between the worker and the objects used in the tasks) that occur in a logistics process. However, the training data of this embodiment and the recognition model generated by training the training data are not limited to this. The recognition model of this embodiment may be any trained model that recognizes interactions between a person and an object. Similarly, the training data of this embodiment may be any data that can be used to train a recognition model that recognizes interactions between a person and an object.

[0021] 1 is a block diagram showing an example of the hardware configuration of a training data generation device 10 of this embodiment. As shown in FIG. 1, the training data generation device 10 includes a control device 11, a main memory device 13, an auxiliary memory device 15, a display device 17, an input device 19, a communication device 21, and various buses 23. The control device 11, the main memory device 13, the auxiliary memory device 15, the display device 17, the input device 19, and the communication device 21 are connected via the various buses 23. As such, the training data generation device 10 of this embodiment has a general hardware configuration using a typical computer.

[0022] The control device 11 controls the overall operation of the training data generation device 10. The control device 11 may be, for example, at least one of a central processing unit (CPU) and a graphics processing unit (GPU), but is not limited to these. There may be any number of CPUs or GPUs as long as they are one or more, and they may be single-core or multi-core.

[0023] Examples of the main storage device 13 include, but are not limited to, a read-only memory (ROM) and a random access memory (RAM). The ROM stores various programs, such as a program for controlling the training data generation device 10 and a program for generating training data according to this embodiment. The RAM is used as a working area when the control device 11 performs various controls based on the programs stored in the ROM.

[0024] The auxiliary storage device 15 stores the various programs described above and data for generating training data in this embodiment. The various programs described above may be stored in at least one of the main storage device 13 and the auxiliary storage device 15. Examples of the auxiliary storage device 15 include, but are not limited to, existing storage devices capable of magnetic, electrical, or optical storage, such as a hard disk drive (HDD), a solid state drive (SSD), and a digital versatile disc (DVD). The auxiliary storage device 15 may be built into the training data generation device 10 or may be externally attached to the training data generation device 10 via an interface such as a universal serial bus (USB). The auxiliary storage device 15 may also be a network-attached storage (NAS) connected via a network such as a local area network (LAN) or a wide area network (WAN).

[0025] The display device 17 displays various screens used when the training data generation device 10 generates training data, and serves as a user interface with the user (operator). Examples of the display device 17 include, but are not limited to, various displays such as a liquid crystal display, an organic electroluminescence (EL) display, and a touch panel display. The display device 17 may be an internal display built into the training data generation device 10, or an external display connected to the training data generation device 10 via a display interface such as HDMI (registered trademark).

[0026] The input device 19 is used for various inputs used when the training data generation device 10 generates training data, and serves as a user interface with the user (operator). Examples of the input device 19 include, but are not limited to, a keyboard, a mouse, and a touch panel. The input device 19 may be built into the training data generation device 10 or may be externally attached to the training data generation device 10 via an interface such as a USB.

[0027] Examples of the communication device 21 include, but are not limited to, a communication device for a wired LAN and a wireless communication device for a wireless LAN. The communication device 21 may be used to externally acquire the program or data for generating the training data of this embodiment, or may be used to externally output the training data generated by the training data generation device 10.

[0028] In addition to the above configuration, the training data generation device 10 may further include hardwired circuits specific to the training data generation device 10, such as an IC (Integrated Circuit), an ASIC (Application Specific Integrated Circuit), and an FPGA (Field-Programmable Gate Array), in order to realize the training data generation function.

[0029] 2 is a block diagram showing an example of the functional configuration of the training data generation device 10 of this embodiment. As shown in FIG. 2, the training data generation device 10 includes a model storage unit 101, a motion storage unit 103, a model registration unit 111, a motion registration unit 113, an action class registration unit 115, an environment setting unit 117, a virtual space control unit 121, a confirmation unit 131, and a training data generation unit 141. The virtual space control unit 121 includes an interaction detection unit 123, a video generation unit 125, a position detection unit 127, and an annotation information generation unit 129. The confirmation unit 131 includes a display control unit 133 and an annotation information correction unit 135.

[0030] The model storage unit 101 and the motion storage unit 103 can be realized by, for example, at least one of the main storage unit 13 and the auxiliary storage unit 15 described with reference to FIG.

[0031] The model registration unit 111, the motion registration unit 113, the action class registration unit 115, the environment setting unit 117, the virtual space control unit 121, the interaction detection unit 123, the video generation unit 125, the position detection unit 127, the annotation information generation unit 129, the confirmation unit 131, the display control unit 133, the annotation information correction unit 135, and the learning data generation unit 141 can be realized, for example, by the control device 11 and the main memory device 13 described in FIG. 1.

[0032] For example, the control device 11 reads out a program for generating learning data according to this embodiment that is stored in the main storage device 13 (ROM) or the auxiliary storage device 15, or that is externally acquired from the communication device 21 via a network, and loads the program into the main storage device 13 (RAM). The control device 11 executes various processes in accordance with the loaded program, thereby realizing each of the above-described functional units. Here, the description has been given taking an example in which each of the above-described functional units is realized as software, but at least a portion of each of the above-described functional units may also be realized as hardware. In this case, the functional units realized as hardware may be realized, for example, by the above-described hardwired circuit. Furthermore, any of the above-described functional units may also be realized by a combination of software and hardware.

[0033] The model storage unit 101 stores character models and object models to be made to appear in the virtual space. In this embodiment, the virtual space is described as a three-dimensional CG space, but the present invention is not limited to this. Examples of character models include three-dimensional model data in which a character is reproduced using CG, but the present invention is not limited to this and may also use two-dimensional model data. Examples of object models include three-dimensional model data in which an object handled by a character is reproduced using CG, but the present invention is not limited to this and may also use two-dimensional model data.

[0034] Both the character model and the object model may be model data created based on a real person or object, or may be model data created based on a fictional person or object. Furthermore, both the character model and the object model may be model data created in a CG environment by a developer or the like of the training data generation device 10 (hereinafter simply referred to as "developer"), or may be model data prepared in advance on a platform that constructs a CG environment.

[0035] The motion storage unit 103 stores motion data for operating the human model or object model stored in the model storage unit 101 in a virtual space. The motion data may be data created in a CG environment by the developer of the training data generation device 10, or may be data prepared in advance on a platform that builds the CG environment. The motion data may also be data generated by capturing the movements of actual human figures or objects using motion capture technology.

[0036] The model registration unit 111 registers character models and object models to be made to appear in the virtual space. Specifically, the model registration unit 111 registers character models and object models selected by a developer from the character models and object models stored in the model storage unit 101 in the virtual space control unit 121.

[0037] For example, when a developer performs an operation input requesting the display of a character model registration screen using the input device 19, the model registration unit 111 displays the character model registration screen on the display device 17. The character model registration screen is a UI (User Interface) screen for registering a character model.

[0038] 3 is a diagram showing an example of a character model registration screen 201 according to this embodiment. As shown in FIG. 3, the character model registration screen 201 includes a pull-down list 211, a display area 213, and a register button 215. The pull-down list 211 is an operation button for selecting a character model. The display area 213 is a display area in which the character model selected in the pull-down list 211 is displayed. The register button 215 is an operation button for registering the character model selected in the pull-down list 211.

[0039] 3 , in order to select a person model 221, which is a CG model of a worker, the developer uses the input device 19 to perform an operational input to select a file with the file name "PersonA.fbx" on the pull-down list 211. As a result, the model registration unit 111 retrieves the person model 221 with the file name "PersonA.fbx" from the model storage unit 101 and displays it in the display area 213. In this situation, when the developer performs an operational input to press the registration button 215 using the input device 19, the model registration unit 111 registers the person model 221 in the virtual space control unit 121 as a person model to appear in the virtual space.

[0040] In this embodiment, the case where the number of character models appearing in the virtual space (to be registered) is one will be described as an example, but the present invention is not limited to this and the number of character models may be multiple. When multiple character models appear in the virtual space, for example, the character model registration process described in the character model registration screen 201 shown in FIG. 3 may be performed for each character model appearing in the virtual space (to be registered).

[0041] Furthermore, for example, when a developer performs an operation input requesting the display of an object model registration screen using the input device 19, the model registration unit 111 displays the object model registration screen on the display device 17. The object model registration screen is a UI screen for registering an object model.

[0042] 4 and 5 are diagrams showing an example of an object model registration screen 251 according to this embodiment. As shown in FIGS. 4 and 5, the object model registration screen 251 includes a pull-down list 261, a display area 263, and a register button 265. The pull-down list 261 is an operation button for selecting an object model. The display area 263 is a display area in which the object model selected in the pull-down list 261 is displayed. The register button 265 is an operation button for registering the object model selected in the pull-down list 261.

[0043] 4 , in order to select object model 271, which is a CG model of scissors, the developer uses the input device 19 to perform an operational input to select a file with the file name "scissors.fbx" on the pull-down list 261. As a result, the model registration unit 111 retrieves the object model 271 with the file name "scissors.fbx" from the model storage unit 101 and displays it in the display area 263. In this situation, when the developer uses the input device 19 to perform an operational input to press the registration button 265, the model registration unit 111 registers the object model 271 in the virtual space control unit 121 as an object model to appear in the virtual space.

[0044] 5, in order to select an object model 273, which is a CG model of paper, the developer selects a file with the file name "paper.fbx" on the pull-down list 261. As a result, the model registration unit 111 retrieves the object model 273 with the file name "paper.fbx" from the model storage unit 101 and displays it in the display area 263. In this situation, when the developer presses the registration button 265, the model registration unit 111 registers the object model 273 in the virtual space control unit 121 as an object model to appear in the virtual space.

[0045] In this embodiment, an example will be described in which there are a plurality of object models (to be registered) appearing in the virtual space, but the present invention is not limited to this and there may be only one object model.

[0046] The motion registration unit 113 registers motion data for operating each of the human model and object model registered by the model registration unit 111. Specifically, the motion registration unit 113 registers, in the virtual space control unit 121, motion data selected by the developer from the motion data stored in the motion storage unit 103 for operating each of the human model and object model registered by the model registration unit 111.

[0047] For example, when a developer performs an operation input using the input device 19 to request the display of a motion registration screen, the motion registration unit 113 displays the motion registration screen on the display device 17. The motion registration screen is a UI screen for registering motion data for each of the human model and object model registered by the model registration unit 111.

[0048] Fig. 6 is a diagram showing an example of a motion registration screen 301 according to this embodiment. Fig. 6 shows the motion registration screen 301 in a case where a person model 221, an object model 271, and an object model 273 have been registered by the model registration unit 111 as person models and object models (to be registered) to appear in the virtual space. As shown in Fig. 6, the motion registration screen 301 includes a pull-down list 311, a register button 313, a pull-down list 321, a register button 323, a pull-down list 331, and a register button 333.

[0049] The pull-down list 311 is an operation button for selecting motion data for the person model 221 (Person A). The register button 313 is an operation button for registering the motion data selected in the pull-down list 311. The pull-down list 321 is an operation button for selecting motion data for the object model 271 (scissors). The register button 323 is an operation button for registering the motion data selected in the pull-down list 321. The pull-down list 331 is an operation button for selecting motion data for the object model 273 (paper). The register button 333 is an operation button for registering the motion data selected in the pull-down list 331.

[0050] 6, the developer uses the input device 19 to perform an operation input to select a file with the file name "PersonA-cut.anim" on the pull-down list 311. In this situation, when the developer uses the input device 19 to perform an operation input to press the registration button 313, the motion registration unit 113 registers the motion data with the file name "PersonA-cut.anim" in the virtual space control unit 121 as motion data for operating the person model 221.

[0051] Similarly, a file with the file name "scissors-cut.anim" is selected by the developer on pull-down list 321. In this situation, when the developer presses register button 323, the motion registration unit 113 registers the motion data with the file name "scissors-cut.anim" in the virtual space control unit 121 as motion data for operating the object model 271. Similarly, a file with the file name "paper-cut.anim" is selected by the developer on pull-down list 331. In this situation, when the developer presses register button 333, the motion registration unit 113 registers the motion data with the file name "paper-cut.anim" in the virtual space control unit 121 as motion data for operating the object model 273.

[0052] The motion data with the file name "PersonA-cut.anim" is data for causing the person model 221 to perform an action of cutting the object model 273 (paper) using the object model 271 (scissors). The file name "scissors-cut.anim" is data for causing the object model 271 (scissors) to perform an action of cutting the object model 273 (paper) when handled by the person model 221. The file name "paper-cut.anim" is data for causing the object model 273 (paper) to perform an action of being cut by the object model 271 (scissors).

[0053] The action class registration unit 115 registers action classes defined in association with the object models registered by the model registration unit 111. An action class indicates the content of an interaction that occurs between an object model for which the action class is defined and a person model. An action class is an example of content information that defines the content that occurs in an object model when the person model comes into contact with the object model. Registering an action class and defining the action class in association with an object model is an example of associating content information with an object model. Note that an action class may indicate the content of an interaction that occurs between a person model and a single object model, or may indicate the content of an interaction that occurs between a person model and multiple object models.

[0054] For example, when a developer performs an operation input requesting the display of an action class registration screen using the input device 19, the action class registration unit 115 displays the action class registration screen on the display device 17. The action class registration screen is a UI screen for registering an action class.

[0055] 7 to 9 are diagrams showing examples of an action class registration screen 401 according to this embodiment. The action class registration screen 401 shown in Fig. 7 shows a case where an action class is defined and registered in association with a single object model, while the action class registration screen 301 shown in Figs. 8 and 9 shows a case where an action class is defined and registered in association with a plurality of object models.

[0056] As shown in FIG. 7, the action class registration screen 301 includes a pull-down list 411 for selecting the number of object models, a pull-down list 421 for selecting an object model, a display area 423, a text box 441, and a registration button 443.

[0057] The pull-down list 411 is an operation button for selecting the number (type) of object models for which an action class is defined. The pull-down list 421 is an operation button for selecting an object model for which an action class is defined. The display area 423 is a display area in which the object model selected in the pull-down list 421 is displayed. The text box 441 is an input field for inputting an action class. The register button 343 is an operation button for defining and registering the action class input in the text box 341 in association with the object model selected in the pull-down list 321.

[0058] In the example shown in Fig. 7, the developer uses the input device 19 to input an operation to set the number of object models for which an action class is defined to "1" on the pull-down list 411. Therefore, on the action class registration screen 401 shown in Fig. 7, the pull-down list and display area for the object models are a single set consisting of a pull-down list 421 and a display area 423. Note that when there is one object model for which an action class is defined, it is assumed that this object model will be possessed and handled (directly handled) by a human model, and therefore on the action class registration screen 401 shown in Fig. 7, this object model is referred to as a "possessive item."

[0059] 7 , it is assumed that an object model 271 registered by the model registration unit 111 is selected as an object model (possessive item) for which an action class is defined. For this purpose, the developer uses the input device 19 to perform an operation input to select a file with the file name "scissors.fbx" on the pull-down list 421. As a result, the action class registration unit 115 obtains the object model 271 with the file name "scissors.fbx" from the model registration unit 111 and displays it in the display area 423.

[0060] 7 , the developer uses the input device 19 to perform an operation input to input "scissors grip" into the text box 441 as an action class defined in association with the object model 271. In this situation, when the developer uses the input device 19 to perform an operation input to press the register button 443, the action class registration unit 115 defines the action class "scissors grip" in association with the object model 271 and registers it in the virtual space control unit 121.

[0061] Next, an example of defining and registering an action class in association with a plurality of object models will be described with reference to Fig. 8. As shown in Fig. 8, an action class registration screen 401 includes a pull-down list 411 for selecting the number of object models, pull-down lists 421 and 425 for selecting an object model, display areas 423 and 427, a text box 441, and a register button 443.

[0062] In the example shown in FIG. 8 , the developer uses the input device 19 to input an operation to set the number of object models for which an action class is defined to "2" on the pull-down list 411. Therefore, on the action class registration screen 401 shown in FIG. 8 , the pull-down list and display area for the object models are divided into two sets: a pull-down list 421 and a display area 423, and a pull-down list 425 and a display area 427. Note that when there are multiple object models for which an action class is defined, the multiple object models are assumed to include the aforementioned possessed items and object models that are handled (indirectly handled) by the person model via the possessed items and are affected by the possessed items. Therefore, on the action class registration screen 401 shown in FIG. 8 , the latter object models are referred to as "acting items."

[0063] 8, as in FIG. 7, a file with the file name "scissors.fbx" is selected as a possessed item in pull-down list 421, and an object model 271 is displayed in display area 423. Furthermore, in the action class registration screen 401 shown in FIG. 8, "2" is selected in pull-down list 411, and therefore, unlike the action class registration screen 401 shown in FIG. 7, a pull-down list 425 and a display area 427 for action items are included. In the example shown in FIG. 8, the developer uses the input device 19 to perform an operation input on pull-down list 425 to select a file with the file name "paper.fbx" as an object model (action item) for which an action class is defined. As a result, the action class registration unit 115 obtains an object model 273 with the file name "paper.fbx" from the model registration unit 111 and displays it in display area 427.

[0064] 8 , the developer uses the input device 19 to perform an operation input to input “disconnect” into the text box 341 as an action class defined in association with the object model 271 and the object model 273. In this situation, when the developer uses the input device 19 to perform an operation input to press the register button 443, the action class registration unit 115 defines the action class “disconnect” in association with the object model 271 and the object model 273, and registers it in the virtual space control unit 121.

[0065] Next, another example of registering an action class defined in association with a plurality of object models will be described with reference to Fig. 9. As shown in Fig. 9, an action class registration screen 401 includes a pull-down list 411 for selecting the number of object models, pull-down lists 421, 425, and 429 for selecting an object model, display areas 423, 427, and 431, a text box 441, and a register button 443.

[0066] 9, an operation input for setting the number of object models for which an action class is defined to "3" is performed on pull-down list 411. Therefore, on action class registration screen 401 shown in FIG. 9, the pull-down list and display area for object models are divided into three sets: pull-down list 421 and display area 423, pull-down list 425 and display area 427, and pull-down list 429 and display area 431.

[0067] 9, a file with the file name "screwdriver.fbx" is selected on the pull-down list 421 as a possessed item, and an object model 281 with the file name "screwdriver.fbx" is displayed in the display area 423. A file with the file name "bolt.fbx" is selected on the pull-down list 425 as an action item, and an object model 283 with the file name "bolt.fbx" is displayed in the display area 427. A file with the file name "PC.fbx" is selected on the pull-down list 429 as another action item that is affected by the action of the object model 283, and an object model 285 with the file name "PC.fbx" is displayed in the display area 431. It is assumed that the object models 281, 283, and 285 are object models that have already been registered by the model registration unit 111.

[0068] 9 , an operation input is performed to input “assembly” into text box 341 as an action class defined in association with object models 281, 283, and 285. In this situation, when the developer performs an operation input to press registration button 443 using input device 19, action class registration unit 115 defines the action class “assembly” in association with object model 281, object model 283, and object model 285, and registers it in the virtual space control unit 121.

[0069] The environment setting unit 117 performs various settings related to the environment of the virtual space. For example, the environment setting unit 117 sets a background image and a background model to be placed in the virtual space to the virtual space control unit 121. In this embodiment, the background image and background model may be, for example, a workbench on which a worker works, but are not limited to this. Furthermore, for example, the environment setting unit 117 sets parameters related to the angle of view of a virtual camera placed in the virtual space and parameters related to lighting placed in the virtual space to the virtual space control unit 121.

[0070] The virtual space control unit 121 performs various controls related to the virtual space. Specifically, the virtual space control unit 121 generates the virtual space, controls the movements of human models and object models in the virtual space, detects interactions between human models and object models, photographs the human models and object models in the virtual space, and generates annotation information.

[0071] For example, the virtual space control unit 121 generates a virtual space and places in the generated virtual space the person model and object model registered by the model registration unit 111, the background image and background model set by the environment setting unit 117, a virtual camera that captures images of the virtual space, and a light source. The virtual space control unit 121 also adjusts the angle of view of the virtual camera using parameters related to the angle of view of the virtual camera set by the environment setting unit 117, and adjusts the brightness of the light source and the like using parameters related to lighting set by the environment setting unit 117.

[0072] For example, in the virtual space where the above-described environmental settings have been applied, the virtual space control unit 121 operates each of the human model and the object model registered by the model registration unit 111 using the motion data registered by the motion registration unit 113. Furthermore, when the virtual space control unit 121 detects contact between a human model and an object model in motion in the virtual space, it detects, as an interaction, an action indicated by the action class registered in association with the object model by the action class registration unit 115. Furthermore, the virtual space control unit 121 generates a video capturing the human model and the object model from a virtual camera and detects the positions of the human model and the object model in the generated video. The virtual space control unit 121 generates annotation information by integrating the detected interaction and the positions of the human model and the object model in the video.

[0073] For example, when a developer uses the input device 19 to perform an operation input requesting the display of a video generation screen, the virtual space control unit 121 displays the video generation screen on the display device 17. The video generation screen is a UI screen for starting the generation of the above-described video and annotation information.

[0074] Fig. 10 is a diagram showing an example of a moving image generation screen 501 of this embodiment. As shown in Fig. 10, the moving image generation screen 501 includes a text box 511 for inputting the moving image name of the moving image to be generated, and a rendering start button 521 for starting generation of the moving image and annotation information. Note that, on the moving image generation screen 501 shown in Fig. 10, "00:30" (30 seconds) is fixedly displayed as the moving image length indicating the length of the moving image, but this is automatically set according to the playback time of the motion data registered by the motion registration unit 113.

[0075] 10 , the developer uses the input device 19 to perform an operation input to input the moving image name "cut.mp4" into the text box 511. In this situation, when the developer uses the input device 19 to perform an operation input to press the rendering start button 521, the virtual space control unit 121 generates a virtual space, performs the above-mentioned environmental settings, and starts generating the moving image and annotation information.

[0076] The following describes how video and annotation information are generated in a virtual space using the interaction detection unit 123, video generation unit 125, position detection unit 127, and annotation information generation unit 129 included in the virtual space control unit 121.

[0077] As described above, the virtual space control unit 121 operates each of the human model and object model registered by the model registration unit 111 in the virtual space using the motion data registered by the motion registration unit 113. The interaction detection unit 123 detects interactions that may occur between a human model and an object model in motion in the virtual space. Specifically, the interaction detection unit 123 detects interactions between a human model and an object model for each frame. Note that, in this embodiment, an example will be described in which the frame rate is 10 fps (frames per second), but the present invention is not limited to this.

[0078] In this embodiment, the interaction detection unit 123 detects the presence or absence of contact between a human model and an object model on a frame-by-frame basis. When the interaction detection unit 123 detects contact between a human model and an object model, the interaction detection unit 123 detects, as an interaction, an action indicated by an action class associated with the object model and registered by the action class registration unit 115. On the other hand, when the interaction detection unit 123 does not detect contact between a human model and an object model, it detects "no interaction."

[0079] The presence or absence of contact between a person model and an object model can be achieved by using a collision function (collision detection function) that is provided in advance on the platform that creates the CG environment. The collision function sets a collider, which is a transparent, three-dimensional object such as a rectangular parallelepiped or sphere that surrounds the person model and the object model, and detects the presence or absence of contact between the person model and the object model based on the presence or absence of contact between the colliders.

[0080] A collider set on a person model is an example of a person model area for interaction detection. A collider set on an object model is an example of an object model area for interaction detection. The interaction detection unit 123 detects an interaction based on whether or not the colliders of the person model and the object model are in contact with each other. Specifically, when the colliders of the person model and the object model are in contact with each other, the interaction detection unit 123 determines that the person model and the object model are in contact with each other, and detects, as an interaction, an action indicated by an action class registered in association with the object model. Note that detecting, as an interaction, an action indicated by an action class registered in association with the object model is an example of detecting, as an interaction, the content indicated by content information associated with the object model. On the other hand, when the colliders of the person model and the object model are not in contact with each other, the interaction detection unit 123 determines that the person model and the object model are not in contact with each other, and detects "no interaction."

[0081] FIG. 11 is an explanatory diagram of an example of the interaction detection method of this embodiment. The example shown in FIG. 11 chronologically illustrates the motion transitions and presence or absence of interactions in the virtual space of the person model 221, the object model 271, the object model 273, and the object model 275. For ease of explanation, in the example shown in FIG. 11 , the person model 221 is displayed with only the hand model 222, not the entire body. The object model 275 is a CG model of a pen. In the example shown in FIG. 11 , an action class "scissor grip" is defined in association with the object model 271 (see FIG. 7 ), an action class "cut" is defined in association with the object models 271 and 273 (see FIG. 8 ), and an action class "pen hold" is defined in association with the object model 275.

[0082] 11 , a collider 522 is set for the hand model 222, a collider 571 is set for the object model 271, a collider 573 is set for the object model 273, and a collider 575 is set for the object model 275. Note that in the example shown in Fig. 11 , for convenience of explanation, the hand model 222, the object model 271, the object model 273, and the object model 275 are shown in two dimensions, but in reality they are three-dimensional models. Similarly, the colliders 522, 571, 573, and collider 575 are shown as visible rectangles, but in reality they are transparent (invisible) rectangular parallelepiped objects.

[0083] First, at time 0 s, the collider 522 of the hand model 222 is not in contact with the colliders of any of the object models. Therefore, the interaction detection unit 123 does not detect contact between the person model 221 and any of the object models, resulting in "no interaction." As a result, the interaction detection unit 123 associates "no interaction" with each object model as the interaction detection result at time 0 s (frame).

[0084] Next, the virtual space control unit 121 moves the hand model 222 so as to approach the object model 271. As a result, at the time point of 10 s, the collider 522 of the hand model 222 comes into contact with the collider 571 of the object model 271, and the interaction detection unit 123 detects contact between the person model 221 and the object model 271, i.e., the presence of an interaction. The interaction detection unit 123 detects "scissors grip" as the interaction between the person model 221 and the object model 271 because the action class defined in association with the object model 271 for which contact was detected is "scissors grip."

[0085] As a result, the interaction detection unit 123 associates "scissors grip" with object model 271 and "no interaction" with object models 273 and 275 as the interaction detection result at the 10 s point (frame). Note that collider 522 of hand model 222 is not in contact with the colliders of any of the object models up until 10 s. Therefore, in the interaction detection results for each frame from 0 s to 9.9 s, "no interaction" is associated with each object model.

[0086] Next, the virtual space control unit 121 moves the hand model 222 and the object model 271 while causing the hand model 222 to hold the object model 271 so that the object model 271 approaches the object model 273. As a result, at the time point of 20 s, the collider 522 of the hand model 222 comes into contact with the collider 571 of the object model 271, and the collider 571 of the object model 271 comes into contact with the collider 573 of the object model 273. Therefore, the interaction detection unit 123 detects contact between the person model 221 and the object model 271, and contact between the object model 271 and the object model 273 (indirect contact between the person model 221 and the object model 273 via the object model 271). The interaction detection unit 123 detects "disconnection" as the interaction between the person model 221, the object model 271, and the object model 273, because the action class defined in association with the object model 271 and the object model 273 where contact was detected is "disconnection."

[0087] As a result, the interaction detection unit 123 associates "disconnection" with object models 271 and 273, and "no interaction" with object model 275, as the interaction detection result at the 20 s point (frame). Note that collider 522 of hand model 222 is in contact only with collider 571 of object model 271 from 10 s to 19.9 s. Therefore, in the interaction detection results for each frame from 10 s to 19.9 s, "scissor grip" is associated with object model 271, and "no interaction" is associated with object models 273 and 275.

[0088] Next, the virtual space control unit 121 causes the hand model 222 to release its grip on the object model 271 and moves the hand model 222 so as to approach the object model 275. As a result, at the time point of 30 s, the collider 522 of the hand model 222 comes into contact with the collider 575 of the object model 275, and the interaction detection unit 123 detects contact between the person model 221 and the object model 275. The action class defined in association with the object model 275 for which contact has been detected is "pen grip," and therefore the interaction detection unit 123 detects "pen grip" as the interaction between the person model 221 and the object model 275.

[0089] As a result, the interaction detection unit 123 associates "pen grip" with the object model 275 and associates "no interaction" with the object models 271 and 273 as the interaction detection result at the time point (frame) of 30 s. Note that the collider 522 of the hand model 222 is not in contact with the colliders of any of the object models from 20 s until the hand model 222 releases its grip on the object model 271. Therefore, in the interaction detection results for each frame from 20 s until the hand model 222 releases its grip on the object model 271, "disconnection" is associated with the object models 271 and 273, and "no interaction" is associated with the object model 275. In addition, in the interaction detection results for each frame from 20 s until 29.9 s after the hand model 222 releases its grip on the object model 271, "no interaction" is associated with each object model.

[0090] The video generation unit 125 generates a video including a person model and one or more object models. Specifically, the video generation unit 125 controls a virtual camera disposed in a virtual space and generates a video capturing a person model and an object model from the virtual camera. For example, the video generation unit 125 generates an image by having the virtual camera project (render) various models, such as person models and object models, and backgrounds, included in the angle of view (field of view) of the virtual camera onto a projection surface (not shown). The rendering converts the world coordinate system, which is a coordinate system within the virtual space, into a camera coordinate system with the viewpoint position, which is the position where the virtual camera is disposed in the virtual space, as the origin, and then converts the various models, backgrounds, etc. represented in the camera coordinate system into a coordinate system on a two-dimensional projection surface by perspective projection transformation. Note that a well-known rendering method, such as the Z-buffer method, may be used. The video generation unit 125 generates a video captured within the virtual space by having the virtual camera perform the above-described image generation for each frame.

[0091] In this embodiment, the virtual camera is fixedly positioned in the virtual space, but is not limited to this. For example, the virtual camera may be positioned above the virtual space, with an angle of view set to overlook the virtual space. In this embodiment, the virtual camera is positioned in the virtual space so as to correspond to the position of a camera that captures images of the worker's work status in the real environment to recognize the worker's work using a recognition model, but is not limited to this. For example, the virtual camera may be positioned in the virtual space so that the relative position between the worker and the camera in the real space corresponds to the relative position between the human model and the camera in the virtual space.

[0092] The position detection unit 127 detects the positions of the person model and the object model on the moving image generated by the moving image generation unit 125. Specifically, the position detection unit 127 detects the positions of the person model and the object model on the image each time an image is generated frame by frame by the moving image generation unit 125. Note that since the person model and the object model on the moving image are converted into two dimensions, the position detection unit 127 detects the positions of the person model and the object model in two-dimensional coordinates. In this embodiment, the position detection unit 127 detects the position of each model to be detected, such as a person model and an object model, by a rectangle surrounding the model, but is not limited to this.

[0093] For example, the position detection unit 127 may detect the position of each model of the detection target by calculating two-dimensional coordinates in the video from the three-dimensional coordinates in the camera coordinate system of each model of the detection target based on coordinate transformation information performed during rendering. Furthermore, for example, the position detection unit 127 may detect the position of each model of the detection target by performing person recognition processing or object recognition processing on the video generated by the video generation unit 125.

[0094] The annotation information generation unit 129 generates annotation information according to interactions between human models and object models. Note that the annotation information generation unit 129 may further generate annotation information according to interactions between object models. In this embodiment, the annotation information generation unit 129 generates annotation information indicating the results of interaction detection by the interaction detection unit 123. Therefore, the annotation information in this embodiment indicates the content of interactions between human models and object models for each frame. The content of the interaction may include, for example, whether or not an interaction exists, and, if an interaction exists, the action content indicated by the action class detected as the interaction.

[0095] Specifically, the annotation information generation unit 129 generates annotation information by integrating the interaction detection result by the interaction detection unit 123 and the position detection result of the human model and the object model by the position detection unit 127. Note that both the interaction detection unit 123 and the position detection unit 127 perform detection on a frame-by-frame basis. Therefore, the annotation information generation unit 129 can integrate the interaction detection result by the interaction detection unit 123 and the position detection result of the human model and the object model by the position detection unit 127 using frames as a key.

[0096] Fig. 12 is a diagram illustrating an example of annotation information according to the present embodiment. The example illustrated in Fig. 12 illustrates, as annotation information, information that integrates the detection result of an interaction in the situation in the virtual space described in Fig. 11 and the position detection result of each model on a video captured of the situation in the virtual space described in Fig. 11.

[0097] In the example shown in FIG. 12 , the annotation information is information that associates a frame, a time, rectangular coordinates of a human model, and rectangular coordinates and interactions of each object model. The interaction of each object model corresponds to the interaction detection result, and the rectangular coordinates of each model correspond to the position detection result of each model in the video. In the example shown in FIG. 12 , the human model (hand) represents the hand model 222 of the human model 221, the object model (scissors) represents the object model 271, the object model (paper) represents the object model 273, and the object model (pen) represents the object model 275. In the example shown in FIG. 12 , the rectangular coordinates represent the upper left x-y coordinates and the lower right x-y coordinates of a rectangle surrounding the model, but are not limited to this.

[0098] 11 , the interaction detection results are associated with each object model. Therefore, when the person model 221 is not in contact with any object model, as at time 0 s, the interaction of each object model is "absent." When the person model 221 is in contact with the object model 271, as at time 10 s, the interaction of the object model 271 is "present" and "scissors grip," and the interactions of the object models 273 and 275 are "absent." When the person model 221 is in contact with the object model 271 and indirectly in contact with the object model 273, as at time 20 s, the interaction of the object models 271 and 273 is "present" and "disconnected," and the interaction of the object model 275 is "absent." Also, when the person model 221 is in contact with the object model 275, as at 30 s, the interaction between the object models 271 and 273 becomes "absent," and the interaction with the object model 275 becomes "present" and "pen gripped."

[0099] When annotation information is generated by the annotation information generating unit 129, the confirmation unit 131 displays a confirmation screen on the display device 17 to prompt the developer to confirm and correct the annotation information. This allows the developer to confirm whether the content of the interaction in the generated annotation information is appropriate or not, and to correct any inappropriate parts, even when annotation information is automatically generated as in this embodiment.

[0100] In particular, in this embodiment, the confirmation unit 131 synchronizes and displays, on a frame-by-frame basis, the animation generated by the animation generation unit 125 in addition to the annotation information generated by the annotation information generation unit 129. This allows the developer to check whether the content of the interaction in the annotation information is appropriate while checking the scene in which the interaction actually occurred on the animation, thereby improving the efficiency of the confirmation work.

[0101] Hereinafter, a method for displaying annotation information and moving images, and modification of annotation information will be described using the above-described display control unit 133 and annotation information modification unit 135 included in the confirmation unit 131. Note that the method for displaying annotation information and moving images, and modification of annotation information will be described using an example different from the examples described in Fig. 11 and Fig. 12. First, the different example will be described with reference to Fig. 13.

[0102] Fig. 13 is a diagram showing an example of a video 601 generated by the video generation unit 125 of this embodiment. The video 601 shown in Fig. 13 is a video of a packing operation performed by a worker in a real environment, reproduced in a virtual space using a human model 221 and filmed. Note that the video data captured to recognize the packing operation of the worker in the real environment is expected to be captured, for example, from above the work site. For this reason, the video 601 shown in Fig. 13 is captured from a virtual camera placed above the work site in the virtual space so as to correspond to the shooting position of the video in the real environment.

[0103] 13 is a video of a scene in which a person model 221 is packing a packing box model 621 in a virtual space using a tape model 631, a tape model 633, and a tape model 635. A workbench model 611 is placed in the virtual space, and a packing box model 621, a tape model 631, a tape model 633, a tape model 635, and a replacement tape model 641 are placed on the workbench model 611.

[0104] The tape model 631, the tape model 633, and the tape model 635 are reproductions of different types of tape, and are attached to the tape cutter. The replacement tape model 641 is a reproduction of the same type of tape as the tape model 631, and is used as a replacement for the tape model 631. For example, when the tape of the tape model 631 runs out, the replacement tape model 641 is attached to the tape cutter of the tape model 631 in place of the tape model 631. The person model 221 packs the packing box model 621 on the workbench model 611 by applying tape to the packing box model 621 while switching between tape models.

[0105] 13, tape model 631, tape model 633, tape model 635, and replacement tape model 641 correspond to the above-mentioned object models, but are not limited to these. Also, in the example shown in Fig. 13, an action class "apply tape A" is defined in association with tape model 631, an action class "apply tape B" is defined in association with tape model 633, an action class "apply tape C" is defined in association with tape model 635, and an action class "refill tape A" is defined in association with tape model 631 and replacement tape model 641.

[0106] Based on the above assumptions, the video shown in Fig. 13 is assumed to include a scene in which the person model 221 applies tape to the packaging box model 621 in the order of tape model 633 (tape B), tape model 631 (tape A), and tape model 635 (tape C), and refills the tape model 631 (tape A). Note that while a detailed description of the annotation information of the video shown in Fig. 13 will be omitted, as described in Fig. 12, the annotation information is information that associates frames, time, rectangular coordinates of the person model, and rectangular coordinates and interactions of each object model.

[0107] 14 is a diagram showing an example of a confirmation screen 701 displayed on the display device 17 by the display control unit 133 of this embodiment. As shown in Fig. 14, the confirmation screen 701 includes a video display area 702, a timeline display area 703, an annotation information display area 704, and a display setting button 771.

[0108] The display control unit 133 displays the video 601 and a view angle change button 715 in the video display area 702, displays a play button 751 and a seek bar 753 in the timeline display area 703, and displays interaction information 761 as annotation information for the video 601 in the annotation information display area 704.

[0109] Furthermore, the display control unit 133 superimposes and displays the annotation information of the video 601 on the video 601 displayed in the video display area 702. Specifically, the display control unit 133 superimposes and displays, on the video 601, as annotation information, a rectangle (an example of a frame) surrounding each model and a line segment indicating the presence or absence of an interaction.

[0110] 14 , the display control unit 133 displays a rectangle 721 superimposed on the person model 221, a rectangle 731 on the tape model 631, a rectangle 733 on the tape model 633, a rectangle 735 on the tape model 635, and a rectangle 741 on the replacement tape model 641 on the video 601. Also, in the example shown in Fig. 14 , the display control unit 133 displays line segments 732, 734, 736, and 742 superimposed on the video 601, which respectively connect the rectangle 721 of the person model 221 to the rectangle 731 of the tape model 631, the rectangle 733 of the tape model 633, the rectangle 735 of the tape model 635, and the rectangle 741 of the replacement tape model 641. Also, in the example shown in Fig. 14 , the display control unit 133 displays the line segment 732 as a dotted line because there is an interaction between the person model 221 and the tape model 631. On the other hand, the display control unit 133 displays the line segments 734, 736, and 742 as solid lines because there is no interaction between the person model 221 and the tape model 633, tape model 635, and replacement tape model 641.

[0111] Specifically, when playing back any frame of the video 601, the display control unit 133 refers to the annotation information and acquires the rectangular coordinates of each model and the interaction of the object model in that frame. Based on the acquired rectangular coordinates of each model, the display control unit 133 displays the above-mentioned rectangle superimposed on the video 601. Furthermore, based on the acquired rectangular coordinates of the human model and the rectangular coordinates of each object model, the display control unit 133 displays the above-mentioned line segment superimposed on the video 601. For example, the display control unit 133 generates a line segment connecting the center coordinate of the rectangular coordinates of the human model and the center coordinate of the rectangular coordinates of each object model, and displays the line segment superimposed on the video 601. Furthermore, the display control unit 133 determines the display mode of the above-mentioned line segment based on the acquired interaction. For example, the display control unit 133 displays the line segment as a dotted line when there is an interaction, and as a solid line when there is no interaction.

[0112] In this way, in this embodiment, annotation information is displayed superimposed on the video, so that developers can check the annotation information on the video, improving the efficiency of the checking work. Furthermore, in this embodiment, the display mode of the annotation information superimposed on the video is changed depending on whether or not there is interaction, making it easier for developers to check the annotation information, and further improving the efficiency of the checking work.

[0113] In the present embodiment, an example has been described in which the display mode of annotation information is changed depending on whether an interaction is present or absent, by switching between displaying a line segment as a solid line and a dotted line. However, the present invention is not limited to this. For example, any display mode may be changed depending on whether an interaction is present or absent, such as by changing the color or thickness of the line segment. Alternatively, a predetermined line segment may be displayed only when an interaction is present.

[0114] A play button 751 displayed in the timeline display area 703 is a button for playing the video 601. A seek bar 753 displayed in the timeline display area 703 is a UI component that displays the playback position (playback frame) of the video 601 using a slider.

[0115] The interaction information 761 displayed in the annotation information display area 704 includes items indicating article models and interactions. The items include tape A corresponding to tape model 631, tape B corresponding to tape model 633, tape C corresponding to tape model 635, and replacement tape A corresponding to replacement tape model 641. The interactions include whether or not an interaction has occurred, and, if an interaction has occurred, the action (the content of the interaction resulting from contact with the human model).

[0116] In this embodiment, the display control unit 133 displays the interaction information 761 in synchronization with the video 601 displayed in the video display area 702 on a frame-by-frame basis. That is, each time a frame being played back in the video 601 is updated, the display control unit 133 refers to the annotation information, acquires the interaction of the object model in that frame, and updates the display content of the interaction information 761 to the acquired interaction content. Therefore, according to this embodiment, the developer can check whether the interaction content in the annotation information is appropriate while checking the scene in which the interaction actually occurred on the video, thereby improving the efficiency of the checking work.

[0117] Furthermore, a correction button is associated with the presence or absence of interaction and the action. When the developer uses the input device 19 to press the correction button and input the correction content, the annotation information correction unit 135 corrects the presence or absence of interaction or the action associated with the correction button to the input content and corrects the corresponding part of the annotation information. For example, the developer uses the input device 19 to press the correction button 763 associated with the presence or absence of interaction of tape A and inputs the correction content. In this case, the annotation information correction unit 135 corrects the presence or absence of interaction to the input content and corrects the corresponding part of the annotation information. Also, for example, the developer uses the input device 19 to press the correction button 765 associated with the action of tape A and inputs the correction content. In this case, the annotation information correction unit 135 corrects the action to the input content and corrects the corresponding part of the annotation information.

[0118] This allows developers to check whether the interaction content of the generated annotation information is appropriate, even when annotation information is automatically generated as in this embodiment, and to correct any parts that are not appropriate.

[0119] The display setting button 771 is a button for displaying a display setting screen for performing display settings for the confirmation screen 701. For example, when a developer performs an operation input of pressing the display setting button 771 using the input device 19, the display control unit 133 displays the display setting screen on the display device 17.

[0120] 15 is a diagram showing an example of a display setting screen 801 according to this embodiment. As shown in Fig. 15, the display setting screen 801 includes check boxes 811, 813, and 815 for configuring display settings for interactions, check boxes 817 and 819 for configuring display settings for the timeline, and an OK button 821.

[0121] In the example shown in Fig. 15, three patterns of interaction display settings can be set: "Show all," "Show only existing," and "Do not display." Note that "Show all" is set by selecting check box 811, "Show only existing" by selecting check box 813, and "Do not display" by selecting check box 815. In the example shown in Fig. 14, all interactions were displayed, so in the example shown in Fig. 15, check box 811 is selected.

[0122] In addition, in the example shown in Fig. 15, two patterns, "simple" and "detailed," can be set as the timeline display setting. "Simple" is set by selecting check box 817, and "detailed" is set by selecting check box 819. In the example shown in Fig. 14, the timeline was displayed in a simple format with only a seek bar, so in the example shown in Fig. 15, check box 817 is selected.

[0123] The OK button 821 is a button for confirming the display settings of the interactions and timeline to the contents selected in the check boxes. For example, on the display setting screen 801 shown in Fig. 15 , assume that the developer uses the input device 19 to deselect the check box 817, select the check box 819, and then press the OK button 821. In this case, the display control unit 133 switches the display content of the timeline display area 703 of the confirmation screen 701 from the simple version to the detailed version.

[0124] Fig. 16 is a diagram showing an example of a confirmation screen 701 after changing the timeline display settings of this embodiment. The confirmation screen 701 shown in Fig. 16 has the same display content as the confirmation screen 701 shown in Fig. 14 except for the display content of the timeline display area 703, although the display size of the video display area 702 and the like are different. For this reason, Fig. 16 will explain the timeline display area 703.

[0125] 16, the timeline display area 703 includes a play button 751, a play bar 754, and a timeline 781. In this way, by changing the timeline display setting from "simple" to "detailed," the seek bar 753 is switched to the timeline 781.

[0126] The timeline 781 is configured for each item (object model). Specifically, the timeline 781 is configured with timelines for tape A (tape model 631), tape B (tape model 633), tape C (tape model 635), and replacement tape A (replacement tape model 641). Each timeline displays chronological changes in interactions with the person model. In the example shown in FIG. 16 , each timeline is displayed without shading during periods without interaction and with shading corresponding to the interaction during periods with interaction, but the display format is not limited to this. Therefore, by checking the timeline 781, it can be seen that in the video 601, the person model 221 applies tape B, tape A, and tape C to the packaging box model 621 in this order, and refills tape A. The playback bar 754 displays the playback position (playback frame) of the video 601.

[0127] The learning data generation unit 141 generates learning data by adding annotation information generated by the annotation information generation unit 129 to the moving image generated by the moving image generation unit 125. Specifically, the learning data generation unit 141 generates learning data by adding annotation information confirmed by the confirmation unit 131 to the moving image generated by the moving image generation unit 125. The learning data generation unit 141 outputs the generated learning data to the auxiliary storage device 15 or to the outside via the communication device 21.

[0128] FIG. 17 is a flowchart showing an example of the training data generation process performed by the training data generation device 10 of this embodiment.

[0129] First, the model registration unit 111 registers a character model selected by the developer from among the character models stored in the model storage unit 101 in the virtual space control unit 121 (step S101).

[0130] Next, the model registration unit 111 registers one or more object models selected by the developer from the object models stored in the model storage unit 101 in the virtual space control unit 121 (step S103).

[0131] Next, the motion registration unit 113 registers in the virtual space control unit 121 the motion data selected by the developer from the motion data stored in the motion memory unit 103 in order to operate the person model and object model registered by the model registration unit 111 (step S105).

[0132] Next, the action class registration unit 115 registers an action class defined in association with the object model registered by the model registration unit 111 (step S107).

[0133] Next, the environment setting unit 117 performs various settings relating to the environment of the virtual space on the virtual space control unit 121 (step S109).

[0134] Next, the virtual space control unit 121 starts taking (generating) a video capturing the human model and object model in the virtual space and generating annotation information (step S111).

[0135] Next, the interaction detection unit 123 detects whether or not there is contact between the human model and the object model in the virtual space (step S113).

[0136] When the interaction detection unit 123 detects contact between a person model and an object model (Yes in step S113), it detects, as an interaction, an action indicated by the action class registered in association with the object model by the action class registration unit 115 (step S115).

[0137] On the other hand, if the interaction detection unit 123 does not detect contact between the human model and the object model (No in step S113), it detects "no" interaction.

[0138] Next, the video generating unit 125 controls a virtual camera placed in the virtual space, and generates an image capturing the human model and the object model from the virtual camera (step S117).

[0139] Next, the position detection unit 127 detects the positions of the human model and the object model on the image generated by the video generation unit 125 (step S119).

[0140] Next, the virtual space control unit 121 checks whether the playback of the motions that move the human model and object model in the virtual space has finished (step S121). If the playback of the motions has not finished (No in step S121), the process returns to step S113. On the other hand, if the playback of the motions has finished (Yes in step S121), the process proceeds to step S123. The processes in steps S113 to S121 are performed on a frame-by-frame basis.

[0141] Next, the annotation information generating unit 129 generates annotation information by integrating the interaction detection result by the interaction detecting unit 123 and the position detection result of the human model and object model by the position detecting unit 127 (step S123).

[0142] Next, the confirmation unit 131 displays a confirmation screen for the annotation information generated by the annotation information generation unit 129 on the display device 17, and allows the developer to confirm and correct the annotation information (step S125).

[0143] Next, the learning data generation unit 141 generates learning data by adding the annotation information confirmed by the confirmation unit 131 to the video generated by the video generation unit 125 (step S127).

[0144] FIG. 18 is a flowchart showing an example of the annotation information checking and correction process shown in step S125 of the flowchart in FIG.

[0145] First, the display control unit 133 displays a confirmation screen 701 including a moving image display area 702, a timeline display area 703, an annotation information display area 704, and a display setting button 771 on the display device 17 (step S201). The display control unit 133 displays the moving image 601 in the moving image display area 702 and the annotation information (interaction information 761) in the annotation information display area 704, and further displays the annotation information superimposed on the moving image 601.

[0146] Next, the display control unit 133 checks whether there is a request to change the display settings of the confirmation screen 701 based on whether there is an operation input to press the display setting button 771 (step S203).

[0147] When the display setting button 771 is pressed to confirm a request to change the display settings of the confirmation screen 701 (Yes in step S203), the display control unit 133 displays the display setting screen 801 on the display device 17 and accepts changes to the display settings of the interactions and the timeline (step S205). As a result, the display settings of the interactions and the timeline on the confirmation screen 701 are changed and displayed. Note that when the display setting button 771 is not pressed and the request to change the display settings of the confirmation screen 701 is not confirmed (No in step S203), the processing of step S205 is not performed.

[0148] Next, when an operation input is made to press the play button 751 on the confirmation screen 701, the display control unit 133 starts playing the video 601, updates the playback frame, and updates the display content of the interaction information 761 to that of the updated playback frame (step S207).

[0149] Next, the annotation information correction unit 135 checks whether there is a request to correct the interaction based on whether there is an interaction in the interaction information 761 and whether there is an operation input to press a correction button associated with the action (step S209). In practice, the annotation information correction unit 135 checks whether there is a request to correct the interaction when the playback of the video 601 is paused by pressing the play button 751 (pause button) during playback of the video 601.

[0150] If the request to modify the interaction is confirmed by pressing the Modify button (Yes in step S209), the annotation information modification unit 135 accepts input of the modification content for the interaction and modifies the interaction 761 (step S211). This modifies the annotation information. If the Modify button is not pressed and the request to modify the interaction is not confirmed (No in step S209), the processing of step S211 is not performed.

[0151] Next, the display control unit 133 checks whether the video 601 has been played to the end (step S213). If the video 601 has not been played to the end (No in step S213), the process returns to step S207. On the other hand, if the video 601 has been played to the end (Yes in step S213), the process ends.

[0152] In this manner, in this embodiment, training data is generated by adding annotation information indicating the content of the interaction to a video containing an interaction between a person model and an object model. Therefore, according to this embodiment, training data useful for training a recognition model for recognizing interactions between a person and an object can be generated. Furthermore, according to this embodiment, the generation and addition of annotation information can be automated, allowing training data to be generated efficiently.

[0153] In this embodiment, an action class is defined in association with an object model, and when a human model comes into contact with an object model, the action content indicated by the action class is detected as an interaction. In this way, in this embodiment, since the detection of interactions can be automated, annotation costs can be reduced even in situations where the annotation costs of generating annotation information indicating the presence or absence of interactions on a frame-by-frame basis would be high.

[0154] Furthermore, in this embodiment, once annotation information is generated, a confirmation screen is displayed to allow the developer to confirm and modify the annotation information. This allows the developer to confirm whether the content of the interaction in the generated annotation information is appropriate, even when annotation information is automatically generated as in this embodiment, and to modify any inappropriate parts.

[0155] In particular, in this embodiment, in addition to annotation information, video is displayed in synchronization with the video on a frame-by-frame basis, allowing developers to check whether the content of the interaction in the annotation information is appropriate while checking the scene in which the interaction actually occurred on the video, thereby improving the efficiency of the checking work.

[0156] In this embodiment, annotation information is also superimposed on the video, allowing developers to check the annotation information on the video, improving the efficiency of the checking work. In addition, in this embodiment, the display mode of the annotation information superimposed on the video changes depending on whether or not there is interaction, making it easier for developers to check the annotation information, further improving the efficiency of the checking work.

[0157] In addition, in this embodiment, in addition to video and annotation information, the chronological changes in interactions with human models can be displayed on a timeline for each object model, making it easier for developers to check annotation information and further improving the efficiency of the checking process.

[0158] (Variation 1) In the above embodiment, an example has been described in which a collider is used to detect contact between a person model and an object model, but the method for detecting contact between a person model and an object model is not limited to this. The interaction detection unit 123 may detect actual contact between a person model and an object model, or may detect an interaction using a positional relationship, such as the distance, between the person model and the object model. In the latter case, the interaction detection unit 123 may, for example, set reference points on the person model and the object model and detect an interaction based on the distance and positional relationship between the reference points.

[0159] (Program) The program executed by the training data generation device 10 of the above embodiment and the above modified example is provided by being stored in an installable or executable file format on a computer-readable storage medium such as a CD-ROM, CD-R, memory card, DVD, or flexible disk (FD).

[0160] The programs executed by the training data generation device 10 of the above embodiment and modified example may be stored on a computer connected to a network such as the Internet and provided by being downloaded via the network. The programs executed by the training data generation device 10 of the above embodiment and modified example may be provided or distributed via a network such as the Internet. The programs executed by the training data generation device 10 of the above embodiment and modified example may be provided by being pre-installed in a ROM or the like.

[0161] The programs executed by the training data generation device 10 of the above embodiment and the above modification have a modular configuration for implementing the above-mentioned units on a computer. In actual hardware, for example, the CPU reads the training program from the HDD onto the RAM and executes it, thereby implementing the above-mentioned units on the computer.

[0162] As described above, according to the above embodiment and the above modification, it is possible to generate learning data that is useful for training a recognition model for recognizing interactions between a person and an object.

[0163] In the above-described conventional technology, annotation information is assigned to images, but not to videos. In contrast, according to the above-described embodiment and the above-described modified example, even when annotation information is assigned to a video, the assigned annotation information can be checked efficiently.

[0164] The above-described embodiment and modifications merely illustrate examples of specific embodiments of the present disclosure, and the technical scope of the present disclosure should not be construed as being limited by these. Therefore, the present disclosure can be implemented in various forms without departing from the spirit or main features thereof. For example, the above-described embodiment and modifications may be appropriately combined in their respective constituent units. Furthermore, for example, some components may be deleted from all components in the above-described embodiment and modifications.

[0165] The present disclosure includes the following aspects.

[0166] (1) A training data generation device including one or more processors, wherein the one or more processors: generate a video including a human model and a first object model; generate annotation information according to an interaction between the human model and the first object model; and generate training data by adding the annotation information to the video.

[0167] In the above configuration (1), training data is generated by adding annotation information generated in response to an interaction to a video containing an interaction between a person model and an object model. Therefore, the above configuration (1) makes it possible to generate training data useful for training a recognition model for recognizing interactions between a person and an object. Furthermore, the above configuration (1) makes it possible to automate the generation and addition of annotation information, thereby enabling efficient generation of training data.

[0168] (2) The training data generation device according to (1), wherein the video further includes a second object model, and the one or more processors further generate the annotation information in response to an interaction between the first object model and the second object model.

[0169] In the above configuration (2), training data is generated by adding annotation information generated in response to interactions to videos containing interactions between object models. Therefore, the above configuration (2) makes it possible to generate training data useful for training a recognition model for recognizing interactions between objects. Furthermore, the above configuration (2) makes it possible to automate the generation and addition of annotation information, thereby enabling efficient generation of training data.

[0170] (3) The training data generation device according to (2), wherein the one or more processors detect the interaction for each frame, and the annotation information indicates, for each frame, the content of the interaction between the human model and the first object model or the second object model.

[0171] The configuration (3) above allows for automated detection of interactions, thereby reducing annotation costs even in situations where the annotation costs of generating annotation information indicating the presence or absence of interactions on a frame-by-frame basis would be high.

[0172] (4) The training data generation device according to (3), wherein the first object model or the second object model is associated with content information that defines content that occurs as a result of contact with the person model, and the one or more processors detect, as the interaction, content indicated by the content information when it is determined that the person model has come into contact with the first object model or the second object model.

[0173] In the above configuration (4), an action class is defined in association with an object model, and when a human model and an object model come into contact, the action content indicated by the action class is detected as an interaction. In this way, the above configuration (4) can automate the detection of interactions, so that annotation costs can be reduced even in situations where the annotation costs of generating annotation information indicating the presence or absence of interactions on a frame-by-frame basis would be high.

[0174] (5) The training data generation device according to (3), wherein a human model area for interaction detection is set in the human model, an object model area for interaction detection is set in the first object model or the second object model, and the one or more processors detect the interaction by utilizing presence or absence of contact between the human model area and the object model area.

[0175] According to the above configuration (5), contact between the human model and the first object model can be detected by simple processing, thereby reducing the processing load.

[0176] (6) The training data generation device according to (3), wherein the one or more processors detect the interaction by utilizing a positional relationship between the human model and the first object model or the second object model.

[0177] (7) An annotation information display device including one or more processors, wherein the one or more processors display a moving image including a human model and an object model in a moving image display area, display annotation information indicating an interaction between the human model and the object model in the annotation information display area, and further display the annotation information superimposed on the moving image.

[0178] In the configuration (7) above, the annotation information is superimposed on the video and displayed, so the developer can check whether or not the annotation information has been corrected on the video, thereby improving the efficiency of the checking work.

[0179] (8) The annotation information display device according to (7), wherein the one or more processors change a display mode of the annotation information superimposed on the video depending on whether or not the interaction occurs.

[0180] In the configuration (8) above, the display mode of the annotation information superimposed on the video differs depending on whether or not there is interaction, making it easier for developers to check the annotation information and further improving the efficiency of the checking work.

[0181] (9) The annotation information display device described in (8) above, wherein the one or more processors display frames surrounding the human model and the object model on the video, display lines connecting the frames on the video as the annotation information, and change the display mode of the lines depending on whether or not the interaction occurs.

[0182] In the configuration (9) above, the display mode of the annotation information superimposed on the video differs depending on whether or not there is interaction, making it easier for developers to check the annotation information and further improving the efficiency of the checking work.

[0183] (10) The annotation information display device according to (7), wherein the annotation information displayed in the annotation information display area indicates, as the interaction, content that occurs in the object model as a result of contact with the human model.

[0184] According to the above configuration (10), annotation information of a content that is difficult to superimpose on a moving image can be confirmed by viewing the annotation information displayed in the annotation information display area.

[0185] (11) The annotation information display device according to (7), wherein the video includes a plurality of the object models, and the one or more processors display a timeline in a timeline display area, the timeline displaying chronological changes in interactions between the human model and each of the plurality of object models.

[0186] According to the configuration (11) above, in addition to the video and annotation information, the time series changes in interactions with the person model can be displayed on a timeline for each object model, making it easier for developers to check the annotation information and further improving the efficiency of the checking work.

[0187] (12) The annotation information display device described in (7) above, wherein the annotation information indicates the content of the interaction between the human model and the object model for each frame, and the one or more processors update the content of the interaction indicated by the annotation information displayed in the annotation information display area to the content of the interaction in the updated frame each time a frame being played in the video is updated.

[0188] In the configuration (12) above, the annotation information and the video are displayed in synchronization on a frame-by-frame basis, so that the developer can check the scene in which the interaction actually occurred on the video and confirm whether the content of the interaction in the annotation information is appropriate, thereby improving the efficiency of the confirmation work.

[0189] (13) A training data generation method, comprising: one or more processors generating a video including a human model and a first object model; generating annotation information according to an interaction between the human model and the first object model; and adding the annotation information to the video to generate training data.

[0190] In the configuration (13), training data is generated by adding annotation information generated in response to an interaction to a video containing an interaction between a person model and an object model. Therefore, the configuration (13) can generate training data useful for training a recognition model for recognizing interactions between a person and an object. Furthermore, the configuration (13) can automate the generation and addition of annotation information, thereby enabling efficient generation of training data.

[0191] (14) An annotation information display method, comprising: one or more processors displaying a moving image including a human model and an object model in a moving image display area; displaying annotation information indicating an interaction between the human model and the object model in an annotation information display area; and further displaying the annotation information superimposed on the moving image.

[0192] In the configuration (14) above, annotation information is superimposed on the video and displayed, so that the developer can check whether or not the annotation information has been corrected on the video, thereby improving the efficiency of the checking work.

[0193] (15) A program causing a computer to execute the following steps: a moving image generating step of generating a moving image including a human model and a first object model; an annotation information generating step of generating annotation information according to an interaction between the human model and the first object model; and a learning data generating step of generating learning data by adding the annotation information to the moving image.

[0194] In the configuration (15) above, training data is generated by adding annotation information generated in response to an interaction to a video containing an interaction between a person model and an object model. Therefore, the configuration (15) above makes it possible to generate training data useful for training a recognition model for recognizing interactions between a person and an object. Furthermore, the configuration (15) above makes it possible to automate the generation and addition of annotation information, thereby enabling efficient generation of training data.

[0195] (16) A program that causes a computer to execute a moving image display step of displaying a moving image including a human model and an object model in a moving image display area, and an annotation information display step of displaying annotation information indicating an interaction between the human model and the object model in an annotation information display area, wherein in the moving image display step, the annotation information is further displayed by being superimposed on the moving image.

[0196] In the above configuration (16), the annotation information is superimposed on the video and displayed, so the developer can check whether or not the annotation information has been corrected on the video, thereby improving the efficiency of the checking work.

[0197] This application is based on the Japanese application of Patent Application No. 2024-054531 filed on March 28, 2024, and the Japanese application of Patent Application No. 2024-054537 filed on March 28, 2024, the contents of which are all incorporated by reference into this application.

[0198] REFERENCE SIGNS LIST 10 Learning data generation device 11 Control device 13 Main memory device 15 Auxiliary memory device 17 Display device 19 Input device 21 Communication device 23 Various buses 101 Model memory unit 103 Motion memory unit 111 Model registration unit 113 Motion registration unit 115 Action class registration unit 117 Environment setting unit 121 Virtual space control unit 123 Interaction detection unit 125 Video generation unit 127 Position detection unit 129 Annotation information generation unit 131 Confirmation unit 133 Display control unit 135 Annotation information correction unit 141 Learning data generation unit

Claims

1. A training data generation device comprising one or more processors, wherein the one or more processors: generate a video including a human model and a first object model; generate annotation information according to an interaction between the human model and the first object model; and add the annotation information to the video to generate training data.

2. The training data generation device according to claim 1, wherein the video further includes a second object model, and the one or more processors further generate the annotation information in response to an interaction between the first object model and the second object model.

3. The training data generation device according to claim 2, wherein the one or more processors detect the interaction for each frame, and the annotation information indicates the content of the interaction between the human model and the first object model or the second object model for each frame.

4. The training data generation device according to claim 3, wherein the first object model or the second object model is associated with content information that defines what occurs as a result of contact with the person model, and the one or more processors detect, when it is deemed that the person model has come into contact with the first object model or the second object model, the content indicated by the content information as the interaction.

5. The training data generation device according to claim 3, wherein a human model area for interaction detection is set in the human model, an object model area for interaction detection is set in the first object model or the second object model, and the one or more processors detect the interaction by utilizing the presence or absence of contact between the human model area and the object model area.

6. The training data generation device according to claim 3, wherein the one or more processors detect the interaction by utilizing a positional relationship between the human model and the first object model or the second object model.

7. An annotation information display device comprising one or more processors, wherein the one or more processors: display a video including a human model and an object model in a video display area; display annotation information showing an interaction between the human model and the object model in the annotation information display area; and further display the annotation information superimposed on the video.

8. The annotation information display device according to claim 7, wherein the one or more processors change the display mode of the annotation information superimposed on the video depending on whether or not there is an interaction.

9. The annotation information display device according to claim 8, wherein the one or more processors display frames surrounding each of the human model and the object model on the video, display lines connecting the frames on the video as the annotation information, and vary the display manner of the lines depending on whether or not the interaction occurs.

10. The annotation information display device according to claim 7, wherein the annotation information displayed in the annotation information display area indicates, as the interaction, what occurs to the object model as a result of contact with the human model.

11. An annotation information display device as described in claim 7, wherein the video includes a plurality of the object models, and the one or more processors display a timeline in a timeline display area, the timeline displaying the chronological changes in interactions between the human model and each of the plurality of object models.

12. An annotation information display device as described in claim 7, wherein the annotation information indicates the content of the interaction between the human model and the object model for each frame, and the one or more processors update the content of the interaction indicated by the annotation information displayed in the annotation information display area to the content of the interaction in the updated frame each time a frame being played in the video is updated.

13. A training data generation method, comprising: one or more processors generating a video including a human model and a first object model; generating annotation information according to an interaction between the human model and the first object model; and adding the annotation information to the video to generate training data.

14. An annotation information display method, in which one or more processors display a video including a human model and an object model in a video display area, display annotation information indicating an interaction between the human model and the object model in an annotation information display area, and further display the annotation information superimposed on the video.

15. A program for causing a computer to execute the following steps: a moving image generation step of generating a moving image including a human model and a first object model; an annotation information generation step of generating annotation information according to an interaction between the human model and the first object model; and a learning data generation step of generating learning data by adding the annotation information to the moving image.

16. A program that causes a computer to execute a moving image display step of displaying a moving image including a human model and an object model in a moving image display area, and an annotation information display step of displaying annotation information indicating an interaction between the human model and the object model in an annotation information display area, wherein in the moving image display step, the annotation information is further displayed superimposed on the moving image.

Citation Information

Patent Citations

  • Information processing apparatus, information processing method, and program

    JP2012249156A

  • Simulation system, simulation program, and simulation method

    JP2018060511A

  • Information processor, information processing method, and program

    JP2019003329A

  • Learning data generation device, learning data generation method, machine learning method, and program

    JP2019023858A

  • Teacher data generation device

    JP2021107981A