Animation generation system, animation generation method, and program
The system uses VR to capture virtual object manipulation, addressing the challenge of creating training data for recognition models by generating videos of human models operating virtual objects, ensuring precise finger movement capture and object interaction.
Patent Information
- Application Number
- JP2024080974
- Authority / Receiving Office
- JP · JP
- Patent Type
- Applications
- Current Assignee / Owner
- Filing Date
- 2024-05-17
- Publication Date
- 2025-11-28
AI Technical Summary
Conventional methods require physical manipulation of objects to generate training data for recognition models, making it difficult to create such data when the object is not available.
A system and method that uses a VR headset to capture an actor's hand movements virtually manipulating virtual objects, generating videos of a human model operating these objects, enabling training data creation without physical objects.
Enables the generation of training data for recognition models that recognize object manipulation, even without the physical presence of the object, allowing precise capture of finger movements and object interaction.
Smart Images

Figure 2025174545000001_ABST
Abstract
Description
[Technical Field]
[0001] The present disclosure relates to a video generation system, a video generation method, and a program. [Background technology]
[0002] A known method of data augmentation, which expands (increases) learning data (teacher data) for machine learning, is to generate learning data using computer graphics (CG).
[0003] For example, the technology disclosed in Patent Document 1 generates 3D coordinated skeletal information from human movements measured by motion capture. The technology disclosed in Patent Document 1 not only generates video (teacher data) of a 3DCG avatar moving using the skeletal information, but also generates video (teacher data) of the 3DCG avatar moving using skeletal information with parameters changed, thereby generating a teacher dataset. The technology disclosed in Patent Document 1 can generate a recognition model that recognizes whether the movement of the recognition target corresponds to the movement of the 3DCG avatar by learning from the generated teacher dataset. [Prior art documents] [Patent documents]
[0004] [Patent Document 1] Japanese Patent Publication No. 2022-140038 Summary of the Invention [Problem to be solved by the invention]
[0005] However, in the above-described conventional technique, when the action to be recognized by the recognition model is an action of manipulating an object, in order to generate the recognition model, it is necessary to capture the action of a person actually manipulating the object and generate the above-described video. Therefore, in the above-described conventional technique, if an object to be manipulated by a person cannot be prepared, it is difficult to generate a video used to train a recognition model that recognizes the action of manipulating an object.
[0006] The present disclosure has been made in consideration of the above circumstances, and aims to provide a video generation system, a video generation method, and a program that are capable of generating videos useful for training a recognition model that recognizes the action of manipulating an object, even without actually preparing an object to be manipulated by a person. [Means for solving the problem]
[0007] One aspect of the video generation system of the present disclosure includes a first display control unit that displays an apparatus model having a moving part, a capture unit that captures the hand movements of an actor virtually operating the moving part, and a video generation unit that generates a video in which the virtual hand of a human model operates the moving part in accordance with the captured hand movements.
[0008] One aspect of the moving image generating method of the present disclosure includes a display step in which a first display control unit displays an apparatus model having a moving part; a capture step in which a capture unit captures hand movements of an actor virtually operating the moving part; and a moving image generating step in which a moving image generating unit generates a moving image in which the virtual hand of a human model operates the moving part in accordance with the captured hand movements.
[0009] One aspect of the program of the present disclosure is for causing a computer to execute a display control step of displaying an apparatus model having a moving part, an acquisition step of acquiring hand movements of an actor virtually operating the moving part, and an animation generation step of generating an animation in which the virtual hand of a human model operates the moving part in accordance with the acquired hand movements. [Effects of the Invention]
[0010] According to the present disclosure, it is possible to provide a video generation system, a video generation method, and a program that can generate videos useful for training a recognition model that recognizes the action of manipulating an object, even if the object to be operated by a person is not actually prepared. [Brief explanation of the drawings]
[0011] [Figure 1] FIG. 1 is a block diagram showing an example of the configuration of a training data generation system according to this embodiment. [Figure 2] FIG. 2 is a block diagram showing an example of the hardware configuration of the information processing device and the VR headset of this embodiment. [Figure 3] FIG. 3 is a block diagram showing an example of the functional configuration of the motion capture system of this embodiment. [Figure 4] FIG. 4 is a diagram showing an example of a capture setting screen according to this embodiment. [Figure 5] FIG. 5 is a diagram showing an example of a person model registration screen according to this embodiment. [Figure 6] FIG. 6 is a diagram showing an example of an object model registration screen according to this embodiment. [Figure 7] FIG. 7 is a diagram showing an example of the capture setting screen after the human model and the object model are registered. [Figure 8] FIG. 8 is a diagram showing an example of a check moving image generated by the check moving image generating unit of this embodiment. [Figure 9] FIG. 9 is a diagram showing an example of a feedback moving image generated by the feedback moving image generating unit of this embodiment. [Figure 10] FIG. 10 is a diagram showing an example of the capture confirmation screen of this embodiment. [Figure 11] FIG. 11 is a diagram showing an example of the capture information stored in the capture information storage unit of this embodiment. [Figure 12] FIG. 12 is a block diagram showing an example of the functional configuration of the training data generating device of this embodiment. [Figure 13] FIG. 13 is a diagram showing an example of the video setting screen of this embodiment. [Figure 14] FIG. 14 is a diagram showing an example of a video confirmation screen according to this embodiment. [Figure 15] FIG. 15 is a diagram showing an example of the video setting screen of this embodiment. [Figure 16] FIG. 16 is a diagram showing an example of a moving image for learning data generated by the moving image generating unit of this embodiment. [Figure 17] FIG. 17 is a flowchart showing an example of the capture process performed in the motion capture system of this embodiment. [Figure 18] FIG. 18 is a flowchart showing an example of the training data generation process performed by the training data generation device of this embodiment. DETAILED DESCRIPTION OF THE INVENTION
[0012] Hereinafter, an embodiment of the present disclosure (hereinafter simply referred to as "the present embodiment") will be described in detail with reference to the drawings. Note that the present disclosure is not limited to the following embodiment. Furthermore, the following embodiment and modified examples can be combined as appropriate.
[0013] In the training data generation system of this embodiment, a VR (Virtual Reality) headset is used to virtually reproduce objects that do not exist in the field as object models, and the movements of an actor wearing the VR headset as he or she virtually manipulates the object models are captured. Furthermore, the training data generation system of this embodiment uses the captured movements to generate a scene in which a human model that reproduces the actor's movements manipulates the object models in a virtual space, and animates the scene to generate training data.
[0014] Therefore, according to this embodiment, even if an object to be operated by an actor is not actually prepared, video useful for training a recognition model that recognizes the movements of manipulating the object can be generated as training data.
[0015] Note that the term "performer" refers to the person wearing the VR headset. Furthermore, since the object model does not exist in real space but is a virtual object presented to the performer via the VR headset, the performer cannot physically manipulate the object model. For this reason, in this embodiment, the performer's manipulation of the object model presented via the VR headset is expressed as "virtually manipulating the object model." Note that when the performer makes a movement to virtually manipulate the object model presented via the VR headset, the performer is actually performing an action (movement) as if manipulating an object that does not exist in the actual location. For this reason, in this embodiment, the wearer of the VR headset is referred to as the "performer" (an example of a performer) as described above.
[0016] Furthermore, the training data generation system of this embodiment uses a camera attached to a VR headset to capture the motion of the performer, making it possible to capture the precise movements of the performer's fingers when virtually manipulating an object model. Therefore, this embodiment makes it possible to generate, as training data, videos that are useful for training a recognition model that recognizes movements that manipulate objects, such as those that require precise finger movements.
[0017] Furthermore, in the learning data generation system of this embodiment, in a scene where motion capture is performed, not only the movement of the performer but also the movement of the object model is captured. Note that in a scene where motion capture is performed, if the performer's hand and the object model are in virtual contact, the object model is moved in conjunction with the movement (operation) of the performer's hand. On the other hand, in a scene where learning data is generated, the object model is moved in conjunction with the virtual hand of the human model based on the captured movement of the object model, regardless of whether or not there is contact between the virtual hand, which is the hand of the human model, and the object model.
[0018] In the training data generation system of this embodiment, in order to expand the data, the above-mentioned scene in which a human model operates an object model may be generated by shifting the position of the human model and the object model while maintaining their relative positions, thereby generating training data that is a variation of the scene. In this embodiment, even in this case, the object model is moved in conjunction with the human model's virtual hand based on the captured movement of the object model, regardless of whether or not there is contact between the human model's virtual hand and the object model. Therefore, according to this embodiment, even if there is no contact between the human model's virtual hand and the object model due to the effect of shifting the position of the human model and the object model, it is possible to reliably generate a video in which a human model operates an object model.
[0019] In the following, the learning data generation system of this embodiment will be described using an example in which a recognition model for recognizing the movements of an operator operating an electronic component mounter is assumed and learning data used to train the recognition model is generated, but the present invention is not limited to this. Note that in the following, the electronic component mounter may be simply referred to as a "mounter."
[0020] Specifically, the training data generation system of this embodiment has an actor wearing a VR headset virtually operate a mounting machine model (an example of a device model with moving parts) that is a CG (Computer Graphics) reproduction of a mounting machine, and captures the actor's movements. Because mounting machines are so-called factory equipment used in electronic device manufacturing sites, it is difficult to prepare actual machines for motion capture. Furthermore, the training data generation system of this embodiment uses the captured movements to generate a scene in which a CG human model of a worker reproduces the actor's movements to operate the mounting machine model in a virtual space, and animates the scene to generate training data.
[0021] However, the training data of this embodiment and the recognition model generated by training the training data are not limited to this. The recognition model of this embodiment may be any trained model for recognizing the movement of a person manipulating an object. Similarly, the training data of this embodiment may be any data that can be used to train a recognition model for recognizing the movement of a person manipulating an object.
[0022] Fig. 1 is a block diagram showing an example of the configuration of a training data generation system 1 according to this embodiment. As shown in Fig. 1, the training data generation system 1 includes an information processing device 10 and a VR headset 20. The information processing device 10 and the VR headset 20 are connected via a communication cable 19.
[0023] When the learning data generation system 1 performs motion capture, the information processing device 10 sets up, controls, and generates moving images in a virtual space for motion capture. When the learning data generation system 1 generates learning data using motion capture results, the information processing device 10 sets up, controls, and generates moving images in a virtual space for generating learning data. Examples of the information processing device 10 include, but are not limited to, a computer capable of running software for setting up, controlling, and generating moving images in the virtual space described above using CG.
[0024] The VR headset 20 is a head mounted device (HMD) that provides a virtual reality (VR) experience to the wearer. When the learning data generation system 1 performs motion capture, the VR headset 20 displays a moving image (video) of the virtual space generated by the information processing device 10. The VR headset 20 also tracks and captures the movements of the wearer who is immersed in the virtual space by viewing the displayed video of the virtual space.
[0025] The communication cable 19 is a communication cable that connects the information processing device 10 and the VR headset 20, and may be, for example, a USB (Universal Serial Bus) cable, but is not limited to this. The communication cable 19 is used for communicating, for example, videos of virtual spaces generated by the information processing device 10, motion capture results of the VR headset 20, etc. Note that the connection between the information processing device 10 and the VR headset 20 may be a wireless connection for wireless communication, rather than a connection via the communication cable 19 for wired communication.
[0026] As described above, the training data generation system 1 of this embodiment includes a process in which the information processing device 10 and the VR headset 20 cooperate to perform motion capture, and a process in which the information processing device 10 alone generates training data using the motion capture results. Therefore, hereinafter, the training data generation system 1 in the scene in which motion capture is performed may be referred to as a "motion capture system," and the training data generation system 1 (information processing device 10) in the scene in which training data is generated may be referred to as a "training data generation device."
[0027] FIG. 2 is a block diagram showing an example of the hardware configuration of the information processing device 10 and the VR headset 20 of this embodiment.
[0028] First, the hardware configuration of the information processing device 10 will be described. As shown in Fig. 2, the information processing device 10 includes a control device 11, a main memory device 12, an auxiliary memory device 13, a display device 14, an input device 15, a communication device 16, and various buses 17. The control device 11, the main memory device 12, the auxiliary memory device 13, the display device 14, the input device 15, and the communication device 16 are connected via the various buses 17. As described above, the information processing device 10 of this embodiment has a general hardware configuration using a normal computer.
[0029] The control device 11 controls the overall operation of the information processing device 10. Examples of the control device 11 include at least one of a CPU (Central Processing Unit) and a GPU (Graphics Processing Unit), but are not limited to these. There may be any number of CPUs or GPUs as long as they are one or more, and they may be single-core or multi-core.
[0030] Examples of the main storage device 12 include, but are not limited to, a ROM (Read Only Memory) and a RAM (Random Access Memory). The ROM stores various programs such as a program for controlling the information processing device 10, a program for controlling motion capture, and a program for generating learning data. The RAM is used as a working area when the control device 11 performs various controls based on the programs stored in the ROM.
[0031] The auxiliary storage device 13 stores the various programs described above, as well as various data such as data for controlling motion capture and data for generating learning data. The various programs described above may be stored in at least one of the main storage device 12 and the auxiliary storage device 13. Examples of the auxiliary storage device 13 include, but are not limited to, at least one of existing storage devices capable of magnetic, electrical, or optical storage, such as a hard disk drive (HDD), a solid state drive (SSD), and a digital versatile disc (DVD). The auxiliary storage device 13 may be built into the information processing device 10 or may be externally attached to the information processing device 10 via an interface such as a universal serial bus (USB). The auxiliary storage device 13 may also be a network-attached storage (NAS) connected via a network such as a local area network (LAN) or a wide area network (WAN).
[0032] The display device 14 displays various screens used when setting up motion capture and generating learning data by the information processing device 10, and serves as a user interface with the user (operator). Examples of the display device 14 include various displays such as a liquid crystal display, an organic electroluminescence (EL) display, and a touch panel display, but are not limited to these. The display device 14 may be a built-in display built into the information processing device 10, or an external display connected to the information processing device 10 via a display interface such as HDMI (registered trademark).
[0033] The input device 15 is used for various inputs used when setting up motion capture or generating learning data by the information processing device 10, and serves as a user interface with a user (operator). Examples of the input device 15 include, but are not limited to, a keyboard, a mouse, and a touch panel. The input device 15 may be built into the information processing device 10 or may be externally attached to the information processing device 10 via an interface such as a USB. In the following, unless otherwise specified, it is assumed that user requests and operations regarding setting up motion capture or generating learning data executed on the information processing device 10 are performed using the input device 15.
[0034] Examples of the communication device 16 include, but are not limited to, a communication device for a wired LAN and a wireless communication device for a wireless LAN. The communication device 16 may be used to acquire various programs and various data of the present embodiment from the outside, or may be used to output various data generated by the learning data generation system 1 to the outside.
[0035] In addition to the above configuration, the information processing device 10 may further include hardwired circuits such as an IC (Integrated Circuit), an ASIC (Application Specific Integrated Circuit), and an FPGA (Field-Programmable Gate Array) specific to the information processing device 10 in order to realize the motion capture setting function and the learning data generation function.
[0036] Next, we will explain the hardware configuration of the VR headset 20. As shown in Fig. 2, the VR headset 20 includes a control device 21, a main memory device 22, an auxiliary memory device 23, a display device 24, and an image capturing device 25. The control device 21, the main memory device 22, the auxiliary memory device 23, the display device 24, and the image capturing device 25 are connected via various buses 27.
[0037] The control device 21 controls the overall operation of the VR headset 20. The method for realizing the control device 21 is similar to that of the control device 11, and therefore a detailed description thereof will be omitted.
[0038] The method for realizing the main memory device 22 is similar to that of the main memory device 12, and therefore detailed description thereof will be omitted. The ROM of the main memory device 22 stores various programs, such as a program for controlling the VR headset 20 and a program for performing motion capture.
[0039] The auxiliary storage device 23 stores the various programs described above and various data such as data for motion capture. The various programs described above may be stored in at least one of the main storage device 22 and the auxiliary storage device 23. The method for realizing the auxiliary storage device 23 is the same as that for the auxiliary storage device 13, and therefore a detailed description thereof will be omitted.
[0040] The display device 24 displays a moving image (video) of a virtual space generated by the information processing device 10, and may be, for example, a non-transparent display. The display device 24 of this embodiment is provided, for example, inside the housing of the VR headset 20, and is configured so that the field of view of the wearer of the VR headset 20 is covered by the inside of the housing and the non-transparent display device 24. Therefore, the display device 24 of this embodiment functions as an immersive display. However, the display device 24 of this embodiment is not limited to the above-described configuration. For example, the display device 24 of this embodiment may be a mixed reality (MR)-compatible display that can be switched between transparent and non-transparent modes, as long as it can be set to non-transparent. Furthermore, for example, the display device 24 of this embodiment may be a stereoscopic display that displays three-dimensional video (images) by displaying different videos (images) with parallax to the left and right eyes of the wearer.
[0041] The image capturing device 25 captures images for tracking the movements of the wearer of the VR headset 20, and can be realized, for example, by one or more cameras. Examples of cameras include, but are not limited to, wide-angle cameras and IR (Infrared Rays) cameras. The image capturing device 25 of this embodiment is realized, for example, by multiple cameras, which are distributed and arranged outside the housing of the VR headset 20. Therefore, the image capturing device 25 of this embodiment is configured to be able to capture images of the wearer wearing the VR headset 20, regardless of the posture of the wearer.
[0042] The control device 21 detects the three-dimensional position and orientation of a specific part of the wearer of the VR headset 20 in real space from the position and orientation of the specific part in images captured frame by frame by the imaging device 25. The control device 21 captures the wearer's movements by tracking the detected three-dimensional position and orientation of the specific part. Note that in this embodiment, an example will be described in which the tracking targets are the head and fingers of the wearer of the VR headset 20. That is, in this embodiment, an example will be described in which the VR headset 20 has a head tracking function for tracking the movement of the wearer's head and a hand tracking function for tracking the movement of the wearer's fingers, but the present invention is not limited to this.
[0043] Note that the head of the wearer of the VR headset 20 is not captured in the image captured by the image capture device 25. Therefore, the control device 21 may realize the head tracking function by using a self-position estimation technology, such as SLAM (Simultaneous Localization and Mapping), that estimates the capture position and orientation of the captured image from the captured image. Furthermore, the control device 21 may also use a technology that infers the movement of a human hand or wrist through machine learning when performing hand tracking. In this way, hand tracking can be continued even if the wearer's fingers are hidden by the torso and not captured by the image capture device 25, or if the wearer's fingers are out of the field of view of the image capture device 25.
[0044] In this embodiment, a case where the position and posture of the fingers of the wearer of the VR headset 20 are directly tracked will be described as an example, but the present invention is not limited to this. For example, the wearer may hold a hand controller in his / her hand, and the position and posture of the held hand controller may be detected to indirectly track the position and posture of the wearer's fingers. In addition, in this embodiment, an inside-out method where tracking is performed by the VR headset 20 alone without external sensors will be described as an example, but the present invention is not limited to this. For example, an outside-in method where tracking is performed using external sensors and the VR headset 20 may be adopted.
[0045] Like the information processing device 10, the VR headset 20 may include a communication device for communicating with the outside.
[0046] The functional configuration of the training data generation system 1 of this embodiment will be described below, divided into the motion capture system and the training data generation device. First, the motion capture system of this embodiment will be described.
[0047] 3 is a block diagram showing an example of the functional configuration of the motion capture system 30 of this embodiment. As described above, the motion capture system 30 is realized by the information processing device 10 and the VR headset 20, and corresponds to the learning data generation system 1 for a scene where motion capture is performed.
[0048] 3, motion capture system 30 includes a model storage unit 101, a capture setting unit 111, a virtual space control unit 121, a capture information saving unit 131, a capture information storage unit 141, a capture unit 201, and a feedback video display control unit 211. The capture setting unit 111 includes a model registration unit 113, a placement determination unit 115, and an additional information setting unit 117. The virtual space control unit 121 includes a confirmation video generation unit 123, a feedback video generation unit 125, and a motion control unit 127.
[0049] The model storage unit 101 and the capture information storage unit 141 can be realized by, for example, at least one of the main storage unit 12 and the auxiliary storage unit 13 described with reference to FIG.
[0050] The capture setting unit 111, model registration unit 113, placement determination unit 115, additional information setting unit 117, virtual space control unit 121, confirmation video generation unit 123, feedback video generation unit 125, operation control unit 127, and capture information storage unit 131 can be realized, for example, by the control device 11 and main memory device 12 described in Figure 2.
[0051] For example, the control device 11 reads out a program for controlling the motion capture of this embodiment, which is stored in the main memory device 12 (ROM) or the auxiliary memory device 13, or which is externally acquired from the communication device 16 via a network, and loads the program into the main memory device 12 (RAM). The control device 11 executes various processes in accordance with the loaded program, thereby realizing each of the above-mentioned functional units. Here, the description has been given taking as an example a case where each of the above-mentioned functional units is realized as software, but at least a part of each of the above-mentioned functional units may also be realized as hardware. In this case, the functional units realized as hardware may be realized, for example, by the above-mentioned hardwired circuit. Furthermore, any of the above-mentioned functional units may also be realized by a combination of software and hardware.
[0052] The capture unit 201 can be realized by, for example, the control device 21, the main storage device 22, and the image capturing device 25 described in Fig. 2. The feedback video display control unit 211 can be realized by, for example, the control device 21 and the main storage device 22 described in Fig. 2.
[0053] The model storage unit 101 stores character models and object models to be made to appear in the virtual space. In this embodiment, the virtual space is described as a three-dimensional CG space, but the present invention is not limited to this. An example of a character model is three-dimensional model data in which a character is reproduced using CG, but the present invention is not limited to this and may also be two-dimensional model data. An example of an object model is three-dimensional model data in which an object operated by a character is reproduced using CG, but the present invention is not limited to this and may also be two-dimensional model data.
[0054] Both the person model and the object model may be model data created based on a real person or object, or may be model data created based on a fictional person or object. Furthermore, both the person model and the object model may be model data created in a CG environment by a developer of the learning data generation system 1 (hereinafter simply referred to as "developer"), or may be model data prepared in advance on a platform that builds a CG environment.
[0055] The capture setting unit 111 performs various settings for capturing the motion of an actor wearing the VR headset 20. Specifically, the capture setting unit 111 performs settings related to the virtual space that allows the actor to virtually operate an object model, and settings for various information to be added to the captured movements of the actor.
[0056] The "developer" of the learning data generation system 1 and the "performer" who wears the VR headset 20 may be the same person or different people. In other words, the developer himself may wear the VR headset 20 and act as the performer, or a person other than the developer (for example, an assistant to the developer) may wear the VR headset 20 and act as the performer.
[0057] The various settings made by the capture setting unit 111 will be specifically described below using the model registration unit 113, the arrangement determination unit 115, and the additional information setting unit 117 included in the capture setting unit 111 as appropriate.
[0058] The model registration unit 113 registers human models and object models to appear in the virtual space. Specifically, the model registration unit 113 registers human models and object models selected by the developer from the human models and object models stored in the model storage unit 101 in the virtual space control unit 121. The placement determination unit 115 determines the initial placement in the virtual space of the human models and object models registered by the model registration unit 113, and registers them in the virtual space control unit 121.
[0059] The registered human model reproduces the movements of the actor captured in the real space in real time in the virtual space, and the registered object model is manipulated in the virtual space by the human model that reproduces the movements of the actor.
[0060] The additional information setting unit 117 sets an ID and a movement class to be added to the captured movements of the performer, etc. The ID is used as an identifier to identify the captured movements of the performer, etc. The movement class is label information indicating the type of the captured movement of the performer, and is used as an annotation when generating learning data using the captured movements of the performer.
[0061] For example, when a developer performs an operation input requesting the display of a capture setting screen (an example of a setting screen) using the input device 15, the capture setting unit 111 (an example of a second display control unit) displays the capture setting screen on the display device 14. The capture setting screen is a UI (User Interface) screen for various settings for performing motion capture of a performer.
[0062] 4 is a diagram showing an example of the capture setting screen 301 of this embodiment. As shown in Fig. 4, the capture setting screen 301 includes a human model registration area 311, an object model registration area 321, a model placement area 331, a text box 341, a text box 351, and a capture start button 361.
[0063] The human model registration area 311 includes a registration button 313 for registering a human model, and placement information 315 indicating the placement and scale of the registered human model in virtual space. The object model registration area 321 includes an add button 323 for additionally registering an object model, a select button 325 for selecting the additionally registered object model, placement information 327 indicating the placement and scale of the selected object model in virtual space, and a delete button 329 for deleting the selected object model.
[0064] For example, when the developer performs an operation input to press the registration button 313 in the human model registration area 311 using the input device 15, the model registration unit 113 displays a human model registration screen on the display device 14. The human model registration screen is a UI screen for registering a human model.
[0065] Fig. 5 is a diagram showing an example of a character model registration screen 401 according to this embodiment. As shown in Fig. 5, the character model registration screen 401 includes a pull-down list 411, a display area 413, and a register button 415. The pull-down list 411 is an operation button for selecting a character model. The display area 413 is a display area in which the character model selected in the pull-down list 411 is displayed. The register button 415 is an operation button for registering the character model selected in the pull-down list 411.
[0066] 5, in order to select a person model 421 (hereinafter sometimes referred to as a "worker model"), which is a CG model of a worker, the developer uses the input device 15 to perform an operational input to select a file with the file name "PersonA.fbx" on the pull-down list 411. As a result, the model registration unit 113 retrieves the person model 421 with the file name "PersonA.fbx" from the model storage unit 101 and displays it in the display area 413. In this situation, when the developer uses the input device 15 to perform an operational input to press the registration button 415, the model registration unit 113 registers the person model 421 in the virtual space control unit 121 as a person model to appear in the virtual space.
[0067] In this embodiment, the case where the number of character models appearing in the virtual space (to be registered) is one will be described as an example, but the present invention is not limited to this and there may be more than one. When multiple character models appear in the virtual space, for example, the character model registration process described in the character model registration screen 401 shown in Fig. 5 may be performed for each character model appearing in the virtual space (to be registered).
[0068] 4, for example, when the developer performs an operation input to press the add button 323 in the object model registration area 321, the model registration unit 113 displays an object model registration screen on the display device 14. The object model registration screen is a UI screen for registering an object model.
[0069] Fig. 6 is a diagram showing an example of an object model registration screen 451 of this embodiment. As shown in Fig. 6, the object model registration screen 451 includes a pull-down list 461, a display area 463, and a register button 465. The pull-down list 461 is an operation button for selecting an object model. The display area 463 is a display area in which the object model selected in the pull-down list 461 is displayed. The register button 465 is an operation button for registering the object model selected in the pull-down list 461.
[0070] 6, in order to select object model 471, which is a CG model of a mounting machine (hereinafter, sometimes referred to as "mounting machine model"), the developer performs an operational input to select a file with the file name "mounter.fbx" on pull-down list 461. As a result, the model registration unit 113 retrieves object model 471 with the file name "mounter.fbx" from the model storage unit 101 and displays it in the display area 463. In this situation, when the developer performs an operational input to press registration button 465, the model registration unit 113 registers object model 471 in the virtual space control unit 121 as an object model to appear in the virtual space.
[0071] In this embodiment, an example will be described in which there are a plurality of object models (to be registered) appearing in the virtual space, but the present invention is not limited to this and there may be only one object model.
[0072] 7 is a diagram showing an example of the capture setting screen 301 after registering the person model 421 and the object model 471. The capture setting screen 301 described with reference to FIG. 4 will be described below with reference to FIG.
[0073] In the model placement area 331, the person model 421 and the object model 471 are placed in a planar view to show the initial placement in the virtual space of the person model 421 and the object model 471 registered by the model registration unit 113 in a planar view. The developer uses the input device 15 to perform an operation input to change the placement of the person model 421 and the object model 471 placed in the model placement area 331. The placement determination unit 115 changes the placement of the person model 421 and the object model 471 in the virtual space in accordance with the operation input.
[0074] The operation input may be an operation input that changes the positions of the person model 421 and the object model 471 in the model placement area 331, thereby changing the placement of the person model 421 and the object model 471. The operation input may also be an operation input that changes the placement of the person model 421 by changing the position, angle, and scale of the placement of the person model 421 displayed in the placement information 315. Similarly, the operation input may be an operation input that changes the placement of the object model 471 by changing the position, angle, and scale of the placement of the object model 471 displayed in the placement information 327.
[0075] Furthermore, as will be described in detail later, the performer wearing the VR headset 20 experiences himself as if he were the person model 421 (worker model) and operates the object model 471 (mounting machine model) as the person model 421. For this reason, it is preferable to determine the positions of the person model 421 and the object model 471 so that they are in the same positional relationship as when an actual worker operates an actual mounting machine.
[0076] 4 and 7 is shown in a planar view, but the display mode of the model placement area 331 is not limited to this. The model placement area 331 can display the model placement area captured from any viewpoint position, and the viewpoint position at this time is displayed as viewpoint position display 333.
[0077] As described above, the placement information 315 displays the position, angle, and scale of the placement of the person model 421 in the virtual space. Similarly, as described above, the placement information 327 displays the position, angle, and scale of the placement of the object model 471 in the virtual space. If multiple object models are registered by the model registration unit 113, the placement information 327 displays the placement of the object model selected by the selection button 325. Furthermore, when an operation input is made to press the delete button 329, the registration of the object model whose placement is displayed in the placement information 327 is deleted.
[0078] The text box 341 is an input field for inputting an ID to be added to the captured movements of the performer, etc. In the example shown in Fig. 7, the developer uses the input device 15 to input the ID "1" into the text box 341.
[0079] The text box 351 is an input field for inputting an action class that indicates the type of captured movement of the performer (the type of operation performed by the human model). In the example shown in Fig. 7, the performer virtually opens the safety cover of the mounting machine model (object model 471) as the target action of motion capture, so the developer inputs the action class "open" into the text box 351.
[0080] The capture start button 361 is an operation button for starting to capture the movements of the performer wearing the VR headset 20 and the movements of the object model virtually operated by the performer. When the developer performs an operation input of pressing the capture start button 361 using the input device 15, the virtual space control unit 121 generates a virtual space in which the human model and object model registered by the model registration unit 113 are placed at positions determined by the placement determination unit 115. The capture unit 201 also starts capturing the motion of the performer wearing the VR headset 20. The virtual space control unit 121 also captures the movements of the object model. This makes it possible to capture the movements of the object model when the human model operates the object model (when the performer virtually operates the object model).
[0081] Below, we will specifically explain the processing contents of the virtual space control unit 121, the capture unit 201, and the feedback video display control unit 211 after the capture start button 361 is pressed. Note that below, as an example of an actor wearing the VR headset 20 virtually manipulating an object model, the actor's movement of virtually opening the safety cover of a mounting machine model (object model 471) will be explained as an example, but the present invention is not limited to this. Other examples of an actor virtually manipulating an object model include, for example, an actor virtually pulling out a feeder from the mounting machine model, or an actor virtually repairing the cover tape sealing electronic components packed in a tape reel of the mounting machine model. Note that the object model virtually manipulated by the actor is not limited to a mounting machine model.
[0082] The safety cover, feeder, and cover tape are examples of movable parts of the device model. In this embodiment, the movable parts of the device model are configured to be movable along a predetermined trajectory, but are configured to be unable to move in a direction other than the predetermined trajectory. For example, the safety cover is configured to be movable along a substantially arc-shaped trajectory with the connection part with the housing of the mounting machine model as the axis, but is configured to be unable to move in a direction other than the trajectory, such as by sliding the safety cover left and right. Also, for example, the feeder is configured to be movable along a linear trajectory back and forth relative to the housing of the mounting machine model, but is configured to be unable to move in a direction other than the trajectory, such as by sliding the feeder left and right or rotating the feeder.
[0083] The virtual space control unit 121 performs various controls related to the virtual space, such as generating the virtual space, generating videos captured within the virtual space, and controlling the movements of human models and object models within the virtual space. Below, the various controls related to the virtual space performed by the virtual space control unit 121 will be specifically explained using the confirmation video generation unit 123, feedback video generation unit 125, and movement control unit 127 included in the virtual space control unit 121 as appropriate.
[0084] The virtual space control unit 121 generates the virtual space described above and places a third-person perspective virtual camera in the virtual space to generate a confirmation video. The confirmation video is a video for confirming the movement of a human model based on the movement of an actor captured by the capture unit 201. By checking the movement of the human model in the confirmation video, the developer can confirm whether the captured movement of the actor can be used to generate learning data for a recognition model to recognize the movement of a worker operating a mounting machine. The placement position of the third-person perspective virtual camera can be, for example, a position that captures the scene of a human model operating an object model from diagonally behind, which corresponds to the camera position when generating a video for learning data (described below). Therefore, the confirmation video is a third-person perspective video. However, the placement position of the third-person perspective virtual camera is not limited to this. In this embodiment, the third-person perspective virtual camera is fixedly placed in the virtual space, but this is not limited to this.
[0085] The virtual space control unit 121 also places a first-person perspective virtual camera for generating a feedback video for the performer at the viewpoint position of the human model in the generated virtual space. Therefore, the feedback video is a first-person perspective video showing the field of view of the human model. The feedback video is a video for feeding back the field of view of the human model to the performer, and is displayed on the display device 24 of the VR headset 20. This allows the performer to virtually operate the object model from the viewpoint of the human model.
[0086] The capture unit 201 captures the movements of the performer wearing the VR headset 20. In this embodiment, the capture unit 201 has the head tracking function and hand tracking function described above. Therefore, in this embodiment, the capture unit 201 captures the three-dimensional position and posture of the performer's head and the three-dimensional position and posture of the performer's fingers in real space on a frame-by-frame basis, thereby capturing the movements of the performer's head and the movements of the performer's fingers. For example, the capture unit 201 captures the movement of the performer's hand virtually opening the safety cover of the mounting machine model (object model 471). However, the parts of the performer that are the subject of capture by the capture unit 201 are not limited to these.
[0087] In this embodiment, the capture unit 201 captures the movement (three-dimensional position and posture) of the performer's head using the three-dimensional position and posture of the performer's head at the start of motion capture as the origin and coordinate axes. Similarly, the capture unit 201 captures the movement (three-dimensional position and posture) of the performer's hand using the three-dimensional position and posture of the performer's hand at the start of motion capture as the origin and coordinate axes. In this way, position adjustments such as offsets can be eliminated when using the captured movement. However, the present invention is not limited to this, and the capture unit 201 may, for example, capture the movement (three-dimensional position and posture) of the performer's head and hand using preset three-dimensional positions and postures as the origin and coordinate axes.
[0088] When motion capture by the capture unit 201 begins, the character model moves in the virtual space in accordance with the captured movements of the performer under the control of the movement control unit 127. The view of the virtual space from the character model is fed back to the performer as a feedback video via the VR headset 20 under the control of the feedback video generation unit 125 and the feedback video display control unit 211, and the view of the character model is displayed to the performer. Since the head movement of the performer is tracked by the capture unit 201, the position and orientation of the character model and the first-person perspective virtual camera are controlled under the control of the movement control unit 127 and the feedback video generation unit 125 so as to be linked to the head movement of the performer. Therefore, the view of the character model displayed to the performer (feedback video) also changes in accordance with the head movement of the performer. For example, if the performer turns to the right, the character model and the first-person perspective virtual camera are also controlled to face right, and the view of the character model as it faces right (feedback video) is displayed to the performer.
[0089] Furthermore, since the movement of the performer's hand is tracked by the capture unit 201, the position and orientation of the virtual hand of the human model are controlled by the movement control unit 127 so as to be linked to the movement of the performer's hand. The virtual hand of the human model is placed in the virtual space so that its positional relationship with the virtual camera for the first-person viewpoint corresponds to the positional relationship between the VR headset 20 and the performer's hand. This allows the performer to move the virtual hand of the human model as if it were his or her own hand.
[0090] As a result, the performer can experience himself as if he were a human model, and can therefore act as if he were manipulating the object model (movements for virtually manipulating the object model) by moving his body in real space in order to operate the object model as a human model in the virtual space. The capture unit 201 can capture the movements for virtually manipulating the object model by capturing the performer's performance.
[0091] The confirmation video generation unit 123 controls a third-person perspective virtual camera fixedly positioned at a predetermined position in the virtual space by the virtual space control unit 121, captures a scene in which a human model manipulates an object model based on the actor's movements captured by the capture unit 201, and generates the confirmation video. For example, the confirmation video generation unit 123 generates an image by having the third-person perspective virtual camera project (render) various models, such as human models and object models, and backgrounds, included in the virtual camera's angle of view (field of view) onto a projection surface (not shown). In rendering, the world coordinate system, which is a coordinate system within the virtual space, is converted into a camera coordinate system whose origin is the viewpoint position, which is the position where the third-person perspective virtual camera is positioned in the virtual space, and the various models, backgrounds, etc. expressed in the camera coordinate system are converted into a coordinate system on a two-dimensional projection surface by perspective projection transformation. Note that a well-known rendering method, such as the Z-buffer method, may be used. The confirmation video generation unit 123 generates the confirmation video by having the third-person perspective virtual camera perform the above-mentioned image generation for each frame.
[0092] Fig. 8 is a diagram showing an example of a confirmation video 501 generated by the confirmation video generating unit 123 of this embodiment. In the example shown in Fig. 8, the confirmation video 501 is described as a video in which a person model 421 (worker model) opens a safety cover 472 of an object model 471 (mounting machine model), but the present invention is not limited to this. Note that the person model 421 makes a movement to open the safety cover 472 based on the movement of the actor captured by the capturing unit 201 (the movement of the actor virtually opening the safety cover 472 of the object model 471).
[0093] In the example shown in FIG. 8 , confirmation video 501 includes image 511 and image 512 as frame images. Image 511 shows a scene before human model 421 opens safety cover 472. Specifically, image 511 shows a scene in which the actor virtually grasps handle portion 473 of safety cover 472, causing human model 421 to grasp handle portion 473 with virtual hand 422. Image 512 shows a scene after human model 421 opens safety cover 472. Specifically, image 512 shows a scene in which the actor virtually grasps handle portion 473 while moving his or her hand upward and slightly toward the back, causing human model 421 to move virtual hand 422 upward and slightly toward the back while grasping handle portion 473, thereby opening safety cover 472. Note that safety cover 472 opens by moving along a substantially arc-shaped trajectory around the connection portion with object model 471 (mounting machine model) as an axis.
[0094] The feedback video generation unit 125 controls the first-person viewpoint virtual camera, which is placed at the viewpoint position of the human model by the virtual space control unit 121, so that it is always located at that viewpoint position, even if the human model moves or changes its posture in the virtual space. The feedback video generation unit 125 also controls the first-person viewpoint virtual camera, captures a scene in which the human model manipulates an object model based on the movement of the performer captured by the capture unit 201, and generates the above-mentioned feedback video. Note that the method of generating the feedback video by the feedback video generation unit 125 is the same as the method of generating the confirmation video by the confirmation video generation unit 123, except that the third-person viewpoint virtual camera is a first-person viewpoint virtual camera, and therefore detailed description thereof will be omitted.
[0095] The feedback video display control unit 211 (an example of a first display control unit) displays the feedback video generated by the feedback video generation unit 125 on the display device 24 of the VR headset 20. Specifically, every time a frame image of the feedback video is generated by the feedback video generation unit 125, the feedback video display control unit 211 displays the frame image on the display device 24. In other words, the feedback video display control unit 211 displays the feedback video generated by the feedback video generation unit 125 on the display device 24 in real time, and feeds back the view of the human model to the performer in real time.
[0096] FIG. 9 is a diagram illustrating an example of a feedback video 521 generated by the feedback video generator 125 of this embodiment. In the example illustrated in FIG. 9, the feedback video 521 is described as a video in which a person model 421 (worker model) opens a safety cover 472 of an object model 471 (mounting machine model), similar to the example illustrated in FIG. 8 ; however, the present invention is not limited to this. Similarly to image 512 of the confirmation video 501 illustrated in FIG. 8 , the feedback video 521 illustrated in FIG. 9 illustrates a scene after the person model 421 opens the safety cover 472. Specifically, the actor virtually grasps the handle portion 473 and moves his / her hand upward and slightly toward the back, thereby causing the person model 421 to move the virtual hand 422 upward and slightly toward the back while grasping the handle portion 473, thereby opening the safety cover 472. Because the feedback video 521 is a field-of-view image illustrating the field of view of the person model as illustrated in FIG. 9 , the actor can virtually operate the object model from the viewpoint of the person model.
[0097] The movement control unit 127 (an example of a movement control unit) acquires the movement of the performer captured by the capture unit 201, and moves the human model in accordance with the acquired movement. Specifically, each time the capture unit 201 captures the movement of the performer's head and fingers, the movement control unit 127 moves the head and fingers of the human model so as to match the captured movement of the head and fingers. In this embodiment, the movement control unit 127 estimates the movement (posture) of parts of the human model excluding the head and fingers, such as the shoulders and lower body, using inverse kinematics (IK), which estimates the posture of the entire body from the position of the head, etc. of the human model, but is not limited to this.
[0098] Furthermore, when the virtual hand of the person model grasps an object model, the operation control unit 127 moves the object model in conjunction with the movement of the virtual hand. For example, when the virtual hand of the person model and the safety cover of the object model (mounting machine model) are within a predetermined distance, the operation control unit 127 moves the safety cover in accordance with the movement of the virtual hand. Specifically, the operation control unit 127 detects for each frame whether or not there is contact between the virtual hand of the person model and the object model, and when contact is detected, determines whether or not the virtual hand has grasped the object model. When the virtual hand grasps the object model, the operation control unit 127 moves the object model in conjunction with the movement of the virtual hand until the virtual hand releases the object model.
[0099] Whether or not the virtual hand has grasped the object model can be determined, for example, by using the movement of the performer's fingers captured by the capture unit 201. When grasping an object with a hand, the distance between the fingertips of the fingers is generally closer after grasping the object than before grasping the object. For this reason, the movement control unit 127 determines, for example, whether or not the distance between the fingertip of the performer's thumb and the fingertips of each finger captured by the capture unit 201 is less than a threshold value. The threshold value may be changed for each finger. Alternatively, whether or not the distance between the fingertip of the thumb and the fingertips of some fingers is less than a threshold value may be determined without using all fingers.
[0100] If the distance between the performer's thumb tip and each finger tip is less than a threshold, the movement control unit 127 determines that the virtual hand has grasped the object model and causes the virtual hand to grasp the object model. On the other hand, if the distance between the performer's thumb tip and each finger tip is equal to or greater than a threshold, the movement control unit 127 determines that the virtual hand has not grasped the object model and does not cause the virtual hand to grasp the object model.
[0101] When the virtual hand grabs an object model, the movement control unit 127 moves the object model so that it moves in the same manner as the virtual hand until the virtual hand releases the object model. For example, when the movement control unit 127 moves the virtual hand a certain distance in a predetermined direction, the movement control unit 127 also moves the object model a certain distance in the same direction. Furthermore, for example, when the virtual hand rotates a certain angle in a predetermined direction, the movement control unit 127 also rotates the object model a certain angle in the same direction.
[0102] If the distance between the tip of the performer's thumb and the tips of each finger remains less than a threshold while the virtual hand is grasping the object model, the movement control unit 127 determines that the virtual hand is still grasping the object model, and continues to grasp the object model with the virtual hand. On the other hand, if the distance between the tip of the performer's thumb and the tips of each finger becomes equal to or greater than a threshold while the virtual hand is grasping the object model, the movement control unit 127 determines that the virtual hand has released the object model, and causes the virtual hand to release the object model. After the virtual hand releases the object model, the movement control unit 127 moves the virtual hand in accordance with the movement of the performer's hand captured by the capture unit 201, without moving the object model.
[0103] When motion capture is started, the movement control unit 127 captures the movement of the object model. In this embodiment, the movement control unit 127 captures the three-dimensional position and posture of the object model in the virtual space on a frame-by-frame basis, thereby capturing the movement of the object model. In this embodiment, it is assumed that the object model is moved by the virtual hand of a human model, and therefore the three-dimensional position and posture remain unchanged from their initial values until the object model is moved by the virtual hand, but this is not limited to this.
[0104] In this embodiment, the movement control unit 127 captures the movement (three-dimensional position and orientation) of the object model using the three-dimensional position and orientation of the object model at the start of motion capture as the origin and coordinate axes. In this way, position adjustment such as offset can be eliminated when using the captured movement. However, without being limited to this, the movement control unit 127 may capture the movement (three-dimensional position and orientation) of the object model using, for example, a preset three-dimensional position and orientation as the origin and coordinate axes.
[0105] For example, when a developer performs an operation input to end motion capture using the input device 15, the capture setting unit 111 displays a capture confirmation screen (an example of a confirmation screen) on the display device 14. The capture confirmation screen is a screen for playing back the confirmation video generated by the confirmation video generation unit 123.
[0106] 10 is a diagram showing an example of a capture confirmation screen 601 of this embodiment. As shown in Fig. 10, the capture confirmation screen 601 includes a text box 341, a text box 351, a back button 611, a confirmation video display area 613, a play button 615, a seek bar 617, a background image superimposition button 619, and a registration button 621.
[0107] The text box 341 and the text box 351 are the same as those described in the capture setting screen 301 shown in FIG. 7, and therefore a description thereof will be omitted.
[0108] The back button 611 is a button for returning to the capture setting screen 301. When the developer performs an operation input to press the back button 611, the capture setting unit 111 terminates the display of the capture confirmation screen 601 and redisplays the capture setting screen 301 shown in Fig. 7. Note that the redisplayed capture setting screen 301 may be provided with an operation button for returning to the capture confirmation screen 601 again. In other words, the capture setting unit 111 may be configured to be able to switch between displaying the capture setting screen 301 and the capture confirmation screen 601.
[0109] Check video display area 613 is an area where the check video generated by check video generation unit 123 is displayed, and check video 501 described in Fig. 8 is displayed. Play button 615 is an operation button for playing check video 501. Seek bar 617 is a UI component that displays the playback position (playback frame) of check video 501 with a slider.
[0110] When the developer performs an operation input of pressing the play button 615, the capture setting unit 111 plays the confirmation video 501. As described above, the developer checks the movement of the person model 421 (worker model) operating the object model 471 (mounting machine model) on the confirmation video. This allows the developer to check whether the captured movements of the actor can be used to generate learning data for a recognition model for recognizing the movements of a worker operating a mounting machine.
[0111] The background image superimposition button 619 is an operation button for superimposing a background image on the confirmation video 501. As will be described in detail later, when the learning data generation system 1 generates learning data, it is preferable to reproduce an environment in which an object model (implementation machine model) exists, and then have a human model operate the object model based on the captured movement to generate learning data. For this reason, the developer can press the background image superimposition button 619 to select a background image to be superimposed on the confirmation video 501, and play the confirmation video 501 with the selected background image superimposed on it, thereby checking the image of the video to be generated as learning data.
[0112] The register button 621 is an operation button for registering capture information. The capture information is information including actor movement information indicating the movement of the actor captured by the capture unit 201, object model movement information indicating the movement of the object model captured by the movement control unit 127, and the ID and movement class set by the additional information setting unit 117.
[0113] When the developer performs an operation input to press registration button 621, additional information setting unit 117 sets the value entered in text box 341 as the ID, sets the information entered in text box 351 as the movement class, and outputs the set ID and movement class to capture information storage unit 131. Furthermore, capture setting unit 111 instructs capture unit 201 to output actor movement information, and capture unit 201 outputs the actor movement information to capture information storage unit 131. Furthermore, capture setting unit 111 instructs movement control unit 127 to output object model movement information, and movement control unit 127 outputs the object model movement information to capture information storage unit 131.
[0114] The capture information saving unit 131 associates the ID and the action class with the actor movement information and the object model movement information, and saves them in the capture information storage unit 141. In this way, the capture information is registered.
[0115] Fig. 11 is a diagram showing an example of the capture information stored in the capture information storage unit 141 of this embodiment. In the example shown in Fig. 11, the capture information is information in which an ID, a capture time, an action class, a target, and movement information are associated with each other.
[0116] The ID is primarily used as an identifier to identify capture information (a set of performer movement information, object model movement information, and action class). In this embodiment, an example will be described in which the ID is composed of a classification number and a sub-number, but this is not limited to this. The classification number is the value to the left of the hyphen, and corresponds to the value set by the additional information setting unit 117. As mentioned above, the classification number is used as an identifier to identify capture information. The sub-number is the value to the right of the hyphen, and is used to identify whether the movement information is performer movement information or object model movement information. In the example shown in Figure 11, when the sub-number is "A", the movement information is performer movement information, and when the sub-number is "B", the movement information is object model movement information, but this is not limited to this.
[0117] The capture time is information indicating the time when the motion information was captured. The action class is label information indicating the type of motion indicated by the motion information, and corresponds to the information set by the additional information setting unit 117. The target is information indicating the target of the motion information. In the example shown in Figure 11, if the motion information is performer motion information, the target is a "person," and if the motion information is object model motion information, the target is an "implementation machine." The motion information corresponds to either performer motion information or object model motion information.
[0118] In the example shown in FIG. 11, the capture information with an ID (classification number) of "1" corresponds to the capture information registered on the capture confirmation screen 601 shown in FIG.
[0119] Next, a training data generation device of this embodiment will be described. Fig. 12 is a block diagram showing an example of the functional configuration of a training data generation device 40 of this embodiment. As described above, the training data generation device 40 is realized by an information processing device 10, and corresponds to a training data generation system 1 in which training data is generated. Note that, as described above, this embodiment will be described taking as an example a case in which the training data generation device 40 is realized by the information processing device 10, but the present invention is not limited to this and may be realized by an information processing device different from the information processing device 10.
[0120] 12 , the training data generation device 40 includes a model storage unit 101, a capture information storage unit 141, a virtual environment setting unit 161, a virtual space control unit 171, and a training data generation unit 181. The virtual environment setting unit 161 includes a model registration unit 163, a placement determination unit 165, a capture information registration unit 167, and an environment setting unit 169. The virtual space control unit 171 includes a video generation unit 173.
[0121] The model storage unit 101 and the capture information storage unit 141 are the same as those in FIG. 3, and therefore detailed description thereof will be omitted.
[0122] The virtual environment setting unit 161, the model registration unit 163, the placement determination unit 165, the capture information registration unit 167, the environment setting unit 169, the virtual space control unit 171, the video generation unit 173, and the learning data generation unit 181 can be realized, for example, by the control device 11 and the main memory device 12 described in Figure 2.
[0123] For example, the control device 11 reads out a program for generating the learning data of this embodiment, which is stored in the main storage device 12 (ROM) or the auxiliary storage device 13, or which is externally acquired from the communication device 16 via a network, and loads the program into the main storage device 12 (RAM). The control device 11 executes various processes in accordance with the loaded program, thereby realizing each of the above-mentioned functional units. Here, the description has been given taking as an example a case where each of the above-mentioned functional units is realized as software, but at least a part of each of the above-mentioned functional units may also be realized as hardware. In this case, the functional units realized as hardware may be realized, for example, by the above-mentioned hardwired circuit. Furthermore, any of the above-mentioned functional units may be realized by a combination of software and hardware.
[0124] The virtual environment setting unit 161 performs settings related to a virtual space in which a human model that reproduces the movements of an actor captured by the motion capture system 30 operates an object model.
[0125] Below, the settings regarding the virtual space performed by the virtual environment setting unit 161 will be explained in detail, appropriately using the model registration unit 163, placement determination unit 165, capture information registration unit 167, and environment setting unit 169 included in the virtual environment setting unit 161.
[0126] The model registration unit 163 and the placement determination unit 165 are similar to the model registration unit 113 and the placement determination unit 115 described in Fig. 3, respectively, and therefore detailed description thereof will be omitted. In this embodiment, a case will be described in which the object model registered by the model registration unit 113 and the object model registered by the model registration unit 163 are the same object model, but the present invention is not limited to this.
[0127] The capture information registration unit 167 registers motion data for moving in a virtual space each of the human model and object model registered by the model registration unit 163. The capture information registration unit 167 also registers an action class to be used as an annotation when generating learning data.
[0128] Specifically, the capture information registration unit 167 receives an operation input from the developer to input the ID of the capture information to be used for generating the learning data. The capture information registration unit 167 uses the received ID as a key to acquire capture information including the ID from the capture information stored in the capture information storage unit 141. The capture information registration unit 167 registers the actor movement information and object model movement information included in the acquired capture information as motion data for operating the human model and object model registered by the model registration unit 163, respectively. The capture information registration unit 167 also registers the movement class included in the acquired capture information as an annotation to be used when generating the learning data.
[0129] The environment setting unit 169 performs various settings related to the environment of the virtual space. For example, the environment setting unit 169 sets a background image and a background model to be placed in the virtual space. In this embodiment, the background image and the background model include, for example, an image or a model for reproducing the environment of an electronic device manufacturing site, such as another mounting machine model placed adjacent to an object model (mounting machine model), but are not limited to these. Also, for example, the environment setting unit 169 sets parameters related to the angle of view of a virtual camera placed in the virtual space and parameters related to lighting placed in the virtual space.
[0130] For example, when a developer uses the input device 15 to perform an operation input requesting the display of a video setting screen, the virtual environment setting unit 161 displays the video setting screen on the display device 14. The video setting screen is a UI screen for various settings for generating a video in which a human model that reproduces the movements of a captured actor operates an object model.
[0131] Fig. 13 is a diagram showing an example of a video setting screen 701 according to this embodiment. As shown in Fig. 13, the video setting screen 701 includes a person model registration area 711, an object model registration area 721, a background selection button 731, a camera setting area 741, a lighting setting area 751, a model placement area 761, a text box 771, and a generation start button 781.
[0132] The human model registration area 711 includes a registration button 713 for registering a human model, and an expansion button 719. The human model registration area 711 and the registration button 713 are similar to the human model registration area 311 and the registration button 713 described in Fig. 7, respectively, and therefore detailed description thereof will be omitted. Note that items corresponding to the placement information 315 described in Fig. 7 are displayed in the human model registration area 711 by pressing the expansion button 719.
[0133] In this embodiment, the model registration unit 163 registers the person model 421 (worker model) as the person model, as in Fig. 7, but is not limited to this. For example, the model registration unit 163 may register a person model that differs from the person model 421 in terms of body type or gender.
[0134] The object model registration area 721 includes an add button 723 for additionally registering an object model, a delete button 725 for deleting an object model, and an expand button 729. The object model registration area 721, the add button 723, and the delete button 725 are similar to the object model registration area 321, the add button 323, and the delete button 329 described in Fig. 7, respectively, and therefore detailed description thereof will be omitted. Note that items corresponding to the select button 325 and the placement information 327 described in Fig. 7 are displayed in the object model registration area 721 by pressing the expand button 729.
[0135] In this embodiment, the model registration unit 163 registers the object model 471 (mounting machine model) as the object model, as in FIG. 7, but is not limited to this.
[0136] The background selection button 731 is an operation button for selecting a background image or background model to be placed in the virtual space. For example, a developer presses the background selection button 731 to select a background image or background model to be placed in the virtual space. This allows the background setting for reproducing movement using the person model 421 and the object model 471. Note that in this embodiment, the environment setting unit 169 selects the background model 481, which is the same implementation model as the object model 471, as the background model, but this is not limited to this.
[0137] The camera setting area 741 includes placement information 743 indicating the placement and zoom of the virtual camera in the virtual space, and a camera effect setting button 745 for setting effects to be applied to the video captured by the virtual camera.
[0138] The placement information 743 displays the placement position and angle of the virtual camera in the virtual space, as well as the zoom ratio. The camera effect setting button 745 is an operation button for setting blur processing or lens distortion processing to be applied as an effect to a video captured by the virtual camera. For example, when a developer presses the camera effect setting button 745 and selects blur processing or lens distortion processing, the environment setting unit 169 sets the selected processing to be applied to the video being captured.
[0139] Because videos captured by a virtual camera are CG images, they typically do not suffer from blurring or lens distortion caused by the camera lens, as occurs in videos captured by a real camera. On the other hand, when using a recognition model to recognize the movements of a worker operating a mounting machine, the video of the recognition target input to the recognition model is video captured by a camera installed at an electronics manufacturing site, and may contain blurring, lens distortion, etc. For this reason, in this embodiment, even if the video of the recognition target contains blurring, lens distortion, etc., it is possible to apply the above-described effects to videos generated as training data so that a recognition model can be generated that can correctly recognize the movements of a worker operating a mounting machine.
[0140] The lighting setting area 751 includes an expand button 759. When the developer presses the expand button 759, a setting UI for parameters related to lighting to be placed in the virtual space is displayed in the lighting setting area 751. The environment setting unit 169 sets the lighting-related parameters based on the operation input by the developer to the parameter setting UI.
[0141] The model placement area 761 shows, in plan view, the initial placement in virtual space of the person model 421 and object model 471 registered by the model registration unit 163, the background model 481 registered by the environment setting unit 169, and the virtual camera 491. In the model placement area 761, the person model 421, the object model 471, the background model 481, and the virtual camera 491 are placed in plan view. Note that the method for changing the initial placement of the person model 421, the object model 471, the background model 481, and the virtual camera 491 in virtual space is the same as that explained for the model placement area 331 in Fig. 7, and therefore a detailed explanation will be omitted.
[0142] However, the person model 421 is made to reproduce the movements of the performer captured by the motion capture system 30 and operate the object model 471. For this reason, the arrangement of the person model 421 and the object model 471 needs to maintain the relative positions of the person model 421 and the object model 471 in their initial arrangement at the time of capture by the motion capture system 30. In other words, positions are set for the person model 421 and the object model 471 to reproduce the movements.
[0143] 13, the virtual camera 491 is placed in the virtual space so as to correspond to the position of a camera in the real environment that captures the movements of the worker operating the mounting machine. For example, the virtual camera 491 is placed in the virtual space so that the positional relationship between the worker and the camera in the real space corresponds to the positional relationship between the person model 421 and the virtual camera 491 in the virtual space, but this is not limited to this. In the example shown in FIG. 13, the initial position of the virtual camera 491 also corresponds to the position of the third-person perspective virtual camera that captured the confirmation video in the motion capture system 30, but this is not limited to this.
[0144] In the example shown in FIG. 13, a background model 481 is placed adjacent to the object model 471 in order to reproduce the environment of an electronic device manufacturing site.
[0145] 13 is shown in a planar view, but the display form of the model placement area 761 is not limited to this. Like the model placement area 331 described in Fig. 7, the model placement area 761 can display the model placement area captured from any viewpoint position, and the viewpoint position at this time is displayed in the model placement area 761 as a viewpoint position display 763.
[0146] Furthermore, as described above, the placement information 743 displays the position and angle of the placement of the virtual camera 491 in the virtual space, as well as the zoom ratio.
[0147] The text box 771 is an input field where the ID of the capture information used to generate learning data is entered. In the example shown in Fig. 13, the developer uses the input device 15 to input the ID "1" into the text box 771.
[0148] The generation start button 781 is an operation button for starting the generation of a video in which a human model reproducing the captured movements of an actor manipulates an object model. When the developer performs an operation input of pressing the generation start button 781 using the input device 15, the capture information registration unit 167 acquires capture information using the ID entered in the text box 771 as a key, and registers actor movement information, object model movement information, and an action class. The virtual space control unit 171 also generates a virtual space, controls the movements of the human model and object model within the virtual space, and captures images of the human model and object model within the virtual space. Hereinafter, various controls related to the virtual space performed by the virtual space control unit 171 will be specifically described, appropriately referring to the video generation unit 173 included in the virtual space control unit 171.
[0149] The virtual space control unit 171 generates a virtual space in which the person model and object model registered by the model registration unit 163, the background model 481 registered by the environment setting unit 169, and the virtual camera 491 are placed at a position determined by the placement determination unit 115.
[0150] Furthermore, the virtual space control unit 171 operates the person model registered by the model registration unit 163 based on the actor movement information registered by the capture information registration unit 167, and operates the object model registered by the model registration unit 163 based on the object model movement information registered by the capture information registration unit 167. In the example shown in FIG. 13 , the virtual space control unit 171 causes the person model 421 to reproduce the movement of the actor based on the actor movement information registered by the capture information registration unit 167. Furthermore, the virtual space control unit 171 causes the object model registered by the model registration unit 113 to reproduce the movement of the object model registered by the model registration unit 163 based on the object model movement information registered by the capture information registration unit 167.
[0151] In this embodiment, the actor's movements indicated by the actor movement information are the movements of the actor's head and fingers, and therefore the actor movement information cannot move parts of the human model other than the head and fingers, such as the shoulders and lower body. For this reason, like the movement control unit 127 of the motion capture system 30, the virtual space control unit 171 estimates the movements (postures) of parts of the human model other than the head and fingers, such as the shoulders and lower body, using inverse kinematics (IK), but is not limited to this.
[0152] As described above, the virtual space control unit 171 moves the object model using the object model movement information. In this embodiment, both the performer movement information and the object model movement information are captured at the start of motion capture, and the movement indicated by the performer movement information and the movement indicated by the object model movement information are synchronized. Therefore, the virtual space control unit 171 can move the object model in conjunction with the movement of the virtual hand simply by moving the object model using the object model movement information. In the example shown in FIG. 13 , by moving the object model 471 using the object model movement information, the virtual space control unit 171 can open the safety cover 472 in conjunction with the human model 421 moving the virtual hand 422 upward and slightly toward the back while grasping the handle portion 473. Therefore, unlike the motion control unit 127 of the motion capture system 30, the virtual space control unit 171 does not need to determine contact between the virtual hand of the human model and the object model or use the movement of the performer's fingers to move the object model in conjunction with the movement of the virtual hand.
[0153] The video generation unit 173 controls the virtual camera 491, which is initially placed at a predetermined position in the virtual space by the virtual space control unit 171, to capture a scene in which a human model operates an object model based on the actor movement information and the object model movement information, thereby generating a video for training data. For example, the video generation unit 173 generates an image by having the virtual camera 491 project (render) various models, such as human models and object models, and backgrounds, etc., included in the angle of view (field of view) of the virtual camera 491 onto a projection surface (not shown). In rendering, the world coordinate system, which is a coordinate system in the virtual space, is converted into a camera coordinate system whose origin is the viewpoint position, which is the position where the virtual camera 491 is placed in the virtual space, and the various models, backgrounds, etc. expressed in the camera coordinate system are converted into a coordinate system on a two-dimensional projection surface by perspective projection transformation. Note that a well-known method, such as the Z-buffer method, may be used as a rendering method. The video generation unit 173 generates a video for training data by having the virtual camera 491 perform the above-mentioned image generation for each frame.
[0154] 13 differs from the time of capture by the motion capture system 30 in that there is a background model 481, but the arrangement of the person model 421 and the object model 471 is the same as that of the time of capture, and the arrangement of the virtual camera 491 is the same as that of the virtual camera for a third-person perspective. Therefore, the video for learning data generated by the video generation unit 173 is the same as the confirmation video 501, except for the presence of the background model 481. For example, the video for learning data generated by the video generation unit 173 is a video in which the virtual hand of the person model operates the safety cover of the object model (mounting machine model) in accordance with the hand movement of the actor captured by the motion capture system 30.
[0155] When the virtual space control unit 171 finishes playing back the actor movement information and the object model movement information and finishes generating the video for the learning data, the virtual environment setting unit 161 displays a video confirmation screen on the display device 14. The video confirmation screen is a screen that is played back to allow the developer to check the video for the learning data generated by the video generation unit 173.
[0156] 14 is a diagram showing an example of a video confirmation screen 901 according to this embodiment. As shown in FIG. 14, the video confirmation screen 901 includes a back button 911, a video display area 913, a play button 915, a seek bar 917, and a register button 919.
[0157] The back button 911 is a button for returning to the video setting screen 701. When the developer performs an operation input of pressing the back button 911, the virtual environment setting unit 161 terminates the display of the video confirmation screen 901 and redisplays the video setting screen 701 shown in Fig. 13. Note that the redisplayed video setting screen 701 may be provided with an operation button for returning to the video confirmation screen 901 again. In other words, the virtual environment setting unit 161 may be configured to be able to switch between displaying the video setting screen 701 and the video confirmation screen 901.
[0158] The video display area 913 is an area where the video for learning data generated by the video generation unit 173 is displayed, and the video for learning data 801 is displayed. As described above, the video for learning data 801 is the same as the confirmation video 501 except that it includes a background model 481. The play button 915 is an operation button for playing the video for learning data 801. The seek bar 917 is a UI component that displays the playback position (playback frame) of the video for learning data 801 with a slider.
[0159] When the developer performs an operation input of pressing the play button 915, the virtual environment setting unit 161 plays the video 801 for learning data. The developer can check the movement of the person model 421 (worker model) operating the object model 471 (mounting machine model) on the video, and confirm whether it can be used for learning data.
[0160] The registration button 919 is an operation button for registering a video for learning data as learning data. When the developer performs an operation input to press the registration button 919, the learning data generation unit 181 generates learning data by annotating the video generated by the video generation unit 173 with the action class registered by the capture information registration unit 167. The learning data generation unit 181 outputs the generated learning data to the auxiliary storage device 13 or to the outside via the communication device 16.
[0161] The training data generation device 40 of this embodiment is capable of generating not one type of training data but multiple types of training data from captured information. For example, the training data generation device 40 can generate multiple types of training data from the same captured information by changing at least one of the type of human model, lighting settings, shooting position and angle of view of the virtual camera, position of the object model, and virtual camera effects.
[0162] 15 and 16, an example of generating a video for learning data different from the video for learning data 801 using the capture information with the above-mentioned ID "1" will be described below. Here, an example will be described in which the arrangement of object models is made different from that in FIG. 13. FIG. 15 is a diagram showing an example of the video setting screen 701 of this embodiment. In the arrangement of each model in the model arrangement area 761 shown in FIG. 15, the positions of the object model 471 and the background model 481 are swapped with respect to the arrangement in the model arrangement area 761 shown in FIG. 13. Furthermore, in order to maintain the positional relationship between the person model 421 and the object model 471 at the time of capture by the motion capture system 30, the arrangement has been changed so that the person model 421 is located in front of the object model 471.
[0163] That is, the object model 471 (second object model) registered by the model registration unit 163 is placed at a position different from that of the object model 471 (first object model) registered by the model registration unit 113. Furthermore, the person model 421 is placed so that the positional relationship with respect to the object model 471 (second object model) registered by the model registration unit 163 becomes the positional relationship of the performer (person model 421) with respect to the object model 471 (first object model) registered by the model registration unit 113.
[0164] In the learning data generation device 40 of this embodiment, the virtual space control unit 171, as described above, uses the performer movement information and the object model movement information to move the person model 421 and the object model 471. Therefore, as shown in the example of Fig. 15, even if the arrangement changes, as long as the positional relationship between the person model 421 and the object model 471 is maintained, the virtual space control unit 171 can open the safety cover 472 in conjunction with the person model 421 moving the virtual hand 422 upward and slightly toward the back while holding the handle portion 473.
[0165] In particular, as described above, the virtual space control unit 171 does not need to determine contact between the virtual hand of the person model and the object model or use the movement of the performer's fingers in order to move the object model in conjunction with the movement of the virtual hand, as does the movement control unit 127 of the motion capture system 30. Therefore, according to this embodiment, even if contact between the virtual hand 422 of the person model 421 and the object model 471 does not occur due to the effect of shifting the positions of the person model 421 and the object model 471, it is possible to reliably generate a moving image in which the person model 421 operates the object model 471.
[0166] 16 is a diagram showing an example of a video 811 for learning data generated by the video generating unit 173 of this embodiment. As shown in Fig. 16, the video 811 for learning data differs from the video 801 for learning data in that it is a video in which a person model 421 operates an object model 471 (mounting machine model) placed at the back.
[0167] FIG. 17 is a flowchart showing an example of the capture process performed in the motion capture system 30 of this embodiment.
[0168] First, the model registration unit 113 registers a human model and an object model selected by the developer from among the human models and object models stored in the model storage unit 101 (step S101).
[0169] Next, the placement determination unit 115 determines the initial placement in the virtual space of the human model and object model registered by the model registration unit 113 (step S103).
[0170] Next, the additional information setting unit 117 sets an ID and a movement class to be added to the captured movement of the performer, etc. (step S105).
[0171] Next, when the developer presses the capture start button 361, the capture setting unit 111 starts the capture process (step S107). Specifically, the virtual space control unit 121 generates a virtual space in which the person model and object model registered by the model registration unit 113 are placed at positions determined by the placement determination unit 115. The virtual space control unit 121 also places a third-person perspective virtual camera for generating a confirmation video and a first-person perspective virtual camera for generating a feedback video to the performer in the generated virtual space.
[0172] Next, the capture unit 201 captures the movement of the performer wearing the VR headset 20, and the virtual space control unit 121 captures the movement of the object model (step S109).
[0173] Next, the operation control unit 127 detects whether or not the virtual hand of the person model has come into contact with the object model (step S111).
[0174] If it is detected that the virtual hand is in contact with the object model (Yes in step S111), the movement control unit 127 determines whether the distance between the tip of the performer's thumb and the tips of each finger captured by the capture unit 201 is less than a threshold value (step S113).
[0175] If the distance between the fingertips is less than the threshold value (Yes in step S113), the movement control unit 127 moves the character model in accordance with the movement of the performer captured by the capture unit 201, and moves the object model in conjunction with the movement of the virtual hand of the character model (step S115).
[0176] On the other hand, if the virtual hand and the object model are not in contact (No in step S111), or if the distance between the fingertips is not less than the threshold (No in step S113), the movement control unit 127 moves the character model independently in accordance with the movement of the performer captured by the capture unit 201 (step S117).
[0177] Next, the confirmation video generation unit 123 controls a virtual camera for a third-person perspective, captures a scene in which a human model manipulates an object model based on the actor's movements captured by the capture unit 201, and generates frame images of the confirmation video (step S119).
[0178] Next, the feedback video generation unit 125 controls a virtual camera for a first-person perspective, captures a scene in which a human model manipulates an object model based on the movements of the actor captured by the capture unit 201, generates frame images of the feedback video, and updates the feedback video displayed on the display device 24 (step S121).
[0179] Subsequently, the processing of steps S111 to S123 is repeated for each frame until the developer performs an operation to end the motion capture (No in step S123).
[0180] Subsequently, when the developer performs an operation to end the motion capture (Yes in step S123), the capture setting unit 111 displays a capture confirmation screen (step S125).
[0181] Next, when the developer checks the check video generated by the check video generation unit 123 and presses the register button 621, the capture information saving unit 131 saves the capture information in the capture information storage unit 141 (step S127). The capture information is information including actor movement information indicating the actor's movement captured by the capture unit 201, object model movement information indicating the object model's movement captured by the movement control unit 127, and the ID and movement class set by the additional information setting unit 117.
[0182] FIG. 18 is a flowchart showing an example of the training data generation process performed by the training data generation device 40 of this embodiment.
[0183] First, the model registration unit 163 registers a human model and an object model selected by the developer from among the human models and object models stored in the model storage unit 101 (step S201).
[0184] Next, the environment setting unit 169 sets a background image and a background model to be placed in the virtual space selected by the developer (step S203).
[0185] Next, the environment setting unit 169 sets the virtual camera effect selected by the developer (step S205).
[0186] Next, the environment setting unit 169 sets parameters related to lighting to be placed in the virtual space selected by the developer (step S207).
[0187] Next, the placement determination unit 165 determines the initial placement of each model and the virtual camera in the virtual space (step S209).
[0188] Next, the virtual environment setting unit 161 sets the ID of the capture information used to generate the learning data (step S211).
[0189] Next, when the developer presses the generation start button 781, the virtual environment setting unit 161 starts generating a video for learning data in which a human model reproducing the captured movements of the actor manipulates an object model (step S213). Specifically, the capture information registration unit 167 acquires capture information using the set ID as a key, and registers actor movement information, object model movement information, and action class. Furthermore, the virtual space control unit 171 generates a virtual space in which the human model and object model registered by the model registration unit 163, the background model 481 registered by the environment setting unit 169, and the virtual camera 491 are placed at the position determined by the placement determination unit 115.
[0190] Next, the virtual space control unit 171 operates the human model registered by the model registration unit 163 according to the actor movement information registered by the capture information registration unit 167, and operates the object model registered by the model registration unit 163 according to the object model movement information registered by the capture information registration unit 167. The video generation unit 173 controls the virtual camera 491 initially placed at a predetermined position in the virtual space by the virtual space control unit 171, captures a scene in which the human model operates the object model based on the actor movement information and the object model movement information, and generates a video for learning data (step S215).
[0191] Next, when the virtual space control unit 171 finishes playing back the actor movement information and the object model movement information and finishes generating the moving image for learning data, the virtual environment setting unit 161 displays a moving image confirmation screen (step S217).
[0192] Next, when the developer checks the video for learning data generated by the video generation unit 173 and presses the registration button 919 (Yes in step S219), the learning data generation unit 181 generates learning data by annotating the video generated by the video generation unit 173 with the action class registered by the capture information registration unit 167 (step S221).
[0193] If the developer does not press the registration button 919 (No in step S219), the process of step S221 is not performed.
[0194] As described above, the training data generation system of this embodiment uses a VR headset to virtually recreate an object that does not exist in the field as an object model, and captures the movements of an actor wearing the VR headset virtually manipulating the object model. Furthermore, the training data generation system of this embodiment uses the captured movements to create a scene in which a human model reproducing the actor's movements manipulates the object model in a virtual space, and animates the scene to generate training data.
[0195] Therefore, according to this embodiment, even if an object to be operated by an actor is not actually prepared, video useful for training a recognition model that recognizes the movements of manipulating the object can be generated as training data.
[0196] Furthermore, the training data generation system of this embodiment uses a camera attached to a VR headset to capture the motion of the performer, making it possible to capture the precise movements of the performer's fingers when virtually manipulating an object model. Therefore, this embodiment makes it possible to generate, as training data, videos that are useful for training a recognition model that recognizes movements that manipulate objects, such as those that require precise finger movements.
[0197] Furthermore, in the learning data generation system of this embodiment, when motion capture is performed, not only the movements of the performer but also the movements of the object model are captured. Note that when motion capture is performed, if the performer's hand and the object model are virtually in contact with each other, the object model is moved in conjunction with the movement of the performer's hand. On the other hand, when learning data is generated, the object model is moved in conjunction with the virtual hand of the human model based on the captured movement of the object model, regardless of whether or not the virtual hand, which is the hand of the human model, is in contact with the object model.
[0198] In the training data generation system of this embodiment, in order to expand the data, the above-mentioned scene in which a human model operates an object model may be generated by shifting the position of the human model and the object model while maintaining their relative positions, thereby generating training data that is a variation of the scene. In this embodiment, even in this case, the object model is moved in conjunction with the human model's virtual hand based on the captured movement of the object model, regardless of whether or not there is contact between the human model's virtual hand and the object model. Therefore, according to this embodiment, even if there is no contact between the human model's virtual hand and the object model due to the effect of shifting the position of the human model and the object model, it is possible to reliably generate a video in which a human model operates an object model.
[0199] (Variation 1) In the above embodiment, an example has been described in which a collider is used to detect contact between a person model and an object model, but the method for detecting contact between a person model and an object model is not limited to this. The movement control unit 127 may detect actual contact between a person model and an object model, or may detect contact using a positional relationship, such as a distance, between the person model and the object model. In the latter case, the movement control unit 127 may, for example, set reference points on the person model and the object model and detect interaction based on the distance and positional relationship between the reference points.
[0200] (Variation 2) In the motion capture system 30 of the above embodiment, when contact occurs between the virtual hand of the human model and the object model, the feedback video generation unit 125 may display a symbol indicating that contact has occurred on at least one of the human model and the virtual hand. In this way, feedback can be provided to the performer that the object model can be grasped with the virtual hand.
[0201] Similarly, in the motion capture system 30 of the above embodiment, when the virtual hand of a human model is grasping an object model, the feedback video generation unit 125 may superimpose a symbol indicating that the virtual hand is grasping the object model on at least one of the human model and the virtual hand. In this way, feedback can be provided to the performer that the virtual hand is grasping the object model.
[0202] (Variation 3) In the above embodiment, an example of realizing the motion capture system 30 using a VR headset has been described, but this is not limited to this, and the motion capture system 30 may also be realized using AR (Augmented Glasses) or an MR headset.
[0203] (Variation 4) In the motion capture system 30 of the above embodiment, the motion control unit 127 (an example of a correction unit) may correct the position of the virtual hand to a predetermined position relative to the safety cover regardless of the position of the performer's hand when the virtual hand of the human model and the safety cover of the object model (mounting machine model) are within a predetermined distance. For example, if the performer's hand virtually grasps the handle of the safety cover and the performer's hand moves off the trajectory for opening the safety cover, the safety cover should not move and the virtual hand and safety cover will no longer be in contact with each other, causing the virtual hand to separate from the safety cover.
[0204] In contrast, in Modification 4, even if the performer's hand deviates from the trajectory for opening the safety cover and moves to a position where the virtual hand would not normally come into contact with the safety cover, the motion control unit 127 may correct the position of the virtual hand so that the virtual hand does not separate from the safety cover but moves along the trajectory. Furthermore, the virtual space control unit 171 of the training data generation device 40 may reproduce this control when generating a video for training data to generate the video. In this way, it is possible to generate the intended video for training data even if there is some deviation when the performer virtually operates the movable part of the object model.
[0205] (Variation 5) In the above embodiment, an example has been described in which the object model (first object model) registered by the model registration unit 113 and the object model (second object model) registered by the model registration unit 163 are the same object model. However, in Modification 5, the first object model and the second object model may be different object models. Specifically, the first object model and the second object model may have the same movable parts but different parts other than the movable parts. For example, the first object model and the second object model may have the same configuration and arrangement of the movable parts but different configurations and arrangements of the parts other than the movable parts. Also, for example, the first object model and the second object model may have the same configuration and arrangement but different colors.
[0206] In this way, learning data can be generated not only using the mounting machine model used during capture by the motion capture system 30, but also using a mounting machine model of a successor model that has undergone minor changes to the mounting machine model, or a mounting machine model of a different color from the mounting machine model. Therefore, according to Modification 5, it becomes possible to generate a wider variety of learning data from captured information. This makes it possible to prepare learning data before equipment such as similar successor models or mounting machines of different colors is actually introduced to the site, so that an action recognition system using the learning data can be put into operation immediately after equipment is replaced on site.
[0207] (Variation 6) In the above embodiment, an example has been described in which an object model is moved based on object model movement information in a scene where learning data is generated. However, in Modification 6, in a scene where motion capture is performed, the movement of the object model virtually operated by the performer may not be captured, and object model movement information may not be generated. In this case, in a scene where learning data is generated, as in a scene where motion capture is performed, when the virtual hand of the human model and the object model are in contact with each other, the object model may be moved in conjunction with the movement of the virtual hand.
[0208] (Variation 7) In the above embodiment, an example has been described in which, in a scene where motion capture is performed, the movements of a performer are captured to generate performer movement information. However, in Modification 7, in a scene where motion capture is performed, instead of capturing the movements of a performer, the movements of a human model that reproduces the movements of the performer in real time may be captured to generate performer movement information.
[0209] (program) The programs executed by the information processing device 10 (motion capture system, learning data generation device) of the above-mentioned embodiment and each of the above-mentioned modified examples are provided as files in an installable or executable format stored on a computer-readable storage medium such as a CD-ROM, CD-R, memory card, DVD, or flexible disk (FD).
[0210] Furthermore, the programs executed by the information processing device 10 (motion capture system, learning data generation device) of the above embodiment and each of the above modified examples may be stored on a computer connected to a network such as the Internet and provided by being downloaded via the network. Furthermore, the programs executed by the information processing device 10 (motion capture system, learning data generation device) of the above embodiment and each of the above modified examples may be provided or distributed via a network such as the Internet. Furthermore, the programs executed by the information processing device 10 (motion capture system, learning data generation device) of the above embodiment and each of the above modified examples may be provided by being pre-installed in a ROM or the like.
[0211] The programs executed by the information processing device 10 (motion capture system, learning data generation device) of the above-described embodiment and each of the above-described modifications have a modular configuration for realizing the above-described units on a computer. As for actual hardware, for example, the CPU reads the learning program from the HDD onto the RAM and executes it, thereby realizing the above-described units on a computer.
[0212] As described above, according to the above embodiment and each of the above variants, it is possible to generate videos that are useful for training a recognition model that recognizes the action of manipulating an object, even if the object to be manipulated by a person is not actually prepared.
[0213] The above-described embodiment and each of the modifications merely illustrate examples of implementations of the present disclosure, and the technical scope of the present disclosure should not be construed as being limited by these. Therefore, the present disclosure can be implemented in various forms without departing from the spirit or main features thereof. For example, the above-described embodiment and each of the modifications may be appropriately combined in their respective constituent units. Furthermore, for example, some components may be deleted from all components in the above-described embodiment and each of the modifications.
[0214] The present disclosure includes the following aspects.
[0215] (1) a first display control unit that displays a device model having a moving part; a capture unit that captures the hand movements of an actor virtually operating the movable unit; a moving image generating unit that generates a moving image in which a virtual hand of a human model operates the movable part in accordance with the captured hand movement; A video generation system comprising:
[0216] (2) The movable part is movable along a predetermined trajectory. The video generation system according to (1) above.
[0217] (3) The robot further includes a movement control unit that moves the movable unit in accordance with the movement of the virtual hand when the virtual hand and the movable unit are within a predetermined distance from each other. The video generation system according to (1) above.
[0218] (4) A correction unit is further provided that corrects the position of the virtual hand to a predetermined position relative to the movable unit when the virtual hand and the movable unit are within a predetermined distance, regardless of the position of the performer's hand. The video generation system according to any one of (1) to (3) above.
[0219] (5) A second display control unit is further provided which registers a character model and the device model that reproduce the hand movements of the actor in real time, sets the positions of the registered character model and the device model, and switchably displays a setting screen for displaying the positioning status of the registered character model and the device model, and a confirmation screen for checking a video of the registered character model operating the device model. The video generation system according to (1) above.
[0220] (6) The setting screen is a screen for setting the type of operation to be performed by the registered human model. (5) The video generation system according to claim 5 above.
[0221] (7) a display step in which the first display control unit displays a device model having a movable part; a capture step in which a capture unit captures hand movements of an actor virtually operating the movable unit; a moving image generating step in which a moving image generating unit generates a moving image in which a virtual hand of a human model operates the movable part in accordance with the captured hand movement; A video generation method including:
[0222] (8) a display control step of displaying a device model having a moving part; an acquisition step of acquiring hand movements of an actor virtually operating the movable part; a moving image generating step of generating a moving image in which a virtual hand of a human model operates the movable part in accordance with the acquired hand movement; A program that causes a computer to execute the following. [Explanation of symbols]
[0223] 1. Training data generation system 10. Information processing equipment 11 Control device 12 Main storage 13 Auxiliary storage device 14 Display device 15 Input Devices 16. Communications equipment 17 Various buses 19 Communication Cable 20 VR headsets 21 Control device 22 Main storage 23 Auxiliary storage device 24 Display device 25 Imaging equipment 27 Various buses 30 Motion Capture System 40 Learning data generation device 101 Model memory section 111 Capture Settings 113 Model Registration Department 115 Placement determination section 117 Additional information setting section 121 Virtual Space Control Unit 123 Confirmation video generation unit 125 Feedback video generation unit 127 Motion control section 131 Capture information storage unit 141 Capture information storage unit 161 Virtual Environment Settings 163 Model Registration Department 165 Placement determination section 167 Capture Information Registration Section 169 Environment Settings 171 Virtual Space Control Unit 173 Video Generation Unit 181 Learning Data Generation Unit 201 Capture section 211 Feedback video display control unit
Claims
1. a first display control unit that displays a device model having a movable part; a capture unit that captures the hand movements of an actor virtually operating the movable unit; a moving image generating unit that generates a moving image in which a virtual hand of a human model operates the movable part in accordance with the captured hand movement; A video generation system comprising:
2. The movable part is movable along a predetermined trajectory. The video generation system of claim 1 .
3. a movement control unit that moves the movable unit in accordance with the movement of the virtual hand when the virtual hand and the movable unit are within a predetermined distance from each other; The video generation system of claim 1 .
4. a correction unit that corrects the position of the virtual hand to a predetermined position relative to the movable unit when the virtual hand and the movable unit are within a predetermined distance, regardless of the position of the performer's hand; The video generation system according to any one of claims 1 to 3.
5. a second display control unit that registers a character model and the device model that reproduce the hand movements of the actor in real time, and sets the positions of the registered character model and the device model, and that switchably displays a setting screen for displaying the positioning status of the registered character model and the device model, and a confirmation screen for checking a video of the registered character model operating the device model; The video generation system of claim 1 .
6. The setting screen is further a screen for setting the type of operation to be performed by the registered human model. The video generation system according to claim 5 .
7. a display step in which a first display control unit displays an apparatus model having a movable part; a capture step in which a capture unit captures hand movements of an actor virtually operating the movable unit; a moving image generating step in which a moving image generating unit generates a moving image in which a virtual hand of a human model operates the movable part in accordance with the captured hand movement; A video generation method including:
8. a display control step of displaying an apparatus model having a movable part; an acquisition step of acquiring hand movements of an actor virtually operating the movable part; a moving image generating step of generating a moving image in which a virtual hand of a human model operates the movable part in accordance with the acquired hand movement; A program that causes a computer to execute the following.
Citation Information
Patent Citations
Information processing device and information processing program
JP2022140038A