System
The system records and optimizes human motion data from virtual reality gameplay to train humanoid robots with human-like movements, addressing the inefficiencies and costs of existing data collection methods.
Patent Information
- Application Number
- JP2024115231
- Authority / Receiving Office
- JP · JP
- Patent Type
- Applications
- Current Assignee / Owner
- Filing Date
- 2024-07-18
- Publication Date
- 2026-01-29
AI Technical Summary
Collecting high-quality human motion data for humanoid robots is time-consuming and costly, and there is a lack of effective methods to incorporate this data into learning models for human-like robot movements.
A system that records user actions in a virtual reality environment, converts the data into a specified format, and transmits it to a server for analysis and machine learning model generation, optimizing the model based on in-game success rates to train humanoid robots.
Efficiently collects and utilizes human motion data to train humanoid robots with human-like movements, improving motion accuracy and reducing costs.
Smart Images

Figure 2026014234000001_ABST
Abstract
Description
[Technical Field]
[0001] The technology of the present disclosure relates to a system. [Background technology]
[0002] Patent document 1 discloses a persona chatbot control method performed by at least one processor, the method including the steps of receiving a user utterance, adding the user utterance to a prompt including an instruction sentence related to a description of the chatbot character, encoding the prompt, and inputting the encoded prompt into a language model to generate a chatbot utterance in response to the user utterance. [Prior art documents] [Patent documents]
[0003] [Patent Document 1] Japanese Patent Publication No. 2022-180282 Summary of the Invention [Problem to be solved by the invention]
[0004] To adapt humanoid robots to human life, high-quality human motion data is necessary, but collecting such data is time-consuming and costly. Furthermore, no appropriate method has been established for effectively incorporating the acquired data into a learning model to make robots' movements more human-like. The present invention aims to solve these problems by providing a means for efficiently and inexpensively collecting human motion data and utilizing it in humanoid robot learning. [Means for solving the problem]
[0005] The present invention provides a means for recording a user's actions while playing a game in a virtual reality environment in real time, converting the recorded actions into a predetermined format, and transmitting the data to a server. The server then analyzes the received data, associates the user's actions with in-game actions, stores the data, and generates a machine learning model based on the data, which can then be used to train a humanoid robot. The system also provides a means for transmitting the recorded action data to a server in real time and optimizing the machine learning model based on the success rate of the user's in-game actions, thereby achieving efficient and effective data collection and model updating.
[0006] "User" refers to a person who uses the system to play games in a virtual reality environment.
[0007] A "virtual reality environment" refers to a virtual space that uses computer technology to provide an experience similar to the real world.
[0008] "Game" refers to content that is provided to users by computer programs for entertainment or for purposefully controlling their actions.
[0009] "Motion data" refers to information regarding physical movements made by a user in a virtual reality environment.
[0010] "Format" refers to the rules for organizing data into a particular form or structure.
[0011] "Server" refers to a computer system that receives, analyzes, stores, and provides services to data over a network.
[0012] "Analysis" refers to the act of examining received data in detail to clarify its meaning and relationships.
[0013] "Action" refers to the operations or actions that a user performs within a game.
[0014] A "machine learning model" refers to an algorithm or program that allows a computer to learn specific patterns and rules based on data and generate future predictions and behaviors.
[0015] A "humanoid robot" refers to a robot that has a form and movements similar to those of a human.
[0016] "Real-time" refers to nearly simultaneous data processing. [Brief explanation of the drawings]
[0017] [Figure 1] 1 is a conceptual diagram showing an example of the configuration of a data processing system according to a first embodiment. [Figure 2] 1 is a conceptual diagram showing an example of main functions of a data processing device and a smart device according to a first embodiment. [Figure 3] FIG. 10 is a conceptual diagram showing an example of the configuration of a data processing system according to a second embodiment. [Figure 4] FIG. 10 is a conceptual diagram showing an example of main functions of a data processing device and smart glasses according to a second embodiment. [Figure 5] FIG. 10 is a conceptual diagram showing an example of the configuration of a data processing system according to a third embodiment. [Figure 6] FIG. 11 is a conceptual diagram showing an example of main functions of a data processing device and a headset-type terminal according to a third embodiment. [Figure 7] FIG. 10 is a conceptual diagram showing an example of the configuration of a data processing system according to a fourth embodiment. [Figure 8] FIG. 10 is a conceptual diagram showing an example of main functions of a data processing device and a robot according to a fourth embodiment. [Figure 9] 1 shows an emotion map onto which multiple emotions are mapped. [Figure 10] 1 shows an emotion map onto which multiple emotions are mapped. [Figure 11] FIG. 3 is a sequence diagram showing a processing flow of the data processing system according to the first embodiment. [Figure 12]FIG. 10 is a sequence diagram showing the flow of processing in the data processing system in Application Example 1. [Figure 13] FIG. 10 is a sequence diagram showing the flow of processing in the data processing system according to the second embodiment when an emotion engine is combined. [Figure 14] FIG. 10 is a sequence diagram showing the flow of processing in the data processing system in Application Example 2 when an emotion engine is combined. DETAILED DESCRIPTION OF THE INVENTION
[0018] An example of an embodiment of a system according to the technology of the present disclosure will be described below with reference to the accompanying drawings.
[0019] First, the terms used in the following description will be explained.
[0020] In the following embodiments, a coded processor (hereinafter simply referred to as a "processor") may be a single arithmetic device or a combination of multiple arithmetic devices. Furthermore, a processor may be a single type of arithmetic device or a combination of multiple types of arithmetic devices. Examples of arithmetic devices include a CPU (Central Processing Unit), a GPU (Graphics Processing Unit), a GPGPU (General-Purpose computing on Graphics Processing Units), and an APU (Accelerated Processing Unit).
[0021] In the following embodiments, a coded RAM (Random Access Memory) is a memory in which information is temporarily stored and is used as a working memory by a processor.
[0022] In the following embodiments, the coded storage is one or more non-volatile storage devices that store various programs, various parameters, etc. Examples of non-volatile storage devices include flash memory (SSD (Solid State Drive)), magnetic disks (e.g., hard disks), and magnetic tapes.
[0023] In the following embodiments, a communication I / F (Interface) with a symbol is an interface including a communication processor, an antenna, etc. The communication I / F controls communication between multiple computers. Examples of communication standards applied to the communication I / F include wireless communication standards including 5G (5th Generation Mobile Communication System), Wi-Fi (registered trademark), Bluetooth (registered trademark), etc.
[0024] In the following embodiments, "A and / or B" is synonymous with "at least one of A and B." In other words, "A and / or B" means that it may be only A, only B, or a combination of A and B. Furthermore, in this specification, the same concept as "A and / or B" is also applied when three or more things are expressed connected by "and / or."
[0025] [First embodiment]
[0026] FIG. 1 shows an example of the configuration of a data processing system 10 according to the first embodiment.
[0027] 1, a data processing system 10 includes a data processing device 12 and a smart device 14. An example of the data processing device 12 is a server.
[0028] The data processing device 12 includes a computer 22, a database 24, and a communication I / F 26. The computer 22 is an example of a "computer" according to the technology of the present disclosure. The computer 22 includes a processor 28, a RAM 30, and a storage 32. The processor 28, the RAM 30, and the storage 32 are connected to a bus 34. The database 24 and the communication I / F 26 are also connected to the bus 34. The communication I / F 26 is connected to a network 54. Examples of the network 54 include a WAN (Wide Area Network) and / or a LAN (Local Area Network).
[0029] The smart device 14 includes a computer 36, a reception device 38, an output device 40, a camera 42, and a communication I / F 44. The computer 36 includes a processor 46, a RAM 48, and a storage 50. The processor 46, the RAM 48, and the storage 50 are connected to a bus 52. The reception device 38, the output device 40, and the camera 42 are also connected to the bus 52.
[0030] The reception device 38 includes a touch panel 38A, a microphone 38B, and the like, and receives user input. The touch panel 38A detects contact with an indicator (for example, a pen or a finger) to receive user input by the touch of the indicator. The microphone 38B detects the user's voice to receive user input by voice. The control unit 46A transmits data indicating the user input received by the touch panel 38A and the microphone 38B to the data processing device 12. In the data processing device 12, the specific processing unit 290 acquires the data indicating the user input.
[0031] The output device 40 includes a display 40A and a speaker 40B, and presents data to the user 20 by outputting the data in a form of expression that the user 20 can perceive (for example, audio and / or text). The display 40A displays visible information such as text and images in accordance with instructions from the processor 46. The speaker 40B outputs audio in accordance with instructions from the processor 46. The camera 42 is a compact digital camera equipped with an optical system including a lens, aperture, and shutter, and an imaging element such as a CMOS (Complementary Metal-Oxide-Semiconductor) image sensor or a CCD (Charge Coupled Device) image sensor.
[0032] The communication I / F 44 is connected to a network 54. The communication I / Fs 44 and 26 control the exchange of various information between the processor 46 and the processor 28 via the network 54.
[0033] FIG. 2 shows an example of the main functions of the data processing device 12 and the smart device 14.
[0034] 2, in the data processing device 12, a specific process is performed by the processor 28. A specific processing program 56 is stored in the storage 32. The specific processing program 56 is an example of a "program" according to the technology of the present disclosure. The processor 28 reads the specific processing program 56 from the storage 32 and executes the read specific processing program 56 on the RAM 30. The specific process is realized by the processor 28 operating as a specific processing unit 290 in accordance with the specific processing program 56 executed on the RAM 30.
[0035] The storage 32 stores a data generation model 58 and an emotion identification model 59. The data generation model 58 and the emotion identification model 59 are used by the identification processing unit 290.
[0036] In the smart device 14, the processor 46 performs the reception output process. The storage 50 stores a reception output program 60. The reception output program 60 is used in conjunction with the specific processing program 56 by the data processing system 10. The processor 46 reads the reception output program 60 from the storage 50 and executes the read reception output program 60 on the RAM 48. The reception output process is realized by the processor 46 operating as the control unit 46A in accordance with the reception output program 60 executed on the RAM 48.
[0037] Next, a description will be given of the specific processing performed by the specific processing unit 290 of the data processing device 12. In the following description, the data processing device 12 will be referred to as a "server" and the smart device 14 will be referred to as a "terminal."
[0038] The present invention is a system that collects user actions while playing a game in a virtual reality (VR) environment and uses the collected actions for motion learning of a humanoid robot. Specific embodiments will be described below.
[0039] Overall system overview
[0040] The system mainly consists of the following components:
[0041] 1. VR games played by users and motion capture suits
[0042] 2. Devices that collect and transmit operational data
[0043] 3. Server that receives and analyzes data and generates machine learning models
[0044] User Preparation
[0045] User: First, the user puts on a VR headset and a motion capture suit, launches a dedicated application, and begins the game. The motion capture suit has built-in sensors that collect data such as the position of each joint and the speed of movement in real time.
[0046] Recording gameplay
[0047] Device: When a user plays a game, it receives real-time motion data from the motion capture suit. For example, when a user plays a shooting game, arm movement and hand position data are collected.
[0048] Terminal: Converts the received data into a specified format (e.g., CSV or JSON) and adds a timestamp and user ID. This data includes the X, Y, Z coordinates and angles of each joint, as well as movement speed.
[0049] Data transmission and storage
[0050] Terminal: Sends formatted motion data to the server, either in real time or in batches.
[0051] Server: Receives data sent from the device and stores it in a database. For example, it stores organized data using MongoDB or an SQL database.
[0052] Data analysis and generation of learning models
[0053] Server: Analyzes the stored data. For example, creates a data frame using a Python library (Pandas, NumPy, etc.) and calculates user behavior characteristics, game success rate, etc.
[0054] Server: Generates a machine learning model based on the analysis results. Machine learning libraries such as TensorFlow and PyTorch are used to generate the model. For example, a deep learning model is built to reproduce the arm movements of a user when shooting an enemy.
[0055] Application in humanoid robots
[0056] Server: Incorporates the generated machine learning model into the control program of the humanoid robot.
[0057] Humanoid robot (user-like): Based on a new model, the robot performs human-like actions. For example, the robot replicates the shooting actions performed by the user in a game.
[0058] As a concrete example, consider the movements of a user playing a shooting game. In this case, a motion capture suit records the user's movements as they move their arms to aim and fire. This data is then sent in real time to the device and then to a server, where it is analyzed and correlated with the user's in-game score. This information is used to generate a machine learning model that allows a humanoid robot to replicate similar shooting movements.
[0059] In this way, the present invention provides a system that can efficiently collect high-quality human motion data and use it to learn the movements of humanoid robots.
[0060] The processing flow will be explained below.
[0061] Step 1:
[0062] User: Put on the VR headset and motion capture suit, launch the VR game application, and start the game.
[0063] Step 2:
[0064] Terminal: Receives real-time user movement data from the motion capture suit, such as the user's arm movements and joint position data.
[0065] Step 3:
[0066] Terminal: Converts the received movement data into a specified format (CSV, JSON, etc.). The converted data includes the X, Y, Z coordinates, angle, and movement speed of each joint.
[0067] Step 4:
[0068] Terminal: Adds identification information such as a timestamp and user ID to the formatted data.
[0069] Step 5:
[0070] Terminal: Sends the converted data to the server in real time or in batch format.
[0071] Step 6:
[0072] Server: Receives data sent from the terminals. The received data includes each user's action data and its time information.
[0073] Step 7:
[0074] Server: Stores the received data in a database, for example, using MongoDB or an SQL database to store structured data.
[0075] Step 8:
[0076] Server: Analyzes the saved data. Specifically, it converts the data into a data frame using Python libraries (Pandas and NumPy) and compares the user's behavioral characteristics and in-game actions with the score.
[0077] Step 9:
[0078] Server: Generates a machine learning model based on the analysis results. For example, a deep learning model is built using TensorFlow or PyTorch to learn the user's behavioral patterns.
[0079] Step 10:
[0080] Server: Incorporates the generated machine learning model into the control program of the humanoid robot.
[0081] Step 11:
[0082] Humanoid robots (user-based): Based on new machine learning models, the robots will be updated to perform human-like actions, such as replicating the actions of a user in a shooting game.
[0083] Example 1
[0084] Next, a description will be given of Example 1. In the following description, the data processing device 12 will be referred to as a "server" and the smart device 14 will be referred to as a "terminal."
[0085] With conventional systems, it has been difficult to efficiently record the movements of a user in a virtual reality environment and use this information to train a humanoid robot. Furthermore, analysis of the movement data and model generation are often done manually, leaving issues in terms of real-time performance and accuracy.
[0086] Furthermore, there was a lack of optimization of the learning model that took into account the correlation between the specific actions performed by the user in the game and the success rate.
[0087] The specific processing by the specific processing unit 290 of the data processing device 12 in the first embodiment is realized by the following means.
[0088] In this invention, the server includes a means for storing and organizing the recorded motion data in a database, a means for programmatically generating and analyzing data frames, and a means for generating a deep learning model based on the analysis results and integrating the model into the control system of the humanoid robot. This makes it possible to efficiently record a user's motions in real time, generate a machine learning model based on the data, and apply it to the humanoid robot. Furthermore, the model can be optimized based on the success rate of the user's in-game actions, thereby improving the robot's motion accuracy.
[0089] A "user" is an entity that plays a game in a virtual reality environment and provides motion data.
[0090] A "virtual reality environment" is a computer-generated, three-dimensional virtual space that a user experiences using special equipment.
[0091] "Motion data" is information that records the position, movement, speed, etc. of each part of the user's body as numerical values.
[0092] A "predetermined format" is a specific data structure or file format (e.g., CSV or JSON format) defined for storing operational data.
[0093] "Conversion" refers to the process of changing the motion data into a predetermined format, and includes adding a timestamp and user identification information.
[0094] "Time information" is information relating to time, such as the date and time when the motion data was acquired.
[0095] "User identification information" is information for identifying the user who provided the motion data.
[0096] "Real time" means that operational data is processed immediately with little delay.
[0097] The "batch format" refers to a format in which operational data is collected at regular intervals or in regular amounts and processed or transmitted all at once.
[0098] A "server" is a remote computer system that stores and analyzes received data and performs model generation.
[0099] A "database" is a digital repository for efficiently storing, searching, and managing a wide range of data.
[0100] A "data frame" is a data structure that holds data retrieved from a database in tabular form for analysis.
[0101] "Analysis" refers to the processing and examination of operational data to extract specific information.
[0102] A "deep learning model" is an algorithm that uses a multi-layer neural network to learn and reproduce user behavior.
[0103] A "humanoid robot" is a machine that is designed to mimic the appearance and behavior of a human being and operates based on a program.
[0104] A "control system" is a set of software and hardware that directs and manages the movements of a humanoid robot.
[0105] "Success rate" is the percentage of times a user successfully performs a game action within a virtual reality environment.
[0106] The present invention is a system that collects the movements of users playing games in a virtual reality (VR) environment and uses that data to train a humanoid robot. Hereinafter, an embodiment of the invention will be described in detail.
[0107] Overall system configuration
[0108] The system consists of the following components:
[0109] 1. VR gaming and motion capture equipment
[0110] 2. Devices that collect and transmit operational data
[0111] 3. Server that receives and analyzes data and generates machine learning models
[0112] User Preparation
[0113] The user first puts on a VR headset and a motion capture suit, then launches a dedicated application and begins the game. The motion capture suit is equipped with sensors that collect the position of each joint and the speed of movement in real time.
[0114] Recording gameplay
[0115] The device receives real-time motion data from the motion capture suit while the user is playing a game. For example, when a user plays a shooting game, data such as arm movements, X, Y, Z coordinates of hand positions, and movement speed are collected.
[0116] The terminal converts the received data into a predetermined format (for example, CSV or JSON format) and adds a timestamp and user ID. This converted data is sent to the server in real time or in batch format.
[0117] Data transmission and storage
[0118] The device sends the formatted data to the server, either in real time or in batches.
[0119] The server receives the data sent from the device and stores it in a database, for example, using MongoDB or an SQL database to organize and store the data.
[0120] Data analysis and generation of learning models
[0121] The server uses Python libraries (Pandas, NumPy, etc.) to analyze the stored data, generating data frames to calculate user behavior characteristics, game success rates, etc.
[0122] The server uses machine learning libraries such as TensorFlow and PyTorch to generate a deep learning model based on the analysis results, which is designed to replicate, for example, the user's arm movements when shooting an enemy.
[0123] Application to humanoid robots
[0124] The server integrates the generated deep learning model into the control program of the humanoid robot.
[0125] Based on the new model, the humanoid robot can perform human-like actions, for example, faithfully reproducing the shooting movements of a user in a game.
[0126] Examples of specific examples and prompts
[0127] As a concrete example, consider the movements of a user playing a shooting game. In this case, a motion capture suit records the user's movements as they move their arms to aim and fire. This data is then transmitted in real time to the device and then to a server, where it is analyzed and correlated with the user's success rate in the game. This information is used to generate a deep learning model that allows a humanoid robot to replicate similar shooting movements.
[0128] Example prompt sentence:
[0129] We would like to generate a machine learning model for a humanoid robot to reproduce the same movements based on the movement data of a user in a VR shooting game. We will provide the following data.
[0130] 1. X, Y, Z coordinates of each joint
[0131] 2. Operating speed
[0132] 3. Timestamp
[0133] 4. User ID
[0134] 5. In-game success rate
[0135] Using this data, a deep learning model is generated that reproduces the arm movements of the user when shooting an enemy.
[0136] The flow of the identification process in the first embodiment will be described with reference to FIG.
[0137] Step 1: Prepare your users
[0138] The user puts on a VR headset and motion capture suit, launches a dedicated application, and begins the game.
[0139] Input: VR headset, motion capture suit, dedicated application
[0140] Specific operation: The user puts on a VR headset and a motion capture suit, then launches a dedicated application on the device and selects a shooting game.
[0141] Output: Initialized operating environment
[0142] Step 2: Record your gameplay
[0143] The device receives the user's movement data in real time from the motion capture suit.
[0144] Input: Motion capture suit sensor data
[0145] Specific actions: When a user raises their arm to aim and fire a bullet at an enemy in the game, this action is recorded in real time by sensors in the motion capture suit.
[0146] Output: Received raw data (position of each joint, speed of movement, etc.)
[0147] Step 3: Transform the data
[0148] The terminal converts the received data into a specified format (CSV or JSON format) and adds a timestamp and user ID.
[0149] Input: Raw data received
[0150] Specific operation: The received data such as the X, Y, Z coordinates of the arm movement and hand position, and movement speed is converted into a format, and time information (timestamp) and a user-specific ID are added to it.
[0151] Output: Transformed data (CSV or JSON format with timestamp)
[0152] Step 4: Sending data
[0153] The terminal transmits the formatted data to the server, either in real time or in batches.
[0154] Input: Transformed data
[0155] Specific operation: The device sends the converted data to the server via Wi-Fi or a wired connection.
[0156] Output: Formatted data sent to the server
[0157] Step 5: Save your data
[0158] The server stores the received data in a database.
[0159] Input: Formatted data sent to the server
[0160] Specific operation: The server stores the received data in a database system such as MongoDB or SQL, and classifies and organizes it by user or game.
[0161] Output: Data stored in the database
[0162] Step 6: Data analysis
[0163] The server analyzes the stored data and generates a data frame using Python tools such as Pandas and NumPy.
[0164] Input: Data stored in a database
[0165] Specific operation: The server uses Pandas to load data from the database into a data frame, and then uses NumPy to calculate various statistical information, such as the average user movement speed and success rate.
[0166] Output: Parsed data frame and statistics
[0167] Step 7: Generate a training model
[0168] The server generates a deep learning model based on the analysis results, using libraries such as TensorFlow and PyTorch.
[0169] Input: Parsed data frame and statistics
[0170] Specific Actions: The server trains a deep learning model with TensorFlow to replicate the user's arm movements when shooting an enemy. This model is trained to predict appropriate actions based on the input motion data.
[0171] Output: A trained deep learning model
[0172] Step 8: Application to humanoid robots
[0173] The server incorporates the generated deep learning model into the control program of the humanoid robot.
[0174] Input: A trained deep learning model
[0175] Specific operation: The server sends the trained model to the robot's control system and applies it to the robot's operation program.
[0176] Output: The deep learning model applied to the robot
[0177] Based on the new model, the humanoid robot performs human-like movements.
[0178] Input: Deep learning model applied to the robot
[0179] Specific Actions: Based on the embedded model, the humanoid robot raises its arm, aims, and shoots, faithfully replicating the shooting actions performed by the user in virtual reality.
[0180] Output: Reproduced human movements
[0181] (Application example 1)
[0182] Next, a description will be given of Application Example 1. In the following description, the data processing device 12 will be referred to as a "server" and the smart device 14 will be referred to as a "terminal."
[0183] In today's industrial environment, training factory robots is expensive and requires advanced expertise. Furthermore, for robots to accurately replicate human movements, a large amount of realistic motion data is required, but collecting such data is not easy. Therefore, there is a need for an efficient, low-cost method for collecting high-quality motion data and training factory robots.
[0184] The specific processing by the specific processing unit 290 of the data processing device 12 in the application example 1 is realized by the following means.
[0185] In this invention, the server includes means for recording a user's actions in a virtual reality environment in real time, means for converting the recorded action data into a predetermined format, means for transmitting the converted data to the server, means for the server to analyze the received data and store the data by correlating the user's actions with actions in the virtual environment, means for generating a machine learning model based on the stored data and using it for motion learning of an operating device, and means for applying the converted action data to an industrial environment. This makes it possible to efficiently collect actions performed by a user in a virtual reality environment and for a factory robot to learn its actions with high accuracy based on the data.
[0186] A "virtual reality environment" is an environment that uses computer technology to allow users to visually and aurally experience a virtual world that is different from the real world.
[0187] "Motion data" refers to information such as the position, angle, and movement speed of each joint that is collected when a user performs a specific movement.
[0188] A "predetermined format" is a predetermined data format used to store and transmit data, such as CSV or JSON.
[0189] A "server" is a part of a computer system that receives, stores, analyzes, and transmits data over a network.
[0190] A "machine learning model" is a collection of algorithms that learn from accumulated data and analyze and predict its patterns.
[0191] "Motion device" refers to a mechanical device that can reproduce human movements, such as a factory robot.
[0192] "Industrial environment" means the physical space where manufacturing and production activities take place, such as a factory or manufacturing floor.
[0193] The present invention is a system that allows a user to perform actions in a virtual reality (VR) environment and works in conjunction with a server to apply the action data to an industrial environment. Specific embodiments of the system will be described below.
[0194] Overall system overview
[0195] Hardware configuration
[0196] VR headset: A device that allows a user to immerse themselves in a virtual reality environment.
[0197] Motion capture suit: A suit with built-in sensors to measure the position of each joint and the speed of movement (e.g., Xsens, Vicon).
[0198] Servers and Network: Computing resources for data storage, analysis, and machine learning model generation (e.g., AWS, Google Cloud).
[0199] Software configuration
[0200] Data collection applications: Applications that collect data from VR headsets and motion capture suits in real time.
[0201] Data transmission module: A module that converts collected data into a specified format and sends it to the server (e.g., Python's requests library).
[0202] Data analysis and machine learning model generation module: A module that analyzes the received data on the server and generates a machine learning model (e.g., TensorFlow, PyTorch).
[0203] Behavior reproduction module: A program for controlling the behavior of factory robots based on the generated machine learning model.
[0204] Operation procedures and examples
[0205] 1. User Preparation
[0206] The user puts on a VR headset and a motion capture suit and launches a dedicated data collection application. As the user performs factory tasks (e.g., assembling parts) in the virtual reality environment, their motion data is collected by the motion capture suit.
[0207] 2. Data collection and transmission
[0208] The data collection application receives the collected operational data in real time and converts it into a specific format (e.g., CSV or JSON format). The converted data is given a timestamp and a user ID. The data transmission module then sends the formatted data to the server.
[0209] 3. Data storage and analysis
[0210] The server stores the received data in a database (e.g., MongoDB or SQL database). It then analyzes the received data using a data analysis module. For analysis, it uses Python libraries (e.g., Pandas and NumPy) to calculate user behavior characteristics, task success rates, and other information.
[0211] 4. Generating a Machine Learning Model
[0212] Based on the results of the data analysis, the server generates a machine learning model using machine learning libraries such as TensorFlow and PyTorch. For example, a deep learning model can be built to learn the actions of a user assembling parts.
[0213] 5. Reproducing the behavior
[0214] The generated machine learning model is transferred from the server to the factory robot and incorporated into the robot's behavior reproduction program, allowing the factory robot to accurately mimic the user's actions and execute factory tasks.
[0215] Examples of specific examples and prompts
[0216] For example, to simulate a part assembly task in a factory in a VR environment, a user puts on a VR headset and a motion capture suit and performs a series of actions, picking up parts and installing them in designated positions. These action data are collected in real time, converted into a specified format, and sent to a server. After analyzing the data on the server, a factory robot can learn similar actions and reproduce the actual factory task.
[0217] Prompt Sentence Examples
[0218] "We are designing a system that simulates factory tasks in a VR environment, collects user motion data, and uses it to train robots. For example, when a user performs an assembly task, the motion capture suit collects that data and sends it to a server in real time. This data can then be analyzed to enable a factory robot to replicate the same assembly tasks."
[0219] The flow of the specific processing in the application example 1 will be described with reference to FIG.
[0220] Program processing flow
[0221] Step 1:
[0222] Input: A user wears a VR headset and motion capture suit and performs movements in a virtual reality environment.
[0223] Action: The user performs a factory task (e.g., assembling parts, moving pallets) within the VR environment.
[0224] Output: Motion data such as the position, angle, and movement speed of each joint is collected from the motion capture suit.
[0225] Step 2:
[0226] Input: Collected behavioral data.
[0227] Operation: The device converts the collected operation data into a specified format (CSV format, JSON format). The data collection application is used for the conversion, and the data is given a timestamp and user ID.
[0228] Output: Formatted behavior data (CSV format, JSON format).
[0229] Step 3:
[0230] Input: Formatted motion data.
[0231] Operation: The terminal transmits the formatted operation data to the server in real time. The data is transmitted via the Internet using the data transmission module.
[0232] Output: The server receives the operation data.
[0233] Step 4:
[0234] Input: The operational data received by the server.
[0235] Operation: The server stores the received data in a database (e.g., MongoDB or SQL database), then analyzes the data using Python libraries (Pandas, NumPy) to calculate user behavior characteristics, task success rates, etc.
[0236] Output: Parsed data and user behavior characteristics.
[0237] Step 5:
[0238] Input: Parsed data and user behavior characteristics.
[0239] How it works: The server generates a machine learning model based on the analyzed data. It trains the model using machine learning libraries such as TensorFlow and PyTorch, and builds a deep learning model to reproduce the user's behavior.
[0240] Output: The generated machine learning model.
[0241] Step 6:
[0242] Input: The generated machine learning model.
[0243] Operation: The server incorporates the generated machine learning model into the factory robot's behavior reproduction program and transfers the model to the factory robot. Data transmission technology is used over the network for the transfer.
[0244] Output: Machine learning models embedded in factory robots.
[0245] Step 7:
[0246] Input: Machine learning models embedded in factory robots.
[0247] Operation: The factory robot replicates the user's actions based on machine learning models and performs real factory tasks, such as assembling parts and moving pallets with high precision.
[0248] Output: Actual factory task completion.
[0249] This series of processes enables the efficient collection of user actions in a virtual reality environment, enabling factory robots to learn and reproduce these actions with high accuracy based on that data.
[0250] Furthermore, an emotion engine that estimates the user's emotion may be combined. That is, the identification processing unit 290 may estimate the user's emotion using the emotion identification model 59 and perform identification processing using the user's emotion.
[0251] The present invention is a system that collects motion and emotion data when a user plays a game in a virtual reality (VR) environment, and uses this data to train a humanoid robot. Specific embodiments are described below.
[0252] Overall system overview
[0253] The system consists of the following main components:
[0254] 1. VR games played by users and motion capture suits
[0255] 2. Device that collects and transmits motion data and emotional data
[0256] 3. Server that receives and analyzes data and generates machine learning models
[0257] 4. Emotion engine that recognizes user emotions
[0258] User Preparation
[0259] User: First, the user puts on the VR headset, motion capture suit, and emotion engine. Then, they launch the dedicated application and start the game. The motion capture suit and emotion engine collect the user's movement data and emotion data, respectively.
[0260] Recording behavioral and emotional data during gameplay
[0261] Terminal: When a user plays a game, the terminal receives the user's movement data in real time from the motion capture suit and acquires emotion data from the emotion engine. For example, when a user plays a shooting game, the terminal collects arm movements, joint position data, and emotional states (joy, anger, sadness, and happiness) in real time.
[0262] Terminal: Converts the received motion data and emotion data into a specified format (CSV, JSON, etc.) and adds a timestamp and user ID to each data.
[0263] Data transmission and storage
[0264] Terminal: Sends formatted motion data and emotion data to the server in real time or in batch format.
[0265] Server: Receives data sent from the device and stores each user's behavioral and emotional data in a database. For example, structured data is stored using MongoDB or an SQL database.
[0266] Data analysis and generation of learning models
[0267] Server: Analyzes the saved motion and emotion data. Specifically, it converts the data into a data frame using Python libraries (such as Pandas and NumPy) and compares the user's motion characteristics, emotional state, and in-game actions with the score.
[0268] Server: Generates a machine learning model based on the analysis results. For example, a deep learning model is built using TensorFlow or PyTorch to learn the user's behavior and emotional patterns.
[0269] Application in humanoid robots
[0270] Server: Incorporates the generated machine learning model into the control program of the humanoid robot.
[0271] Humanoid robot (user-like): Based on a new machine learning model, it performs human-like movements and emotional responses, for example, reproducing the movements and emotional states of a user in a shooting game.
[0272] As a concrete example, consider the movements and emotions of a user playing a shooting game. In this case, the user's arm movements to aim and fire are recorded. At the same time, an emotion engine records the user's emotional state (e.g., excitement or tension). This data is transmitted in real time to the device and then to a server. The server analyzes the data and associates the user's movements and emotional state with the in-game score. This information is used to generate a machine learning model that enables a humanoid robot to reproduce similar shooting movements and emotions.
[0273] In this way, the present invention provides a system that can efficiently collect human motion data and emotional states and use them for motion learning and emotional expression of humanoid robots.
[0274] The processing flow will be explained below.
[0275] Step 1:
[0276] User: Puts on the VR headset, motion capture suit, and emotion engine, launches the VR game application, and begins the game.
[0277] Step 2:
[0278] Terminal: Receives user movement data (e.g., joint positions, arm movements) in real time from the motion capture suit. Obtains user emotion data (e.g., joy, anger, sadness, and happiness) in real time from the emotion engine.
[0279] Step 3:
[0280] Terminal: Converts the received movement and emotion data into a specified format (CSV, JSON, etc.). The converted data includes the X, Y, and Z coordinates of each joint, angle, movement speed, and emotional state.
[0281] Step 4:
[0282] Terminal: Adds identification information such as a timestamp and user ID to the formatted data.
[0283] Step 5:
[0284] Terminal: Transmits the converted motion data and emotion data to the server in real time or in batch format.
[0285] Step 6:
[0286] Server: Receives data sent from the devices. The received data includes each user's motion data, emotion data, and time information.
[0287] Step 7:
[0288] Server: Stores the received data in a database, for example, using MongoDB or an SQL database to store structured data.
[0289] Step 8:
[0290] Server: Analyzes the saved behavioral and emotional data. Specifically, it converts the data into a data frame using Python libraries (Pandas and NumPy) and compares the user's behavioral characteristics, emotional state, in-game actions, and score.
[0291] Step 9:
[0292] Server: Generates a machine learning model based on the analysis results. For example, a deep learning model is built using TensorFlow or PyTorch to learn the user's behavior and emotional patterns.
[0293] Step 10:
[0294] Server: Incorporates the generated machine learning model into the control program of the humanoid robot.
[0295] Step 11:
[0296] Humanoid robot (user-like): Based on a new machine learning model, the robot can perform human-like actions and emotional responses, such as reproducing the actions and emotional states (e.g., excitement and tension) of a user in a shooting game.
[0297] Example 2
[0298] Next, a description will be given of Example 2. In the following description, the data processing device 12 will be referred to as a "server" and the smart device 14 will be referred to as a "terminal."
[0299] While there are systems that record only the user's motion data during gameplay in conventional virtual reality environments, there are no systems that simultaneously record the user's emotional data in real time and generate highly accurate machine learning models based on that data. As a result, it is difficult for humanoid robots to reproduce human-like motions and emotional expressions.
[0300] The identification process by the identification processing unit 290 of the data processing device 12 in the second embodiment is realized by the following means. In this invention, the server includes means for recording movement and emotion data in real time, means for converting the recorded movement and emotion data into a predetermined format, and means for adding a timestamp and user ID to the converted data and transmitting it to the server. This makes it possible to collect and analyze user movement and emotion data with high accuracy.
[0301] "User" refers to a person who plays a game in a virtual reality environment.
[0302] A "virtual reality environment" refers to an environment in which you can experience a virtual space that is different from reality using devices such as a VR headset.
[0303] "Motion data" refers to data relating to the physical movements of a user when playing a game in a virtual reality environment.
[0304] "Emotion data" refers to data relating to the user's emotional state (for example, joy, anger, sadness, happiness, etc.).
[0305] "Real-time" refers to the near-simultaneous collection, processing, and transmission of data.
[0306] "Specified format" refers to converting data into a prescribed format, such as CSV or JSON.
[0307] "Time stamp" refers to the time information of data collection.
[0308] "User ID" refers to an identifier that uniquely identifies a user.
[0309] "Server" refers to the computer system that stores and analyzes the received data and generates the machine learning model.
[0310] A "machine learning model" is a model that learns patterns based on collected data and makes predictions and classifications for unknown data.
[0311] A "humanoid robot" refers to a robot that has the ability to imitate human movements and emotional expressions.
[0312] MODE FOR CARRYING OUT THE INVENTION
[0313] The present invention is a system for collecting motion and emotion data when a user plays a game in a virtual reality environment, and for using the collected data to train a humanoid robot. Specific embodiments of the system are described below.
[0314] User Preparation
[0315] User: First, the user puts on the VR headset, motion capture suit, and emotion engine. Next, they launch a dedicated application and start the game. This prepares the motion capture suit and emotion engine to collect the user's movement and emotion data.
[0316] Real-time recording and conversion of motion and emotion data
[0317] Device: When a user plays a game, the device receives the user's movement data in real time from the motion capture suit and obtains emotion data from the emotion engine. For example, when playing a shooting game, the device collects the user's arm movements, joint position data, and emotional state (joy, anger, sadness, and happiness) in real time. This data is converted into CSV or JSON format, and a timestamp and user ID are attached to each data point.
[0318] Data transmission and storage
[0319] Terminal: Sends formatted motion data and emotion data to the server in real time or in batch format.
[0320] Server: Receives data sent from the devices and stores each user's motion data and emotion data in a database. MongoDB or an SQL database is used for the database. For example, using MongoDB makes it possible to efficiently store and manage motion data and emotion data.
[0321] Data analysis and generation of learning models
[0322] Server: Analyzes the saved movement and emotion data. Specifically, it uses Python libraries (such as Pandas and NumPy) to create data frames and compares the user's movement characteristics, emotional state, and in-game movements with their score. For example, by combining and analyzing the user's joint position data and emotion data, it is possible to gain a detailed understanding of the user's movement and emotion patterns.
[0323] Server: Generates a machine learning model based on the analysis results. A deep learning model is built using TensorFlow or PyTorch to learn the user's behavior and emotional patterns. For example, TensorFlow can be used to build a multi-layer perceptron or recurrent neural network (RNN) to achieve a highly accurate prediction model.
[0324] Application in humanoid robots
[0325] Server: The generated machine learning model is incorporated into the control program of the humanoid robot, enabling it to reproduce human-like behavior and emotions based on the behavior and emotion patterns that the robot has learned.
[0326] Humanoid robots: Based on new machine learning models, these robots can perform human-like actions and emotional responses. For example, they can replicate the actions and emotional states of a user in a shooting game. Specifically, the robot can aim at a target, pull the trigger, and simultaneously express emotional states such as excitement or tension.
[0327] Specific examples
[0328] Consider the motion and emotional data of a user playing a shooting game. The user's arm movements to aim and fire are recorded. At the same time, an emotion engine records the user's emotional state (e.g., excitement or tension). This data is transmitted in real time to the device and then to a server. The server analyzes the data and associates the user's motion and emotional state with the in-game score. A machine learning model is generated based on this information, enabling a humanoid robot to reproduce similar shooting motions and emotions.
[0329] Prompt Sentence Examples
[0330] "Describe a system that collects motion and emotion data from a user playing a shooting game and generates a learning model that allows a humanoid robot to reproduce those motions and emotions."
[0331] The flow of the identification process in the second embodiment will be described with reference to FIG.
[0332] Step 1: Prepare your users
[0333] Input: The user provides a VR headset, a motion capture suit, and an emotion engine.
[0334] Specific operation: The user puts on a VR headset and a motion capture suit. They then wear or install an emotion engine (heart rate sensor and facial recognition camera). They then launch the dedicated application on their device, select a game, and start playing.
[0335] Output: The user's behavioral and emotional data is ready to be collected.
[0336] Step 2: Recording behavioral and emotional data during gameplay
[0337] Input: The user begins playing the game.
[0338] Specific movements: When the user moves in the game, sensors in the motion capture suit capture the position and movement of the joints, and the emotion engine monitors facial expressions and heart rate in real time.
[0339] Output: Movement data from the motion capture suit and emotion data from the emotion engine are sent to the terminal in real time.
[0340] Step 3: Transforming behavior and emotion data
[0341] Input: Real-time received behavioral and emotional data.
[0342] Specific operation: The terminal converts the data received in real time into CSV or JSON format and adds a timestamp and user ID to each data.
[0343] Output: Data formatted in a given format.
[0344] Step 4: Sending data
[0345] Input: Formatted motion and emotion data.
[0346] Specific operation: The terminal sends the formatted data packets to the server in real time or in batch format.
[0347] Output: Data packets sent to the server.
[0348] Step 5: Save your data
[0349] Input: The data packet sent to the server.
[0350] Specific operation: The server unpacks the received data packets and stores the data in a MongoDB or SQL database.
[0351] Output: Behavioral and emotional data stored in a database.
[0352] Step 6: Analyze the data
[0353] Input: Stored behavioral and emotional data.
[0354] Specific operation: The server uses Python libraries (such as Pandas and NumPy) to convert the data into a data frame, process missing values, and analyze the user's behavioral characteristics and emotional state.
[0355] Output: Analysis results (data about the user's behavioral characteristics, emotional state, and in-game score).
[0356] Step 7: Generate a training model
[0357] Input: Analysis results.
[0358] Specific operation: Based on the analysis results, the server uses TensorFlow or PyTorch to build a deep learning model to learn the user's behavior and emotional patterns.
[0359] Output: The generated machine learning model.
[0360] Step 8: Application in humanoid robots
[0361] Input: The generated machine learning model.
[0362] Specific behavior: The server deploys the trained model to the control system of the humanoid robot, and the robot reproduces human-like behavior and emotions based on the learned behavior and emotion patterns.
[0363] Output: The humanoid robot performs actions that reflect the user's actions and emotions.
[0364] (Application example 2)
[0365] Next, a description will be given of Application Example 2. In the following description, the data processing device 12 will be referred to as a "server" and the smart device 14 will be referred to as a "terminal."
[0366] The present invention relates to a system for providing effective customer service in brick-and-mortar stores based on motion and emotion data acquired when users play games in a virtual reality environment. Conventional technologies have made it difficult for customer service staff to grasp customers' emotions and motions in real time and provide appropriate service. This has resulted in reduced customer satisfaction and lost sales opportunities.
[0367] The specific processing by the specific processing unit 290 of the data processing device 12 in Application Example 2 is realized by the following means. In this invention, the server includes means for recording in real time the actions of a user playing a game in a virtual reality environment, means for converting the recorded action data into a predetermined format, means for transmitting the converted data to the server, means for analyzing the data received by the server and storing the data in association with the user's actions and in-game actions, means for generating a machine learning model based on the stored data and using it for motion training of a humanoid robot, means for capturing the customer's facial expressions and body movements in real time, collecting and analyzing emotional data and action data, and means for providing customer service advice to the customer based on the analysis results. This enables the customer service staff to grasp the customer's emotional state in real time and provide appropriate service.
[0368] A "virtual reality environment" is a technology that allows users to experience virtual spaces and situations that feel real using a computer or special equipment.
[0369] "Motion data" is data that records motion information such as the user's body movements and gestures as numerical values.
[0370] "Emotion data" is data collected to express the user's emotional state, and is inferred from facial expressions, voice, behavior, etc.
[0371] A "motion capture suit" is a device that records a user's body movements with high precision.
[0372] The "emotion engine" is software that analyzes collected data to estimate and recognize the user's emotions.
[0373] A "server" is a computer system that receives, stores, and analyzes data over a network.
[0374] A "machine learning model" is an algorithm that is trained to perform a specific task based on collected and analyzed data.
[0375] A "humanoid robot" is a robot designed to resemble a human and capable of performing human-like movements.
[0376] "Customer service staff" are employees whose role is to provide products and services to customers in physical stores.
[0377] "Response advice" refers to instructions for action or suggestions for improvement that the system provides to customer service staff based on the analysis results.
[0378] A "brick and mortar store" is a physical store where customers can visit in person to purchase goods or services.
[0379] Overall system overview
[0380] The system of the present invention analyzes the facial expressions and movements of customers in real time when the user is serving customers in a physical store, and provides corresponding advice to the staff. The system includes the following main components and software:
[0381] 1. Smart glasses (equipped with a camera for capturing facial expressions and movements)
[0382] 2. Device that collects behavioral and emotional data
[0383] 3. Server that receives and analyzes data and generates machine learning models
[0384] 4. Emotion engine that recognizes user emotions
[0385] Preparing the smart glasses and devices
[0386] Customer service staff: Customer service staff wear smart glasses and launch a dedicated application. The smart glasses are equipped with a camera that captures facial expressions and movements. The images captured by the camera are sent to a terminal in real time.
[0387] Recording facial and movement data
[0388] Terminal: When a wait staff member meets with a customer, the smart glasses camera receives real-time facial and behavior data, including face detection and image processing using OpenCV and emotion prediction using TensorFlow.
[0389] Data transmission and storage
[0390] Terminal: The collected facial and movement data is converted into JSON format and sent to the server in real time or in batch format. The data is accompanied by a timestamp and the staff member ID.
[0391] Server: Receives data sent from the devices and stores the behavioral and emotional data of each customer service staff member in a database, typically using MongoDB or an SQL database.
[0392] Data analysis and generation of learning models
[0393] Server: Analyzes the stored behavioral and emotional data. Python libraries such as Pandas and NumPy are used to create data frames, and the behavioral characteristics, emotional states, and reactions of the staff to customer interactions are compared.
[0394] Server: Generates a machine learning model based on the analysis results. A deep learning model is built using TensorFlow and PyTorch, and trained to learn the behavior and emotional patterns of the customer service staff.
[0395] Providing customer service assistants
[0396] Server: Based on the generated machine learning model, the server provides real-time advice to the customer service staff based on the analysis results. For example, if a customer looks dissatisfied, the server displays the advice, "The customer seems dissatisfied. Please suggest additional services."
[0397] Specific examples
[0398] Consider a scenario where a customer enters a store and looks at products. At this time, smart glasses capture the customer's facial expressions and identify emotions such as "excitement" or "tension." The analysis results are sent to the device in real time and then stored and analyzed on a server. Based on this information, the customer service staff can receive advice such as, "The customer is showing interest. Please explain the product to them."
[0399] Prompt Sentence Examples
[0400] "Enter customer facial expression data and predict their emotion. Generate text that suggests appropriate customer interactions based on the predicted emotion."
[0401] The flow of the specific processing in the application example 2 will be described with reference to FIG.
[0402] Program processing steps and detailed explanations
[0403] Step 1:
[0404] Input: Smart glasses camera image
[0405] How it works: Waiting staff wear smart glasses and capture camera footage in real time as they interact with customers.
[0406] Output: Real-time video data
[0407] explanation:
[0408] The user (a customer service staff member) puts on the smart glasses and begins to interact with the customer. The camera in the smart glasses captures the customer's facial expressions and movements in real time and sends the video data to the terminal.
[0409] Step 2:
[0410] Input: Real-time video data
[0411] How it works: The device processes video data using OpenCV, performs face detection and image processing, and then performs facial recognition using TensorFlow.
[0412] Output: facial expression data and movement data
[0413] explanation:
[0414] The device uses OpenCV to analyze real-time video data received from the smart glasses. First, it performs face detection, then extracts the facial region, converts it to grayscale, and inputs it into a facial expression recognition model (TensorFlow). This generates facial expression data (emotional state) and movement data of the customer.
[0415] Step 3:
[0416] Input: facial expression data and movement data
[0417] How it works: The device converts the data into JSON format and adds a timestamp and the customer service staff ID.
[0418] Output: Formatted data (JSON format)
[0419] explanation:
[0420] The device converts the acquired facial expression and movement data into a specified format (JSON) and adds a timestamp and the customer service staff ID. This data is used for subsequent analysis.
[0421] Step 4:
[0422] Input: Formatted data (JSON format)
[0423] How it works: The device sends formatted data to the server in real time or in batches.
[0424] Output: Transmitted data
[0425] explanation:
[0426] The terminal sends formatted JSON data to the server in real time or in batch format, and communication protocols such as HTTP and WebSocket can be used.
[0427] Step 5:
[0428] Input: Send data
[0429] Operation: The server stores the received data in a database.
[0430] Output: Stored data
[0431] explanation:
[0432] The server receives the JSON data sent from the device and stores it in a MongoDB or SQL database, making it easy to store and analyze the data later.
[0433] Step 6:
[0434] Input: Stored data
[0435] How it works: The server uses Pandas or NumPy to create a data frame and perform analysis.
[0436] Output: Analysis results
[0437] explanation:
[0438] The server then converts the stored data into a data frame using Pandas or NumPy to analyze the customer's behavioral characteristics and emotional state, making it easier to analyze the data and extract specific patterns.
[0439] Step 7:
[0440] Input: Analysis results
[0441] How it works: The server generates a machine learning model using TensorFlow or PyTorch.
[0442] Output: Machine learning model
[0443] explanation:
[0444] The server uses TensorFlow and PyTorch to generate machine learning models based on the analysis results, which improves the accuracy of predictions of customer behavior and emotions.
[0445] Step 8:
[0446] Input: Machine learning models and real-time sentiment data
[0447] Operation: The server generates response advice based on the analysis results and sends it to the device.
[0448] Output: Response advice
[0449] explanation:
[0450] The server generates analysis results based on the generated machine learning model and real-time emotion data, and provides real-time advice to the customer service staff, which is displayed to the customer service staff via their terminal.
[0451] Examples of specific examples and prompts
[0452] If the customer looks unhappy, the advice displayed will be "The customer looks unhappy. Please suggest additional services."
[0453] Prompt Sentence Examples
[0454] "Enter customer facial expression data and predict their emotion. Generate text that suggests appropriate customer interactions based on the predicted emotion."
[0455] The specific processing unit 290 transmits the result of the specific processing to the smart device 14. In the smart device 14, the control unit 46A causes the output device 40 to output the result of the specific processing. The microphone 38B acquires audio indicating a user input regarding the result of the specific processing. The control unit 46A transmits audio data indicating the user input acquired by the microphone 38B to the data processing device 12. In the data processing device 12, the specific processing unit 290 acquires the audio data.
[0456] The data generation model 58 is a so-called generative AI (Artificial Intelligence). An example of the data generation model 58 is ChatGPT (Internet Search<URL: https: / / openai.com / blog / chatgpt> ), Gemini (Internet search <url: https: gemini.google.com ?hl="ja">) and other generation AIs. The data generation model 58 is obtained by performing deep learning on a neural network. A prompt including an instruction is input to the data generation model 58, and inference data such as voice data indicating voice, text data indicating text, and image data indicating an image is also input. The data generation model 58 performs inference on the input inference data in accordance with the instruction indicated by the prompt, and outputs the inference result in a data format such as voice data and text data. Here, inference refers to, for example, analysis, classification, prediction, and / or summarization.
[0457] In the above embodiment, an example in which the specific process is performed by the data processing device 12 has been given, but the technology of the present disclosure is not limited to this, and the specific process may be performed by the smart device 14.
[0458] [Second embodiment]
[0459] FIG. 3 shows an example of the configuration of a data processing system 210 according to the second embodiment.
[0460] 3, the data processing system 210 includes the data processing device 12 and smart glasses 214. An example of the data processing device 12 is a server.
[0461] The data processing device 12 includes a computer 22, a database 24, and a communication I / F 26. The computer 22 is an example of a "computer" according to the technology of the present disclosure. The computer 22 includes a processor 28, a RAM 30, and a storage 32. The processor 28, the RAM 30, and the storage 32 are connected to a bus 34. The database 24 and the communication I / F 26 are also connected to the bus 34. The communication I / F 26 is connected to a network 54. Examples of the network 54 include a WAN (Wide Area Network) and / or a LAN (Local Area Network).
[0462] The smart glasses 214 include a computer 36, a microphone 238, a speaker 240, a camera 42, and a communication I / F 44. The computer 36 includes a processor 46, a RAM 48, and a storage 50. The processor 46, the RAM 48, and the storage 50 are connected to a bus 52. The microphone 238, the speaker 240, and the camera 42 are also connected to the bus 52.
[0463] The microphone 238 receives instructions and the like from the user 20 by receiving voice uttered by the user 20. The microphone 238 captures the voice uttered by the user 20, converts the captured voice into audio data, and outputs it to the processor 46. The speaker 240 outputs audio in accordance with instructions from the processor 46.
[0464] Camera 42 is a small digital camera equipped with an optical system including a lens, aperture, and shutter, and an imaging element such as a CMOS (Complementary Metal-Oxide-Semiconductor) image sensor or a CCD (Charge Coupled Device) image sensor, and captures images of the surroundings of user 20 (for example, an imaging range defined by an angle of view equivalent to the field of vision of a typical healthy person).
[0465] The communication I / F 44 is connected to a network 54. The communication I / Fs 44 and 26 are responsible for the exchange of various information between the processor 46 and the processor 28 via the network 54. The exchange of various information between the processor 46 and the processor 28 using the communication I / Fs 44 and 26 is carried out in a secure state.
[0466] Fig. 4 shows an example of the main functions of the data processing device 12 and the smart glasses 214. As shown in Fig. 4, in the data processing device 12, a specific process is performed by the processor 28. A specific process program 56 is stored in the storage 32.
[0467] The specific processing program 56 is an example of a "program" according to the technology of the present disclosure. The processor 28 reads the specific processing program 56 from the storage 32 and executes the read specific processing program 56 on the RAM 30. The specific processing is realized by the processor 28 operating as a specific processing unit 290 in accordance with the specific processing program 56 executed on the RAM 30.
[0468] The storage 32 stores a data generation model 58 and an emotion identification model 59. The data generation model 58 and the emotion identification model 59 are used by the identification processing unit 290.
[0469] In the smart glasses 214, the reception output process is performed by the processor 46. A reception output program 60 is stored in the storage 50. The processor 46 reads the reception output program 60 from the storage 50 and executes the read reception output program 60 on the RAM 48. The reception output process is realized by the processor 46 operating as the control unit 46A in accordance with the reception output program 60 executed on the RAM 48.
[0470] Next, a description will be given of the identification process performed by the identification processing unit 290 of the data processing device 12. In the following description, the data processing device 12 will be referred to as the "server" and the smart glasses 214 will be referred to as the "terminal."
[0471] The present invention is a system that collects user actions while playing a game in a virtual reality (VR) environment and uses the collected actions for motion learning of a humanoid robot. Specific embodiments will be described below.
[0472] Overall system overview
[0473] The system mainly consists of the following components:
[0474] 1. VR games played by users and motion capture suits
[0475] 2. Devices that collect and transmit operational data
[0476] 3. Server that receives and analyzes data and generates machine learning models
[0477] User Preparation
[0478] User: First, the user puts on a VR headset and a motion capture suit, launches a dedicated application, and begins the game. The motion capture suit has built-in sensors that collect data such as the position of each joint and the speed of movement in real time.
[0479] Recording gameplay
[0480] Device: When a user plays a game, it receives real-time motion data from the motion capture suit. For example, when a user plays a shooting game, arm movement and hand position data are collected.
[0481] Terminal: Converts the received data into a specified format (e.g., CSV or JSON) and adds a timestamp and user ID. This data includes the X, Y, Z coordinates and angles of each joint, as well as movement speed.
[0482] Data transmission and storage
[0483] Terminal: Sends formatted motion data to the server, either in real time or in batches.
[0484] Server: Receives data sent from the device and stores it in a database. For example, it stores organized data using MongoDB or an SQL database.
[0485] Data analysis and generation of learning models
[0486] Server: Analyzes the stored data. For example, creates a data frame using a Python library (Pandas, NumPy, etc.) and calculates user behavior characteristics, game success rate, etc.
[0487] Server: Generates a machine learning model based on the analysis results. Machine learning libraries such as TensorFlow and PyTorch are used to generate the model. For example, a deep learning model is built to reproduce the arm movements of a user when shooting an enemy.
[0488] Application in humanoid robots
[0489] Server: Incorporates the generated machine learning model into the control program of the humanoid robot.
[0490] Humanoid robot (user-like): Based on a new model, the robot performs human-like actions. For example, the robot replicates the shooting actions performed by the user in a game.
[0491] As a concrete example, consider the movements of a user playing a shooting game. In this case, a motion capture suit records the user's movements as they move their arms to aim and fire. This data is then sent in real time to the device and then to a server, where it is analyzed and correlated with the user's in-game score. This information is used to generate a machine learning model that allows a humanoid robot to replicate similar shooting movements.
[0492] In this way, the present invention provides a system that can efficiently collect high-quality human motion data and use it to learn the movements of humanoid robots.
[0493] The processing flow will be explained below.
[0494] Step 1:
[0495] User: Put on the VR headset and motion capture suit, launch the VR game application, and start the game.
[0496] Step 2:
[0497] Terminal: Receives real-time user movement data from the motion capture suit, such as the user's arm movements and joint position data.
[0498] Step 3:
[0499] Terminal: Converts the received movement data into a specified format (CSV, JSON, etc.). The converted data includes the X, Y, Z coordinates, angle, and movement speed of each joint.
[0500] Step 4:
[0501] Terminal: Adds identification information such as a timestamp and user ID to the formatted data.
[0502] Step 5:
[0503] Terminal: Sends the converted data to the server in real time or in batch format.
[0504] Step 6:
[0505] Server: Receives data sent from the terminals. The received data includes each user's action data and its time information.
[0506] Step 7:
[0507] Server: Stores the received data in a database, for example, using MongoDB or an SQL database to store structured data.
[0508] Step 8:
[0509] Server: Analyzes the saved data. Specifically, it converts the data into a data frame using Python libraries (Pandas and NumPy) and compares the user's behavioral characteristics and in-game actions with the score.
[0510] Step 9:
[0511] Server: Generates a machine learning model based on the analysis results. For example, a deep learning model is built using TensorFlow or PyTorch to learn the user's behavioral patterns.
[0512] Step 10:
[0513] Server: Incorporates the generated machine learning model into the control program of the humanoid robot.
[0514] Step 11:
[0515] Humanoid robots (user-based): Based on new machine learning models, the robots will be updated to perform human-like actions, such as replicating the actions of a user in a shooting game.
[0516] Example 1
[0517] Next, a description will be given of Example 1. In the following description, the data processing device 12 will be referred to as a "server" and the smart glasses 214 will be referred to as a "terminal."
[0518] With conventional systems, it has been difficult to efficiently record the movements of a user in a virtual reality environment and use this information to train a humanoid robot. Furthermore, analysis of the movement data and model generation are often done manually, leaving issues in terms of real-time performance and accuracy.
[0519] Furthermore, there was a lack of optimization of the learning model that took into account the correlation between the specific actions performed by the user in the game and the success rate.
[0520] The specific processing by the specific processing unit 290 of the data processing device 12 in the first embodiment is realized by the following means.
[0521] In this invention, the server includes a means for storing and organizing the recorded motion data in a database, a means for programmatically generating and analyzing data frames, and a means for generating a deep learning model based on the analysis results and integrating the model into the control system of the humanoid robot. This makes it possible to efficiently record a user's motions in real time, generate a machine learning model based on the data, and apply it to the humanoid robot. Furthermore, the model can be optimized based on the success rate of the user's in-game actions, thereby improving the robot's motion accuracy.
[0522] A "user" is an entity that plays a game in a virtual reality environment and provides motion data.
[0523] A "virtual reality environment" is a computer-generated, three-dimensional virtual space that a user experiences using special equipment.
[0524] "Motion data" is information that records the position, movement, speed, etc. of each part of the user's body as numerical values.
[0525] A "predetermined format" is a specific data structure or file format (e.g., CSV or JSON format) defined for storing operational data.
[0526] "Conversion" refers to the process of changing the motion data into a predetermined format, and includes adding a timestamp and user identification information.
[0527] "Time information" is information relating to time, such as the date and time when the motion data was acquired.
[0528] "User identification information" is information for identifying the user who provided the motion data.
[0529] "Real time" means that operational data is processed immediately with little delay.
[0530] The "batch format" refers to a format in which operational data is collected at regular intervals or in regular amounts and processed or transmitted all at once.
[0531] A "server" is a remote computer system that stores and analyzes received data and performs model generation.
[0532] A "database" is a digital repository for efficiently storing, searching, and managing a wide range of data.
[0533] A "data frame" is a data structure that holds data retrieved from a database in tabular form for analysis.
[0534] "Analysis" refers to the processing and examination of operational data to extract specific information.
[0535] A "deep learning model" is an algorithm that uses a multi-layer neural network to learn and reproduce user behavior.
[0536] A "humanoid robot" is a machine that is designed to mimic the appearance and behavior of a human being and operates based on a program.
[0537] A "control system" is a set of software and hardware that directs and manages the movements of a humanoid robot.
[0538] "Success rate" is the percentage of times a user successfully performs a game action within a virtual reality environment.
[0539] The present invention is a system that collects the movements of users playing games in a virtual reality (VR) environment and uses that data to train a humanoid robot. Hereinafter, an embodiment of the invention will be described in detail.
[0540] Overall system configuration
[0541] The system consists of the following components:
[0542] 1. VR gaming and motion capture equipment
[0543] 2. Devices that collect and transmit operational data
[0544] 3. Server that receives and analyzes data and generates machine learning models
[0545] User Preparation
[0546] The user first puts on a VR headset and a motion capture suit, then launches a dedicated application and begins the game. The motion capture suit is equipped with sensors that collect the position of each joint and the speed of movement in real time.
[0547] Recording gameplay
[0548] The device receives real-time motion data from the motion capture suit while the user is playing a game. For example, when a user plays a shooting game, data such as arm movements, X, Y, Z coordinates of hand positions, and movement speed are collected.
[0549] The terminal converts the received data into a predetermined format (for example, CSV or JSON format) and adds a timestamp and user ID. This converted data is sent to the server in real time or in batch format.
[0550] Data transmission and storage
[0551] The device sends the formatted data to the server, either in real time or in batches.
[0552] The server receives the data sent from the device and stores it in a database, for example, using MongoDB or an SQL database to organize and store the data.
[0553] Data analysis and generation of learning models
[0554] The server uses Python libraries (Pandas, NumPy, etc.) to analyze the stored data, generating data frames to calculate user behavior characteristics, game success rates, etc.
[0555] The server uses machine learning libraries such as TensorFlow and PyTorch to generate a deep learning model based on the analysis results, which is designed to replicate, for example, the user's arm movements when shooting an enemy.
[0556] Application to humanoid robots
[0557] The server integrates the generated deep learning model into the control program of the humanoid robot.
[0558] Based on the new model, the humanoid robot can perform human-like actions, for example, faithfully reproducing the shooting movements of a user in a game.
[0559] Examples of specific examples and prompts
[0560] As a concrete example, consider the movements of a user playing a shooting game. In this case, a motion capture suit records the user's movements as they move their arms to aim and fire. This data is then transmitted in real time to the device and then to a server, where it is analyzed and correlated with the user's success rate in the game. This information is used to generate a deep learning model that allows a humanoid robot to replicate similar shooting movements.
[0561] Example prompt sentence:
[0562] We would like to generate a machine learning model for a humanoid robot to reproduce the same movements based on the movement data of a user in a VR shooting game. We will provide the following data.
[0563] 1. X, Y, Z coordinates of each joint
[0564] 2. Operating speed
[0565] 3. Timestamp
[0566] 4. User ID
[0567] 5. In-game success rate
[0568] Using this data, a deep learning model is generated that reproduces the arm movements of the user when shooting an enemy.
[0569] The flow of the identification process in the first embodiment will be described with reference to FIG.
[0570] Step 1: Prepare your users
[0571] The user puts on a VR headset and motion capture suit, launches a dedicated application, and begins the game.
[0572] Input: VR headset, motion capture suit, dedicated application
[0573] Specific operation: The user puts on a VR headset and a motion capture suit, then launches a dedicated application on the device and selects a shooting game.
[0574] Output: Initialized operating environment
[0575] Step 2: Record your gameplay
[0576] The device receives the user's movement data in real time from the motion capture suit.
[0577] Input: Motion capture suit sensor data
[0578] Specific actions: When a user raises their arm to aim and fire a bullet at an enemy in the game, this action is recorded in real time by sensors in the motion capture suit.
[0579] Output: Received raw data (position of each joint, speed of movement, etc.)
[0580] Step 3: Transform the data
[0581] The terminal converts the received data into a specified format (CSV or JSON format) and adds a timestamp and user ID.
[0582] Input: Raw data received
[0583] Specific operation: The received data such as the X, Y, Z coordinates of the arm movement and hand position, and movement speed is converted into a format, and time information (timestamp) and a user-specific ID are added to it.
[0584] Output: Transformed data (CSV or JSON format with timestamp)
[0585] Step 4: Sending data
[0586] The terminal transmits the formatted data to the server, either in real time or in batches.
[0587] Input: Transformed data
[0588] Specific operation: The device sends the converted data to the server via Wi-Fi or a wired connection.
[0589] Output: Formatted data sent to the server
[0590] Step 5: Save your data
[0591] The server stores the received data in a database.
[0592] Input: Formatted data sent to the server
[0593] Specific operation: The server stores the received data in a database system such as MongoDB or SQL, and classifies and organizes it by user or game.
[0594] Output: Data stored in the database
[0595] Step 6: Data analysis
[0596] The server analyzes the stored data and generates a data frame using Python tools such as Pandas and NumPy.
[0597] Input: Data stored in a database
[0598] Specific operation: The server uses Pandas to load data from the database into a data frame, and then uses NumPy to calculate various statistical information, such as the average user movement speed and success rate.
[0599] Output: Parsed data frame and statistics
[0600] Step 7: Generate a training model
[0601] The server generates a deep learning model based on the analysis results, using libraries such as TensorFlow and PyTorch.
[0602] Input: Parsed data frame and statistics
[0603] Specific Actions: The server trains a deep learning model with TensorFlow to replicate the user's arm movements when shooting an enemy. This model is trained to predict appropriate actions based on the input motion data.
[0604] Output: A trained deep learning model
[0605] Step 8: Application to humanoid robots
[0606] The server incorporates the generated deep learning model into the control program of the humanoid robot.
[0607] Input: A trained deep learning model
[0608] Specific operation: The server sends the trained model to the robot's control system and applies it to the robot's operation program.
[0609] Output: The deep learning model applied to the robot
[0610] Based on the new model, the humanoid robot performs human-like movements.
[0611] Input: Deep learning model applied to the robot
[0612] Specific Actions: Based on the embedded model, the humanoid robot raises its arm, aims, and shoots, faithfully replicating the shooting actions performed by the user in virtual reality.
[0613] Output: Reproduced human movements
[0614] (Application example 1)
[0615] Next, a description will be given of Application Example 1. In the following description, the data processing device 12 will be referred to as a "server" and the smart glasses 214 will be referred to as a "terminal."
[0616] In today's industrial environment, training factory robots is expensive and requires advanced expertise. Furthermore, for robots to accurately replicate human movements, a large amount of realistic motion data is required, but collecting such data is not easy. Therefore, there is a need for an efficient, low-cost method for collecting high-quality motion data and training factory robots.
[0617] The specific processing by the specific processing unit 290 of the data processing device 12 in the application example 1 is realized by the following means.
[0618] In this invention, the server includes means for recording a user's actions in a virtual reality environment in real time, means for converting the recorded action data into a predetermined format, means for transmitting the converted data to the server, means for the server to analyze the received data and store the data by correlating the user's actions with actions in the virtual environment, means for generating a machine learning model based on the stored data and using it for motion learning of an operating device, and means for applying the converted action data to an industrial environment. This makes it possible to efficiently collect actions performed by a user in a virtual reality environment and for a factory robot to learn its actions with high accuracy based on the data.
[0619] A "virtual reality environment" is an environment that uses computer technology to allow users to visually and aurally experience a virtual world that is different from the real world.
[0620] "Motion data" refers to information such as the position, angle, and movement speed of each joint that is collected when a user performs a specific movement.
[0621] A "predetermined format" is a predetermined data format used to store and transmit data, such as CSV or JSON.
[0622] A "server" is a part of a computer system that receives, stores, analyzes, and transmits data over a network.
[0623] A "machine learning model" is a collection of algorithms that learn from accumulated data and analyze and predict its patterns.
[0624] "Motion device" refers to a mechanical device that can reproduce human movements, such as a factory robot.
[0625] "Industrial environment" means the physical space where manufacturing and production activities take place, such as a factory or manufacturing floor.
[0626] The present invention is a system that allows a user to perform actions in a virtual reality (VR) environment and works in conjunction with a server to apply the action data to an industrial environment. Specific embodiments of the system will be described below.
[0627] Overall system overview
[0628] Hardware configuration
[0629] VR headset: A device that allows a user to immerse themselves in a virtual reality environment.
[0630] Motion capture suit: A suit with built-in sensors to measure the position of each joint and the speed of movement (e.g., Xsens, Vicon).
[0631] Servers and Network: Computing resources for data storage, analysis, and machine learning model generation (e.g., AWS, Google Cloud).
[0632] Software configuration
[0633] Data collection applications: Applications that collect data from VR headsets and motion capture suits in real time.
[0634] Data transmission module: A module that converts collected data into a specified format and sends it to the server (e.g., Python's requests library).
[0635] Data analysis and machine learning model generation module: A module that analyzes the received data on the server and generates a machine learning model (e.g., TensorFlow, PyTorch).
[0636] Behavior reproduction module: A program for controlling the behavior of factory robots based on the generated machine learning model.
[0637] Operation procedures and examples
[0638] 1. User Preparation
[0639] The user puts on a VR headset and a motion capture suit and launches a dedicated data collection application. As the user performs factory tasks (e.g., assembling parts) in the virtual reality environment, their motion data is collected by the motion capture suit.
[0640] 2. Data collection and transmission
[0641] The data collection application receives the collected operational data in real time and converts it into a specific format (e.g., CSV or JSON format). The converted data is given a timestamp and a user ID. The data transmission module then sends the formatted data to the server.
[0642] 3. Data storage and analysis
[0643] The server stores the received data in a database (e.g., MongoDB or SQL database). It then analyzes the received data using a data analysis module. For analysis, it uses Python libraries (e.g., Pandas and NumPy) to calculate user behavior characteristics, task success rates, and other information.
[0644] 4. Generating a Machine Learning Model
[0645] Based on the results of the data analysis, the server generates a machine learning model using machine learning libraries such as TensorFlow and PyTorch. For example, a deep learning model can be built to learn the actions of a user assembling parts.
[0646] 5. Reproducing the behavior
[0647] The generated machine learning model is transferred from the server to the factory robot and incorporated into the robot's behavior reproduction program, allowing the factory robot to accurately mimic the user's actions and execute factory tasks.
[0648] Examples of specific examples and prompts
[0649] For example, to simulate a part assembly task in a factory in a VR environment, a user puts on a VR headset and a motion capture suit and performs a series of actions, picking up parts and installing them in designated positions. These action data are collected in real time, converted into a specified format, and sent to a server. After analyzing the data on the server, a factory robot can learn similar actions and reproduce the actual factory task.
[0650] Prompt Sentence Examples
[0651] "We are designing a system that simulates factory tasks in a VR environment, collects user motion data, and uses it to train robots. For example, when a user performs an assembly task, the motion capture suit collects that data and sends it to a server in real time. This data can then be analyzed to enable a factory robot to replicate the same assembly tasks."
[0652] The flow of the specific processing in the application example 1 will be described with reference to FIG.
[0653] Program processing flow
[0654] Step 1:
[0655] Input: A user wears a VR headset and motion capture suit and performs movements in a virtual reality environment.
[0656] Action: The user performs a factory task (e.g., assembling parts, moving pallets) within the VR environment.
[0657] Output: Motion data such as the position, angle, and movement speed of each joint is collected from the motion capture suit.
[0658] Step 2:
[0659] Input: Collected behavioral data.
[0660] Operation: The device converts the collected operation data into a specified format (CSV format, JSON format). The data collection application is used for the conversion, and the data is given a timestamp and user ID.
[0661] Output: Formatted behavior data (CSV format, JSON format).
[0662] Step 3:
[0663] Input: Formatted motion data.
[0664] Operation: The terminal transmits the formatted operation data to the server in real time. The data is transmitted via the Internet using the data transmission module.
[0665] Output: The server receives the operation data.
[0666] Step 4:
[0667] Input: The operational data received by the server.
[0668] Operation: The server stores the received data in a database (e.g., MongoDB or SQL database), then analyzes the data using Python libraries (Pandas, NumPy) to calculate user behavior characteristics, task success rates, etc.
[0669] Output: Parsed data and user behavior characteristics.
[0670] Step 5:
[0671] Input: Parsed data and user behavior characteristics.
[0672] How it works: The server generates a machine learning model based on the analyzed data. It trains the model using machine learning libraries such as TensorFlow and PyTorch, and builds a deep learning model to reproduce the user's behavior.
[0673] Output: The generated machine learning model.
[0674] Step 6:
[0675] Input: The generated machine learning model.
[0676] Operation: The server incorporates the generated machine learning model into the factory robot's behavior reproduction program and transfers the model to the factory robot. Data transmission technology is used over the network for the transfer.
[0677] Output: Machine learning models embedded in factory robots.
[0678] Step 7:
[0679] Input: Machine learning models embedded in factory robots.
[0680] Operation: The factory robot replicates the user's actions based on machine learning models and performs real factory tasks, such as assembling parts and moving pallets with high precision.
[0681] Output: Actual factory task completion.
[0682] This series of processes enables the efficient collection of user actions in a virtual reality environment, enabling factory robots to learn and reproduce these actions with high accuracy based on that data.
[0683] Furthermore, an emotion engine that estimates the user's emotion may be further combined. That is, the identification processing unit 290 may estimate the user's emotion using the emotion identification model 59, and perform identification processing using the user's emotion.
[0684] The present invention is a system that collects motion and emotion data when a user plays a game in a virtual reality (VR) environment, and uses this data to train a humanoid robot. Specific embodiments are described below.
[0685] Overall system overview
[0686] The system consists of the following main components:
[0687] 1. VR games played by users and motion capture suits
[0688] 2. Device that collects and transmits motion data and emotional data
[0689] 3. Server that receives and analyzes data and generates machine learning models
[0690] 4. Emotion engine that recognizes user emotions
[0691] User Preparation
[0692] User: First, the user puts on the VR headset, motion capture suit, and emotion engine. Then, they launch the dedicated application and start the game. The motion capture suit and emotion engine collect the user's movement data and emotion data, respectively.
[0693] Recording behavioral and emotional data during gameplay
[0694] Terminal: When a user plays a game, the terminal receives the user's movement data in real time from the motion capture suit and acquires emotion data from the emotion engine. For example, when a user plays a shooting game, the terminal collects arm movements, joint position data, and emotional states (joy, anger, sadness, and happiness) in real time.
[0695] Terminal: Converts the received motion data and emotion data into a specified format (CSV, JSON, etc.) and adds a timestamp and user ID to each data.
[0696] Data transmission and storage
[0697] Terminal: Sends formatted motion data and emotion data to the server in real time or in batch format.
[0698] Server: Receives data sent from the device and stores each user's behavioral and emotional data in a database. For example, structured data is stored using MongoDB or an SQL database.
[0699] Data analysis and generation of learning models
[0700] Server: Analyzes the saved motion and emotion data. Specifically, it converts the data into a data frame using Python libraries (such as Pandas and NumPy) and compares the user's motion characteristics, emotional state, and in-game actions with the score.
[0701] Server: Generates a machine learning model based on the analysis results. For example, a deep learning model is built using TensorFlow or PyTorch to learn the user's behavior and emotional patterns.
[0702] Application in humanoid robots
[0703] Server: Incorporates the generated machine learning model into the control program of the humanoid robot.
[0704] Humanoid robot (user-like): Based on a new machine learning model, it performs human-like movements and emotional responses, for example, reproducing the movements and emotional states of a user in a shooting game.
[0705] As a concrete example, consider the movements and emotions of a user playing a shooting game. In this case, the user's arm movements to aim and fire are recorded. At the same time, an emotion engine records the user's emotional state (e.g., excitement or tension). This data is transmitted in real time to the device and then to a server. The server analyzes the data and associates the user's movements and emotional state with the in-game score. This information is used to generate a machine learning model that enables a humanoid robot to reproduce similar shooting movements and emotions.
[0706] In this way, the present invention provides a system that can efficiently collect human motion data and emotional states and use them for motion learning and emotional expression of humanoid robots.
[0707] The processing flow will be explained below.
[0708] Step 1:
[0709] User: Puts on the VR headset, motion capture suit, and emotion engine, launches the VR game application, and begins the game.
[0710] Step 2:
[0711] Terminal: Receives user movement data (e.g., joint positions, arm movements) in real time from the motion capture suit. Obtains user emotion data (e.g., joy, anger, sadness, and happiness) in real time from the emotion engine.
[0712] Step 3:
[0713] Terminal: Converts the received movement and emotion data into a specified format (CSV, JSON, etc.). The converted data includes the X, Y, and Z coordinates of each joint, angle, movement speed, and emotional state.
[0714] Step 4:
[0715] Terminal: Adds identification information such as a timestamp and user ID to the formatted data.
[0716] Step 5:
[0717] Terminal: Transmits the converted motion data and emotion data to the server in real time or in batch format.
[0718] Step 6:
[0719] Server: Receives data sent from the devices. The received data includes each user's motion data, emotion data, and time information.
[0720] Step 7:
[0721] Server: Stores the received data in a database, for example, using MongoDB or an SQL database to store structured data.
[0722] Step 8:
[0723] Server: Analyzes the saved behavioral and emotional data. Specifically, it converts the data into a data frame using Python libraries (Pandas and NumPy) and compares the user's behavioral characteristics, emotional state, in-game actions, and score.
[0724] Step 9:
[0725] Server: Generates a machine learning model based on the analysis results. For example, a deep learning model is built using TensorFlow or PyTorch to learn the user's behavior and emotional patterns.
[0726] Step 10:
[0727] Server: Incorporates the generated machine learning model into the control program of the humanoid robot.
[0728] Step 11:
[0729] Humanoid robot (user-like): Based on a new machine learning model, the robot can perform human-like actions and emotional responses, such as reproducing the actions and emotional states (e.g., excitement and tension) of a user in a shooting game.
[0730] Example 2
[0731] Next, a description will be given of Example 2. In the following description, the data processing device 12 will be referred to as a "server" and the smart glasses 214 will be referred to as a "terminal."
[0732] While there are systems that record only the user's motion data during gameplay in conventional virtual reality environments, there are no systems that simultaneously record the user's emotional data in real time and generate highly accurate machine learning models based on that data. As a result, it is difficult for humanoid robots to reproduce human-like motions and emotional expressions.
[0733] The identification process by the identification processing unit 290 of the data processing device 12 in the second embodiment is realized by the following means. In this invention, the server includes means for recording movement and emotion data in real time, means for converting the recorded movement and emotion data into a predetermined format, and means for adding a timestamp and user ID to the converted data and transmitting it to the server. This makes it possible to collect and analyze user movement and emotion data with high accuracy.
[0734] "User" refers to a person who plays a game in a virtual reality environment.
[0735] A "virtual reality environment" refers to an environment in which you can experience a virtual space that is different from reality using devices such as a VR headset.
[0736] "Motion data" refers to data relating to the physical movements of a user when playing a game in a virtual reality environment.
[0737] "Emotion data" refers to data relating to the user's emotional state (for example, joy, anger, sadness, happiness, etc.).
[0738] "Real-time" refers to the near-simultaneous collection, processing, and transmission of data.
[0739] "Specified format" refers to converting data into a prescribed format, such as CSV or JSON.
[0740] "Time stamp" refers to the time information of data collection.
[0741] "User ID" refers to an identifier that uniquely identifies a user.
[0742] "Server" refers to the computer system that stores and analyzes the received data and generates the machine learning model.
[0743] A "machine learning model" is a model that learns patterns based on collected data and makes predictions and classifications for unknown data.
[0744] A "humanoid robot" refers to a robot that has the ability to imitate human movements and emotional expressions.
[0745] MODE FOR CARRYING OUT THE INVENTION
[0746] The present invention is a system for collecting motion and emotion data when a user plays a game in a virtual reality environment, and for using the collected data to train a humanoid robot. Specific embodiments of the system are described below.
[0747] User Preparation
[0748] User: First, the user puts on the VR headset, motion capture suit, and emotion engine. Next, they launch a dedicated application and start the game. This prepares the motion capture suit and emotion engine to collect the user's movement and emotion data.
[0749] Real-time recording and conversion of motion and emotion data
[0750] Device: When a user plays a game, the device receives the user's movement data in real time from the motion capture suit and obtains emotion data from the emotion engine. For example, when playing a shooting game, the device collects the user's arm movements, joint position data, and emotional state (joy, anger, sadness, and happiness) in real time. This data is converted into CSV or JSON format, and a timestamp and user ID are attached to each data point.
[0751] Data transmission and storage
[0752] Terminal: Sends formatted motion data and emotion data to the server in real time or in batch format.
[0753] Server: Receives data sent from the devices and stores each user's motion data and emotion data in a database. MongoDB or an SQL database is used for the database. For example, using MongoDB makes it possible to efficiently store and manage motion data and emotion data.
[0754] Data analysis and generation of learning models
[0755] Server: Analyzes the saved movement and emotion data. Specifically, it uses Python libraries (such as Pandas and NumPy) to create data frames and compares the user's movement characteristics, emotional state, and in-game movements with their score. For example, by combining and analyzing the user's joint position data and emotion data, it is possible to gain a detailed understanding of the user's movement and emotion patterns.
[0756] Server: Generates a machine learning model based on the analysis results. A deep learning model is built using TensorFlow or PyTorch to learn the user's behavior and emotional patterns. For example, TensorFlow can be used to build a multi-layer perceptron or recurrent neural network (RNN) to achieve a highly accurate prediction model.
[0757] Application in humanoid robots
[0758] Server: The generated machine learning model is incorporated into the control program of the humanoid robot, enabling it to reproduce human-like behavior and emotions based on the behavior and emotion patterns that the robot has learned.
[0759] Humanoid robots: Based on new machine learning models, these robots can perform human-like actions and emotional responses. For example, they can replicate the actions and emotional states of a user in a shooting game. Specifically, the robot can aim at a target, pull the trigger, and simultaneously express emotional states such as excitement or tension.
[0760] Specific examples
[0761] Consider the motion and emotional data of a user playing a shooting game. The user's arm movements to aim and fire are recorded. At the same time, an emotion engine records the user's emotional state (e.g., excitement or tension). This data is transmitted in real time to the device and then to a server. The server analyzes the data and associates the user's motion and emotional state with the in-game score. A machine learning model is generated based on this information, enabling a humanoid robot to reproduce similar shooting motions and emotions.
[0762] Prompt Sentence Examples
[0763] "Describe a system that collects motion and emotion data from a user playing a shooting game and generates a learning model that allows a humanoid robot to reproduce those motions and emotions."
[0764] The flow of the identification process in the second embodiment will be described with reference to FIG.
[0765] Step 1: Prepare your users
[0766] Input: The user provides a VR headset, a motion capture suit, and an emotion engine.
[0767] Specific operation: The user puts on a VR headset and a motion capture suit. They then wear or install an emotion engine (heart rate sensor and facial recognition camera). They then launch the dedicated application on their device, select a game, and start playing.
[0768] Output: The user's behavioral and emotional data is ready to be collected.
[0769] Step 2: Recording behavioral and emotional data during gameplay
[0770] Input: The user begins playing the game.
[0771] Specific movements: When the user moves in the game, sensors in the motion capture suit capture the position and movement of the joints, and the emotion engine monitors facial expressions and heart rate in real time.
[0772] Output: Movement data from the motion capture suit and emotion data from the emotion engine are sent to the terminal in real time.
[0773] Step 3: Transforming behavior and emotion data
[0774] Input: Real-time received behavioral and emotional data.
[0775] Specific operation: The terminal converts the data received in real time into CSV or JSON format and adds a timestamp and user ID to each data.
[0776] Output: Data formatted in a given format.
[0777] Step 4: Sending data
[0778] Input: Formatted motion and emotion data.
[0779] Specific operation: The terminal sends the formatted data packets to the server in real time or in batch format.
[0780] Output: Data packets sent to the server.
[0781] Step 5: Save your data
[0782] Input: The data packet sent to the server.
[0783] Specific operation: The server unpacks the received data packets and stores the data in a MongoDB or SQL database.
[0784] Output: Behavioral and emotional data stored in a database.
[0785] Step 6: Analyze the data
[0786] Input: Stored behavioral and emotional data.
[0787] Specific operation: The server uses Python libraries (such as Pandas and NumPy) to convert the data into a data frame, process missing values, and analyze the user's behavioral characteristics and emotional state.
[0788] Output: Analysis results (data about the user's behavioral characteristics, emotional state, and in-game score).
[0789] Step 7: Generate a training model
[0790] Input: Analysis results.
[0791] Specific operation: Based on the analysis results, the server uses TensorFlow or PyTorch to build a deep learning model to learn the user's behavior and emotional patterns.
[0792] Output: The generated machine learning model.
[0793] Step 8: Application in humanoid robots
[0794] Input: The generated machine learning model.
[0795] Specific behavior: The server deploys the trained model to the control system of the humanoid robot, and the robot reproduces human-like behavior and emotions based on the learned behavior and emotion patterns.
[0796] Output: The humanoid robot performs actions that reflect the user's actions and emotions.
[0797] (Application example 2)
[0798] Next, a description will be given of Application Example 2. In the following description, the data processing device 12 will be referred to as a "server" and the smart glasses 214 will be referred to as a "terminal."
[0799] The present invention relates to a system for providing effective customer service in brick-and-mortar stores based on motion and emotion data acquired when users play games in a virtual reality environment. Conventional technologies have made it difficult for customer service staff to grasp customers' emotions and motions in real time and provide appropriate service. This has resulted in reduced customer satisfaction and lost sales opportunities.
[0800] The specific processing by the specific processing unit 290 of the data processing device 12 in Application Example 2 is realized by the following means. In this invention, the server includes means for recording in real time the actions of a user playing a game in a virtual reality environment, means for converting the recorded action data into a predetermined format, means for transmitting the converted data to the server, means for analyzing the data received by the server and storing the data in association with the user's actions and in-game actions, means for generating a machine learning model based on the stored data and using it for motion training of a humanoid robot, means for capturing the customer's facial expressions and body movements in real time, collecting and analyzing emotional data and action data, and means for providing customer service advice to the customer based on the analysis results. This enables the customer service staff to grasp the customer's emotional state in real time and provide appropriate service.
[0801] A "virtual reality environment" is a technology that allows users to experience virtual spaces and situations that feel real using a computer or special equipment.
[0802] "Motion data" is data that records motion information such as the user's body movements and gestures as numerical values.
[0803] "Emotion data" is data collected to express the user's emotional state, and is inferred from facial expressions, voice, behavior, etc.
[0804] A "motion capture suit" is a device that records a user's body movements with high precision.
[0805] The "emotion engine" is software that analyzes collected data to estimate and recognize the user's emotions.
[0806] A "server" is a computer system that receives, stores, and analyzes data over a network.
[0807] A "machine learning model" is an algorithm that is trained to perform a specific task based on collected and analyzed data.
[0808] A "humanoid robot" is a robot designed to resemble a human and capable of performing human-like movements.
[0809] "Customer service staff" are employees whose role is to provide products and services to customers in physical stores.
[0810] "Response advice" refers to instructions for action or suggestions for improvement that the system provides to customer service staff based on the analysis results.
[0811] A "brick and mortar store" is a physical store where customers can visit in person to purchase goods or services.
[0812] Overall system overview
[0813] The system of the present invention analyzes the facial expressions and movements of customers in real time when the user is serving customers in a physical store, and provides corresponding advice to the staff. The system includes the following main components and software:
[0814] 1. Smart glasses (equipped with a camera for capturing facial expressions and movements)
[0815] 2. Device that collects behavioral and emotional data
[0816] 3. Server that receives and analyzes data and generates machine learning models
[0817] 4. Emotion engine that recognizes user emotions
[0818] Preparing the smart glasses and devices
[0819] Customer service staff: Customer service staff wear smart glasses and launch a dedicated application. The smart glasses are equipped with a camera that captures facial expressions and movements. The images captured by the camera are sent to a terminal in real time.
[0820] Recording facial and movement data
[0821] Terminal: When a wait staff member meets with a customer, the smart glasses camera receives real-time facial and behavior data, including face detection and image processing using OpenCV and emotion prediction using TensorFlow.
[0822] Data transmission and storage
[0823] Terminal: The collected facial and movement data is converted into JSON format and sent to the server in real time or in batch format. The data is accompanied by a timestamp and the staff member ID.
[0824] Server: Receives data sent from the devices and stores the behavioral and emotional data of each customer service staff member in a database, typically using MongoDB or an SQL database.
[0825] Data analysis and generation of learning models
[0826] Server: Analyzes the stored behavioral and emotional data. Python libraries such as Pandas and NumPy are used to create data frames, and the behavioral characteristics, emotional states, and reactions of the staff to customer interactions are compared.
[0827] Server: Generates a machine learning model based on the analysis results. A deep learning model is built using TensorFlow and PyTorch, and trained to learn the behavior and emotional patterns of the customer service staff.
[0828] Providing customer service assistants
[0829] Server: Based on the generated machine learning model, the server provides real-time advice to the customer service staff based on the analysis results. For example, if a customer looks dissatisfied, the server displays the advice, "The customer seems dissatisfied. Please suggest additional services."
[0830] Specific examples
[0831] Consider a scenario where a customer enters a store and looks at products. At this time, smart glasses capture the customer's facial expressions and identify emotions such as "excitement" or "tension." The analysis results are sent to the device in real time and then stored and analyzed on a server. Based on this information, the customer service staff can receive advice such as, "The customer is showing interest. Please explain the product to them."
[0832] Prompt Sentence Examples
[0833] "Enter customer facial expression data and predict their emotion. Generate text that suggests appropriate customer interactions based on the predicted emotion."
[0834] The flow of the specific processing in the application example 2 will be described with reference to FIG.
[0835] Program processing steps and detailed explanations
[0836] Step 1:
[0837] Input: Smart glasses camera image
[0838] How it works: Waiting staff wear smart glasses and capture camera footage in real time as they interact with customers.
[0839] Output: Real-time video data
[0840] explanation:
[0841] The user (a customer service staff member) puts on the smart glasses and begins to interact with the customer. The camera in the smart glasses captures the customer's facial expressions and movements in real time and sends the video data to the terminal.
[0842] Step 2:
[0843] Input: Real-time video data
[0844] How it works: The device processes video data using OpenCV, performs face detection and image processing, and then performs facial recognition using TensorFlow.
[0845] Output: facial expression data and movement data
[0846] explanation:
[0847] The device uses OpenCV to analyze real-time video data received from the smart glasses. First, it performs face detection, then extracts the facial region, converts it to grayscale, and inputs it into a facial expression recognition model (TensorFlow). This generates facial expression data (emotional state) and movement data of the customer.
[0848] Step 3:
[0849] Input: facial expression data and movement data
[0850] How it works: The device converts the data into JSON format and adds a timestamp and the customer service staff ID.
[0851] Output: Formatted data (JSON format)
[0852] explanation:
[0853] The device converts the acquired facial expression and movement data into a specified format (JSON) and adds a timestamp and the customer service staff ID. This data is used for subsequent analysis.
[0854] Step 4:
[0855] Input: Formatted data (JSON format)
[0856] How it works: The device sends formatted data to the server in real time or in batches.
[0857] Output: Transmitted data
[0858] explanation:
[0859] The terminal sends formatted JSON data to the server in real time or in batch format, and communication protocols such as HTTP and WebSocket can be used.
[0860] Step 5:
[0861] Input: Send data
[0862] Operation: The server stores the received data in a database.
[0863] Output: Stored data
[0864] explanation:
[0865] The server receives the JSON data sent from the device and stores it in a MongoDB or SQL database, making it easy to store and analyze the data later.
[0866] Step 6:
[0867] Input: Stored data
[0868] How it works: The server uses Pandas or NumPy to create a data frame and perform analysis.
[0869] Output: Analysis results
[0870] explanation:
[0871] The server then converts the stored data into a data frame using Pandas or NumPy to analyze the customer's behavioral characteristics and emotional state, making it easier to analyze the data and extract specific patterns.
[0872] Step 7:
[0873] Input: Analysis results
[0874] How it works: The server generates a machine learning model using TensorFlow or PyTorch.
[0875] Output: Machine learning model
[0876] explanation:
[0877] The server uses TensorFlow and PyTorch to generate machine learning models based on the analysis results, which improves the accuracy of predictions of customer behavior and emotions.
[0878] Step 8:
[0879] Input: Machine learning models and real-time sentiment data
[0880] Operation: The server generates response advice based on the analysis results and sends it to the device.
[0881] Output: Response advice
[0882] explanation:
[0883] The server generates analysis results based on the generated machine learning model and real-time emotion data, and provides real-time advice to the customer service staff, which is displayed to the customer service staff via their terminal.
[0884] Examples of specific examples and prompts
[0885] If the customer looks unhappy, the advice displayed will be "The customer looks unhappy. Please suggest additional services."
[0886] Prompt Sentence Examples
[0887] "Enter customer facial expression data and predict their emotion. Generate text that suggests appropriate customer interactions based on the predicted emotion."
[0888] The specific processing unit 290 transmits the result of the specific processing to the smart glasses 214. In the smart glasses 214, the control unit 46A causes the speaker 240 to output the result of the specific processing. The microphone 238 acquires audio indicating a user input regarding the result of the specific processing. The control unit 46A transmits audio data indicating the user input acquired by the microphone 238 to the data processing device 12. In the data processing device 12, the specific processing unit 290 acquires the audio data.
[0889] The data generation model 58 is a so-called generative AI (Artificial Intelligence). An example of the data generation model 58 is ChatGPT (Internet Search<URL: https: / / openai.com / blog / chatgpt> ), Gemini (Internet search <url: https: gemini.google.com ?hl="ja">) and other generation AIs. The data generation model 58 is obtained by performing deep learning on a neural network. A prompt including an instruction is input to the data generation model 58, and inference data such as voice data indicating voice, text data indicating text, and image data indicating an image is also input. The data generation model 58 performs inference on the input inference data in accordance with the instruction indicated by the prompt, and outputs the inference result in a data format such as voice data and text data. Here, inference refers to, for example, analysis, classification, prediction, and / or summarization.
[0890] In the above embodiment, an example in which the specific processing is performed by the data processing device 12 has been given, but the technology of the present disclosure is not limited to this, and the specific processing may be performed by the smart glasses 214.
[0891] [Third embodiment]
[0892] FIG. 5 shows an example of the configuration of a data processing system 310 according to the third embodiment.
[0893] 5, the data processing system 310 includes the data processing device 12 and a headset type terminal 314. An example of the data processing device 12 is a server.
[0894] The data processing device 12 includes a computer 22, a database 24, and a communication I / F 26. The computer 22 is an example of a "computer" according to the technology of the present disclosure. The computer 22 includes a processor 28, a RAM 30, and a storage 32. The processor 28, the RAM 30, and the storage 32 are connected to a bus 34. The database 24 and the communication I / F 26 are also connected to the bus 34. The communication I / F 26 is connected to a network 54. Examples of the network 54 include a WAN (Wide Area Network) and / or a LAN (Local Area Network).
[0895] The headset type terminal 314 includes a computer 36, a microphone 238, a speaker 240, a camera 42, a communication I / F 44, and a display 343. The computer 36 includes a processor 46, a RAM 48, and a storage 50. The processor 46, the RAM 48, and the storage 50 are connected to a bus 52. The microphone 238, the speaker 240, the camera 42, and the display 343 are also connected to the bus 52.
[0896] The microphone 238 receives instructions and the like from the user 20 by receiving voice uttered by the user 20. The microphone 238 captures the voice uttered by the user 20, converts the captured voice into audio data, and outputs it to the processor 46. The speaker 240 outputs audio in accordance with instructions from the processor 46.
[0897] Camera 42 is a small digital camera equipped with an optical system including a lens, aperture, and shutter, and an imaging element such as a CMOS (Complementary Metal-Oxide-Semiconductor) image sensor or a CCD (Charge Coupled Device) image sensor, and captures images of the surroundings of user 20 (for example, an imaging range defined by an angle of view equivalent to the field of vision of a typical healthy person).
[0898] The communication I / F 44 is connected to a network 54. The communication I / Fs 44 and 26 are responsible for the exchange of various information between the processor 46 and the processor 28 via the network 54. The exchange of various information between the processor 46 and the processor 28 using the communication I / Fs 44 and 26 is carried out in a secure state.
[0899] Fig. 6 shows an example of the main functions of the data processing device 12 and the headset type terminal 314. As shown in Fig. 6, in the data processing device 12, a specific process is performed by the processor 28. A specific process program 56 is stored in the storage 32.
[0900] The specific processing program 56 is an example of a "program" according to the technology of the present disclosure. The processor 28 reads the specific processing program 56 from the storage 32 and executes the read specific processing program 56 on the RAM 30. The specific processing is realized by the processor 28 operating as a specific processing unit 290 in accordance with the specific processing program 56 executed on the RAM 30.
[0901] The storage 32 stores a data generation model 58 and an emotion identification model 59. The data generation model 58 and the emotion identification model 59 are used by the identification processing unit 290.
[0902] In the headset type terminal 314, a reception output process is performed by the processor 46. A reception output program 60 is stored in the storage 50. The processor 46 reads the reception output program 60 from the storage 50 and executes the read reception output program 60 on the RAM 48. The reception output process is realized by the processor 46 operating as the control unit 46A in accordance with the reception output program 60 executed on the RAM 48.
[0903] Next, a description will be given of the identification process performed by the identification processing unit 290 of the data processing device 12. In the following description, the data processing device 12 will be referred to as the "server" and the headset type terminal 314 will be referred to as the "terminal."
[0904] The present invention is a system that collects user actions while playing a game in a virtual reality (VR) environment and uses the collected actions for motion learning of a humanoid robot. Specific embodiments will be described below.
[0905] Overall system overview
[0906] The system mainly consists of the following components:
[0907] 1. VR games played by users and motion capture suits
[0908] 2. Devices that collect and transmit operational data
[0909] 3. Server that receives and analyzes data and generates machine learning models
[0910] User Preparation
[0911] User: First, the user puts on a VR headset and a motion capture suit, launches a dedicated application, and begins the game. The motion capture suit has built-in sensors that collect data such as the position of each joint and the speed of movement in real time.
[0912] Recording gameplay
[0913] Device: When a user plays a game, it receives real-time motion data from the motion capture suit. For example, when a user plays a shooting game, arm movement and hand position data are collected.
[0914] Terminal: Converts the received data into a specified format (e.g., CSV or JSON) and adds a timestamp and user ID. This data includes the X, Y, Z coordinates and angles of each joint, as well as movement speed.
[0915] Data transmission and storage
[0916] Terminal: Sends formatted motion data to the server, either in real time or in batches.
[0917] Server: Receives data sent from the device and stores it in a database. For example, it stores organized data using MongoDB or an SQL database.
[0918] Data analysis and generation of learning models
[0919] Server: Analyzes the stored data. For example, creates a data frame using a Python library (Pandas, NumPy, etc.) and calculates user behavior characteristics, game success rate, etc.
[0920] Server: Generates a machine learning model based on the analysis results. Machine learning libraries such as TensorFlow and PyTorch are used to generate the model. For example, a deep learning model is built to reproduce the arm movements of a user when shooting an enemy.
[0921] Application in humanoid robots
[0922] Server: Incorporates the generated machine learning model into the control program of the humanoid robot.
[0923] Humanoid robot (user-like): Based on a new model, the robot performs human-like actions. For example, the robot replicates the shooting actions performed by the user in a game.
[0924] As a concrete example, consider the movements of a user playing a shooting game. In this case, a motion capture suit records the user's movements as they move their arms to aim and fire. This data is then sent in real time to the device and then to a server, where it is analyzed and correlated with the user's in-game score. This information is used to generate a machine learning model that allows a humanoid robot to replicate similar shooting movements.
[0925] In this way, the present invention provides a system that can efficiently collect high-quality human motion data and use it to learn the movements of humanoid robots.
[0926] The processing flow will be explained below.
[0927] Step 1:
[0928] User: Put on the VR headset and motion capture suit, launch the VR game application, and start the game.
[0929] Step 2:
[0930] Terminal: Receives real-time user movement data from the motion capture suit, such as the user's arm movements and joint position data.
[0931] Step 3:
[0932] Terminal: Converts the received movement data into a specified format (CSV, JSON, etc.). The converted data includes the X, Y, Z coordinates, angle, and movement speed of each joint.
[0933] Step 4:
[0934] Terminal: Adds identification information such as a timestamp and user ID to the formatted data.
[0935] Step 5:
[0936] Terminal: Sends the converted data to the server in real time or in batch format.
[0937] Step 6:
[0938] Server: Receives data sent from the terminals. The received data includes each user's action data and its time information.
[0939] Step 7:
[0940] Server: Stores the received data in a database, for example, using MongoDB or an SQL database to store structured data.
[0941] Step 8:
[0942] Server: Analyzes the saved data. Specifically, it converts the data into a data frame using Python libraries (Pandas and NumPy) and compares the user's behavioral characteristics and in-game actions with the score.
[0943] Step 9:
[0944] Server: Generates a machine learning model based on the analysis results. For example, a deep learning model is built using TensorFlow or PyTorch to learn the user's behavioral patterns.
[0945] Step 10:
[0946] Server: Incorporates the generated machine learning model into the control program of the humanoid robot.
[0947] Step 11:
[0948] Humanoid robots (user-based): Based on new machine learning models, the robots will be updated to perform human-like actions, such as replicating the actions of a user in a shooting game.
[0949] Example 1
[0950] Next, a description will be given of Example 1. In the following description, the data processing device 12 will be referred to as a "server" and the headset type terminal 314 will be referred to as a "terminal."
[0951] With conventional systems, it has been difficult to efficiently record the movements of a user in a virtual reality environment and use this information to train a humanoid robot. Furthermore, analysis of the movement data and model generation are often done manually, leaving issues in terms of real-time performance and accuracy.
[0952] Furthermore, there was a lack of optimization of the learning model that took into account the correlation between the specific actions performed by the user in the game and the success rate.
[0953] The specific processing by the specific processing unit 290 of the data processing device 12 in the first embodiment is realized by the following means.
[0954] In this invention, the server includes a means for storing and organizing the recorded motion data in a database, a means for programmatically generating and analyzing data frames, and a means for generating a deep learning model based on the analysis results and integrating the model into the control system of the humanoid robot. This makes it possible to efficiently record a user's motions in real time, generate a machine learning model based on the data, and apply it to the humanoid robot. Furthermore, the model can be optimized based on the success rate of the user's in-game actions, thereby improving the robot's motion accuracy.
[0955] A "user" is an entity that plays a game in a virtual reality environment and provides motion data.
[0956] A "virtual reality environment" is a computer-generated, three-dimensional virtual space that a user experiences using special equipment.
[0957] "Motion data" is information that records the position, movement, speed, etc. of each part of the user's body as numerical values.
[0958] A "predetermined format" is a specific data structure or file format (e.g., CSV or JSON format) defined for storing operational data.
[0959] "Conversion" refers to the process of changing the motion data into a predetermined format, and includes adding a timestamp and user identification information.
[0960] "Time information" is information relating to time, such as the date and time when the motion data was acquired.
[0961] "User identification information" is information for identifying the user who provided the motion data.
[0962] "Real time" means that operational data is processed immediately with little delay.
[0963] The "batch format" refers to a format in which operational data is collected at regular intervals or in regular amounts and processed or transmitted all at once.
[0964] A "server" is a remote computer system that stores and analyzes received data and performs model generation.
[0965] A "database" is a digital repository for efficiently storing, searching, and managing a wide range of data.
[0966] A "data frame" is a data structure that holds data retrieved from a database in tabular form for analysis.
[0967] "Analysis" refers to the processing and examination of operational data to extract specific information.
[0968] A "deep learning model" is an algorithm that uses a multi-layer neural network to learn and reproduce user behavior.
[0969] A "humanoid robot" is a machine that is designed to mimic the appearance and behavior of a human being and operates based on a program.
[0970] A "control system" is a set of software and hardware that directs and manages the movements of a humanoid robot.
[0971] "Success rate" is the percentage of times a user successfully performs a game action within a virtual reality environment.
[0972] The present invention is a system that collects the movements of users playing games in a virtual reality (VR) environment and uses that data to train a humanoid robot. Hereinafter, an embodiment of the invention will be described in detail.
[0973] Overall system configuration
[0974] The system consists of the following components:
[0975] 1. VR gaming and motion capture equipment
[0976] 2. Devices that collect and transmit operational data
[0977] 3. Server that receives and analyzes data and generates machine learning models
[0978] User Preparation
[0979] The user first puts on a VR headset and a motion capture suit, then launches a dedicated application and begins the game. The motion capture suit is equipped with sensors that collect the position of each joint and the speed of movement in real time.
[0980] Recording gameplay
[0981] The device receives real-time motion data from the motion capture suit while the user is playing a game. For example, when a user plays a shooting game, data such as arm movements, X, Y, Z coordinates of hand positions, and movement speed are collected.
[0982] The terminal converts the received data into a predetermined format (for example, CSV or JSON format) and adds a timestamp and user ID. This converted data is sent to the server in real time or in batch format.
[0983] Data transmission and storage
[0984] The device sends the formatted data to the server, either in real time or in batches.
[0985] The server receives the data sent from the device and stores it in a database, for example, using MongoDB or an SQL database to organize and store the data.
[0986] Data analysis and generation of learning models
[0987] The server uses Python libraries (Pandas, NumPy, etc.) to analyze the stored data, generating data frames to calculate user behavior characteristics, game success rates, etc.
[0988] The server uses machine learning libraries such as TensorFlow and PyTorch to generate a deep learning model based on the analysis results, which is designed to replicate, for example, the user's arm movements when shooting an enemy.
[0989] Application to humanoid robots
[0990] The server integrates the generated deep learning model into the control program of the humanoid robot.
[0991] Based on the new model, the humanoid robot can perform human-like actions, for example, faithfully reproducing the shooting movements of a user in a game.
[0992] Examples of specific examples and prompts
[0993] As a concrete example, consider the movements of a user playing a shooting game. In this case, a motion capture suit records the user's movements as they move their arms to aim and fire. This data is then transmitted in real time to the device and then to a server, where it is analyzed and correlated with the user's success rate in the game. This information is used to generate a deep learning model that allows a humanoid robot to replicate similar shooting movements.
[0994] Example prompt sentence:
[0995] We would like to generate a machine learning model for a humanoid robot to reproduce the same movements based on the movement data of a user in a VR shooting game. We will provide the following data.
[0996] 1. X, Y, Z coordinates of each joint
[0997] 2. Operating speed
[0998] 3. Timestamp
[0999] 4. User ID
[1000] 5. In-game success rate
[1001] Using this data, a deep learning model is generated that reproduces the arm movements of the user when shooting an enemy.
[1002] The flow of the identification process in the first embodiment will be described with reference to FIG.
[1003] Step 1: Prepare your users
[1004] The user puts on a VR headset and motion capture suit, launches a dedicated application, and begins the game.
[1005] Input: VR headset, motion capture suit, dedicated application
[1006] Specific operation: The user puts on a VR headset and a motion capture suit, then launches a dedicated application on the device and selects a shooting game.
[1007] Output: Initialized operating environment
[1008] Step 2: Record your gameplay
[1009] The device receives the user's movement data in real time from the motion capture suit.
[1010] Input: Motion capture suit sensor data
[1011] Specific actions: When a user raises their arm to aim and fire a bullet at an enemy in the game, this action is recorded in real time by sensors in the motion capture suit.
[1012] Output: Received raw data (position of each joint, speed of movement, etc.)
[1013] Step 3: Transform the data
[1014] The terminal converts the received data into a specified format (CSV or JSON format) and adds a timestamp and user ID.
[1015] Input: Raw data received
[1016] Specific operation: The received data such as the X, Y, Z coordinates of the arm movement and hand position, and movement speed is converted into a format, and time information (timestamp) and a user-specific ID are added to it.
[1017] Output: Transformed data (CSV or JSON format with timestamp)
[1018] Step 4: Sending data
[1019] The terminal transmits the formatted data to the server, either in real time or in batches.
[1020] Input: Transformed data
[1021] Specific operation: The device sends the converted data to the server via Wi-Fi or a wired connection.
[1022] Output: Formatted data sent to the server
[1023] Step 5: Save your data
[1024] The server stores the received data in a database.
[1025] Input: Formatted data sent to the server
[1026] Specific operation: The server stores the received data in a database system such as MongoDB or SQL, and classifies and organizes it by user or game.
[1027] Output: Data stored in the database
[1028] Step 6: Data analysis
[1029] The server analyzes the stored data and generates a data frame using Python tools such as Pandas and NumPy.
[1030] Input: Data stored in a database
[1031] Specific operation: The server uses Pandas to load data from the database into a data frame, and then uses NumPy to calculate various statistical information, such as the average user movement speed and success rate.
[1032] Output: Parsed data frame and statistics
[1033] Step 7: Generate a training model
[1034] The server generates a deep learning model based on the analysis results, using libraries such as TensorFlow and PyTorch.
[1035] Input: Parsed data frame and statistics
[1036] Specific Actions: The server trains a deep learning model with TensorFlow to replicate the user's arm movements when shooting an enemy. This model is trained to predict appropriate actions based on the input motion data.
[1037] Output: A trained deep learning model
[1038] Step 8: Application to humanoid robots
[1039] The server incorporates the generated deep learning model into the control program of the humanoid robot.
[1040] Input: A trained deep learning model
[1041] Specific operation: The server sends the trained model to the robot's control system and applies it to the robot's operation program.
[1042] Output: The deep learning model applied to the robot
[1043] Based on the new model, the humanoid robot performs human-like movements.
[1044] Input: Deep learning model applied to the robot
[1045] Specific Actions: Based on the embedded model, the humanoid robot raises its arm, aims, and shoots, faithfully replicating the shooting actions performed by the user in virtual reality.
[1046] Output: Reproduced human movements
[1047] (Application example 1)
[1048] Next, a description will be given of Application Example 1. In the following description, the data processing device 12 will be referred to as a "server" and the headset type terminal 314 will be referred to as a "terminal."
[1049] In today's industrial environment, training factory robots is expensive and requires advanced expertise. Furthermore, for robots to accurately replicate human movements, a large amount of realistic motion data is required, but collecting such data is not easy. Therefore, there is a need for an efficient, low-cost method for collecting high-quality motion data and training factory robots.
[1050] The specific processing by the specific processing unit 290 of the data processing device 12 in the application example 1 is realized by the following means.
[1051] In this invention, the server includes means for recording a user's actions in a virtual reality environment in real time, means for converting the recorded action data into a predetermined format, means for transmitting the converted data to the server, means for the server to analyze the received data and store the data by correlating the user's actions with actions in the virtual environment, means for generating a machine learning model based on the stored data and using it for motion learning of an operating device, and means for applying the converted action data to an industrial environment. This makes it possible to efficiently collect actions performed by a user in a virtual reality environment and for a factory robot to learn its actions with high accuracy based on the data.
[1052] A "virtual reality environment" is an environment that uses computer technology to allow users to visually and aurally experience a virtual world that is different from the real world.
[1053] "Motion data" refers to information such as the position, angle, and movement speed of each joint that is collected when a user performs a specific movement.
[1054] A "predetermined format" is a predetermined data format used to store and transmit data, such as CSV or JSON.
[1055] A "server" is a part of a computer system that receives, stores, analyzes, and transmits data over a network.
[1056] A "machine learning model" is a collection of algorithms that learn from accumulated data and analyze and predict its patterns.
[1057] "Motion device" refers to a mechanical device that can reproduce human movements, such as a factory robot.
[1058] "Industrial environment" means the physical space where manufacturing and production activities take place, such as a factory or manufacturing floor.
[1059] The present invention is a system that allows a user to perform actions in a virtual reality (VR) environment and works in conjunction with a server to apply the action data to an industrial environment. Specific embodiments of the system will be described below.
[1060] Overall system overview
[1061] Hardware configuration
[1062] VR headset: A device that allows a user to immerse themselves in a virtual reality environment.
[1063] Motion capture suit: A suit with built-in sensors to measure the position of each joint and the speed of movement (e.g., Xsens, Vicon).
[1064] Servers and Network: Computing resources for data storage, analysis, and machine learning model generation (e.g., AWS, Google Cloud).
[1065] Software configuration
[1066] Data collection applications: Applications that collect data from VR headsets and motion capture suits in real time.
[1067] Data transmission module: A module that converts collected data into a specified format and sends it to the server (e.g., Python's requests library).
[1068] Data analysis and machine learning model generation module: A module that analyzes the received data on the server and generates a machine learning model (e.g., TensorFlow, PyTorch).
[1069] Behavior reproduction module: A program for controlling the behavior of factory robots based on the generated machine learning model.
[1070] Operation procedures and examples
[1071] 1. User Preparation
[1072] The user puts on a VR headset and a motion capture suit and launches a dedicated data collection application. As the user performs factory tasks (e.g., assembling parts) in the virtual reality environment, their motion data is collected by the motion capture suit.
[1073] 2. Data collection and transmission
[1074] The data collection application receives the collected operational data in real time and converts it into a specific format (e.g., CSV or JSON format). The converted data is given a timestamp and a user ID. The data transmission module then sends the formatted data to the server.
[1075] 3. Data storage and analysis
[1076] The server stores the received data in a database (e.g., MongoDB or SQL database). It then analyzes the received data using a data analysis module. For analysis, it uses Python libraries (e.g., Pandas and NumPy) to calculate user behavior characteristics, task success rates, and other information.
[1077] 4. Generating a Machine Learning Model
[1078] Based on the results of the data analysis, the server generates a machine learning model using machine learning libraries such as TensorFlow and PyTorch. For example, a deep learning model can be built to learn the actions of a user assembling parts.
[1079] 5. Reproducing the behavior
[1080] The generated machine learning model is transferred from the server to the factory robot and incorporated into the robot's behavior reproduction program, allowing the factory robot to accurately mimic the user's actions and execute factory tasks.
[1081] Examples of specific examples and prompts
[1082] For example, to simulate a part assembly task in a factory in a VR environment, a user puts on a VR headset and a motion capture suit and performs a series of actions, picking up parts and installing them in designated positions. These action data are collected in real time, converted into a specified format, and sent to a server. After analyzing the data on the server, a factory robot can learn similar actions and reproduce the actual factory task.
[1083] Prompt Sentence Examples
[1084] "We are designing a system that simulates factory tasks in a VR environment, collects user motion data, and uses it to train robots. For example, when a user performs an assembly task, the motion capture suit collects that data and sends it to a server in real time. This data can then be analyzed to enable a factory robot to replicate the same assembly tasks."
[1085] The flow of the specific processing in the application example 1 will be described with reference to FIG.
[1086] Program processing flow
[1087] Step 1:
[1088] Input: A user wears a VR headset and motion capture suit and performs movements in a virtual reality environment.
[1089] Action: The user performs a factory task (e.g., assembling parts, moving pallets) within the VR environment.
[1090] Output: Motion data such as the position, angle, and movement speed of each joint is collected from the motion capture suit.
[1091] Step 2:
[1092] Input: Collected behavioral data.
[1093] Operation: The device converts the collected operation data into a specified format (CSV format, JSON format). The data collection application is used for the conversion, and the data is given a timestamp and user ID.
[1094] Output: Formatted behavior data (CSV format, JSON format).
[1095] Step 3:
[1096] Input: Formatted motion data.
[1097] Operation: The terminal transmits the formatted operation data to the server in real time. The data is transmitted via the Internet using the data transmission module.
[1098] Output: The server receives the operation data.
[1099] Step 4:
[1100] Input: The operational data received by the server.
[1101] Operation: The server stores the received data in a database (e.g., MongoDB or SQL database), then analyzes the data using Python libraries (Pandas, NumPy) to calculate user behavior characteristics, task success rates, etc.
[1102] Output: Parsed data and user behavior characteristics.
[1103] Step 5:
[1104] Input: Parsed data and user behavior characteristics.
[1105] How it works: The server generates a machine learning model based on the analyzed data. It trains the model using machine learning libraries such as TensorFlow and PyTorch, and builds a deep learning model to reproduce the user's behavior.
[1106] Output: The generated machine learning model.
[1107] Step 6:
[1108] Input: The generated machine learning model.
[1109] Operation: The server incorporates the generated machine learning model into the factory robot's behavior reproduction program and transfers the model to the factory robot. Data transmission technology is used over the network for the transfer.
[1110] Output: Machine learning models embedded in factory robots.
[1111] Step 7:
[1112] Input: Machine learning models embedded in factory robots.
[1113] Operation: The factory robot replicates the user's actions based on machine learning models and performs real factory tasks, such as assembling parts and moving pallets with high precision.
[1114] Output: Actual factory task completion.
[1115] This series of processes enables the efficient collection of user actions in a virtual reality environment, enabling factory robots to learn and reproduce these actions with high accuracy based on that data.
[1116] Furthermore, an emotion engine that estimates the user's emotion may be further combined. That is, the identification processing unit 290 may estimate the user's emotion using the emotion identification model 59, and perform identification processing using the user's emotion.
[1117] The present invention is a system that collects motion and emotion data when a user plays a game in a virtual reality (VR) environment, and uses this data to train a humanoid robot. Specific embodiments are described below.
[1118] Overall system overview
[1119] The system consists of the following main components:
[1120] 1. VR games played by users and motion capture suits
[1121] 2. Device that collects and transmits motion data and emotional data
[1122] 3. Server that receives and analyzes data and generates machine learning models
[1123] 4. Emotion engine that recognizes user emotions
[1124] User Preparation
[1125] User: First, the user puts on the VR headset, motion capture suit, and emotion engine. Then, they launch the dedicated application and start the game. The motion capture suit and emotion engine collect the user's movement data and emotion data, respectively.
[1126] Recording behavioral and emotional data during gameplay
[1127] Terminal: When a user plays a game, the terminal receives the user's movement data in real time from the motion capture suit and acquires emotion data from the emotion engine. For example, when a user plays a shooting game, the terminal collects arm movements, joint position data, and emotional states (joy, anger, sadness, and happiness) in real time.
[1128] Terminal: Converts the received motion data and emotion data into a specified format (CSV, JSON, etc.) and adds a timestamp and user ID to each data.
[1129] Data transmission and storage
[1130] Terminal: Sends formatted motion data and emotion data to the server in real time or in batch format.
[1131] Server: Receives data sent from the device and stores each user's behavioral and emotional data in a database. For example, structured data is stored using MongoDB or an SQL database.
[1132] Data analysis and generation of learning models
[1133] Server: Analyzes the saved motion and emotion data. Specifically, it converts the data into a data frame using Python libraries (such as Pandas and NumPy) and compares the user's motion characteristics, emotional state, and in-game actions with the score.
[1134] Server: Generates a machine learning model based on the analysis results. For example, a deep learning model is built using TensorFlow or PyTorch to learn the user's behavior and emotional patterns.
[1135] Application in humanoid robots
[1136] Server: Incorporates the generated machine learning model into the control program of the humanoid robot.
[1137] Humanoid robot (user-like): Based on a new machine learning model, it performs human-like movements and emotional responses, for example, reproducing the movements and emotional states of a user in a shooting game.
[1138] As a concrete example, consider the movements and emotions of a user playing a shooting game. In this case, the user's arm movements to aim and fire are recorded. At the same time, an emotion engine records the user's emotional state (e.g., excitement or tension). This data is transmitted in real time to the device and then to a server. The server analyzes the data and associates the user's movements and emotional state with the in-game score. This information is used to generate a machine learning model that enables a humanoid robot to reproduce similar shooting movements and emotions.
[1139] In this way, the present invention provides a system that can efficiently collect human motion data and emotional states and use them for motion learning and emotional expression of humanoid robots.
[1140] The processing flow will be explained below.
[1141] Step 1:
[1142] User: Puts on the VR headset, motion capture suit, and emotion engine, launches the VR game application, and begins the game.
[1143] Step 2:
[1144] Terminal: Receives user movement data (e.g., joint positions, arm movements) in real time from the motion capture suit. Obtains user emotion data (e.g., joy, anger, sadness, and happiness) in real time from the emotion engine.
[1145] Step 3:
[1146] Terminal: Converts the received movement and emotion data into a specified format (CSV, JSON, etc.). The converted data includes the X, Y, and Z coordinates of each joint, angle, movement speed, and emotional state.
[1147] Step 4:
[1148] Terminal: Adds identification information such as a timestamp and user ID to the formatted data.
[1149] Step 5:
[1150] Terminal: Transmits the converted motion data and emotion data to the server in real time or in batch format.
[1151] Step 6:
[1152] Server: Receives data sent from the devices. The received data includes each user's motion data, emotion data, and time information.
[1153] Step 7:
[1154] Server: Stores the received data in a database, for example, using MongoDB or an SQL database to store structured data.
[1155] Step 8:
[1156] Server: Analyzes the saved behavioral and emotional data. Specifically, it converts the data into a data frame using Python libraries (Pandas and NumPy) and compares the user's behavioral characteristics, emotional state, in-game actions, and score.
[1157] Step 9:
[1158] Server: Generates a machine learning model based on the analysis results. For example, a deep learning model is built using TensorFlow or PyTorch to learn the user's behavior and emotional patterns.
[1159] Step 10:
[1160] Server: Incorporates the generated machine learning model into the control program of the humanoid robot.
[1161] Step 11:
[1162] Humanoid robot (user-like): Based on a new machine learning model, the robot can perform human-like actions and emotional responses, such as reproducing the actions and emotional states (e.g., excitement and tension) of a user in a shooting game.
[1163] Example 2
[1164] Next, a description will be given of Example 2. In the following description, the data processing device 12 will be referred to as a "server" and the headset type terminal 314 will be referred to as a "terminal."
[1165] While there are systems that record only the user's motion data during gameplay in conventional virtual reality environments, there are no systems that simultaneously record the user's emotional data in real time and generate highly accurate machine learning models based on that data. As a result, it is difficult for humanoid robots to reproduce human-like motions and emotional expressions.
[1166] The identification process by the identification processing unit 290 of the data processing device 12 in the second embodiment is realized by the following means. In this invention, the server includes means for recording movement and emotion data in real time, means for converting the recorded movement and emotion data into a predetermined format, and means for adding a timestamp and user ID to the converted data and transmitting it to the server. This makes it possible to collect and analyze user movement and emotion data with high accuracy.
[1167] "User" refers to a person who plays a game in a virtual reality environment.
[1168] A "virtual reality environment" refers to an environment in which you can experience a virtual space that is different from reality using devices such as a VR headset.
[1169] "Motion data" refers to data relating to the physical movements of a user when playing a game in a virtual reality environment.
[1170] "Emotion data" refers to data relating to the user's emotional state (for example, joy, anger, sadness, happiness, etc.).
[1171] "Real-time" refers to the near-simultaneous collection, processing, and transmission of data.
[1172] "Specified format" refers to converting data into a prescribed format, such as CSV or JSON.
[1173] "Time stamp" refers to the time information of data collection.
[1174] "User ID" refers to an identifier that uniquely identifies a user.
[1175] "Server" refers to the computer system that stores and analyzes the received data and generates the machine learning model.
[1176] A "machine learning model" is a model that learns patterns based on collected data and makes predictions and classifications for unknown data.
[1177] A "humanoid robot" refers to a robot that has the ability to imitate human movements and emotional expressions.
[1178] MODE FOR CARRYING OUT THE INVENTION
[1179] The present invention is a system for collecting motion and emotion data when a user plays a game in a virtual reality environment, and for using the collected data to train a humanoid robot. Specific embodiments of the system are described below.
[1180] User Preparation
[1181] User: First, the user puts on the VR headset, motion capture suit, and emotion engine. Next, they launch a dedicated application and start the game. This prepares the motion capture suit and emotion engine to collect the user's movement and emotion data.
[1182] Real-time recording and conversion of motion and emotion data
[1183] Device: When a user plays a game, the device receives the user's movement data in real time from the motion capture suit and obtains emotion data from the emotion engine. For example, when playing a shooting game, the device collects the user's arm movements, joint position data, and emotional state (joy, anger, sadness, and happiness) in real time. This data is converted into CSV or JSON format, and a timestamp and user ID are attached to each data point.
[1184] Data transmission and storage
[1185] Terminal: Sends formatted motion data and emotion data to the server in real time or in batch format.
[1186] Server: Receives data sent from the devices and stores each user's motion data and emotion data in a database. MongoDB or an SQL database is used for the database. For example, using MongoDB makes it possible to efficiently store and manage motion data and emotion data.
[1187] Data analysis and generation of learning models
[1188] Server: Analyzes the saved movement and emotion data. Specifically, it uses Python libraries (such as Pandas and NumPy) to create data frames and compares the user's movement characteristics, emotional state, and in-game movements with their score. For example, by combining and analyzing the user's joint position data and emotion data, it is possible to gain a detailed understanding of the user's movement and emotion patterns.
[1189] Server: Generates a machine learning model based on the analysis results. A deep learning model is built using TensorFlow or PyTorch to learn the user's behavior and emotional patterns. For example, TensorFlow can be used to build a multi-layer perceptron or recurrent neural network (RNN) to achieve a highly accurate prediction model.
[1190] Application in humanoid robots
[1191] Server: The generated machine learning model is incorporated into the control program of the humanoid robot, enabling it to reproduce human-like behavior and emotions based on the behavior and emotion patterns that the robot has learned.
[1192] Humanoid robots: Based on new machine learning models, these robots can perform human-like actions and emotional responses. For example, they can replicate the actions and emotional states of a user in a shooting game. Specifically, the robot can aim at a target, pull the trigger, and simultaneously express emotional states such as excitement or tension.
[1193] Specific examples
[1194] Consider the motion and emotional data of a user playing a shooting game. The user's arm movements to aim and fire are recorded. At the same time, an emotion engine records the user's emotional state (e.g., excitement or tension). This data is transmitted in real time to the device and then to a server. The server analyzes the data and associates the user's motion and emotional state with the in-game score. A machine learning model is generated based on this information, enabling a humanoid robot to reproduce similar shooting motions and emotions.
[1195] Prompt Sentence Examples
[1196] "Describe a system that collects motion and emotion data from a user playing a shooting game and generates a learning model that allows a humanoid robot to reproduce those motions and emotions."
[1197] The flow of the identification process in the second embodiment will be described with reference to FIG.
[1198] Step 1: Prepare your users
[1199] Input: The user provides a VR headset, a motion capture suit, and an emotion engine.
[1200] Specific operation: The user puts on a VR headset and a motion capture suit. They then wear or install an emotion engine (heart rate sensor and facial recognition camera). They then launch the dedicated application on their device, select a game, and start playing.
[1201] Output: The user's behavioral and emotional data is ready to be collected.
[1202] Step 2: Recording behavioral and emotional data during gameplay
[1203] Input: The user begins playing the game.
[1204] Specific movements: When the user moves in the game, sensors in the motion capture suit capture the position and movement of the joints, and the emotion engine monitors facial expressions and heart rate in real time.
[1205] Output: Movement data from the motion capture suit and emotion data from the emotion engine are sent to the terminal in real time.
[1206] Step 3: Transforming behavior and emotion data
[1207] Input: Real-time received behavioral and emotional data.
[1208] Specific operation: The terminal converts the data received in real time into CSV or JSON format and adds a timestamp and user ID to each data.
[1209] Output: Data formatted in a given format.
[1210] Step 4: Sending data
[1211] Input: Formatted motion and emotion data.
[1212] Specific operation: The terminal sends the formatted data packets to the server in real time or in batch format.
[1213] Output: Data packets sent to the server.
[1214] Step 5: Save your data
[1215] Input: The data packet sent to the server.
[1216] Specific operation: The server unpacks the received data packets and stores the data in a MongoDB or SQL database.
[1217] Output: Behavioral and emotional data stored in a database.
[1218] Step 6: Analyze the data
[1219] Input: Stored behavioral and emotional data.
[1220] Specific operation: The server uses Python libraries (such as Pandas and NumPy) to convert the data into a data frame, process missing values, and analyze the user's behavioral characteristics and emotional state.
[1221] Output: Analysis results (data about the user's behavioral characteristics, emotional state, and in-game score).
[1222] Step 7: Generate a training model
[1223] Input: Analysis results.
[1224] Specific operation: Based on the analysis results, the server uses TensorFlow or PyTorch to build a deep learning model to learn the user's behavior and emotional patterns.
[1225] Output: The generated machine learning model.
[1226] Step 8: Application in humanoid robots
[1227] Input: The generated machine learning model.
[1228] Specific behavior: The server deploys the trained model to the control system of the humanoid robot, and the robot reproduces human-like behavior and emotions based on the learned behavior and emotion patterns.
[1229] Output: The humanoid robot performs actions that reflect the user's actions and emotions.
[1230] (Application example 2)
[1231] Next, a description will be given of Application Example 2. In the following description, the data processing device 12 will be referred to as a "server" and the headset type terminal 314 will be referred to as a "terminal."
[1232] The present invention relates to a system for providing effective customer service in brick-and-mortar stores based on motion and emotion data acquired when users play games in a virtual reality environment. Conventional technologies have made it difficult for customer service staff to grasp customers' emotions and motions in real time and provide appropriate service. This has resulted in reduced customer satisfaction and lost sales opportunities.
[1233] The specific processing by the specific processing unit 290 of the data processing device 12 in Application Example 2 is realized by the following means. In this invention, the server includes means for recording in real time the actions of a user playing a game in a virtual reality environment, means for converting the recorded action data into a predetermined format, means for transmitting the converted data to the server, means for analyzing the data received by the server and storing the data in association with the user's actions and in-game actions, means for generating a machine learning model based on the stored data and using it for motion training of a humanoid robot, means for capturing the customer's facial expressions and body movements in real time, collecting and analyzing emotional data and action data, and means for providing customer service advice to the customer based on the analysis results. This enables the customer service staff to grasp the customer's emotional state in real time and provide appropriate service.
[1234] A "virtual reality environment" is a technology that allows users to experience virtual spaces and situations that feel real using a computer or special equipment.
[1235] "Motion data" is data that records motion information such as the user's body movements and gestures as numerical values.
[1236] "Emotion data" is data collected to express the user's emotional state, and is inferred from facial expressions, voice, behavior, etc.
[1237] A "motion capture suit" is a device that records a user's body movements with high precision.
[1238] The "emotion engine" is software that analyzes collected data to estimate and recognize the user's emotions.
[1239] A "server" is a computer system that receives, stores, and analyzes data over a network.
[1240] A "machine learning model" is an algorithm that is trained to perform a specific task based on collected and analyzed data.
[1241] A "humanoid robot" is a robot designed to resemble a human and capable of performing human-like movements.
[1242] "Customer service staff" are employees whose role is to provide products and services to customers in physical stores.
[1243] "Response advice" refers to instructions for action or suggestions for improvement that the system provides to customer service staff based on the analysis results.
[1244] A "brick and mortar store" is a physical store where customers can visit in person to purchase goods or services.
[1245] Overall system overview
[1246] The system of the present invention analyzes the facial expressions and movements of customers in real time when the user is serving customers in a physical store, and provides corresponding advice to the staff. The system includes the following main components and software:
[1247] 1. Smart glasses (equipped with a camera for capturing facial expressions and movements)
[1248] 2. Device that collects behavioral and emotional data
[1249] 3. Server that receives and analyzes data and generates machine learning models
[1250] 4. Emotion engine that recognizes user emotions
[1251] Preparing the smart glasses and devices
[1252] Customer service staff: Customer service staff wear smart glasses and launch a dedicated application. The smart glasses are equipped with a camera that captures facial expressions and movements. The images captured by the camera are sent to a terminal in real time.
[1253] Recording facial and movement data
[1254] Terminal: When a wait staff member meets with a customer, the smart glasses camera receives real-time facial and behavior data, including face detection and image processing using OpenCV and emotion prediction using TensorFlow.
[1255] Data transmission and storage
[1256] Terminal: The collected facial and movement data is converted into JSON format and sent to the server in real time or in batch format. The data is accompanied by a timestamp and the staff member ID.
[1257] Server: Receives data sent from the devices and stores the behavioral and emotional data of each customer service staff member in a database, typically using MongoDB or an SQL database.
[1258] Data analysis and generation of learning models
[1259] Server: Analyzes the stored behavioral and emotional data. Python libraries such as Pandas and NumPy are used to create data frames, and the behavioral characteristics, emotional states, and reactions of the staff to customer interactions are compared.
[1260] Server: Generates a machine learning model based on the analysis results. A deep learning model is built using TensorFlow and PyTorch, and trained to learn the behavior and emotional patterns of the customer service staff.
[1261] Providing customer service assistants
[1262] Server: Based on the generated machine learning model, the server provides real-time advice to the customer service staff based on the analysis results. For example, if a customer looks dissatisfied, the server displays the advice, "The customer seems dissatisfied. Please suggest additional services."
[1263] Specific examples
[1264] Consider a scenario where a customer enters a store and looks at products. At this time, smart glasses capture the customer's facial expressions and identify emotions such as "excitement" or "tension." The analysis results are sent to the device in real time and then stored and analyzed on a server. Based on this information, the customer service staff can receive advice such as, "The customer is showing interest. Please explain the product to them."
[1265] Prompt Sentence Examples
[1266] "Enter customer facial expression data and predict their emotion. Generate text that suggests appropriate customer interactions based on the predicted emotion."
[1267] The flow of the specific processing in the application example 2 will be described with reference to FIG.
[1268] Program processing steps and detailed explanations
[1269] Step 1:
[1270] Input: Smart glasses camera image
[1271] How it works: Waiting staff wear smart glasses and capture camera footage in real time as they interact with customers.
[1272] Output: Real-time video data
[1273] explanation:
[1274] The user (a customer service staff member) puts on the smart glasses and begins to interact with the customer. The camera in the smart glasses captures the customer's facial expressions and movements in real time and sends the video data to the terminal.
[1275] Step 2:
[1276] Input: Real-time video data
[1277] How it works: The device processes video data using OpenCV, performs face detection and image processing, and then performs facial recognition using TensorFlow.
[1278] Output: facial expression data and movement data
[1279] explanation:
[1280] The device uses OpenCV to analyze real-time video data received from the smart glasses. First, it performs face detection, then extracts the facial region, converts it to grayscale, and inputs it into a facial expression recognition model (TensorFlow). This generates facial expression data (emotional state) and movement data of the customer.
[1281] Step 3:
[1282] Input: facial expression data and movement data
[1283] How it works: The device converts the data into JSON format and adds a timestamp and the customer service staff ID.
[1284] Output: Formatted data (JSON format)
[1285] explanation:
[1286] The device converts the acquired facial expression and movement data into a specified format (JSON) and adds a timestamp and the customer service staff ID. This data is used for subsequent analysis.
[1287] Step 4:
[1288] Input: Formatted data (JSON format)
[1289] How it works: The device sends formatted data to the server in real time or in batches.
[1290] Output: Transmitted data
[1291] explanation:
[1292] The terminal sends formatted JSON data to the server in real time or in batch format, and communication protocols such as HTTP and WebSocket can be used.
[1293] Step 5:
[1294] Input: Send data
[1295] Operation: The server stores the received data in a database.
[1296] Output: Stored data
[1297] explanation:
[1298] The server receives the JSON data sent from the device and stores it in a MongoDB or SQL database, making it easy to store and analyze the data later.
[1299] Step 6:
[1300] Input: Stored data
[1301] How it works: The server uses Pandas or NumPy to create a data frame and perform analysis.
[1302] Output: Analysis results
[1303] explanation:
[1304] The server then converts the stored data into a data frame using Pandas or NumPy to analyze the customer's behavioral characteristics and emotional state, making it easier to analyze the data and extract specific patterns.
[1305] Step 7:
[1306] Input: Analysis results
[1307] How it works: The server generates a machine learning model using TensorFlow or PyTorch.
[1308] Output: Machine learning model
[1309] explanation:
[1310] The server uses TensorFlow and PyTorch to generate machine learning models based on the analysis results, which improves the accuracy of predictions of customer behavior and emotions.
[1311] Step 8:
[1312] Input: Machine learning models and real-time sentiment data
[1313] Operation: The server generates response advice based on the analysis results and sends it to the device.
[1314] Output: Response advice
[1315] explanation:
[1316] The server generates analysis results based on the generated machine learning model and real-time emotion data, and provides real-time advice to the customer service staff, which is displayed to the customer service staff via their terminal.
[1317] Examples of specific examples and prompts
[1318] If the customer looks unhappy, the advice displayed will be "The customer looks unhappy. Please suggest additional services."
[1319] Prompt Sentence Examples
[1320] "Enter customer facial expression data and predict their emotion. Generate text that suggests appropriate customer interactions based on the predicted emotion."
[1321] The specific processing unit 290 transmits the result of the specific processing to the headset type terminal 314. In the headset type terminal 314, the control unit 46A causes the speaker 240 and the display 343 to output the result of the specific processing. The microphone 238 acquires audio indicating a user input regarding the result of the specific processing. The control unit 46A transmits audio data indicating the user input acquired by the microphone 238 to the data processing device 12. In the data processing device 12, the specific processing unit 290 acquires the audio data.
[1322] The data generation model 58 is a so-called generative AI (Artificial Intelligence). An example of the data generation model 58 is ChatGPT (Internet Search<URL: https: / / openai.com / blog / chatgpt> ), Gemini (Internet search <url: https: gemini.google.com ?hl="ja">) and other generation AIs. The data generation model 58 is obtained by performing deep learning on a neural network. A prompt including an instruction is input to the data generation model 58, and inference data such as voice data indicating voice, text data indicating text, and image data indicating an image is also input. The data generation model 58 performs inference on the input inference data in accordance with the instruction indicated by the prompt, and outputs the inference result in a data format such as voice data and text data. Here, inference refers to, for example, analysis, classification, prediction, and / or summarization.
[1323] In the above embodiment, an example was given in which the specific processing is performed by the data processing device 12, but the technology of the present disclosure is not limited to this, and the specific processing may be performed by the headset type terminal 314.
[1324] [Fourth embodiment]
[1325] FIG. 7 shows an example of the configuration of a data processing system 410 according to the fourth embodiment.
[1326] 7, a data processing system 410 includes a data processing device 12 and a robot 414. An example of the data processing device 12 is a server.
[1327] The data processing device 12 includes a computer 22, a database 24, and a communication I / F 26. The computer 22 is an example of a "computer" according to the technology of the present disclosure. The computer 22 includes a processor 28, a RAM 30, and a storage 32. The processor 28, the RAM 30, and the storage 32 are connected to a bus 34. The database 24 and the communication I / F 26 are also connected to the bus 34. The communication I / F 26 is connected to a network 54. Examples of the network 54 include a WAN (Wide Area Network) and / or a LAN (Local Area Network).
[1328] The robot 414 includes a computer 36, a microphone 238, a speaker 240, a camera 42, a communication I / F 44, and a control target 443. The computer 36 includes a processor 46, a RAM 48, and a storage 50. The processor 46, the RAM 48, and the storage 50 are connected to a bus 52. The microphone 238, the speaker 240, the camera 42, and the control target 443 are also connected to the bus 52.
[1329] The microphone 238 receives instructions and the like from the user 20 by receiving voice uttered by the user 20. The microphone 238 captures the voice uttered by the user 20, converts the captured voice into audio data, and outputs it to the processor 46. The speaker 240 outputs audio in accordance with instructions from the processor 46.
[1330] Camera 42 is a small digital camera equipped with an optical system including a lens, aperture, and shutter, and an imaging element such as a CMOS (Complementary Metal-Oxide-Semiconductor) image sensor or a CCD (Charge Coupled Device) image sensor, and captures images of the surroundings of user 20 (for example, an imaging range defined by an angle of view equivalent to the field of vision of a typical healthy person).
[1331] The communication I / F 44 is connected to a network 54. The communication I / Fs 44 and 26 are responsible for the exchange of various information between the processor 46 and the processor 28 via the network 54. The exchange of various information between the processor 46 and the processor 28 using the communication I / Fs 44 and 26 is carried out in a secure state.
[1332] The control object 443 includes a display device, LEDs in the eyes, and motors for driving the arms, hands, and feet. The posture and gestures of the robot 414 are controlled by controlling the motors of the arms, hands, and feet. Some of the emotions of the robot 414 can be expressed by controlling these motors. In addition, the facial expressions of the robot 414 can also be expressed by controlling the light emission state of the LEDs in the eyes of the robot 414.
[1333] Fig. 8 shows an example of the main functions of the data processing device 12 and the robot 414. As shown in Fig. 8, in the data processing device 12, a specific process is performed by the processor 28. A specific process program 56 is stored in the storage 32.
[1334] The specific processing program 56 is an example of a "program" according to the technology of the present disclosure. The processor 28 reads the specific processing program 56 from the storage 32 and executes the read specific processing program 56 on the RAM 30. The specific processing is realized by the processor 28 operating as a specific processing unit 290 in accordance with the specific processing program 56 executed on the RAM 30.
[1335] The storage 32 stores a data generation model 58 and an emotion identification model 59. The data generation model 58 and the emotion identification model 59 are used by the identification processing unit 290.
[1336] In the robot 414, the processor 46 performs the reception output process. A reception output program 60 is stored in the storage 50. The processor 46 reads the reception output program 60 from the storage 50 and executes the read reception output program 60 on the RAM 48. The reception output process is realized by the processor 46 operating as the control unit 46A in accordance with the reception output program 60 executed on the RAM 48.
[1337] Next, a description will be given of the specific processing performed by the specific processing unit 290 of the data processing device 12. In the following description, the data processing device 12 will be referred to as a "server" and the robot 414 will be referred to as a "terminal."
[1338] The present invention is a system that collects user actions while playing a game in a virtual reality (VR) environment and uses the collected actions for motion learning of a humanoid robot. Specific embodiments will be described below.
[1339] Overall system overview
[1340] The system mainly consists of the following components:
[1341] 1. VR games played by users and motion capture suits
[1342] 2. Devices that collect and transmit operational data
[1343] 3. Server that receives and analyzes data and generates machine learning models
[1344] User Preparation
[1345] User: First, the user puts on a VR headset and a motion capture suit, launches a dedicated application, and begins the game. The motion capture suit has built-in sensors that collect data such as the position of each joint and the speed of movement in real time.
[1346] Recording gameplay
[1347] Device: When a user plays a game, it receives real-time motion data from the motion capture suit. For example, when a user plays a shooting game, arm movement and hand position data are collected.
[1348] Terminal: Converts the received data into a specified format (e.g., CSV or JSON) and adds a timestamp and user ID. This data includes the X, Y, Z coordinates and angles of each joint, as well as movement speed.
[1349] Data transmission and storage
[1350] Terminal: Sends formatted motion data to the server, either in real time or in batches.
[1351] Server: Receives data sent from the device and stores it in a database. For example, it stores organized data using MongoDB or an SQL database.
[1352] Data analysis and generation of learning models
[1353] Server: Analyzes the stored data. For example, creates a data frame using a Python library (Pandas, NumPy, etc.) and calculates user behavior characteristics, game success rate, etc.
[1354] Server: Generates a machine learning model based on the analysis results. Machine learning libraries such as TensorFlow and PyTorch are used to generate the model. For example, a deep learning model is built to reproduce the arm movements of a user when shooting an enemy.
[1355] Application in humanoid robots
[1356] Server: Incorporates the generated machine learning model into the control program of the humanoid robot.
[1357] Humanoid robot (user-like): Based on a new model, the robot performs human-like actions. For example, the robot replicates the shooting actions performed by the user in a game.
[1358] As a concrete example, consider the movements of a user playing a shooting game. In this case, a motion capture suit records the user's movements as they move their arms to aim and fire. This data is then sent in real time to the device and then to a server, where it is analyzed and correlated with the user's in-game score. This information is used to generate a machine learning model that allows a humanoid robot to replicate similar shooting movements.
[1359] In this way, the present invention provides a system that can efficiently collect high-quality human motion data and use it to learn the movements of humanoid robots.
[1360] The processing flow will be explained below.
[1361] Step 1:
[1362] User: Put on the VR headset and motion capture suit, launch the VR game application, and start the game.
[1363] Step 2:
[1364] Terminal: Receives real-time user movement data from the motion capture suit, such as the user's arm movements and joint position data.
[1365] Step 3:
[1366] Terminal: Converts the received movement data into a specified format (CSV, JSON, etc.). The converted data includes the X, Y, Z coordinates, angle, and movement speed of each joint.
[1367] Step 4:
[1368] Terminal: Adds identification information such as a timestamp and user ID to the formatted data.
[1369] Step 5:
[1370] Terminal: Sends the converted data to the server in real time or in batch format.
[1371] Step 6:
[1372] Server: Receives data sent from the terminals. The received data includes each user's action data and its time information.
[1373] Step 7:
[1374] Server: Stores the received data in a database, for example, using MongoDB or an SQL database to store structured data.
[1375] Step 8:
[1376] Server: Analyzes the saved data. Specifically, it converts the data into a data frame using Python libraries (Pandas and NumPy) and compares the user's behavioral characteristics and in-game actions with the score.
[1377] Step 9:
[1378] Server: Generates a machine learning model based on the analysis results. For example, a deep learning model is built using TensorFlow or PyTorch to learn the user's behavioral patterns.
[1379] Step 10:
[1380] Server: Incorporates the generated machine learning model into the control program of the humanoid robot.
[1381] Step 11:
[1382] Humanoid robots (user-based): Based on new machine learning models, the robots will be updated to perform human-like actions, such as replicating the actions of a user in a shooting game.
[1383] Example 1
[1384] Next, a description will be given of Example 1. In the following description, the data processing device 12 will be referred to as a "server" and the robot 414 will be referred to as a "terminal."
[1385] With conventional systems, it has been difficult to efficiently record the movements of a user in a virtual reality environment and use this information to train a humanoid robot. Furthermore, analysis of the movement data and model generation are often done manually, leaving issues in terms of real-time performance and accuracy.
[1386] Furthermore, there was a lack of optimization of the learning model that took into account the correlation between the specific actions performed by the user in the game and the success rate.
[1387] The specific processing by the specific processing unit 290 of the data processing device 12 in the first embodiment is realized by the following means.
[1388] In this invention, the server includes a means for storing and organizing the recorded motion data in a database, a means for programmatically generating and analyzing data frames, and a means for generating a deep learning model based on the analysis results and integrating the model into the control system of the humanoid robot. This makes it possible to efficiently record a user's motions in real time, generate a machine learning model based on the data, and apply it to the humanoid robot. Furthermore, the model can be optimized based on the success rate of the user's in-game actions, thereby improving the robot's motion accuracy.
[1389] A "user" is an entity that plays a game in a virtual reality environment and provides motion data.
[1390] A "virtual reality environment" is a computer-generated, three-dimensional virtual space that a user experiences using special equipment.
[1391] "Motion data" is information that records the position, movement, speed, etc. of each part of the user's body as numerical values.
[1392] A "predetermined format" is a specific data structure or file format (e.g., CSV or JSON format) defined for storing operational data.
[1393] "Conversion" refers to the process of changing the motion data into a predetermined format, and includes adding a timestamp and user identification information.
[1394] "Time information" is information relating to time, such as the date and time when the motion data was acquired.
[1395] "User identification information" is information for identifying the user who provided the motion data.
[1396] "Real time" means that operational data is processed immediately with little delay.
[1397] The "batch format" refers to a format in which operational data is collected at regular intervals or in regular amounts and processed or transmitted all at once.
[1398] A "server" is a remote computer system that stores and analyzes received data and performs model generation.
[1399] A "database" is a digital repository for efficiently storing, searching, and managing a wide range of data.
[1400] A "data frame" is a data structure that holds data retrieved from a database in tabular form for analysis.
[1401] "Analysis" refers to the processing and examination of operational data to extract specific information.
[1402] A "deep learning model" is an algorithm that uses a multi-layer neural network to learn and reproduce user behavior.
[1403] A "humanoid robot" is a machine that is designed to mimic the appearance and behavior of a human being and operates based on a program.
[1404] A "control system" is a set of software and hardware that directs and manages the movements of a humanoid robot.
[1405] "Success rate" is the percentage of times a user successfully performs a game action within a virtual reality environment.
[1406] The present invention is a system that collects the movements of users playing games in a virtual reality (VR) environment and uses that data to train a humanoid robot. Hereinafter, an embodiment of the invention will be described in detail.
[1407] Overall system configuration
[1408] The system consists of the following components:
[1409] 1. VR gaming and motion capture equipment
[1410] 2. Devices that collect and transmit operational data
[1411] 3. Server that receives and analyzes data and generates machine learning models
[1412] User Preparation
[1413] The user first puts on a VR headset and a motion capture suit, then launches a dedicated application and begins the game. The motion capture suit is equipped with sensors that collect the position of each joint and the speed of movement in real time.
[1414] Recording gameplay
[1415] The device receives real-time motion data from the motion capture suit while the user is playing a game. For example, when a user plays a shooting game, data such as arm movements, X, Y, Z coordinates of hand positions, and movement speed are collected.
[1416] The terminal converts the received data into a predetermined format (for example, CSV or JSON format) and adds a timestamp and user ID. This converted data is sent to the server in real time or in batch format.
[1417] Data transmission and storage
[1418] The device sends the formatted data to the server, either in real time or in batches.
[1419] The server receives the data sent from the device and stores it in a database, for example, using MongoDB or an SQL database to organize and store the data.
[1420] Data analysis and generation of learning models
[1421] The server uses Python libraries (Pandas, NumPy, etc.) to analyze the stored data, generating data frames to calculate user behavior characteristics, game success rates, etc.
[1422] The server uses machine learning libraries such as TensorFlow and PyTorch to generate a deep learning model based on the analysis results, which is designed to replicate, for example, the user's arm movements when shooting an enemy.
[1423] Application to humanoid robots
[1424] The server integrates the generated deep learning model into the control program of the humanoid robot.
[1425] Based on the new model, the humanoid robot can perform human-like actions, for example, faithfully reproducing the shooting movements of a user in a game.
[1426] Examples of specific examples and prompts
[1427] As a concrete example, consider the movements of a user playing a shooting game. In this case, a motion capture suit records the user's movements as they move their arms to aim and fire. This data is then transmitted in real time to the device and then to a server, where it is analyzed and correlated with the user's success rate in the game. This information is used to generate a deep learning model that allows a humanoid robot to replicate similar shooting movements.
[1428] Example prompt sentence:
[1429] We would like to generate a machine learning model for a humanoid robot to reproduce the same movements based on the movement data of a user in a VR shooting game. We will provide the following data.
[1430] 1. X, Y, Z coordinates of each joint
[1431] 2. Operating speed
[1432] 3. Timestamp
[1433] 4. User ID
[1434] 5. In-game success rate
[1435] Using this data, a deep learning model is generated that reproduces the arm movements of the user when shooting an enemy.
[1436] The flow of the identification process in the first embodiment will be described with reference to FIG.
[1437] Step 1: Prepare your users
[1438] The user puts on a VR headset and motion capture suit, launches a dedicated application, and begins the game.
[1439] Input: VR headset, motion capture suit, dedicated application
[1440] Specific operation: The user puts on a VR headset and a motion capture suit, then launches a dedicated application on the device and selects a shooting game.
[1441] Output: Initialized operating environment
[1442] Step 2: Record your gameplay
[1443] The device receives the user's movement data in real time from the motion capture suit.
[1444] Input: Motion capture suit sensor data
[1445] Specific actions: When a user raises their arm to aim and fire a bullet at an enemy in the game, this action is recorded in real time by sensors in the motion capture suit.
[1446] Output: Received raw data (position of each joint, speed of movement, etc.)
[1447] Step 3: Transform the data
[1448] The terminal converts the received data into a specified format (CSV or JSON format) and adds a timestamp and user ID.
[1449] Input: Raw data received
[1450] Specific operation: The received data such as the X, Y, Z coordinates of the arm movement and hand position, and movement speed is converted into a format, and time information (timestamp) and a user-specific ID are added to it.
[1451] Output: Transformed data (CSV or JSON format with timestamp)
[1452] Step 4: Sending data
[1453] The terminal transmits the formatted data to the server, either in real time or in batches.
[1454] Input: Transformed data
[1455] Specific operation: The device sends the converted data to the server via Wi-Fi or a wired connection.
[1456] Output: Formatted data sent to the server
[1457] Step 5: Save your data
[1458] The server stores the received data in a database.
[1459] Input: Formatted data sent to the server
[1460] Specific operation: The server stores the received data in a database system such as MongoDB or SQL, and classifies and organizes it by user or game.
[1461] Output: Data stored in the database
[1462] Step 6: Data analysis
[1463] The server analyzes the stored data and generates a data frame using Python tools such as Pandas and NumPy.
[1464] Input: Data stored in a database
[1465] Specific operation: The server uses Pandas to load data from the database into a data frame, and then uses NumPy to calculate various statistical information, such as the average user movement speed and success rate.
[1466] Output: Parsed data frame and statistics
[1467] Step 7: Generate a training model
[1468] The server generates a deep learning model based on the analysis results, using libraries such as TensorFlow and PyTorch.
[1469] Input: Parsed data frame and statistics
[1470] Specific Actions: The server trains a deep learning model with TensorFlow to replicate the user's arm movements when shooting an enemy. This model is trained to predict appropriate actions based on the input motion data.
[1471] Output: A trained deep learning model
[1472] Step 8: Application to humanoid robots
[1473] The server incorporates the generated deep learning model into the control program of the humanoid robot.
[1474] Input: A trained deep learning model
[1475] Specific operation: The server sends the trained model to the robot's control system and applies it to the robot's operation program.
[1476] Output: The deep learning model applied to the robot
[1477] Based on the new model, the humanoid robot performs human-like movements.
[1478] Input: Deep learning model applied to the robot
[1479] Specific Actions: Based on the embedded model, the humanoid robot raises its arm, aims, and shoots, faithfully replicating the shooting actions performed by the user in virtual reality.
[1480] Output: Reproduced human movements
[1481] (Application example 1)
[1482] Next, a description will be given of Application Example 1. In the following description, the data processing device 12 will be referred to as a "server" and the robot 414 will be referred to as a "terminal."
[1483] In today's industrial environment, training factory robots is expensive and requires advanced expertise. Furthermore, for robots to accurately replicate human movements, a large amount of realistic motion data is required, but collecting such data is not easy. Therefore, there is a need for an efficient, low-cost method for collecting high-quality motion data and training factory robots.
[1484] The specific processing by the specific processing unit 290 of the data processing device 12 in the application example 1 is realized by the following means.
[1485] In this invention, the server includes means for recording a user's actions in a virtual reality environment in real time, means for converting the recorded action data into a predetermined format, means for transmitting the converted data to the server, means for the server to analyze the received data and store the data by correlating the user's actions with actions in the virtual environment, means for generating a machine learning model based on the stored data and using it for motion learning of an operating device, and means for applying the converted action data to an industrial environment. This makes it possible to efficiently collect actions performed by a user in a virtual reality environment and for a factory robot to learn its actions with high accuracy based on the data.
[1486] A "virtual reality environment" is an environment that uses computer technology to allow users to visually and aurally experience a virtual world that is different from the real world.
[1487] "Motion data" refers to information such as the position, angle, and movement speed of each joint that is collected when a user performs a specific movement.
[1488] A "predetermined format" is a predetermined data format used to store and transmit data, such as CSV or JSON.
[1489] A "server" is a part of a computer system that receives, stores, analyzes, and transmits data over a network.
[1490] A "machine learning model" is a collection of algorithms that learn from accumulated data and analyze and predict its patterns.
[1491] "Motion device" refers to a mechanical device that can reproduce human movements, such as a factory robot.
[1492] "Industrial environment" means the physical space where manufacturing and production activities take place, such as a factory or manufacturing floor.
[1493] The present invention is a system that allows a user to perform actions in a virtual reality (VR) environment and works in conjunction with a server to apply the action data to an industrial environment. Specific embodiments of the system will be described below.
[1494] Overall system overview
[1495] Hardware configuration
[1496] VR headset: A device that allows a user to immerse themselves in a virtual reality environment.
[1497] Motion capture suit: A suit with built-in sensors to measure the position of each joint and the speed of movement (e.g., Xsens, Vicon).
[1498] Servers and Network: Computing resources for data storage, analysis, and machine learning model generation (e.g., AWS, Google Cloud).
[1499] Software configuration
[1500] Data collection applications: Applications that collect data from VR headsets and motion capture suits in real time.
[1501] Data transmission module: A module that converts collected data into a specified format and sends it to the server (e.g., Python's requests library).
[1502] Data analysis and machine learning model generation module: A module that analyzes the received data on the server and generates a machine learning model (e.g., TensorFlow, PyTorch).
[1503] Behavior reproduction module: A program for controlling the behavior of factory robots based on the generated machine learning model.
[1504] Operation procedures and examples
[1505] 1. User Preparation
[1506] The user puts on a VR headset and a motion capture suit and launches a dedicated data collection application. As the user performs factory tasks (e.g., assembling parts) in the virtual reality environment, their motion data is collected by the motion capture suit.
[1507] 2. Data collection and transmission
[1508] The data collection application receives the collected operational data in real time and converts it into a specific format (e.g., CSV or JSON format). The converted data is given a timestamp and a user ID. The data transmission module then sends the formatted data to the server.
[1509] 3. Data storage and analysis
[1510] The server stores the received data in a database (e.g., MongoDB or SQL database). It then analyzes the received data using a data analysis module. For analysis, it uses Python libraries (e.g., Pandas and NumPy) to calculate user behavior characteristics, task success rates, and other information.
[1511] 4. Generating a Machine Learning Model
[1512] Based on the results of the data analysis, the server generates a machine learning model using machine learning libraries such as TensorFlow and PyTorch. For example, a deep learning model can be built to learn the actions of a user assembling parts.
[1513] 5. Reproducing the behavior
[1514] The generated machine learning model is transferred from the server to the factory robot and incorporated into the robot's behavior reproduction program, allowing the factory robot to accurately mimic the user's actions and execute factory tasks.
[1515] Examples of specific examples and prompts
[1516] For example, to simulate a part assembly task in a factory in a VR environment, a user puts on a VR headset and a motion capture suit and performs a series of actions, picking up parts and installing them in designated positions. These action data are collected in real time, converted into a specified format, and sent to a server. After analyzing the data on the server, a factory robot can learn similar actions and reproduce the actual factory task.
[1517] Prompt Sentence Examples
[1518] "We are designing a system that simulates factory tasks in a VR environment, collects user motion data, and uses it to train robots. For example, when a user performs an assembly task, the motion capture suit collects that data and sends it to a server in real time. This data can then be analyzed to enable a factory robot to replicate the same assembly tasks."
[1519] The flow of the specific processing in the application example 1 will be described with reference to FIG.
[1520] Program processing flow
[1521] Step 1:
[1522] Input: A user wears a VR headset and motion capture suit and performs movements in a virtual reality environment.
[1523] Action: The user performs a factory task (e.g., assembling parts, moving pallets) within the VR environment.
[1524] Output: Motion data such as the position, angle, and movement speed of each joint is collected from the motion capture suit.
[1525] Step 2:
[1526] Input: Collected behavioral data.
[1527] Operation: The device converts the collected operation data into a specified format (CSV format, JSON format). The data collection application is used for the conversion, and the data is given a timestamp and user ID.
[1528] Output: Formatted behavior data (CSV format, JSON format).
[1529] Step 3:
[1530] Input: Formatted motion data.
[1531] Operation: The terminal transmits the formatted operation data to the server in real time. The data is transmitted via the Internet using the data transmission module.
[1532] Output: The server receives the operation data.
[1533] Step 4:
[1534] Input: The operational data received by the server.
[1535] Operation: The server stores the received data in a database (e.g., MongoDB or SQL database), then analyzes the data using Python libraries (Pandas, NumPy) to calculate user behavior characteristics, task success rates, etc.
[1536] Output: Parsed data and user behavior characteristics.
[1537] Step 5:
[1538] Input: Parsed data and user behavior characteristics.
[1539] How it works: The server generates a machine learning model based on the analyzed data. It trains the model using machine learning libraries such as TensorFlow and PyTorch, and builds a deep learning model to reproduce the user's behavior.
[1540] Output: The generated machine learning model.
[1541] Step 6:
[1542] Input: The generated machine learning model.
[1543] Operation: The server incorporates the generated machine learning model into the factory robot's behavior reproduction program and transfers the model to the factory robot. Data transmission technology is used over the network for the transfer.
[1544] Output: Machine learning models embedded in factory robots.
[1545] Step 7:
[1546] Input: Machine learning models embedded in factory robots.
[1547] Operation: The factory robot replicates the user's actions based on machine learning models and performs real factory tasks, such as assembling parts and moving pallets with high precision.
[1548] Output: Actual factory task completion.
[1549] This series of processes enables the efficient collection of user actions in a virtual reality environment, enabling factory robots to learn and reproduce these actions with high accuracy based on that data.
[1550] Furthermore, an emotion engine that estimates the user's emotion may be further combined. That is, the identification processing unit 290 may estimate the user's emotion using the emotion identification model 59, and perform identification processing using the user's emotion.
[1551] The present invention is a system that collects motion and emotion data when a user plays a game in a virtual reality (VR) environment, and uses this data to train a humanoid robot. Specific embodiments are described below.
[1552] Overall system overview
[1553] The system consists of the following main components:
[1554] 1. VR games played by users and motion capture suits
[1555] 2. Device that collects and transmits motion data and emotional data
[1556] 3. Server that receives and analyzes data and generates machine learning models
[1557] 4. Emotion engine that recognizes user emotions
[1558] User Preparation
[1559] User: First, the user puts on the VR headset, motion capture suit, and emotion engine. Then, they launch the dedicated application and start the game. The motion capture suit and emotion engine collect the user's movement data and emotion data, respectively.
[1560] Recording behavioral and emotional data during gameplay
[1561] Terminal: When a user plays a game, the terminal receives the user's movement data in real time from the motion capture suit and acquires emotion data from the emotion engine. For example, when a user plays a shooting game, the terminal collects arm movements, joint position data, and emotional states (joy, anger, sadness, and happiness) in real time.
[1562] Terminal: Converts the received motion data and emotion data into a specified format (CSV, JSON, etc.) and adds a timestamp and user ID to each data.
[1563] Data transmission and storage
[1564] Terminal: Sends formatted motion data and emotion data to the server in real time or in batch format.
[1565] Server: Receives data sent from the device and stores each user's behavioral and emotional data in a database. For example, structured data is stored using MongoDB or an SQL database.
[1566] Data analysis and generation of learning models
[1567] Server: Analyzes the saved motion and emotion data. Specifically, it converts the data into a data frame using Python libraries (such as Pandas and NumPy) and compares the user's motion characteristics, emotional state, and in-game actions with the score.
[1568] Server: Generates a machine learning model based on the analysis results. For example, a deep learning model is built using TensorFlow or PyTorch to learn the user's behavior and emotional patterns.
[1569] Application in humanoid robots
[1570] Server: Incorporates the generated machine learning model into the control program of the humanoid robot.
[1571] Humanoid robot (user-like): Based on a new machine learning model, it performs human-like movements and emotional responses, for example, reproducing the movements and emotional states of a user in a shooting game.
[1572] As a concrete example, consider the movements and emotions of a user playing a shooting game. In this case, the user's arm movements to aim and fire are recorded. At the same time, an emotion engine records the user's emotional state (e.g., excitement or tension). This data is transmitted in real time to the device and then to a server. The server analyzes the data and associates the user's movements and emotional state with the in-game score. This information is used to generate a machine learning model that enables a humanoid robot to reproduce similar shooting movements and emotions.
[1573] In this way, the present invention provides a system that can efficiently collect human motion data and emotional states and use them for motion learning and emotional expression of humanoid robots.
[1574] The processing flow will be explained below.
[1575] Step 1:
[1576] User: Puts on the VR headset, motion capture suit, and emotion engine, launches the VR game application, and begins the game.
[1577] Step 2:
[1578] Terminal: Receives user movement data (e.g., joint positions, arm movements) in real time from the motion capture suit. Obtains user emotion data (e.g., joy, anger, sadness, and happiness) in real time from the emotion engine.
[1579] Step 3:
[1580] Terminal: Converts the received movement and emotion data into a specified format (CSV, JSON, etc.). The converted data includes the X, Y, and Z coordinates of each joint, angle, movement speed, and emotional state.
[1581] Step 4:
[1582] Terminal: Adds identification information such as a timestamp and user ID to the formatted data.
[1583] Step 5:
[1584] Terminal: Transmits the converted motion data and emotion data to the server in real time or in batch format.
[1585] Step 6:
[1586] Server: Receives data sent from the devices. The received data includes each user's motion data, emotion data, and time information.
[1587] Step 7:
[1588] Server: Stores the received data in a database, for example, using MongoDB or an SQL database to store structured data.
[1589] Step 8:
[1590] Server: Analyzes the saved behavioral and emotional data. Specifically, it converts the data into a data frame using Python libraries (Pandas and NumPy) and compares the user's behavioral characteristics, emotional state, in-game actions, and score.
[1591] Step 9:
[1592] Server: Generates a machine learning model based on the analysis results. For example, a deep learning model is built using TensorFlow or PyTorch to learn the user's behavior and emotional patterns.
[1593] Step 10:
[1594] Server: Incorporates the generated machine learning model into the control program of the humanoid robot.
[1595] Step 11:
[1596] Humanoid robot (user-like): Based on a new machine learning model, the robot can perform human-like actions and emotional responses, such as reproducing the actions and emotional states (e.g., excitement and tension) of a user in a shooting game.
[1597] Example 2
[1598] Next, a description will be given of Example 2. In the following description, the data processing device 12 will be referred to as a "server" and the robot 414 will be referred to as a "terminal."
[1599] While there are systems that record only the user's motion data during gameplay in conventional virtual reality environments, there are no systems that simultaneously record the user's emotional data in real time and generate highly accurate machine learning models based on that data. As a result, it is difficult for humanoid robots to reproduce human-like motions and emotional expressions.
[1600] The identification process by the identification processing unit 290 of the data processing device 12 in the second embodiment is realized by the following means. In this invention, the server includes means for recording movement and emotion data in real time, means for converting the recorded movement and emotion data into a predetermined format, and means for adding a timestamp and user ID to the converted data and transmitting it to the server. This makes it possible to collect and analyze user movement and emotion data with high accuracy.
[1601] "User" refers to a person who plays a game in a virtual reality environment.
[1602] A "virtual reality environment" refers to an environment in which you can experience a virtual space that is different from reality using devices such as a VR headset.
[1603] "Motion data" refers to data relating to the physical movements of a user when playing a game in a virtual reality environment.
[1604] "Emotion data" refers to data relating to the user's emotional state (for example, joy, anger, sadness, happiness, etc.).
[1605] "Real-time" refers to the near-simultaneous collection, processing, and transmission of data.
[1606] "Specified format" refers to converting data into a prescribed format, such as CSV or JSON.
[1607] "Time stamp" refers to the time information of data collection.
[1608] "User ID" refers to an identifier that uniquely identifies a user.
[1609] "Server" refers to the computer system that stores and analyzes the received data and generates the machine learning model.
[1610] A "machine learning model" is a model that learns patterns based on collected data and makes predictions and classifications for unknown data.
[1611] A "humanoid robot" refers to a robot that has the ability to imitate human movements and emotional expressions.
[1612] MODE FOR CARRYING OUT THE INVENTION
[1613] The present invention is a system for collecting motion and emotion data when a user plays a game in a virtual reality environment, and for using the collected data to train a humanoid robot. Specific embodiments of the system are described below.
[1614] User Preparation
[1615] User: First, the user puts on the VR headset, motion capture suit, and emotion engine. Next, they launch a dedicated application and start the game. This prepares the motion capture suit and emotion engine to collect the user's movement and emotion data.
[1616] Real-time recording and conversion of motion and emotion data
[1617] Device: When a user plays a game, the device receives the user's movement data in real time from the motion capture suit and obtains emotion data from the emotion engine. For example, when playing a shooting game, the device collects the user's arm movements, joint position data, and emotional state (joy, anger, sadness, and happiness) in real time. This data is converted into CSV or JSON format, and a timestamp and user ID are attached to each data point.
[1618] Data transmission and storage
[1619] Terminal: Sends formatted motion data and emotion data to the server in real time or in batch format.
[1620] Server: Receives data sent from the devices and stores each user's motion data and emotion data in a database. MongoDB or an SQL database is used for the database. For example, using MongoDB makes it possible to efficiently store and manage motion data and emotion data.
[1621] Data analysis and generation of learning models
[1622] Server: Analyzes the saved movement and emotion data. Specifically, it uses Python libraries (such as Pandas and NumPy) to create data frames and compares the user's movement characteristics, emotional state, and in-game movements with their score. For example, by combining and analyzing the user's joint position data and emotion data, it is possible to gain a detailed understanding of the user's movement and emotion patterns.
[1623] Server: Generates a machine learning model based on the analysis results. A deep learning model is built using TensorFlow or PyTorch to learn the user's behavior and emotional patterns. For example, TensorFlow can be used to build a multi-layer perceptron or recurrent neural network (RNN) to achieve a highly accurate prediction model.
[1624] Application in humanoid robots
[1625] Server: The generated machine learning model is incorporated into the control program of the humanoid robot, enabling it to reproduce human-like behavior and emotions based on the behavior and emotion patterns that the robot has learned.
[1626] Humanoid robots: Based on new machine learning models, these robots can perform human-like actions and emotional responses. For example, they can replicate the actions and emotional states of a user in a shooting game. Specifically, the robot can aim at a target, pull the trigger, and simultaneously express emotional states such as excitement or tension.
[1627] Specific examples
[1628] Consider the motion and emotional data of a user playing a shooting game. The user's arm movements to aim and fire are recorded. At the same time, an emotion engine records the user's emotional state (e.g., excitement or tension). This data is transmitted in real time to the device and then to a server. The server analyzes the data and associates the user's motion and emotional state with the in-game score. A machine learning model is generated based on this information, enabling a humanoid robot to reproduce similar shooting motions and emotions.
[1629] Prompt Sentence Examples
[1630] "Describe a system that collects motion and emotion data from a user playing a shooting game and generates a learning model that allows a humanoid robot to reproduce those motions and emotions."
[1631] The flow of the identification process in the second embodiment will be described with reference to FIG.
[1632] Step 1: Prepare your users
[1633] Input: The user provides a VR headset, a motion capture suit, and an emotion engine.
[1634] Specific operation: The user puts on a VR headset and a motion capture suit. They then wear or install an emotion engine (heart rate sensor and facial recognition camera). They then launch the dedicated application on their device, select a game, and start playing.
[1635] Output: The user's behavioral and emotional data is ready to be collected.
[1636] Step 2: Recording behavioral and emotional data during gameplay
[1637] Input: The user begins playing the game.
[1638] Specific movements: When the user moves in the game, sensors in the motion capture suit capture the position and movement of the joints, and the emotion engine monitors facial expressions and heart rate in real time.
[1639] Output: Movement data from the motion capture suit and emotion data from the emotion engine are sent to the terminal in real time.
[1640] Step 3: Transforming behavior and emotion data
[1641] Input: Real-time received behavioral and emotional data.
[1642] Specific operation: The terminal converts the data received in real time into CSV or JSON format and adds a timestamp and user ID to each data.
[1643] Output: Data formatted in a given format.
[1644] Step 4: Sending data
[1645] Input: Formatted motion and emotion data.
[1646] Specific operation: The terminal sends the formatted data packets to the server in real time or in batch format.
[1647] Output: Data packets sent to the server.
[1648] Step 5: Save your data
[1649] Input: The data packet sent to the server.
[1650] Specific operation: The server unpacks the received data packets and stores the data in a MongoDB or SQL database.
[1651] Output: Behavioral and emotional data stored in a database.
[1652] Step 6: Analyze the data
[1653] Input: Stored behavioral and emotional data.
[1654] Specific operation: The server uses Python libraries (such as Pandas and NumPy) to convert the data into a data frame, process missing values, and analyze the user's behavioral characteristics and emotional state.
[1655] Output: Analysis results (data about the user's behavioral characteristics, emotional state, and in-game score).
[1656] Step 7: Generate a training model
[1657] Input: Analysis results.
[1658] Specific operation: Based on the analysis results, the server uses TensorFlow or PyTorch to build a deep learning model to learn the user's behavior and emotional patterns.
[1659] Output: The generated machine learning model.
[1660] Step 8: Application in humanoid robots
[1661] Input: The generated machine learning model.
[1662] Specific behavior: The server deploys the trained model to the control system of the humanoid robot, and the robot reproduces human-like behavior and emotions based on the learned behavior and emotion patterns.
[1663] Output: The humanoid robot performs actions that reflect the user's actions and emotions.
[1664] (Application example 2)
[1665] Next, a description will be given of Application Example 2. In the following description, the data processing device 12 will be referred to as a "server" and the robot 414 will be referred to as a "terminal."
[1666] The present invention relates to a system for providing effective customer service in brick-and-mortar stores based on motion and emotion data acquired when users play games in a virtual reality environment. Conventional technologies have made it difficult for customer service staff to grasp customers' emotions and motions in real time and provide appropriate service. This has resulted in reduced customer satisfaction and lost sales opportunities.
[1667] The specific processing by the specific processing unit 290 of the data processing device 12 in Application Example 2 is realized by the following means. In this invention, the server includes means for recording in real time the actions of a user playing a game in a virtual reality environment, means for converting the recorded action data into a predetermined format, means for transmitting the converted data to the server, means for analyzing the data received by the server and storing the data in association with the user's actions and in-game actions, means for generating a machine learning model based on the stored data and using it for motion training of a humanoid robot, means for capturing the customer's facial expressions and body movements in real time, collecting and analyzing emotional data and action data, and means for providing customer service advice to the customer based on the analysis results. This enables the customer service staff to grasp the customer's emotional state in real time and provide appropriate service.
[1668] A "virtual reality environment" is a technology that allows users to experience virtual spaces and situations that feel real using a computer or special equipment.
[1669] "Motion data" is data that records motion information such as the user's body movements and gestures as numerical values.
[1670] "Emotion data" is data collected to express the user's emotional state, and is inferred from facial expressions, voice, behavior, etc.
[1671] A "motion capture suit" is a device that records a user's body movements with high precision.
[1672] The "emotion engine" is software that analyzes collected data to estimate and recognize the user's emotions.
[1673] A "server" is a computer system that receives, stores, and analyzes data over a network.
[1674] A "machine learning model" is an algorithm that is trained to perform a specific task based on collected and analyzed data.
[1675] A "humanoid robot" is a robot designed to resemble a human and capable of performing human-like movements.
[1676] "Customer service staff" are employees whose role is to provide products and services to customers in physical stores.
[1677] "Response advice" refers to instructions for action or suggestions for improvement that the system provides to customer service staff based on the analysis results.
[1678] A "brick and mortar store" is a physical store where customers can visit in person to purchase goods or services.
[1679] Overall system overview
[1680] The system of the present invention analyzes the facial expressions and movements of customers in real time when the user is serving customers in a physical store, and provides corresponding advice to the staff. The system includes the following main components and software:
[1681] 1. Smart glasses (equipped with a camera for capturing facial expressions and movements)
[1682] 2. Device that collects behavioral and emotional data
[1683] 3. Server that receives and analyzes data and generates machine learning models
[1684] 4. Emotion engine that recognizes user emotions
[1685] Preparing the smart glasses and devices
[1686] Customer service staff: Customer service staff wear smart glasses and launch a dedicated application. The smart glasses are equipped with a camera that captures facial expressions and movements. The images captured by the camera are sent to a terminal in real time.
[1687] Recording facial and movement data
[1688] Terminal: When a wait staff member meets with a customer, the smart glasses camera receives real-time facial and behavior data, including face detection and image processing using OpenCV and emotion prediction using TensorFlow.
[1689] Data transmission and storage
[1690] Terminal: The collected facial and movement data is converted into JSON format and sent to the server in real time or in batch format. The data is accompanied by a timestamp and the staff member ID.
[1691] Server: Receives data sent from the devices and stores the behavioral and emotional data of each customer service staff member in a database, typically using MongoDB or an SQL database.
[1692] Data analysis and generation of learning models
[1693] Server: Analyzes the stored behavioral and emotional data. Python libraries such as Pandas and NumPy are used to create data frames, and the behavioral characteristics, emotional states, and reactions of the staff to customer interactions are compared.
[1694] Server: Generates a machine learning model based on the analysis results. A deep learning model is built using TensorFlow and PyTorch, and trained to learn the behavior and emotional patterns of the customer service staff.
[1695] Providing customer service assistants
[1696] Server: Based on the generated machine learning model, the server provides real-time advice to the customer service staff based on the analysis results. For example, if a customer looks dissatisfied, the server displays the advice, "The customer seems dissatisfied. Please suggest additional services."
[1697] Specific examples
[1698] Consider a scenario where a customer enters a store and looks at products. At this time, smart glasses capture the customer's facial expressions and identify emotions such as "excitement" or "tension." The analysis results are sent to the device in real time and then stored and analyzed on a server. Based on this information, the customer service staff can receive advice such as, "The customer is showing interest. Please explain the product to them."
[1699] Prompt Sentence Examples
[1700] "Enter customer facial expression data and predict their emotion. Generate text that suggests appropriate customer interactions based on the predicted emotion."
[1701] The flow of the specific processing in the application example 2 will be described with reference to FIG.
[1702] Program processing steps and detailed explanations
[1703] Step 1:
[1704] Input: Smart glasses camera image
[1705] How it works: Waiting staff wear smart glasses and capture camera footage in real time as they interact with customers.
[1706] Output: Real-time video data
[1707] explanation:
[1708] The user (a customer service staff member) puts on the smart glasses and begins to interact with the customer. The camera in the smart glasses captures the customer's facial expressions and movements in real time and sends the video data to the terminal.
[1709] Step 2:
[1710] Input: Real-time video data
[1711] How it works: The device processes video data using OpenCV, performs face detection and image processing, and then performs facial recognition using TensorFlow.
[1712] Output: facial expression data and movement data
[1713] explanation:
[1714] The device uses OpenCV to analyze real-time video data received from the smart glasses. First, it performs face detection, then extracts the facial region, converts it to grayscale, and inputs it into a facial expression recognition model (TensorFlow). This generates facial expression data (emotional state) and movement data of the customer.
[1715] Step 3:
[1716] Input: facial expression data and movement data
[1717] How it works: The device converts the data into JSON format and adds a timestamp and the customer service staff ID.
[1718] Output: Formatted data (JSON format)
[1719] explanation:
[1720] The device converts the acquired facial expression and movement data into a specified format (JSON) and adds a timestamp and the customer service staff ID. This data is used for subsequent analysis.
[1721] Step 4:
[1722] Input: Formatted data (JSON format)
[1723] How it works: The device sends formatted data to the server in real time or in batches.
[1724] Output: Transmitted data
[1725] explanation:
[1726] The terminal sends formatted JSON data to the server in real time or in batch format, and communication protocols such as HTTP and WebSocket can be used.
[1727] Step 5:
[1728] Input: Send data
[1729] Operation: The server stores the received data in a database.
[1730] Output: Stored data
[1731] explanation:
[1732] The server receives the JSON data sent from the device and stores it in a MongoDB or SQL database, making it easy to store and analyze the data later.
[1733] Step 6:
[1734] Input: Stored data
[1735] How it works: The server uses Pandas or NumPy to create a data frame and perform analysis.
[1736] Output: Analysis results
[1737] explanation:
[1738] The server then converts the stored data into a data frame using Pandas or NumPy to analyze the customer's behavioral characteristics and emotional state, making it easier to analyze the data and extract specific patterns.
[1739] Step 7:
[1740] Input: Analysis results
[1741] How it works: The server generates a machine learning model using TensorFlow or PyTorch.
[1742] Output: Machine learning model
[1743] explanation:
[1744] The server uses TensorFlow and PyTorch to generate machine learning models based on the analysis results, which improves the accuracy of predictions of customer behavior and emotions.
[1745] Step 8:
[1746] Input: Machine learning models and real-time sentiment data
[1747] Operation: The server generates response advice based on the analysis results and sends it to the device.
[1748] Output: Response advice
[1749] explanation:
[1750] The server generates analysis results based on the generated machine learning model and real-time emotion data, and provides real-time advice to the customer service staff, which is displayed to the customer service staff via their terminal.
[1751] Examples of specific examples and prompts
[1752] If the customer looks unhappy, the advice displayed will be "The customer looks unhappy. Please suggest additional services."
[1753] Prompt Sentence Examples
[1754] "Enter customer facial expression data and predict their emotion. Generate text that suggests appropriate customer interactions based on the predicted emotion."
[1755] The specific processing unit 290 transmits the result of the specific processing to the robot 414. In the robot 414, the control unit 46A causes the speaker 240 and the control target 443 to output the result of the specific processing. The microphone 238 acquires voice indicating a user input regarding the result of the specific processing. The control unit 46A transmits voice data indicating the user input acquired by the microphone 238 to the data processing device 12. In the data processing device 12, the specific processing unit 290 acquires the voice data.
[1756] The data generation model 58 is a so-called generative AI (Artificial Intelligence). An example of the data generation model 58 is ChatGPT (Internet Search<URL: https: / / openai.com / blog / chatgpt> ), Gemini (Internet search <url: https: gemini.google.com ?hl="ja">) and other generation AIs. The data generation model 58 is obtained by performing deep learning on a neural network. A prompt including an instruction is input to the data generation model 58, and inference data such as voice data indicating voice, text data indicating text, and image data indicating an image is also input. The data generation model 58 performs inference on the input inference data in accordance with the instruction indicated by the prompt, and outputs the inference result in a data format such as voice data and text data. Here, inference refers to, for example, analysis, classification, prediction, and / or summarization.
[1757] In the above embodiment, an example was given in which the specific processing is performed by the data processing device 12, but the technology of the present disclosure is not limited to this, and the specific processing may be performed by the robot 414.
[1758] The emotion identification model 59 as an emotion engine may determine the user's emotion according to a specific mapping. Specifically, the emotion identification model 59 may determine the user's emotion according to an emotion map (see FIG. 9), which is a specific mapping. Similarly, the emotion identification model 59 may determine the robot's emotion, and the identification processing unit 290 may perform identification processing using the robot's emotion.
[1759] FIG. 9 is a diagram illustrating an emotion map 400 on which multiple emotions are mapped. In the emotion map 400, emotions are arranged in concentric circles radiating from the center. Emotions closer to the center of the concentric circles are more primitive. Emotions representing states and actions arising from a state of mind are arranged on the outer edges of the concentric circles. The concept of emotion includes both affect and mental states. Emotions generally generated from reactions occurring in the brain are arranged on the left side of the concentric circles. Emotions generally induced by situational judgment are arranged on the right side of the concentric circles. Emotions generally generated from reactions occurring in the brain and induced by situational judgment are arranged on the upper and lower sides of the concentric circles. Furthermore, the emotion of "pleasure" is arranged on the upper side of the concentric circles, and the emotion of "discomfort" is arranged on the lower side. In this way, in the emotion map 400, multiple emotions are mapped based on the structure by which emotions are generated, and emotions that tend to occur simultaneously are mapped close to each other.
[1760] These emotions are distributed in the 3 o'clock direction on emotion map 400, and typically fluctuate between relief and anxiety. In the right half of emotion map 400, situational awareness dominates over internal sensations, resulting in a sense of calm.
[1761] The inside of emotion map 400 represents what is going on in the mind, and the outside of emotion map 400 represents behavior, so the further you go outside emotion map 400, the more visible the emotions become (the more they are expressed in behavior).
[1762] Human emotions are based on various balances, such as posture and blood sugar levels. When these balances deviate from the ideal, a state of discomfort is indicated, and when they approach the ideal, a state of pleasure is indicated. Emotions can also be created for robots, automobiles, and motorcycles, based on various balances, such as posture and remaining battery life. When these balances deviate from the ideal, a state of discomfort is indicated, and when they approach the ideal, a state of pleasure is indicated. An emotion map can be generated, for example, based on Dr. Mitsuyoshi's emotion map (Research on Voice Emotion Recognition and Emotional Brain Physiological Signal Analysis Systems, Tokushima University, Doctoral Dissertation: https: / / ci.nii.ac.jp / naid / 500000375379). The left half of the emotion map lists emotions belonging to the "reaction" domain, where sensation is dominant. The right half of the emotion map lists emotions belonging to the "situation" domain, where situational awareness is dominant.
[1763] The emotion map defines two emotions that promote learning. One is a negative emotion on the situation side, around the middle of "repentance" or "reflection." In other words, this occurs when the robot experiences negative emotions such as "I never want to feel this way again" or "I don't want to be scolded again." The other is a positive emotion on the response side, around "desire." In other words, this occurs when the robot experiences positive feelings such as "I want more" or "I want to know more."
[1764] The emotion identification model 59 inputs user input into a pre-trained neural network, obtains emotion values indicating each emotion shown in the emotion map 400, and determines the user's emotion. This neural network is pre-trained based on multiple pieces of training data that are combinations of user input and emotion values indicating each emotion shown in the emotion map 400. Furthermore, this neural network is trained so that emotions that are located close to each other have similar values, as in the emotion map 900 shown in FIG. 10. FIG. 10 shows an example in which multiple emotions, "relieved," "calm," and "reassuring," have similar emotion values.
[1765] The system according to the present disclosure has been described above mainly with respect to the functions of the data processing device 12, but the system according to the present disclosure is not necessarily implemented on a server. The system according to the present disclosure may be implemented as a general information processing system. The present disclosure may be implemented, for example, as a software program running on a personal computer or an application running on a smartphone, etc. The method according to the present disclosure may be provided to users in the form of SaaS (Software as a Service).
[1766] In the above embodiment, an example was given in which the specific processing is performed by one computer 22, but the technology of the present disclosure is not limited to this, and the specific processing may be distributed and performed by a plurality of computers including the computer 22. For example, the data generation model 58 may be provided in an external device of the data processing device 12, and data may be generated in the external device in accordance with input data.
[1767] In the above embodiment, an example in which the specific processing program 56 is stored in the storage 32 has been described, but the technology of the present disclosure is not limited to this. For example, the specific processing program 56 may be stored in a portable, computer-readable, non-transitory storage medium such as a USB (Universal Serial Bus) memory. The specific processing program 56 stored in the non-transitory storage medium is installed in the computer 22 of the data processing device 12. The processor 28 executes the specific processing in accordance with the specific processing program 56.
[1768] Alternatively, the specific processing program 56 may be stored in a storage device such as a server connected to the data processing device 12 via the network 54, and the specific processing program 56 may be downloaded and installed on the computer 22 in response to a request from the data processing device 12.
[1769] It is not necessary to store all of the specific processing program 56 in a storage device such as a server connected to the data processing device 12 via the network 54, or to store all of the specific processing program 56 in the storage 32; only a portion of the specific processing program 56 may be stored.
[1770] The hardware resource for executing a specific process can be any of the following processors: An example of a processor is a CPU, which is a general-purpose processor that functions as a hardware resource for executing a specific process by executing software, i.e., a program. Another example of a processor is a dedicated electrical circuit, such as an FPGA (Field-Programmable Gate Array), a PLD (Programmable Logic Device), or an ASIC (Application Specific Integrated Circuit), which is a processor with a circuit configuration designed specifically for executing a specific process. Each processor has built-in or connected memory, and each processor uses the memory to execute the specific process.
[1771] The hardware resource that executes the specific processing may be configured with one of these various processors, or may be configured with a combination of two or more processors of the same or different types (for example, a combination of multiple FPGAs, or a combination of a CPU and an FPGA). Also, the hardware resource that executes the specific processing may be a single processor.
[1772] As an example of a system configured with a single processor, first, one processor is configured by combining one or more CPUs and software, and this processor functions as a hardware resource that executes a specific process. Second, there is a system that uses a processor that realizes the functions of an entire system including multiple hardware resources that execute a specific process on a single IC chip, as typified by SoC (System-on-a-chip). In this way, a specific process is realized using one or more of the above-mentioned various processors as hardware resources.
[1773] Furthermore, the hardware structure of these various processors can be, more specifically, an electric circuit that combines circuit elements such as semiconductor devices. The specific processing described above is merely an example. Therefore, it goes without saying that unnecessary steps may be deleted, new steps may be added, or the processing order may be rearranged, without departing from the spirit of the invention.
[1774] The above-described description and illustrations are a detailed explanation of the parts related to the technology of the present disclosure and are merely an example of the technology of the present disclosure. For example, the above description of the configuration, functions, actions, and effects is an explanation of an example of the configuration, functions, actions, and effects of the parts related to the technology of the present disclosure. Therefore, it goes without saying that unnecessary parts may be deleted, new elements may be added, or replacements may be made to the above-described description and illustrations within the scope of the gist of the technology of the present disclosure. Furthermore, to avoid confusion and facilitate understanding of the parts related to the technology of the present disclosure, the above-described description and illustrations omit explanations of common technical knowledge that do not require particular explanation to enable the implementation of the technology of the present disclosure.
[1775] All publications, patent applications, and technical standards mentioned in this specification are herein incorporated by reference to the same extent as if each individual publication, patent application, or technical standard was specifically and individually indicated to be incorporated by reference.
[1776] The following is further disclosed regarding the above embodiment.
[1777] (Claim 1)
[1778] means for recording in real time the actions of a user playing a game in a virtual reality environment;
[1779] means for converting the recorded motion data into a predetermined format;
[1780] means for transmitting the converted data to a server;
[1781] A means for analyzing the data received by the server and storing the data in association with the user's actions in the game;
[1782] A means to generate a machine learning model based on the accumulated data and use it to train the movements of a humanoid robot;
[1783] A system including:
[1784] (Claim 2)
[1785] 10. The system of claim 1, further comprising means for transmitting the recorded motion data to a server in real time.
[1786] (Claim 3)
[1787] 10. The system of claim 1, further comprising: means for optimizing the machine learning model based on a success rate of a user's in-game actions.
[1788] "Example 1"
[1789] (Claim 1)
[1790] means for recording in real time the actions of a user playing a game in a virtual reality environment;
[1791] means for converting the recorded motion data into a predetermined format and adding time information and user identification information;
[1792] means for transmitting the converted data to a server in real time or in batch;
[1793] a means for storing and organizing the data received by the server in a database and programmatically generating and analyzing data frames;
[1794] A means for generating a deep learning model based on the analysis results and integrating the model into the control system of the humanoid robot;
[1795] A system including:
[1796] (Claim 2)
[1797] 10. The system of claim 1, further comprising means for transmitting the recorded motion data to a server in real time.
[1798] (Claim 3)
[1799] 10. The system of claim 1, further comprising: means for optimizing the deep learning model based on a success rate of a user's in-game actions.
[1800] "Application Example 1"
[1801] (Claim 1)
[1802] means for a user to record movements in a virtual reality environment in real time;
[1803] means for converting the recorded motion data into a predetermined format;
[1804] means for transmitting the converted data to a server;
[1805] A means for analyzing the data received by the server and associating the user's actions with actions in the virtual environment and storing the results;
[1806] A means for generating a machine learning model based on the accumulated data and using the model for motion learning of the motion device;
[1807] means for applying the converted operational data to an industrial environment;
[1808] A system including:
[1809] (Claim 2)
[1810] 10. The system of claim 1, further comprising means for transmitting the recorded motion data to a server in real time.
[1811] (Claim 3)
[1812] 10. The system of claim 1, further comprising means for optimizing the machine learning model based on a success rate of a user's actions in the virtual environment.
[1813] "Example 2: Combining Emotion Engines"
[1814] New Claims
[1815] (Claim 1)
[1816] means for recording motion and emotion data in real time as a user plays a game in a virtual reality environment;
[1817] means for converting the recorded motion data and emotion data into a predetermined format;
[1818] a means for adding a timestamp and a user ID to the converted data and transmitting the data to a server;
[1819] A means for analyzing the data received by the server and associating the user's movements and emotional state with in-game actions and storing the associated data;
[1820] A means to generate a machine learning model based on the accumulated data and use it to learn the movements of a humanoid robot;
[1821] A system including:
[1822] (Claim 2)
[1823] 10. The system of claim 1, further comprising means for transmitting the recorded motion data and emotion data to a server in real time.
[1824] (Claim 3)
[1825] 10. The system of claim 1, further comprising: means for optimizing the machine learning model based on a success rate of a user's in-game actions.
[1826] "Application example 2 when combining emotion engines"
[1827] (Claim 1)
[1828] means for recording in real time the actions of a user playing a game in a virtual reality environment;
[1829] means for converting the recorded motion data into a predetermined format;
[1830] means for transmitting the converted data to a server;
[1831] A means for analyzing the data received by the server and storing the data in association with the user's actions in the game;
[1832] A means to generate a machine learning model based on the accumulated data and use it to train the movements of a humanoid robot;
[1833] A means of capturing customer facial expressions and body movements in real time, and collecting and analyzing emotional and behavioral data;
[1834] A means for providing response advice to customer service staff based on the analysis results;
[1835] A system including:
[1836] (Claim 2)
[1837] 10. The system of claim 1, further comprising means for transmitting the recorded motion data to a server in real time.
[1838] (Claim 3)
[1839] 10. The system of claim 1, further comprising: means for optimizing the machine learning model based on a success rate of a user's in-game actions. [Explanation of symbols]
[1840] 10, 210, 310, 410 Data Processing Systems 12 Data Processing Device 14 Smart Devices 214 Smart Glasses 314 Headset-type terminal 414 Robot< / url:> < / url:> < / url:> < / url:>
Claims
1. means for recording in real time the actions of a user playing a game in a virtual reality environment; means for converting the recorded motion data into a predetermined format; means for transmitting the converted data to a server; A means for analyzing the data received by the server and storing the data in association with the user's actions in the game; A means to generate a machine learning model based on the accumulated data and use it to train the movements of a humanoid robot; A system including:
2. 10. The system of claim 1, further comprising means for transmitting the recorded motion data to a server in real time.
3. The system of claim 1 , further comprising: means for optimizing the machine learning model based on the success rate of a user's in-game actions.
Citation Information
Patent Citations
Persona chatbot control method and system
JP2022180282A