system

The system addresses inefficient educational methods by using motion capture and AR/VR feedback to provide real-time, personalized learning experiences that enhance skill acquisition.

JP2026069171APending Publication Date: 2026-04-23SOFTBANK GROUP CORP
View PDF 1 Cites 0 Cited by

Patent Information

Authority / Receiving Office
JP · JP
Patent Type
Applications
Current Assignee / Owner
SOFTBANK GROUP CORP
Filing Date
2024-10-11
Publication Date
2026-04-23

AI Technical Summary

Technical Problem

Existing educational methods rely heavily on direct instructor guidance and self-study, leading to insufficient individual feedback and inefficient technology acquisition, particularly in terms of pace and tailored learning.

Method used

A system utilizing a motion capture device to compare user and teacher movements, providing real-time feedback through AR/VR devices, and incorporating an emotion engine for personalized guidance.

Benefits of technology

Enhances learning efficiency by offering immediate, personalized feedback that aligns with individual learning styles and emotional states, improving skill acquisition.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure 2026069171000001_ABST
    Figure 2026069171000001_ABST
Patent Text Reader

Abstract

We provide the system. [Solution] A motion capture device for capturing movements in real time, A processing unit that compares and analyzes user behavior data and teacher behavior data, A display device that overlays and displays the user's actions in a virtual space based on the analysis results, A feedback device that provides users with feedback on areas for improvement in operation, A system that includes this.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The technology of the present disclosure relates to a system.

Background Art

[0002] Patent Document 1 discloses a method for controlling a persona chatbot, which is performed by at least one processor, and includes steps of receiving a user utterance, adding the user utterance to a prompt including an instruction sentence related to an explanation of a character of the chatbot, encoding the prompt, and inputting the encoded prompt into a language model to generate a chatbot utterance in response to the user utterance.

Prior Art Documents

Patent Documents

[0003]

Patent Document 1

Summary of the Invention

Problems to be Solved by the Invention

[0004] In modern education, the acquisition of technologies where operations are particularly important still largely depends on direct guidance from instructors and self-study by students. As a result, there are many situations where individual feedback is limited and efficient technology acquisition is difficult. In addition, there are problems that learning and feedback tailored to the pace of each student are not sufficient and acquisition takes time. The present invention aims to solve these problems and provide an environment in which students can acquire technologies more effectively and efficiently.

Means for Solving the Problems

[0005] This invention includes a motion capture device for capturing movements in real time, thereby acquiring motion data from both the user and the teacher. A processing device is then used to compare and analyze the user's motion data and the teacher's motion data, performing a highly accurate analysis of their respective motion characteristics. Based on the analysis results, a display device is introduced that overlays the user's actual movements and the teacher's ideal movements onto a virtual space, allowing the user to intuitively understand areas for improvement in their own movements. Furthermore, a dedicated feedback device is provided to give the user real-time feedback on areas for improvement, helping to accelerate student learning.

[0006] A "motion capture device" is a device that captures movement in real time and collects positional information of each joint and skeleton.

[0007] "User motion data" refers to data collected by a motion capture device, including positional information and motion characteristics related to the user's body movements.

[0008] "Teacher motion data" refers to data collected by motion capture devices, including positional information and motion characteristics related to the body movements of instructors demonstrating ideal movements.

[0009] "Comparative analysis" is a process that compares user behavior data with training data, analyzes the differences, and derives information for improvement.

[0010] A "processing device" is a computing device used to analyze collected data and compare the movements of the user and the teacher.

[0011] A "display device" is a device that visually presents the results of overlaying the user's actions with those of the teacher.

[0012] A "feedback device" is a device that directly and in real time communicates to the user areas for improvement and the degree of achievement in their operation. [Brief explanation of the drawing]

[0013] [Figure 1] This is a conceptual diagram showing an example of the configuration of a data processing system according to the first embodiment. [Figure 2] This is a conceptual diagram showing an example of the essential functions of a data processing device and a smart device according to the first embodiment. [Figure 3] This is a conceptual diagram showing an example of the configuration of a data processing system according to the second embodiment. [Figure 4] This is a conceptual diagram showing an example of the main functions of a data processing device and smart glasses according to the second embodiment. [Figure 5] This is a conceptual diagram showing an example of the configuration of a data processing system according to the third embodiment. [Figure 6] This is a conceptual diagram showing an example of the main functions of a data processing device and a headset-type terminal according to the third embodiment. [Figure 7] This is a conceptual diagram showing an example of the configuration of a data processing system according to the fourth embodiment. [Figure 8] This is a conceptual diagram showing an example of the main functions of a data processing device and a robot according to the fourth embodiment. [Figure 9] This shows an emotion map where multiple emotions are mapped. [Figure 10] This shows an emotion map where multiple emotions are mapped. [Figure 11] This is a sequence diagram showing the processing flow of the data processing system in Example 1. [Figure 12] This is a sequence diagram showing the processing flow of the data processing system in Application Example 1. [Figure 13] This is a sequence diagram showing the processing flow of the data processing system in Example 2, which incorporates an emotion engine. [Figure 14] This is a sequence diagram showing the processing flow of the data processing system in Application Example 2, which combines an emotion engine. [Modes for carrying out the invention]

[0014] An example of an embodiment of the system according to the technology of the present disclosure will be described below with reference to the accompanying drawings.

[0015] First, the terms used in the following description will be explained.

[0016] In the following embodiments, the numbered processor (hereinafter simply referred to as "processor") may be a single arithmetic unit or a combination of multiple arithmetic units. Also, the processor may be a single type of arithmetic unit or a combination of multiple types of arithmetic units. Examples of arithmetic units include a CPU (Central Processing Unit), a GPU (Graphics Processing Unit), a GPGPU (General-Purpose computing on Graphics Processing Units), an APU (Accelerated Processing Unit), and the like.

[0017] In the following embodiments, the numbered RAM (Random Access Memory) is a memory in which information is temporarily stored and is used as a work memory by the processor.

[0018] In the following embodiments, the numbered storage is one or more non-volatile storage devices that store various programs, various parameters, and the like. Examples of non-volatile storage devices include flash memory (SSD (Solid State Drive)), magnetic disks (e.g., hard disks), or magnetic tapes, and the like.

[0019] In the following embodiments, the signed communication interface (I / F) is an interface that includes a communication processor and an antenna, etc. The communication interface manages communication between multiple computers. Examples of communication standards applicable to the communication interface include wireless communication standards such as 5G (5th Generation Mobile Communication System), Wi-Fi (registered trademark), or Bluetooth (registered trademark).

[0020] In the following embodiments, "A and / or B" is synonymous with "at least one of A and B." That is, "A and / or B" means that it may be A alone, or B alone, or a combination of A and B. Furthermore, in this specification, the same concept as "A and / or B" applies when expressing three or more things linked by "and / or."

[0021] [First Embodiment]

[0022] Figure 1 shows an example of the configuration of the data processing system 10 according to the first embodiment.

[0023] As shown in Figure 1, the data processing system 10 includes a data processing device 12 and a smart device 14. An example of the data processing device 12 is a server.

[0024] The data processing device 12 comprises a computer 22, a database 24, and a communication interface 26. The computer 22 is an example of a "computer" related to the technology of this disclosure. The computer 22 comprises a processor 28, RAM 30, and storage 32. The processor 28, RAM 30, and storage 32 are connected to a bus 34. The database 24 and the communication interface 26 are also connected to the bus 34. The communication interface 26 is connected to a network 54. An example of the network 54 is a WAN (Wide Area Network) and / or a LAN (Local Area Network).

[0025] The smart device 14 comprises a computer 36, a reception device 38, an output device 40, a camera 42, and a communication interface 44. The computer 36 comprises a processor 46, RAM 48, and storage 50. The processor 46, RAM 48, and storage 50 are connected to a bus 52. The reception device 38, output device 40, and camera 42 are also connected to the bus 52.

[0026] The reception device 38 is equipped with a touch panel 38A and a microphone 38B, etc., and receives user input. The touch panel 38A receives user input by detecting contact with an object (e.g., a pen or finger). The microphone 38B receives user input by detecting the user's voice. The control unit 46A transmits data indicating the user input received by the touch panel 38A and microphone 38B to the data processing device 12. In the data processing device 12, the specific processing unit 290 acquires the data indicating the user input.

[0027] The output device 40 includes a display 40A and a speaker 40B, and presents data to the user 20 by outputting the data in a form perceptible to the user 20 (e.g., audio and / or text). The display 40A displays visible information such as text and images according to instructions from the processor 46. The speaker 40B outputs audio according to instructions from the processor 46. The camera 42 is a small digital camera equipped with an optical system such as a lens, aperture, and shutter, and an image sensor such as a CMOS (Complementary Metal-Oxide-Semiconductor) image sensor or a CCD (Charge Coupled Device) image sensor.

[0028] Communication interface 44 is connected to network 54. Communication interfaces 44 and 26 are responsible for the exchange of various types of information between processor 46 and processor 28 via network 54.

[0029] Figure 2 shows an example of the main functions of the data processing device 12 and the smart device 14.

[0030] As shown in Figure 2, in the data processing device 12, a specific processing is performed by the processor 28. A specific processing program 56 is stored in the storage 32. The specific processing program 56 is an example of a "program" related to the technology of this disclosure. The processor 28 reads the specific processing program 56 from the storage 32 and executes the read specific processing program 56 on the RAM 30. The specific processing is realized by the processor 28 operating as a specific processing unit 290 according to the specific processing program 56 executed on the RAM 30.

[0031] The storage 32 stores the data generation model 58 and the emotion identification model 59. The data generation model 58 and the emotion identification model 59 are used by the identification processing unit 290.

[0032] In the smart device 14, the processor 46 performs the reception output processing. The storage 50 stores the reception output program 60. The reception output program 60 is used in conjunction with a specific processing program 56 by the data processing system 10. The processor 46 reads the reception output program 60 from the storage 50 and executes the read reception output program 60 on the RAM 48. The reception output processing is realized by the processor 46 operating as a control unit 46A according to the reception output program 60 executed on the RAM 48.

[0033] Next, the specific processing performed by the specific processing unit 290 of the data processing device 12 will be described. In the following description, the data processing device 12 will be referred to as the "server" and the smart device 14 as the "terminal".

[0034] The AR / VR teaching system of the present invention utilizes real-time motion capture and feedback functions to support users in acquiring skills. The specific implementation of this system is described below.

[0035] First, the device uses a motion capture system to capture the user's movements. The motion capture system acquires the positional information and movement of each segment of the user's body in real time. This makes it possible to collect accurate data at the moment the user is performing actions such as playing music or participating in sports. The collected data is transmitted to the device through sensors and camera devices worn by the user.

[0036] Next, the terminal sends the captured user motion data to the server. The server runs a dedicated analysis algorithm to analyze the received data and compare it with the training motion data. The algorithm used here is designed to analyze elements such as the timing, speed, and angle of each motion in detail. The server identifies which parts of the user's motion can be improved and generates the necessary feedback information.

[0037] After the analysis is complete, the server generates feedback on the user's actions and sends it to the terminal. The terminal then presents this feedback information to the user through an AR / VR device. Specifically, the user's actions, overlaid on the teacher's ideal actions, are displayed in real time on the head-mounted display or AR glasses worn by the user. This allows the user to intuitively understand which parts of the actions are different visually.

[0038] Another distinctive feature of the system is that the device can provide feedback information to the user in the form of audio and video. For example, if a user is learning to play the violin, the device can overlay the teacher's hand movements, showing in detail what each finger is doing. With the addition of audio assistance, users can receive feedback in a format that suits their individual learning style.

[0039] This entire process aims to dramatically improve user learning efficiency and provide a personalized, interactive learning experience. In the future, this technology is expected to be applied beyond education to support motor-based learning in various fields.

[0040] The following describes the processing flow.

[0041] Step 1:

[0042] The terminal activates the motion capture device and prepares to collect user movement data in real time. The motion capture device begins acquiring location information for each segment.

[0043] Step 2:

[0044] The terminal packets the user's activity data it has acquired and sends that data to the server using a low-latency communication protocol. The data is streamed continuously and in real time.

[0045] Step 3:

[0046] The server receives the user's behavior data and performs a comparative analysis with the teacher's behavior data that has already been registered. Using artificial intelligence, it meticulously analyzes the differences between the user's and the teacher's behavior to identify specific areas for improvement and discrepancies.

[0047] Step 4:

[0048] Based on the improvements identified in the analysis and the ideal operating model, the server generates visual data that overlays the user's movements onto the virtual space. This visual data includes guidelines on how the user should adjust their movements.

[0049] Step 5:

[0050] The server sends the generated visual data to the terminal, which then displays it on the AR / VR device. Users can adjust their actions while visually confirming their own actions and the teacher's actions.

[0051] Step 6:

[0052] In addition to visual data, the device provides users with audio and video feedback through a feedback system. This allows users to understand areas for improvement from multiple perspectives and learn more efficiently.

[0053] Step 7:

[0054] The user modifies their movements based on the feedback provided and records their movements again using the motion capture device. This creates a cycle of continuous improvement through feedback.

[0055] (Example 1)

[0056] Next, we will describe Example 1. In the following description, the data processing device 12 will be referred to as the "server," and the smart device 14 will be referred to as the "terminal."

[0057] Conventional motion instruction systems have a problem where users cannot intuitively understand the difference between their own movements and ideal movements, leading to decreased instruction efficiency. Furthermore, the technical hurdles to ensuring real-time analysis of motion data are high, resulting in delays in feedback, which is a challenge.

[0058] The identification process performed by the identification processing unit 290 of the data processing device 12 in Example 1 is realized by the following means.

[0059] In this invention, the server includes an information processing device that compares and analyzes the user's movement information with that of the instructor, a presentation device that overlays the user's movements onto an abstract space based on the analysis results, and a reporting device that provides the user with areas for improvement in their movements. This allows the user to intuitively and in real time understand areas for improvement in their own movements, enabling effective learning.

[0060] A "motion detection device" is a device that measures the user's body movements and posture in real time and collects that information.

[0061] An "information processing device" is a device used to analyze collected behavioral information, and in particular, it has the function of comparing the actions of the user and the instructor.

[0062] A "presentation device" is a device that visually displays operational information to the user based on analysis results, and is a device that enables the superimposition and display of operations in an abstract space.

[0063] A "reporting device" is a device that presents the user with instructions for improvements and operational modifications based on the analysis results.

[0064] An "intelligent algorithm" is an automated analysis method that analyzes motion information and effectively compares the actions of the user and the instructor.

[0065] "Communication means" refers to technical means for sending and receiving information between a terminal and an information management device with low latency and efficiency.

[0066] This invention is an AR / VR teaching system for assisting users in acquiring skills. The system includes a motion detection device, an information processing device, a presentation device, and a reporting device, and provides real-time analysis and feedback of motion.

[0067] The device detects the user's movements using a motion capture device attached to the user's body. This device digitizes the movement and position of each part of the user's body in real time, providing precise motion information. The motion capture device includes multiple sensors and high-resolution cameras.

[0068] The server receives motion information transmitted from the terminal and processes it using intelligent algorithms. Software on the server compares the user's motion data with the instructor's ideal motion data and performs analysis. This analysis includes elements such as motion speed, angle, and timing. The analysis results are then processed into feedback for presentation to the user.

[0069] The terminal visually presents the feedback received from the server to the user using an AR / VR device. The display device overlays the analyzed movements with the ideal movements through the head-mounted display or AR glasses worn by the user. This allows the user to intuitively understand their own movements and make necessary adjustments.

[0070] Users can improve their performance based on feedback from the device. The reporting device provides users with detailed instructions, via audio and video, on which parts of their performance should be improved and how. For example, if a user is learning to play the violin, improvements would be suggested via audio guidance, along with a digital overlay demonstrating the correct finger movements.

[0071] An example of a prompt for a generative AI model is, "Design a system to help a user learn the correct way to play the violin." Based on this prompt, the model will compare the actions of the user and the instructor in real time, select how to generate effective feedback, and provide specific suggestions.

[0072] The flow of the specific processing in Example 1 will be explained using Figure 11.

[0073] Step 1:

[0074] The user prepares to begin the process. They attach the motion capture device to their body and verify its connection to the system. The device prepares to detect the user's body movements in real time. The user's action (e.g., playing the violin) is set as input.

[0075] Step 2:

[0076] The terminal acquires user movement information from the motion capture device. The device records the user's movements and position, converts this information into digital data format, and transmits it to the terminal. As output, the user's movement data is acquired and temporarily stored on the terminal.

[0077] Step 3:

[0078] The terminal sends acquired motion data to the server. The data is transmitted over the network, allowing the server to perform real-time analysis. The input is the user's real-time motion data, and this data is sent to the server as output.

[0079] Step 4:

[0080] The server analyzes the received user motion data using intelligent algorithms. It compares the user's motion information with the instructor's ideal motion and performs data calculations to evaluate the accuracy, speed, angle, etc. of the motion. The input to the analysis is the user's motion information, and the output is the analysis result.

[0081] Step 5:

[0082] The server generates feedback information based on the analysis results. It prepares suggestions for improving the user's actions in text, audio, and video formats. This process clarifies to the user which actions they should correct and how. Detailed feedback information is generated as output.

[0083] Step 6:

[0084] The server sends the generated feedback information to the terminal. The feedback is formatted so that it is presented appropriately on the AR / VR device used by the user. The input is the feedback information, and the output is the transmission of data to the terminal.

[0085] Step 7:

[0086] The terminal displays the received feedback on the user's AR / VR device. The display device allows the user to overlay and re-examine the ideal action against their own action. This enables the user to receive direct visual feedback. Visual feedback is provided to the user as output.

[0087] Step 8:

[0088] The user modifies their movements based on the feedback provided. Specifically, they adjust their body movements by referring to the ideal movements displayed on the device. The improved movements are recorded again by the motion capture device, and the process is repeated. In this step, the input is the feedback information, and the output is the user's improved movements.

[0089] (Application Example 1)

[0090] Next, we will explain Application Example 1. In the following explanation, the data processing device 12 will be referred to as the "server," and the smart device 14 will be referred to as the "terminal."

[0091] In manufacturing environments, there is a problem with inefficient training methods when workers learn complex robot operation procedures. Furthermore, a lack of flexibility to accommodate individual learning styles and the absence of real-time feedback can delay the learning process. This can potentially impact the accuracy and productivity of operations.

[0092] The specific processing performed by the specific processing unit 290 of the data processing device 12 in Application Example 1 is realized by the following means.

[0093] In this invention, the server includes a motion measurement device for measuring movements in real time, an information processing device for comparing and analyzing the user's motion information and the instructor's motion information, and a display means for overlaying the user's movements onto a virtual environment based on the analysis results. This allows the worker to receive immediate visual and auditory feedback, enabling efficient acquisition of operating procedures based on their individual learning style.

[0094] A "motion measurement device" is a device that measures a user's physical movements in real time and acquires motion information.

[0095] An "information processing device" is a device used to compare and analyze acquired motion information with the motion information of an instructor.

[0096] "Display means" refers to a means for visually displaying the user's actions overlaid on a virtual environment based on the analysis results.

[0097] A "notification method" is a means of notifying the user of areas where functionality has been improved as feedback.

[0098] An "interface means" is a means of providing instructions to a worker visually and aurally, and supporting efficient learning.

[0099] "Machine learning technology" is an artificial intelligence technology used to analyze motion information and identify differences between the user's actions and the instructor's actions.

[0100] "Exchange means" refers to a means for exchanging information between a terminal and an information processing device with low latency and efficiency.

[0101] The server manages a device that uses motion measurement equipment to measure the user's physical movements in real time and acquire that motion information. This device is equipped with multiple sensors and has the function of accurately capturing data on the user's limbs and posture. The acquired motion data is immediately transmitted to the server via communication means.

[0102] The server analyzes this operational information through an information processing device and compares it with standard operational data from pre-registered instructors. The analysis algorithm uses machine learning techniques to identify differences in operation and determine which areas need improvement. Based on the results, the server generates feedback data specifically indicating the areas for improvement.

[0103] The terminal receives feedback data sent from the server and notifies the user through a display device. A head-mounted display or smart glasses are used as the display device, allowing the user to visually understand areas for improvement by overlaying their own actions with ideal actions within a virtual environment. Furthermore, the terminal can enhance learning efficiency by utilizing interface devices to present a combination of visual and auditory feedback.

[0104] A typical prompt for the entire system would be: "We have installed a new robotic arm in the factory. Please tell us how to build an AR / VR system that provides real-time feedback on the correct operating procedures." This allows for a specific inquiry about improving motion management.

[0105] This system enables workers in the factory work environment to efficiently acquire advanced skills and improve the precision and safety of their operations.

[0106] The flow of a specific process in Application Example 1 will be explained using Figure 12.

[0107] Step 1:

[0108] A motion measurement device captures the user's physical movements as they operate the robot using sensors. The acquired motion information is transmitted to a terminal as detailed data such as the user's position and speed. This allows for real-time input of motion data.

[0109] Step 2:

[0110] The terminal transmits the received operational data to the server via a communication means. The server converts this data into the format required for analysis and generates a dataset for analysis. The operational data obtained as input is output to the server's analysis system.

[0111] Step 3:

[0112] The server uses machine learning algorithms to compare user movement data with instructor movement data. Data processing analyzes movement timing, speed, angle, etc., identifying areas for improvement. Analysis results are generated and output as specific areas for improvement.

[0113] Step 4:

[0114] The server generates feedback data on user actions based on the analysis results. This data includes visual improvement instructions and procedures. The generated feedback data is output to a display device.

[0115] Step 5:

[0116] The terminal receives feedback data sent from the server and presents it to the user through a display device. Using a head-mounted display or smart glasses, the user is shown a virtual overlay of the ideal behavior. This allows the user to intuitively understand which parts need improvement.

[0117] Step 6:

[0118] Based on the feedback provided, the user modifies their actions and operates the robot again. The interface provides additional visual and auditory guidance, further enhancing the user's learning. This allows the user to input new actions in response to the feedback.

[0119] Furthermore, an emotion engine that estimates the user's emotions may be incorporated. That is, the identification processing unit 290 may use the emotion identification model 59 to estimate the user's emotions and perform identification processing using the user's emotions.

[0120] The present invention is a system that enables interactive feedback that takes into account the learner's emotional state by combining an emotion engine with an AR / VR teaching system for supporting motion learning. This system includes a motion capture device, a processing device, a display device, a feedback device, a communication means, and an emotion engine.

[0121] The device uses a motion capture system to record the user's body movements in detail in order to collect user motion data in real time. In addition, the emotion engine acquires the user's facial expression data with a camera and analyzes this data to recognize the user's emotions. The emotion engine uses a facial expression analysis algorithm to identify multiple emotions in real time, such as joy, sadness, surprise, anger, fear, disgust, and neutral expression.

[0122] Next, the device sends the collected motion data and emotion data to the server. The server uses the received motion data to perform a comparative analysis with the teacher's actions. Artificial intelligence is used in the analysis to perform a precise comparison of actions and identify specific areas for improvement. Emotion data is also analyzed simultaneously, and feedback is adjusted according to the user's emotions.

[0123] The server then generates feedback on the user's actions based on the analysis results and sends it to the terminal. For visual feedback, an AR / VR device overlays the teacher's ideal actions with the user's actions in the user's field of view. Additionally, voice feedback optimized for the user's emotions is provided. For example, if motivation is low, positive voice assistance is enhanced to increase confidence in improving actions.

[0124] For example, if a user is learning dance moves and the emotion engine detects a decrease in the user's concentration, the device can provide visual feedback along with voice messages of encouragement and relaxation, urging the user to take their time.

[0125] In this way, by considering the user's actions and emotions as an integrated whole, this system provides a more personalized learning experience and greatly supports the user's skill acquisition process.

[0126] The following describes the processing flow.

[0127] Step 1:

[0128] The device acquires user movement data in real time using a motion capture system. This includes collecting positional information of joints and skeleton using sensors or cameras worn by the user.

[0129] Step 2:

[0130] The device acquires user facial expression data through its camera and analyzes it using an emotion engine. It also determines the user's emotional state in real time.

[0131] Step 3:

[0132] The device sends user behavioral data and emotional data to the server. This transmission is performed with low latency, ensuring that data is transferred without compromising real-time capabilities.

[0133] Step 4:

[0134] The server uses the received behavioral data to begin an analysis that compares it to the teacher's behavioral data. This analysis uses artificial intelligence to identify differences and areas for improvement in the behavior in detail.

[0135] Step 5:

[0136] In parallel, the server analyzes data from the emotion engine and determines the tone and content of feedback based on the user's emotion index. If the user's emotions are unstable, it uses a reinforcement learning algorithm to select encouraging or adapted guidance.

[0137] Step 6:

[0138] The server generates feedback information based on the analysis results and sends it to the terminal. The feedback is provided visually and audibly and is customized to enhance user motivation.

[0139] Step 7:

[0140] The device provides feedback to the user via an AR / VR device. The user can view a video in which the teacher's movement model and their own movements are overlaid, and make corrections as needed. Voice assistance is also provided in parallel, allowing the user to receive guidance on specific ways to improve.

[0141] (Example 2)

[0142] Next, we will describe Example 2. In the following description, the data processing device 12 will be referred to as the "server" and the smart device 14 as the "terminal".

[0143] When learners acquire a skill, simply comparing the accuracy of the skill makes it difficult to provide appropriate instruction tailored to each individual's learning progress and emotional state. Furthermore, standardized feedback can make it difficult for learners to maintain motivation.

[0144] The identification process performed by the identification processing unit 290 of the data processing device 12 in Example 2 is realized by the following means.

[0145] In this invention, the server includes an action acquisition device, a processing device, and an emotion recognition device. This allows for more effective action learning by not only comparing the user's action information with standard actions, but also providing feedback that takes into account the user's emotional state.

[0146] A "motion acquisition device" is a device that collects the user's body movements in real time and plays a role in acquiring specific motion information of the user.

[0147] A "calculation unit" is a device that compares and analyzes user operation information with standard operation information, performing calculations to identify deviations and areas for improvement in operation.

[0148] A "display means" is a means of providing information visually and displaying the user's actions overlaid on a virtual environment based on the analysis results.

[0149] An "information provision device" is a device that provides users with feedback on areas for improvement in their operation, issuing specific instructions and information to support their learning.

[0150] An "emotion recognition device" is a device that identifies a user's emotional state from their facial expressions, and it acquires emotional information using an emotion analysis algorithm.

[0151] A "feedback adjustment device" is a device that optimizes the content and format of the feedback provided based on the user's emotional state.

[0152] A "communication device" is a device that enables low-latency information communication between a terminal and an information processing device, allowing for high-speed and stable data exchange.

[0153] This system is designed to effectively support the user's motion learning and is configured as follows: First, the terminal collects the user's motion data in real time using a motion acquisition device. This motion acquisition device includes cameras and sensors to accurately capture the user's body movements. Furthermore, the terminal utilizes an emotion recognition device to acquire the user's facial expression data via the camera and analyzes their emotional state from it.

[0154] The collected data is transmitted to the server via a low-latency communication device. The server uses a computing unit to perform calculations to compare the user's behavioral data with standard behavioral data. Machine learning techniques are applied to these calculations to identify deviations and areas for improvement. Simultaneously, the emotional data obtained by the emotion recognition device is analyzed by a feedback adjustment device to generate feedback tailored to the user's emotional state.

[0155] The generated feedback is transmitted to the terminal via an information provider and provided to the user. Visual feedback utilizes display means, overlaying the user's actions with ideal actions in a virtual environment. Emotion-based voice feedback is also provided, playing a role in increasing user motivation.

[0156] For example, if a user practicing dance wants to improve the timing of their movements, the system can collect motion data for each step, analyze the differences, and provide feedback. At the same time, if the user's concentration is waning, it can provide a voice message encouraging them to relax.

[0157] Examples of prompts for a generative AI model:

[0158] "Please describe a system that generates appropriate feedback when a user is losing focus while practicing dance moves."

[0159] In this way, this system not only promotes the improvement of users' operational skills but also provides emotional support, resulting in a more fulfilling learning experience.

[0160] The flow of the specific processing in Example 2 will be explained using Figure 13.

[0161] Step 1:

[0162] The terminal uses a motion acquisition device to collect user motion data in real time. The input is the user's physical movements, which are captured by cameras and sensors and converted into numerical data such as the position and angle of the movements. The output is detailed motion data for analysis. During this process, the distance the user moves and the angles of their joints are specifically recorded.

[0163] Step 2:

[0164] The device uses an emotion recognition device to acquire and analyze the user's facial expression data. The input is image data of the user's face, which is applied to an analysis algorithm to extract emotional states such as joy and sadness. The output is an emotion label and data indicating the intensity of that emotion. Specifically, the degree of the user's smile and eyebrow movements are analyzed.

[0165] Step 3:

[0166] The terminal sends the acquired behavioral and emotional data to the server. The input consists of the previously collected behavioral and emotional data, and the output is the generation of data packets to be sent to the server. In this step, the communication device plays the role of transmitting the data with low latency.

[0167] Step 4:

[0168] The server uses a computing device to compare and analyze operational data with standard operational data. The input consists of user operational data and pre-prepared standard operational data. Machine learning is used to calculate the differences in operation and identify areas for improvement. The output is an analysis showing the deviations in operation. Specifically, timing delays and movement precision are quantified.

[0169] Step 5:

[0170] The server generates feedback using a feedback adjustment device based on emotional data. The input is an emotional label and its intensity, which is used to construct a feedback message optimized for the user's emotional state. The output is the adjusted feedback message, which includes positive content and, where necessary, words of encouragement.

[0171] Step 6:

[0172] The server sends the generated feedback to the terminal. The terminal provides this feedback to the user using visual and auditory means. The input is feedback data from the server, and the output is a visual overlay display and audio instructions. Specifically, this involves visualization of actions using AR / VR devices and audio guidance using speakers.

[0173] Step 7:

[0174] The user receives feedback and attempts to improve the operation. The input is feedback from the device, and the output is the improved operation. In this process, the user can learn by following specific instructions and correcting the operation.

[0175] (Application Example 2)

[0176] Next, we will explain application example 2. In the following explanation, the data processing device 12 will be referred to as the "server," and the smart device 14 will be referred to as the "terminal."

[0177] In the industrial sector, efficient and safe operation of workers and machinery requires both advanced skills acquisition and emotional support. However, conventional technologies struggle to provide personalized feedback in real time that considers the emotional state of individual workers, in addition to evaluating their movements. Therefore, there is a need for more effective and motivating methods.

[0178] The specific processing performed by the specific processing unit 290 of the data processing device 12 in Application Example 2 is realized by the following means.

[0179] In this invention, the server includes processing means for comparing and examining user behavior information and instructor behavior information, emotion analysis means for analyzing user emotional information, and adjustment means for optimizing output means considering the results of the emotional information analysis. This enables feedback that combines highly accurate evaluation of behavior with emotional support.

[0180] A "detection device" is a device used to record operations and changes in real time.

[0181] The "processing means" refers to a function that analyzes acquired motion information and compares it with a standard motion.

[0182] A "display means" is a device that visually superimposes actions in a virtual space based on the results of the study.

[0183] An "output means" is a device that provides the user with suggestions for improving the operation and feedback.

[0184] An "emotional analysis tool" is a function that analyzes the user's emotions and identifies their state.

[0185] A "modification mechanism" is a system for optimizing the content of feedback based on analyzed emotional information.

[0186] A "communication method" is a system for sending and receiving information with low latency between a terminal device and an information processing device.

[0187] The system in this invention primarily analyzes the actions and emotions of workers in real time and provides personalized feedback. In this system, the server, terminal, and user function as follows:

[0188] First, the terminal acquires the worker's movements in real time using a motion capture device. Examples of motion capture devices include Arduino and Kinect. The terminal also acquires facial expression data via a camera, collecting basic data for analyzing the user's emotions.

[0189] Next, the terminal sends the acquired motion data and emotion data to the server. On the server side, the motion data is analyzed using processing equipment. Specifically, a machine learning algorithm is used to compare the user's motion information with the instructor's motion information to identify accuracy and areas for improvement. In addition, emotion analysis equipment is used to analyze facial expression data and evaluate the user's emotional state.

[0190] Based on these analysis results, the server generates optimized feedback for the output method. This feedback is provided visually on the AR / VR device through the display method. Furthermore, the adjustment method generates audio feedback tailored to the user's emotional state, effectively supporting their motivation.

[0191] To give a concrete example, imagine a scenario in a training session for assembling a new automobile model, where workers are learning precise movements. In this case, it would be possible to capture the workers' movements in detail, evaluate the accuracy of those movements, and then provide visual and audio guidance for the actual movements.

[0192] When using a generative AI model, examples of prompt statements include the following:

[0193] "Propose a program for training new car assembly that provides optimal movement guidance and generates feedback tailored to the worker's emotions. Consider worker movement and emotional data to specifically identify feedback that enhances motivation."

[0194] The flow of a specific process in Application Example 2 will be explained using Figure 14.

[0195] Step 1:

[0196] The device uses a motion capture system to acquire user movement data in real time. The input is raw data from the motion capture sensor, and the output is a stream of movement data. The movement involves recording the user's body movements in detail and converting them into a format that can be transferred as data.

[0197] Step 2:

[0198] The device uses a camera to acquire facial expression data and performs analysis using an emotion analysis tool. The input is image data acquired from the camera, and the output is the analyzed emotional state. Specifically, the device passes the image data through a facial expression analysis algorithm to identify emotions such as joy, anger, sadness, and happiness in real time.

[0199] Step 3:

[0200] The terminal transmits acquired behavioral and emotional data to the server. The input is a data stream related to behavior and emotion, and the output is a successful transmission of data to the server. The terminal efficiently compresses the data and transfers it to the server using a low-latency communication protocol.

[0201] Step 4:

[0202] The server uses processing tools to compare and analyze the user's motion data with the instructor's motion data. The input consists of the user's motion data and existing baseline data, and the output is the extraction of motion similarities and areas for improvement. The server applies machine learning algorithms to perform motion evaluation.

[0203] Step 5:

[0204] The server evaluates facial expression data using emotion analysis tools and determines the emotional state. The input is processed facial expression data, and the output is the user's emotional state and its intensity. The server uses an emotion recognition algorithm to identify emotional patterns.

[0205] Step 6:

[0206] The server optimizes output methods based on analysis results to generate feedback. Inputs are the results of behavioral comparisons and sentiment evaluations, while output is detailed guidance in the form of visual and auditory feedback. The server customizes the content to ensure the feedback is delivered to the user most effectively.

[0207] Step 7:

[0208] The terminal provides feedback received from the server to the user using display means. Input is feedback data, and output is visual guidance via AR / VR devices and audio instructions via speakers. The terminal overlays the instructor's actions with the user's actions to highlight areas for improvement.

[0209] The specific processing unit 290 transmits the result of the specific processing to the smart device 14. In the smart device 14, the control unit 46A causes the output device 40 to output the result of the specific processing. The microphone 38B acquires audio indicating user input for the result of the specific processing. The control unit 46A transmits the audio data indicating user input acquired by the microphone 38B to the data processing device 12. In the data processing device 12, the specific processing unit 290 acquires the audio data.

[0210] Data generation model 58 is a so-called generative AI (Artificial Intelligence). An example of data generation model 58 is ChatGPT (registered trademark) (Internet search).<URL: https: / / openai.com / blog / chatgpt> ), Gemini (registered trademark) (Internet search) <url: https: gemini.google.com ?hl="ja">Examples of generative AI include the following. The data generation model 58 is obtained by performing deep learning on a neural network. The data generation model 58 is input with prompts containing instructions, and with inference data such as audio data representing speech, text data representing text, and image data representing images. The data generation model 58 infers from the input inference data according to the instructions indicated by the prompts, and outputs the inference results in data formats such as audio data and text data. Here, inference refers to, for example, analysis, classification, prediction, and / or summarization.

[0211] In the above embodiment, an example was given in which specific processing is performed by the data processing device 12, but the technology of this disclosure is not limited thereto, and the specific processing may also be performed by the smart device 14.

[0212] [Second Embodiment]

[0213] Figure 3 shows an example of the configuration of the data processing system 210 according to the second embodiment.

[0214] As shown in Figure 3, the data processing system 210 includes a data processing device 12 and smart glasses 214. An example of the data processing device 12 is a server.

[0215] The data processing device 12 comprises a computer 22, a database 24, and a communication interface 26. The computer 22 is an example of a "computer" related to the technology of this disclosure. The computer 22 comprises a processor 28, RAM 30, and storage 32. The processor 28, RAM 30, and storage 32 are connected to a bus 34. The database 24 and the communication interface 26 are also connected to the bus 34. The communication interface 26 is connected to a network 54. An example of the network 54 is a WAN (Wide Area Network) and / or a LAN (Local Area Network).

[0216] The smart glasses 214 include a computer 36, a microphone 238, a speaker 240, a camera 42, and a communication interface 44. The computer 36 includes a processor 46, RAM 48, and storage 50. The processor 46, RAM 48, and storage 50 are connected to a bus 52. The microphone 238, speaker 240, and camera 42 are also connected to the bus 52.

[0217] The microphone 238 receives voice signals from the user 20 and receives instructions from the user 20. The microphone 238 captures the voice signals from the user 20, converts the captured voice into audio data, and outputs it to the processor 46. The speaker 240 outputs audio according to the instructions from the processor 46.

[0218] Camera 42 is a small digital camera equipped with an optical system including a lens, aperture, and shutter, and an image sensor such as a CMOS (Complementary Metal-Oxide-Semiconductor) image sensor or a CCD (Charge Coupled Device) image sensor, and captures images of the area around the user 20 (for example, an imaging range defined by a field of view equivalent to the width of a typical healthy person's field of vision).

[0219] Communication interface 44 is connected to network 54. Communication interfaces 44 and 26 are responsible for the exchange of various information between processor 46 and processor 28 via network 54. The exchange of various information between processor 46 and processor 28 using communication interfaces 44 and 26 is performed in a secure manner.

[0220] Figure 4 shows an example of the main functions of the data processing device 12 and the smart glasses 214. As shown in Figure 4, the data processing device 12 performs specific processing using the processor 28. The storage 32 stores the specific processing program 56.

[0221] The specific processing program 56 is an example of a "program" relating to the technology of this disclosure. The processor 28 reads the specific processing program 56 from the storage 32 and executes the read specific processing program 56 on the RAM 30. The specific processing is realized by the processor 28 operating as a specific processing unit 290 in accordance with the specific processing program 56 executed on the RAM 30.

[0222] The storage 32 stores the data generation model 58 and the emotion identification model 59. The data generation model 58 and the emotion identification model 59 are used by the identification processing unit 290.

[0223] In the smart glasses 214, the processor 46 performs the reception output processing. The storage 50 stores the reception output program 60. The processor 46 reads the reception output program 60 from the storage 50 and executes the read reception output program 60 on the RAM 48. The reception output processing is realized by the processor 46 operating as a control unit 46A according to the reception output program 60 executed on the RAM 48.

[0224] Next, the identification processing performed by the identification processing unit 290 of the data processing device 12 will be described. In the following description, the data processing device 12 will be referred to as the "server" and the smart glasses 214 will be referred to as the "terminal".

[0225] The AR / VR teaching system of the present invention utilizes real-time motion capture and feedback functions to support users in acquiring skills. The specific implementation of this system is described below.

[0226] First, the device uses a motion capture system to capture the user's movements. The motion capture system acquires the positional information and movement of each segment of the user's body in real time. This makes it possible to collect accurate data at the moment the user is performing actions such as playing music or participating in sports. The collected data is transmitted to the device through sensors and camera devices worn by the user.

[0227] Next, the terminal sends the captured user motion data to the server. The server runs a dedicated analysis algorithm to analyze the received data and compare it with the training motion data. The algorithm used here is designed to analyze elements such as the timing, speed, and angle of each motion in detail. The server identifies which parts of the user's motion can be improved and generates the necessary feedback information.

[0228] After the analysis is complete, the server generates feedback on the user's actions and sends it to the terminal. The terminal then presents this feedback information to the user through an AR / VR device. Specifically, the user's actions, overlaid on the teacher's ideal actions, are displayed in real time on the head-mounted display or AR glasses worn by the user. This allows the user to intuitively understand which parts of the actions are different visually.

[0229] Another distinctive feature of the system is that the device can provide feedback information to the user in the form of audio and video. For example, if a user is learning to play the violin, the device can overlay the teacher's hand movements, showing in detail what each finger is doing. With the addition of audio assistance, users can receive feedback in a format that suits their individual learning style.

[0230] This entire process aims to dramatically improve user learning efficiency and provide a personalized, interactive learning experience. In the future, this technology is expected to be applied beyond education to support motor-based learning in various fields.

[0231] The following describes the processing flow.

[0232] Step 1:

[0233] The terminal activates the motion capture device and prepares to collect user movement data in real time. The motion capture device begins acquiring location information for each segment.

[0234] Step 2:

[0235] The terminal packets the user's activity data it has acquired and sends that data to the server using a low-latency communication protocol. The data is streamed continuously and in real time.

[0236] Step 3:

[0237] The server receives the user's behavior data and performs a comparative analysis with the teacher's behavior data that has already been registered. Using artificial intelligence, it meticulously analyzes the differences between the user's and the teacher's behavior to identify specific areas for improvement and discrepancies.

[0238] Step 4:

[0239] Based on the improvements identified in the analysis and the ideal operating model, the server generates visual data that overlays the user's movements onto the virtual space. This visual data includes guidelines on how the user should adjust their movements.

[0240] Step 5:

[0241] The server sends the generated visual data to the terminal, which then displays it on the AR / VR device. Users can adjust their actions while visually confirming their own actions and the teacher's actions.

[0242] Step 6:

[0243] In addition to visual data, the device provides users with audio and video feedback through a feedback system. This allows users to understand areas for improvement from multiple perspectives and learn more efficiently.

[0244] Step 7:

[0245] The user modifies their movements based on the feedback provided and records their movements again using the motion capture device. This creates a cycle of continuous improvement through feedback.

[0246] (Example 1)

[0247] Next, we will describe Example 1. In the following description, the data processing device 12 will be referred to as the "server," and the smart glasses 214 will be referred to as the "terminal."

[0248] Conventional motion instruction systems have a problem where users cannot intuitively understand the difference between their own movements and ideal movements, leading to decreased instruction efficiency. Furthermore, the technical hurdles to ensuring real-time analysis of motion data are high, resulting in delays in feedback, which is a challenge.

[0249] The identification process performed by the identification processing unit 290 of the data processing device 12 in Example 1 is realized by the following means.

[0250] In this invention, the server includes an information processing device that compares and analyzes the user's movement information with that of the instructor, a presentation device that overlays the user's movements onto an abstract space based on the analysis results, and a reporting device that provides the user with areas for improvement in their movements. This allows the user to intuitively and in real time understand areas for improvement in their own movements, enabling effective learning.

[0251] A "motion detection device" is a device that measures the user's body movements and posture in real time and collects that information.

[0252] An "information processing device" is a device used to analyze collected behavioral information, and in particular, it has the function of comparing the actions of the user and the instructor.

[0253] A "presentation device" is a device that visually displays operational information to the user based on analysis results, and is a device that enables the superimposition and display of operations in an abstract space.

[0254] A "reporting device" is a device that presents the user with instructions for improvements and operational modifications based on the analysis results.

[0255] An "intelligent algorithm" is an automated analysis method that analyzes motion information and effectively compares the actions of the user and the instructor.

[0256] "Communication means" refers to technical means for sending and receiving information between a terminal and an information management device with low latency and efficiency.

[0257] This invention is an AR / VR teaching system for assisting users in acquiring skills. The system includes a motion detection device, an information processing device, a presentation device, and a reporting device, and provides real-time analysis and feedback of motion.

[0258] The device detects the user's movements using a motion capture device attached to the user's body. This device digitizes the movement and position of each part of the user's body in real time, providing precise motion information. The motion capture device includes multiple sensors and high-resolution cameras.

[0259] The server receives motion information transmitted from the terminal and processes it using intelligent algorithms. Software on the server compares the user's motion data with the instructor's ideal motion data and performs analysis. This analysis includes elements such as motion speed, angle, and timing. The analysis results are then processed into feedback for presentation to the user.

[0260] The terminal visually presents the feedback received from the server to the user using an AR / VR device. The display device overlays the analyzed movements with the ideal movements through the head-mounted display or AR glasses worn by the user. This allows the user to intuitively understand their own movements and make necessary adjustments.

[0261] Users can improve their performance based on feedback from the device. The reporting device provides users with detailed instructions, via audio and video, on which parts of their performance should be improved and how. For example, if a user is learning to play the violin, improvements would be suggested via audio guidance, along with a digital overlay demonstrating the correct finger movements.

[0262] An example of a prompt for a generative AI model is, "Design a system to help a user learn the correct way to play the violin." Based on this prompt, the model will compare the actions of the user and the instructor in real time, select how to generate effective feedback, and provide specific suggestions.

[0263] The flow of the specific processing in Example 1 will be explained using Figure 11.

[0264] Step 1:

[0265] The user prepares to begin the process. They attach the motion capture device to their body and verify its connection to the system. The device prepares to detect the user's body movements in real time. The user's action (e.g., playing the violin) is set as input.

[0266] Step 2:

[0267] The terminal acquires user movement information from the motion capture device. The device records the user's movements and position, converts this information into digital data format, and transmits it to the terminal. As output, the user's movement data is acquired and temporarily stored on the terminal.

[0268] Step 3:

[0269] The terminal sends acquired motion data to the server. The data is transmitted over the network, allowing the server to perform real-time analysis. The input is the user's real-time motion data, and this data is sent to the server as output.

[0270] Step 4:

[0271] The server analyzes the received user motion data using intelligent algorithms. It compares the user's motion information with the instructor's ideal motion and performs data calculations to evaluate the accuracy, speed, angle, etc. of the motion. The input to the analysis is the user's motion information, and the output is the analysis result.

[0272] Step 5:

[0273] The server generates feedback information based on the analysis results. It prepares suggestions for improving the user's actions in text, audio, and video formats. This process clarifies to the user which actions they should correct and how. Detailed feedback information is generated as output.

[0274] Step 6:

[0275] The server sends the generated feedback information to the terminal. The feedback is formatted so that it is presented appropriately on the AR / VR device used by the user. The input is the feedback information, and the output is the transmission of data to the terminal.

[0276] Step 7:

[0277] The terminal displays the received feedback on the user's AR / VR device. The display device allows the user to overlay and re-examine the ideal action against their own action. This enables the user to receive direct visual feedback. Visual feedback is provided to the user as output.

[0278] Step 8:

[0279] The user modifies their movements based on the feedback provided. Specifically, they adjust their body movements by referring to the ideal movements displayed on the device. The improved movements are recorded again by the motion capture device, and the process is repeated. In this step, the input is the feedback information, and the output is the user's improved movements.

[0280] (Application Example 1)

[0281] Next, we will explain Application Example 1. In the following explanation, the data processing device 12 will be referred to as the "server," and the smart glasses 214 will be referred to as the "terminal."

[0282] In manufacturing environments, there is a problem with inefficient training methods when workers learn complex robot operation procedures. Furthermore, a lack of flexibility to accommodate individual learning styles and the absence of real-time feedback can delay the learning process. This can potentially impact the accuracy and productivity of operations.

[0283] The specific processing by the specific processing unit 290 of the data processing device 12 in Application Example 1 is realized by the following means.

[0284] In this invention, the server includes an operation measurement device for measuring operations in real time, an information processing device for comparing and analyzing the operation information of the user and the operation information of the instructor, and display means for superimposing and displaying the user's operation on a virtual environment based on the analysis result. Thereby, the operator can receive immediate feedback visually and auditorily, and it becomes possible to acquire an efficient operation procedure based on an individual learning style.

[0285] The "operation measurement device" is a device for measuring the physical operations of the user in real time and acquiring operation information.

[0286] The "information processing device" is a device for comparing and analyzing the acquired operation information with the operation information of the instructor.

[0287] The "display means" is means for visually superimposing and displaying the user's operation on a virtual environment based on the analysis result.

[0288] The "notification means" is means for notifying the user of the operation improvement part as feedback.

[0289] The "interface means" is means for providing procedures visually and auditorily to the operator and supporting efficient learning.

[0290] The "machine learning technology" is an artificial intelligence technology used for analyzing operation information and identifying the differences between the user's operations and the instructor's operations.

[0291] The "exchange means" is means for efficiently exchanging information with low latency between the terminal and the information processing device.

[0292] The server manages a device that uses motion measurement equipment to measure the user's physical movements in real time and acquire that motion information. This device is equipped with multiple sensors and has the function of accurately capturing data on the user's limbs and posture. The acquired motion data is immediately transmitted to the server via communication means.

[0293] The server analyzes this operational information through an information processing device and compares it with standard operational data from pre-registered instructors. The analysis algorithm uses machine learning techniques to identify differences in operation and determine which areas need improvement. Based on the results, the server generates feedback data specifically indicating the areas for improvement.

[0294] The terminal receives feedback data sent from the server and notifies the user through a display device. A head-mounted display or smart glasses are used as the display device, allowing the user to visually understand areas for improvement by overlaying their own actions with ideal actions within a virtual environment. Furthermore, the terminal can enhance learning efficiency by utilizing interface devices to present a combination of visual and auditory feedback.

[0295] A typical prompt for the entire system would be: "We have installed a new robotic arm in the factory. Please tell us how to build an AR / VR system that provides real-time feedback on the correct operating procedures." This allows for a specific inquiry about improving motion management.

[0296] This system enables workers in the factory work environment to efficiently acquire advanced skills and improve the precision and safety of their operations.

[0297] The flow of a specific process in Application Example 1 will be explained using Figure 12.

[0298] Step 1:

[0299] When a user operates a robot using a motion measurement device, the body movements are captured by sensors. The acquired motion information is transmitted to a terminal as detailed data such as the user's position and speed. This enables real-time input of motion data.

[0300] Step 2:

[0301] The terminal transmits the received motion data to a server via communication means. The server converts this data into a format required for analysis and generates a data set for analysis. The motion data obtained as input is output to the server's analysis system.

[0302] Step 3:

[0303] The server uses a machine learning algorithm to compare the user's motion data with the motion data of an instructor. Through data processing, aspects such as the timing, speed, and angle of the motion are analyzed to identify areas that need improvement. An analysis result is generated and output as specific improvement points.

[0304] Step 4:

[0305] Based on the analysis result, the server generates feedback data for the user's motion. This data includes visual improvement instructions and procedures. The generated feedback data is output to display means.

[0306] Step 5:

[0307] The terminal receives the feedback data sent from the server and presents it to the user through display means. Using a head-mounted display or smart glasses, an ideal motion virtually superimposed on the user is displayed. This allows the user to intuitively understand which parts need improvement.

[0308] Step 6:

[0309] Based on the feedback provided, the user modifies their actions and operates the robot again. The interface provides additional visual and auditory guidance, further enhancing the user's learning. This allows the user to input new actions in response to the feedback.

[0310] Furthermore, an emotion engine that estimates the user's emotions may be incorporated. That is, the identification processing unit 290 may use the emotion identification model 59 to estimate the user's emotions and perform identification processing using the user's emotions.

[0311] The present invention is a system that enables interactive feedback that takes into account the learner's emotional state by combining an emotion engine with an AR / VR teaching system for supporting motion learning. This system includes a motion capture device, a processing device, a display device, a feedback device, a communication means, and an emotion engine.

[0312] The device uses a motion capture system to record the user's body movements in detail in order to collect user motion data in real time. In addition, the emotion engine acquires the user's facial expression data with a camera and analyzes this data to recognize the user's emotions. The emotion engine uses a facial expression analysis algorithm to identify multiple emotions in real time, such as joy, sadness, surprise, anger, fear, disgust, and neutral expression.

[0313] Next, the device sends the collected motion data and emotion data to the server. The server uses the received motion data to perform a comparative analysis with the teacher's actions. Artificial intelligence is used in the analysis to perform a precise comparison of actions and identify specific areas for improvement. Emotion data is also analyzed simultaneously, and feedback is adjusted according to the user's emotions.

[0314] The server then generates feedback on the user's actions based on the analysis results and sends it to the terminal. For visual feedback, an AR / VR device overlays the teacher's ideal actions with the user's actions in the user's field of view. Additionally, voice feedback optimized for the user's emotions is provided. For example, if motivation is low, positive voice assistance is enhanced to increase confidence in improving actions.

[0315] For example, if a user is learning dance moves and the emotion engine detects a decrease in the user's concentration, the device can provide visual feedback along with voice messages of encouragement and relaxation, urging the user to take their time.

[0316] In this way, by considering the user's actions and emotions as an integrated whole, this system provides a more personalized learning experience and greatly supports the user's skill acquisition process.

[0317] The following describes the processing flow.

[0318] Step 1:

[0319] The device acquires user movement data in real time using a motion capture system. This includes collecting positional information of joints and skeleton using sensors or cameras worn by the user.

[0320] Step 2:

[0321] The device acquires user facial expression data through its camera and analyzes it using an emotion engine. It also determines the user's emotional state in real time.

[0322] Step 3:

[0323] The device sends user behavioral data and emotional data to the server. This transmission is performed with low latency, ensuring that data is transferred without compromising real-time capabilities.

[0324] Step 4:

[0325] The server uses the received behavioral data to begin an analysis that compares it to the teacher's behavioral data. This analysis uses artificial intelligence to identify differences and areas for improvement in the behavior in detail.

[0326] Step 5:

[0327] In parallel, the server analyzes data from the emotion engine and determines the tone and content of feedback based on the user's emotion index. If the user's emotions are unstable, it uses a reinforcement learning algorithm to select encouraging or adapted guidance.

[0328] Step 6:

[0329] The server generates feedback information based on the analysis results and sends it to the terminal. The feedback is provided visually and audibly and is customized to enhance user motivation.

[0330] Step 7:

[0331] The device provides feedback to the user via an AR / VR device. The user can view a video in which the teacher's movement model and their own movements are overlaid, and make corrections as needed. Voice assistance is also provided in parallel, allowing the user to receive guidance on specific ways to improve.

[0332] (Example 2)

[0333] Next, we will describe Example 2. In the following description, the data processing device 12 will be referred to as the "server" and the smart glasses 214 will be referred to as the "terminal".

[0334] When learners acquire a skill, simply comparing the accuracy of the skill makes it difficult to provide appropriate instruction tailored to each individual's learning progress and emotional state. Furthermore, standardized feedback can make it difficult for learners to maintain motivation.

[0335] The identification process performed by the identification processing unit 290 of the data processing device 12 in Example 2 is realized by the following means.

[0336] In this invention, the server includes an action acquisition device, a processing device, and an emotion recognition device. This allows for more effective action learning by not only comparing the user's action information with standard actions, but also providing feedback that takes into account the user's emotional state.

[0337] A "motion acquisition device" is a device that collects the user's body movements in real time and plays a role in acquiring specific motion information of the user.

[0338] A "calculation unit" is a device that compares and analyzes user operation information with standard operation information, performing calculations to identify deviations and areas for improvement in operation.

[0339] A "display means" is a means of providing information visually and displaying the user's actions overlaid on a virtual environment based on the analysis results.

[0340] An "information provision device" is a device that provides users with feedback on areas for improvement in their operation, issuing specific instructions and information to support their learning.

[0341] An "emotion recognition device" is a device that identifies a user's emotional state from their facial expressions, and it acquires emotional information using an emotion analysis algorithm.

[0342] A "feedback adjustment device" is a device that optimizes the content and format of the feedback provided based on the user's emotional state.

[0343] A "communication device" is a device that enables low-latency information communication between a terminal and an information processing device, allowing for high-speed and stable data exchange.

[0344] This system is designed to effectively support the user's motion learning and is configured as follows: First, the terminal collects the user's motion data in real time using a motion acquisition device. This motion acquisition device includes cameras and sensors to accurately capture the user's body movements. Furthermore, the terminal utilizes an emotion recognition device to acquire the user's facial expression data via the camera and analyzes their emotional state from it.

[0345] The collected data is transmitted to the server via a low-latency communication device. The server uses a computing unit to perform calculations to compare the user's behavioral data with standard behavioral data. Machine learning techniques are applied to these calculations to identify deviations and areas for improvement. Simultaneously, the emotional data obtained by the emotion recognition device is analyzed by a feedback adjustment device to generate feedback tailored to the user's emotional state.

[0346] The generated feedback is transmitted to the terminal via an information provider and provided to the user. Visual feedback utilizes display means, overlaying the user's actions with ideal actions in a virtual environment. Emotion-based voice feedback is also provided, playing a role in increasing user motivation.

[0347] For example, if a user practicing dance wants to improve the timing of their movements, the system can collect motion data for each step, analyze the differences, and provide feedback. At the same time, if the user's concentration is waning, it can provide a voice message encouraging them to relax.

[0348] Examples of prompts for a generative AI model:

[0349] "Please describe a system that generates appropriate feedback when a user is losing focus while practicing dance moves."

[0350] In this way, this system not only promotes the improvement of users' operational skills but also provides emotional support, resulting in a more fulfilling learning experience.

[0351] The flow of the specific processing in Example 2 will be explained using Figure 13.

[0352] Step 1:

[0353] The terminal uses a motion acquisition device to collect user motion data in real time. The input is the user's physical movements, which are captured by cameras and sensors and converted into numerical data such as the position and angle of the movements. The output is detailed motion data for analysis. During this process, the distance the user moves and the angles of their joints are specifically recorded.

[0354] Step 2:

[0355] The device uses an emotion recognition device to acquire and analyze the user's facial expression data. The input is image data of the user's face, which is applied to an analysis algorithm to extract emotional states such as joy and sadness. The output is an emotion label and data indicating the intensity of that emotion. Specifically, the degree of the user's smile and eyebrow movements are analyzed.

[0356] Step 3:

[0357] The terminal sends the acquired behavioral and emotional data to the server. The input consists of the previously collected behavioral and emotional data, and the output is the generation of data packets to be sent to the server. In this step, the communication device plays the role of transmitting the data with low latency.

[0358] Step 4:

[0359] The server uses a computing device to compare and analyze operational data with standard operational data. The input consists of user operational data and pre-prepared standard operational data. Machine learning is used to calculate the differences in operation and identify areas for improvement. The output is an analysis showing the deviations in operation. Specifically, timing delays and movement precision are quantified.

[0360] Step 5:

[0361] The server generates feedback using a feedback adjustment device based on emotional data. The input is an emotional label and its intensity, which is used to construct a feedback message optimized for the user's emotional state. The output is the adjusted feedback message, which includes positive content and, where necessary, words of encouragement.

[0362] Step 6:

[0363] The server sends the generated feedback to the terminal. The terminal provides this feedback to the user using visual and auditory means. The input is feedback data from the server, and the output is a visual overlay display and audio instructions. Specifically, this involves visualization of actions using AR / VR devices and audio guidance using speakers.

[0364] Step 7:

[0365] The user receives feedback and attempts to improve the operation. The input is feedback from the device, and the output is the improved operation. In this process, the user can learn by following specific instructions and correcting the operation.

[0366] (Application Example 2)

[0367] Next, we will explain application example 2. In the following explanation, the data processing device 12 will be referred to as the "server," and the smart glasses 214 will be referred to as the "terminal."

[0368] In the industrial sector, efficient and safe operation of workers and machinery requires both advanced skills acquisition and emotional support. However, conventional technologies struggle to provide personalized feedback in real time that considers the emotional state of individual workers, in addition to evaluating their movements. Therefore, there is a need for more effective and motivating methods.

[0369] The specific processing performed by the specific processing unit 290 of the data processing device 12 in Application Example 2 is realized by the following means.

[0370] In this invention, the server includes processing means for comparing and examining user behavior information and instructor behavior information, emotion analysis means for analyzing user emotional information, and adjustment means for optimizing output means considering the results of the emotional information analysis. This enables feedback that combines highly accurate evaluation of behavior with emotional support.

[0371] A "detection device" is a device used to record operations and changes in real time.

[0372] The "processing means" refers to a function that analyzes acquired motion information and compares it with a standard motion.

[0373] A "display means" is a device that visually superimposes actions in a virtual space based on the results of the study.

[0374] An "output means" is a device that provides the user with suggestions for improving the operation and feedback.

[0375] An "emotional analysis tool" is a function that analyzes the user's emotions and identifies their state.

[0376] A "modification mechanism" is a system for optimizing the content of feedback based on analyzed emotional information.

[0377] A "communication method" is a system for sending and receiving information with low latency between a terminal device and an information processing device.

[0378] The system in this invention primarily analyzes the actions and emotions of workers in real time and provides personalized feedback. In this system, the server, terminal, and user function as follows:

[0379] First, the terminal acquires the worker's movements in real time using a motion capture device. Examples of motion capture devices include Arduino and Kinect. The terminal also acquires facial expression data via a camera, collecting basic data for analyzing the user's emotions.

[0380] Next, the terminal sends the acquired motion data and emotion data to the server. On the server side, the motion data is analyzed using processing equipment. Specifically, a machine learning algorithm is used to compare the user's motion information with the instructor's motion information to identify accuracy and areas for improvement. In addition, emotion analysis equipment is used to analyze facial expression data and evaluate the user's emotional state.

[0381] Based on these analysis results, the server generates optimized feedback for the output method. This feedback is provided visually on the AR / VR device through the display method. Furthermore, the adjustment method generates audio feedback tailored to the user's emotional state, effectively supporting their motivation.

[0382] To give a concrete example, imagine a scenario in a training session for assembling a new automobile model, where workers are learning precise movements. In this case, it would be possible to capture the workers' movements in detail, evaluate the accuracy of those movements, and then provide visual and audio guidance for the actual movements.

[0383] When using a generative AI model, examples of prompt statements include the following:

[0384] "Propose a program for training new car assembly that provides optimal movement guidance and generates feedback tailored to the worker's emotions. Consider worker movement and emotional data to specifically identify feedback that enhances motivation."

[0385] The flow of a specific process in Application Example 2 will be explained using Figure 14.

[0386] Step 1:

[0387] The device uses a motion capture system to acquire user movement data in real time. The input is raw data from the motion capture sensor, and the output is a stream of movement data. The movement involves recording the user's body movements in detail and converting them into a format that can be transferred as data.

[0388] Step 2:

[0389] The device uses a camera to acquire facial expression data and performs analysis using an emotion analysis tool. The input is image data acquired from the camera, and the output is the analyzed emotional state. Specifically, the device passes the image data through a facial expression analysis algorithm to identify emotions such as joy, anger, sadness, and happiness in real time.

[0390] Step 3:

[0391] The terminal transmits acquired behavioral and emotional data to the server. The input is a data stream related to behavior and emotion, and the output is a successful transmission of data to the server. The terminal efficiently compresses the data and transfers it to the server using a low-latency communication protocol.

[0392] Step 4:

[0393] The server uses processing tools to compare and analyze the user's motion data with the instructor's motion data. The input consists of the user's motion data and existing baseline data, and the output is the extraction of motion similarities and areas for improvement. The server applies machine learning algorithms to perform motion evaluation.

[0394] Step 5:

[0395] The server evaluates facial expression data using emotion analysis tools and determines the emotional state. The input is processed facial expression data, and the output is the user's emotional state and its intensity. The server uses an emotion recognition algorithm to identify emotional patterns.

[0396] Step 6:

[0397] The server optimizes output methods based on analysis results to generate feedback. Inputs are the results of behavioral comparisons and sentiment evaluations, while output is detailed guidance in the form of visual and auditory feedback. The server customizes the content to ensure the feedback is delivered to the user most effectively.

[0398] Step 7:

[0399] The terminal provides feedback received from the server to the user using display means. Input is feedback data, and output is visual guidance via AR / VR devices and audio instructions via speakers. The terminal overlays the instructor's actions with the user's actions to highlight areas for improvement.

[0400] The specific processing unit 290 transmits the result of the specific processing to the smart glasses 214. In the smart glasses 214, the control unit 46A causes the speaker 240 to output the result of the specific processing. The microphone 238 acquires audio indicating user input for the result of the specific processing. The control unit 46A transmits the audio data indicating user input acquired by the microphone 238 to the data processing unit 12. In the data processing unit 12, the specific processing unit 290 acquires the audio data.

[0401] Data generation model 58 is a type of so-called generative AI (Artificial Intelligence). One example of data generation model 58 is ChatGPT (Internet Search).<URL: https: / / openai.com / blog / chatgpt> ), Gemini (Internet search) <url: https: gemini.google.com ?hl="ja">Examples of generative AI include the following. The data generation model 58 is obtained by performing deep learning on a neural network. The data generation model 58 is input with prompts containing instructions, and with inference data such as audio data representing speech, text data representing text, and image data representing images. The data generation model 58 infers from the input inference data according to the instructions indicated by the prompts, and outputs the inference results in data formats such as audio data and text data. Here, inference refers to, for example, analysis, classification, prediction, and / or summarization.

[0402] In the above embodiment, an example was given in which specific processing is performed by the data processing device 12, but the technology of this disclosure is not limited thereto, and the specific processing may also be performed by the smart glasses 214.

[0403] [Third Embodiment]

[0404] Figure 5 shows an example of the configuration of the data processing system 310 according to the third embodiment.

[0405] As shown in Figure 5, the data processing system 310 includes a data processing device 12 and a headset terminal 314. An example of the data processing device 12 is a server.

[0406] The data processing device 12 comprises a computer 22, a database 24, and a communication interface 26. The computer 22 is an example of a "computer" related to the technology of this disclosure. The computer 22 comprises a processor 28, RAM 30, and storage 32. The processor 28, RAM 30, and storage 32 are connected to a bus 34. The database 24 and the communication interface 26 are also connected to the bus 34. The communication interface 26 is connected to a network 54. An example of the network 54 is a WAN (Wide Area Network) and / or a LAN (Local Area Network).

[0407] The headset terminal 314 includes a computer 36, a microphone 238, a speaker 240, a camera 42, a communication interface 44, and a display 343. The computer 36 includes a processor 46, RAM 48, and storage 50. The processor 46, RAM 48, and storage 50 are connected to a bus 52. The microphone 238, speaker 240, camera 42, and display 343 are also connected to the bus 52.

[0408] The microphone 238 receives voice signals from the user 20 and receives instructions from the user 20. The microphone 238 captures the voice signals from the user 20, converts the captured voice into audio data, and outputs it to the processor 46. The speaker 240 outputs audio according to the instructions from the processor 46.

[0409] Camera 42 is a small digital camera equipped with an optical system including a lens, aperture, and shutter, and an image sensor such as a CMOS (Complementary Metal-Oxide-Semiconductor) image sensor or a CCD (Charge Coupled Device) image sensor, and captures images of the area around the user 20 (for example, an imaging range defined by a field of view equivalent to the width of a typical healthy person's field of vision).

[0410] Communication interface 44 is connected to network 54. Communication interfaces 44 and 26 are responsible for the exchange of various information between processor 46 and processor 28 via network 54. The exchange of various information between processor 46 and processor 28 using communication interfaces 44 and 26 is performed in a secure manner.

[0411] Figure 6 shows an example of the main functions of the data processing device 12 and the headset terminal 314. As shown in Figure 6, the data processing device 12 performs specific processing using the processor 28. The storage 32 stores the specific processing program 56.

[0412] The specific processing program 56 is an example of a "program" relating to the technology of this disclosure. The processor 28 reads the specific processing program 56 from the storage 32 and executes the read specific processing program 56 on the RAM 30. The specific processing is realized by the processor 28 operating as a specific processing unit 290 in accordance with the specific processing program 56 executed on the RAM 30.

[0413] The storage 32 stores the data generation model 58 and the emotion identification model 59. The data generation model 58 and the emotion identification model 59 are used by the identification processing unit 290.

[0414] In the headset terminal 314, the processor 46 performs the reception output processing. The storage 50 stores the reception output program 60. The processor 46 reads the reception output program 60 from the storage 50 and executes the read reception output program 60 on the RAM 48. The reception output processing is realized by the processor 46 operating as a control unit 46A according to the reception output program 60 executed on the RAM 48.

[0415] Next, the specific processing performed by the specific processing unit 290 of the data processing device 12 will be described. In the following description, the data processing device 12 will be referred to as the "server" and the headset terminal 314 will be referred to as the "terminal".

[0416] The AR / VR teaching system of the present invention utilizes real-time motion capture and feedback functions to support users in acquiring skills. The specific implementation of this system is described below.

[0417] First, the device uses a motion capture system to capture the user's movements. The motion capture system acquires the positional information and movement of each segment of the user's body in real time. This makes it possible to collect accurate data at the moment the user is performing actions such as playing music or participating in sports. The collected data is transmitted to the device through sensors and camera devices worn by the user.

[0418] Next, the terminal sends the captured user motion data to the server. The server runs a dedicated analysis algorithm to analyze the received data and compare it with the training motion data. The algorithm used here is designed to analyze elements such as the timing, speed, and angle of each motion in detail. The server identifies which parts of the user's motion can be improved and generates the necessary feedback information.

[0419] After the analysis is complete, the server generates feedback on the user's actions and sends it to the terminal. The terminal then presents this feedback information to the user through an AR / VR device. Specifically, the user's actions, overlaid on the teacher's ideal actions, are displayed in real time on the head-mounted display or AR glasses worn by the user. This allows the user to intuitively understand which parts of the actions are different visually.

[0420] Another distinctive feature of the system is that the device can provide feedback information to the user in the form of audio and video. For example, if a user is learning to play the violin, the device can overlay the teacher's hand movements, showing in detail what each finger is doing. With the addition of audio assistance, users can receive feedback in a format that suits their individual learning style.

[0421] This entire process aims to dramatically improve user learning efficiency and provide a personalized, interactive learning experience. In the future, this technology is expected to be applied beyond education to support motor-based learning in various fields.

[0422] The following describes the processing flow.

[0423] Step 1:

[0424] The terminal activates the motion capture device and prepares to collect user movement data in real time. The motion capture device begins acquiring location information for each segment.

[0425] Step 2:

[0426] The terminal packets the user's activity data it has acquired and sends that data to the server using a low-latency communication protocol. The data is streamed continuously and in real time.

[0427] Step 3:

[0428] The server receives the user's behavior data and performs a comparative analysis with the teacher's behavior data that has already been registered. Using artificial intelligence, it meticulously analyzes the differences between the user's and the teacher's behavior to identify specific areas for improvement and discrepancies.

[0429] Step 4:

[0430] Based on the improvements identified in the analysis and the ideal operating model, the server generates visual data that overlays the user's movements onto the virtual space. This visual data includes guidelines on how the user should adjust their movements.

[0431] Step 5:

[0432] The server sends the generated visual data to the terminal, which then displays it on the AR / VR device. Users can adjust their actions while visually confirming their own actions and the teacher's actions.

[0433] Step 6:

[0434] In addition to visual data, the device provides users with audio and video feedback through a feedback system. This allows users to understand areas for improvement from multiple perspectives and learn more efficiently.

[0435] Step 7:

[0436] The user modifies their movements based on the feedback provided and records their movements again using the motion capture device. This creates a cycle of continuous improvement through feedback.

[0437] (Example 1)

[0438] Next, we will describe Example 1. In the following description, the data processing device 12 will be referred to as the "server," and the headset-type terminal 314 will be referred to as the "terminal."

[0439] Conventional motion instruction systems have a problem where users cannot intuitively understand the difference between their own movements and ideal movements, leading to decreased instruction efficiency. Furthermore, the technical hurdles to ensuring real-time analysis of motion data are high, resulting in delays in feedback, which is a challenge.

[0440] The identification process performed by the identification processing unit 290 of the data processing device 12 in Example 1 is realized by the following means.

[0441] In this invention, the server includes an information processing device that compares and analyzes the user's movement information with that of the instructor, a presentation device that overlays the user's movements onto an abstract space based on the analysis results, and a reporting device that provides the user with areas for improvement in their movements. This allows the user to intuitively and in real time understand areas for improvement in their own movements, enabling effective learning.

[0442] A "motion detection device" is a device that measures the user's body movements and posture in real time and collects that information.

[0443] An "information processing device" is a device used to analyze collected behavioral information, and in particular, it has the function of comparing the actions of the user and the instructor.

[0444] A "presentation device" is a device that visually displays operational information to the user based on analysis results, and is a device that enables the superimposition and display of operations in an abstract space.

[0445] A "reporting device" is a device that presents the user with instructions for improvements and operational modifications based on the analysis results.

[0446] An "intelligent algorithm" is an automated analysis method that analyzes motion information and effectively compares the actions of the user and the instructor.

[0447] "Communication means" refers to technical means for sending and receiving information between a terminal and an information management device with low latency and efficiency.

[0448] This invention is an AR / VR teaching system for assisting users in acquiring skills. The system includes a motion detection device, an information processing device, a presentation device, and a reporting device, and provides real-time analysis and feedback of motion.

[0449] The device detects the user's movements using a motion capture device attached to the user's body. This device digitizes the movement and position of each part of the user's body in real time, providing precise motion information. The motion capture device includes multiple sensors and high-resolution cameras.

[0450] The server receives motion information transmitted from the terminal and processes it using intelligent algorithms. Software on the server compares the user's motion data with the instructor's ideal motion data and performs analysis. This analysis includes elements such as motion speed, angle, and timing. The analysis results are then processed into feedback for presentation to the user.

[0451] The terminal visually presents the feedback received from the server to the user using an AR / VR device. The display device overlays the analyzed movements with the ideal movements through the head-mounted display or AR glasses worn by the user. This allows the user to intuitively understand their own movements and make necessary adjustments.

[0452] Users can improve their performance based on feedback from the device. The reporting device provides users with detailed instructions, via audio and video, on which parts of their performance should be improved and how. For example, if a user is learning to play the violin, improvements would be suggested via audio guidance, along with a digital overlay demonstrating the correct finger movements.

[0453] An example of a prompt for a generative AI model is, "Design a system to help a user learn the correct way to play the violin." Based on this prompt, the model will compare the actions of the user and the instructor in real time, select how to generate effective feedback, and provide specific suggestions.

[0454] The flow of the specific processing in Example 1 will be explained using Figure 11.

[0455] Step 1:

[0456] The user prepares to begin the process. They attach the motion capture device to their body and verify its connection to the system. The device prepares to detect the user's body movements in real time. The user's action (e.g., playing the violin) is set as input.

[0457] Step 2:

[0458] The terminal acquires user movement information from the motion capture device. The device records the user's movements and position, converts this information into digital data format, and transmits it to the terminal. As output, the user's movement data is acquired and temporarily stored on the terminal.

[0459] Step 3:

[0460] The terminal sends acquired motion data to the server. The data is transmitted over the network, allowing the server to perform real-time analysis. The input is the user's real-time motion data, and this data is sent to the server as output.

[0461] Step 4:

[0462] The server analyzes the received user motion data using intelligent algorithms. It compares the user's motion information with the instructor's ideal motion and performs data calculations to evaluate the accuracy, speed, angle, etc. of the motion. The input to the analysis is the user's motion information, and the output is the analysis result.

[0463] Step 5:

[0464] The server generates feedback information based on the analysis results. It prepares suggestions for improving the user's actions in text, audio, and video formats. This process clarifies to the user which actions they should correct and how. Detailed feedback information is generated as output.

[0465] Step 6:

[0466] The server sends the generated feedback information to the terminal. The feedback is formatted so that it is presented appropriately on the AR / VR device used by the user. The input is the feedback information, and the output is the transmission of data to the terminal.

[0467] Step 7:

[0468] The terminal displays the received feedback on the user's AR / VR device. The display device allows the user to overlay and re-examine the ideal action against their own action. This enables the user to receive direct visual feedback. Visual feedback is provided to the user as output.

[0469] Step 8:

[0470] The user modifies their movements based on the feedback provided. Specifically, they adjust their body movements by referring to the ideal movements displayed on the device. The improved movements are recorded again by the motion capture device, and the process is repeated. In this step, the input is the feedback information, and the output is the user's improved movements.

[0471] (Application Example 1)

[0472] Next, we will explain Application Example 1. In the following explanation, the data processing device 12 will be referred to as the "server," and the headset-type terminal 314 will be referred to as the "terminal."

[0473] In manufacturing environments, there is a problem with inefficient training methods when workers learn complex robot operation procedures. Furthermore, a lack of flexibility to accommodate individual learning styles and the absence of real-time feedback can delay the learning process. This can potentially impact the accuracy and productivity of operations.

[0474] The specific processing performed by the specific processing unit 290 of the data processing device 12 in Application Example 1 is realized by the following means.

[0475] In this invention, the server includes a motion measurement device for measuring movements in real time, an information processing device for comparing and analyzing the user's motion information and the instructor's motion information, and a display means for overlaying the user's movements onto a virtual environment based on the analysis results. This allows the worker to receive immediate visual and auditory feedback, enabling efficient acquisition of operating procedures based on their individual learning style.

[0476] A "motion measurement device" is a device that measures a user's physical movements in real time and acquires motion information.

[0477] An "information processing device" is a device used to compare and analyze acquired motion information with the motion information of an instructor.

[0478] "Display means" refers to a means for visually displaying the user's actions overlaid on a virtual environment based on the analysis results.

[0479] A "notification method" is a means of notifying the user of areas where functionality has been improved as feedback.

[0480] An "interface means" is a means of providing instructions to a worker visually and aurally, and supporting efficient learning.

[0481] "Machine learning technology" is an artificial intelligence technology used to analyze motion information and identify differences between the user's actions and the instructor's actions.

[0482] "Exchange means" refers to a means for exchanging information between a terminal and an information processing device with low latency and efficiency.

[0483] The server manages a device that uses motion measurement equipment to measure the user's physical movements in real time and acquire that motion information. This device is equipped with multiple sensors and has the function of accurately capturing data on the user's limbs and posture. The acquired motion data is immediately transmitted to the server via communication means.

[0484] The server analyzes this operational information through an information processing device and compares it with standard operational data from pre-registered instructors. The analysis algorithm uses machine learning techniques to identify differences in operation and determine which areas need improvement. Based on the results, the server generates feedback data specifically indicating the areas for improvement.

[0485] The terminal receives feedback data sent from the server and notifies the user through a display device. A head-mounted display or smart glasses are used as the display device, allowing the user to visually understand areas for improvement by overlaying their own actions with ideal actions within a virtual environment. Furthermore, the terminal can enhance learning efficiency by utilizing interface devices to present a combination of visual and auditory feedback.

[0486] A typical prompt for the entire system would be: "We have installed a new robotic arm in the factory. Please tell us how to build an AR / VR system that provides real-time feedback on the correct operating procedures." This allows for a specific inquiry about improving motion management.

[0487] This system enables workers in the factory work environment to efficiently acquire advanced skills and improve the precision and safety of their operations.

[0488] The flow of a specific process in Application Example 1 will be explained using Figure 12.

[0489] Step 1:

[0490] A motion measurement device captures the user's physical movements as they operate the robot using sensors. The acquired motion information is transmitted to a terminal as detailed data such as the user's position and speed. This allows for real-time input of motion data.

[0491] Step 2:

[0492] The terminal transmits the received operational data to the server via a communication means. The server converts this data into the format required for analysis and generates a dataset for analysis. The operational data obtained as input is output to the server's analysis system.

[0493] Step 3:

[0494] The server uses machine learning algorithms to compare user movement data with instructor movement data. Data processing analyzes movement timing, speed, angle, etc., identifying areas for improvement. Analysis results are generated and output as specific areas for improvement.

[0495] Step 4:

[0496] The server generates feedback data on user actions based on the analysis results. This data includes visual improvement instructions and procedures. The generated feedback data is output to a display device.

[0497] Step 5:

[0498] The terminal receives feedback data sent from the server and presents it to the user through a display device. Using a head-mounted display or smart glasses, the user is shown a virtual overlay of the ideal behavior. This allows the user to intuitively understand which parts need improvement.

[0499] Step 6:

[0500] Based on the feedback provided, the user modifies their actions and operates the robot again. The interface provides additional visual and auditory guidance, further enhancing the user's learning. This allows the user to input new actions in response to the feedback.

[0501] Furthermore, an emotion engine that estimates the user's emotions may be incorporated. That is, the identification processing unit 290 may use the emotion identification model 59 to estimate the user's emotions and perform identification processing using the user's emotions.

[0502] The present invention is a system that enables interactive feedback that takes into account the learner's emotional state by combining an emotion engine with an AR / VR teaching system for supporting motion learning. This system includes a motion capture device, a processing device, a display device, a feedback device, a communication means, and an emotion engine.

[0503] The device uses a motion capture system to record the user's body movements in detail in order to collect user motion data in real time. In addition, the emotion engine acquires the user's facial expression data with a camera and analyzes this data to recognize the user's emotions. The emotion engine uses a facial expression analysis algorithm to identify multiple emotions in real time, such as joy, sadness, surprise, anger, fear, disgust, and neutral expression.

[0504] Next, the device sends the collected motion data and emotion data to the server. The server uses the received motion data to perform a comparative analysis with the teacher's actions. Artificial intelligence is used in the analysis to perform a precise comparison of actions and identify specific areas for improvement. Emotion data is also analyzed simultaneously, and feedback is adjusted according to the user's emotions.

[0505] The server then generates feedback on the user's actions based on the analysis results and sends it to the terminal. For visual feedback, an AR / VR device overlays the teacher's ideal actions with the user's actions in the user's field of view. Additionally, voice feedback optimized for the user's emotions is provided. For example, if motivation is low, positive voice assistance is enhanced to increase confidence in improving actions.

[0506] For example, if a user is learning dance moves and the emotion engine detects a decrease in the user's concentration, the device can provide visual feedback along with voice messages of encouragement and relaxation, urging the user to take their time.

[0507] In this way, by considering the user's actions and emotions as an integrated whole, this system provides a more personalized learning experience and greatly supports the user's skill acquisition process.

[0508] The following describes the processing flow.

[0509] Step 1:

[0510] The device acquires user movement data in real time using a motion capture system. This includes collecting positional information of joints and skeleton using sensors or cameras worn by the user.

[0511] Step 2:

[0512] The device acquires user facial expression data through its camera and analyzes it using an emotion engine. It also determines the user's emotional state in real time.

[0513] Step 3:

[0514] The device sends user behavioral data and emotional data to the server. This transmission is performed with low latency, ensuring that data is transferred without compromising real-time capabilities.

[0515] Step 4:

[0516] The server uses the received behavioral data to begin an analysis that compares it to the teacher's behavioral data. This analysis uses artificial intelligence to identify differences and areas for improvement in the behavior in detail.

[0517] Step 5:

[0518] In parallel, the server analyzes data from the emotion engine and determines the tone and content of feedback based on the user's emotion index. If the user's emotions are unstable, it uses a reinforcement learning algorithm to select encouraging or adapted guidance.

[0519] Step 6:

[0520] The server generates feedback information based on the analysis results and sends it to the terminal. The feedback is provided visually and audibly and is customized to enhance user motivation.

[0521] Step 7:

[0522] The device provides feedback to the user via an AR / VR device. The user can view a video in which the teacher's movement model and their own movements are overlaid, and make corrections as needed. Voice assistance is also provided in parallel, allowing the user to receive guidance on specific ways to improve.

[0523] (Example 2)

[0524] Next, we will describe Example 2. In the following description, the data processing device 12 will be referred to as the "server," and the headset-type terminal 314 will be referred to as the "terminal."

[0525] When learners acquire a skill, simply comparing the accuracy of the skill makes it difficult to provide appropriate instruction tailored to each individual's learning progress and emotional state. Furthermore, standardized feedback can make it difficult for learners to maintain motivation.

[0526] The identification process performed by the identification processing unit 290 of the data processing device 12 in Example 2 is realized by the following means.

[0527] In this invention, the server includes an action acquisition device, a processing device, and an emotion recognition device. This allows for more effective action learning by not only comparing the user's action information with standard actions, but also providing feedback that takes into account the user's emotional state.

[0528] A "motion acquisition device" is a device that collects the user's body movements in real time and plays a role in acquiring specific motion information of the user.

[0529] A "calculation unit" is a device that compares and analyzes user operation information with standard operation information, performing calculations to identify deviations and areas for improvement in operation.

[0530] A "display means" is a means of providing information visually and displaying the user's actions overlaid on a virtual environment based on the analysis results.

[0531] An "information provision device" is a device that provides users with feedback on areas for improvement in their operation, issuing specific instructions and information to support their learning.

[0532] An "emotion recognition device" is a device that identifies a user's emotional state from their facial expressions, and it acquires emotional information using an emotion analysis algorithm.

[0533] A "feedback adjustment device" is a device that optimizes the content and format of the feedback provided based on the user's emotional state.

[0534] A "communication device" is a device that enables low-latency information communication between a terminal and an information processing device, allowing for high-speed and stable data exchange.

[0535] This system is designed to effectively support the user's motion learning and is configured as follows: First, the terminal collects the user's motion data in real time using a motion acquisition device. This motion acquisition device includes cameras and sensors to accurately capture the user's body movements. Furthermore, the terminal utilizes an emotion recognition device to acquire the user's facial expression data via the camera and analyzes their emotional state from it.

[0536] The collected data is transmitted to the server via a low-latency communication device. The server uses a computing unit to perform calculations to compare the user's behavioral data with standard behavioral data. Machine learning techniques are applied to these calculations to identify deviations and areas for improvement. Simultaneously, the emotional data obtained by the emotion recognition device is analyzed by a feedback adjustment device to generate feedback tailored to the user's emotional state.

[0537] The generated feedback is transmitted to the terminal via an information provider and provided to the user. Visual feedback utilizes display means, overlaying the user's actions with ideal actions in a virtual environment. Emotion-based voice feedback is also provided, playing a role in increasing user motivation.

[0538] For example, if a user practicing dance wants to improve the timing of their movements, the system can collect motion data for each step, analyze the differences, and provide feedback. At the same time, if the user's concentration is waning, it can provide a voice message encouraging them to relax.

[0539] Examples of prompts for a generative AI model:

[0540] "Please describe a system that generates appropriate feedback when a user is losing focus while practicing dance moves."

[0541] In this way, this system not only promotes the improvement of users' operational skills but also provides emotional support, resulting in a more fulfilling learning experience.

[0542] The flow of the specific processing in Example 2 will be explained using Figure 13.

[0543] Step 1:

[0544] The terminal uses a motion acquisition device to collect user motion data in real time. The input is the user's physical movements, which are captured by cameras and sensors and converted into numerical data such as the position and angle of the movements. The output is detailed motion data for analysis. During this process, the distance the user moves and the angles of their joints are specifically recorded.

[0545] Step 2:

[0546] The device uses an emotion recognition device to acquire and analyze the user's facial expression data. The input is image data of the user's face, which is applied to an analysis algorithm to extract emotional states such as joy and sadness. The output is an emotion label and data indicating the intensity of that emotion. Specifically, the degree of the user's smile and eyebrow movements are analyzed.

[0547] Step 3:

[0548] The terminal sends the acquired behavioral and emotional data to the server. The input consists of the previously collected behavioral and emotional data, and the output is the generation of data packets to be sent to the server. In this step, the communication device plays the role of transmitting the data with low latency.

[0549] Step 4:

[0550] The server uses a computing device to compare and analyze operational data with standard operational data. The input consists of user operational data and pre-prepared standard operational data. Machine learning is used to calculate the differences in operation and identify areas for improvement. The output is an analysis showing the deviations in operation. Specifically, timing delays and movement precision are quantified.

[0551] Step 5:

[0552] The server generates feedback using a feedback adjustment device based on emotional data. The input is an emotional label and its intensity, which is used to construct a feedback message optimized for the user's emotional state. The output is the adjusted feedback message, which includes positive content and, where necessary, words of encouragement.

[0553] Step 6:

[0554] The server sends the generated feedback to the terminal. The terminal provides this feedback to the user using visual and auditory means. The input is feedback data from the server, and the output is a visual overlay display and audio instructions. Specifically, this involves visualization of actions using AR / VR devices and audio guidance using speakers.

[0555] Step 7:

[0556] The user receives feedback and attempts to improve the operation. The input is feedback from the device, and the output is the improved operation. In this process, the user can learn by following specific instructions and correcting the operation.

[0557] (Application Example 2)

[0558] Next, we will explain application example 2. In the following explanation, the data processing device 12 will be referred to as the "server," and the headset-type terminal 314 will be referred to as the "terminal."

[0559] In the industrial sector, efficient and safe operation of workers and machinery requires both advanced skills acquisition and emotional support. However, conventional technologies struggle to provide personalized feedback in real time that considers the emotional state of individual workers, in addition to evaluating their movements. Therefore, there is a need for more effective and motivating methods.

[0560] The specific processing performed by the specific processing unit 290 of the data processing device 12 in Application Example 2 is realized by the following means.

[0561] In this invention, the server includes processing means for comparing and examining user behavior information and instructor behavior information, emotion analysis means for analyzing user emotional information, and adjustment means for optimizing output means considering the results of the emotional information analysis. This enables feedback that combines highly accurate evaluation of behavior with emotional support.

[0562] A "detection device" is a device used to record operations and changes in real time.

[0563] The "processing means" refers to a function that analyzes acquired motion information and compares it with a standard motion.

[0564] A "display means" is a device that visually superimposes actions in a virtual space based on the results of the study.

[0565] An "output means" is a device that provides the user with suggestions for improving the operation and feedback.

[0566] An "emotional analysis tool" is a function that analyzes the user's emotions and identifies their state.

[0567] A "modification mechanism" is a system for optimizing the content of feedback based on analyzed emotional information.

[0568] A "communication method" is a system for sending and receiving information with low latency between a terminal device and an information processing device.

[0569] The system in this invention primarily analyzes the actions and emotions of workers in real time and provides personalized feedback. In this system, the server, terminal, and user function as follows:

[0570] First, the terminal acquires the worker's movements in real time using a motion capture device. Examples of motion capture devices include Arduino and Kinect. The terminal also acquires facial expression data via a camera, collecting basic data for analyzing the user's emotions.

[0571] Next, the terminal sends the acquired motion data and emotion data to the server. On the server side, the motion data is analyzed using processing equipment. Specifically, a machine learning algorithm is used to compare the user's motion information with the instructor's motion information to identify accuracy and areas for improvement. In addition, emotion analysis equipment is used to analyze facial expression data and evaluate the user's emotional state.

[0572] Based on these analysis results, the server generates optimized feedback for the output method. This feedback is provided visually on the AR / VR device through the display method. Furthermore, the adjustment method generates audio feedback tailored to the user's emotional state, effectively supporting their motivation.

[0573] To give a concrete example, imagine a scenario in a training session for assembling a new automobile model, where workers are learning precise movements. In this case, it would be possible to capture the workers' movements in detail, evaluate the accuracy of those movements, and then provide visual and audio guidance for the actual movements.

[0574] When using a generative AI model, examples of prompt statements include the following:

[0575] "Propose a program for training new car assembly that provides optimal movement guidance and generates feedback tailored to the worker's emotions. Consider worker movement and emotional data to specifically identify feedback that enhances motivation."

[0576] The flow of a specific process in Application Example 2 will be explained using Figure 14.

[0577] Step 1:

[0578] The device uses a motion capture system to acquire user movement data in real time. The input is raw data from the motion capture sensor, and the output is a stream of movement data. The movement involves recording the user's body movements in detail and converting them into a format that can be transferred as data.

[0579] Step 2:

[0580] The device uses a camera to acquire facial expression data and performs analysis using an emotion analysis tool. The input is image data acquired from the camera, and the output is the analyzed emotional state. Specifically, the device passes the image data through a facial expression analysis algorithm to identify emotions such as joy, anger, sadness, and happiness in real time.

[0581] Step 3:

[0582] The terminal transmits acquired behavioral and emotional data to the server. The input is a data stream related to behavior and emotion, and the output is a successful transmission of data to the server. The terminal efficiently compresses the data and transfers it to the server using a low-latency communication protocol.

[0583] Step 4:

[0584] The server uses processing tools to compare and analyze the user's motion data with the instructor's motion data. The input consists of the user's motion data and existing baseline data, and the output is the extraction of motion similarities and areas for improvement. The server applies machine learning algorithms to perform motion evaluation.

[0585] Step 5:

[0586] The server evaluates facial expression data using emotion analysis tools and determines the emotional state. The input is processed facial expression data, and the output is the user's emotional state and its intensity. The server uses an emotion recognition algorithm to identify emotional patterns.

[0587] Step 6:

[0588] The server optimizes output methods based on analysis results to generate feedback. Inputs are the results of behavioral comparisons and sentiment evaluations, while output is detailed guidance in the form of visual and auditory feedback. The server customizes the content to ensure the feedback is delivered to the user most effectively.

[0589] Step 7:

[0590] The terminal provides feedback received from the server to the user using display means. Input is feedback data, and output is visual guidance via AR / VR devices and audio instructions via speakers. The terminal overlays the instructor's actions with the user's actions to highlight areas for improvement.

[0591] The specific processing unit 290 transmits the result of the specific processing to the headset terminal 314. In the headset terminal 314, the control unit 46A causes the speaker 240 and display 343 to output the result of the specific processing. The microphone 238 acquires audio indicating user input for the result of the specific processing. The control unit 46A transmits the audio data indicating user input acquired by the microphone 238 to the data processing unit 12. In the data processing unit 12, the specific processing unit 290 acquires the audio data.

[0592] Data generation model 58 is a type of so-called generative AI (Artificial Intelligence). One example of data generation model 58 is ChatGPT (Internet Search).<URL: https: / / openai.com / blog / chatgpt> ), Gemini (Internet search) <url: https: gemini.google.com ?hl="ja">Examples of generative AI include the following. The data generation model 58 is obtained by performing deep learning on a neural network. The data generation model 58 is input with prompts containing instructions, and with inference data such as audio data representing speech, text data representing text, and image data representing images. The data generation model 58 infers from the input inference data according to the instructions indicated by the prompts, and outputs the inference results in data formats such as audio data and text data. Here, inference refers to, for example, analysis, classification, prediction, and / or summarization.

[0593] In the above embodiment, an example was given in which specific processing is performed by the data processing device 12, but the technology of this disclosure is not limited thereto, and specific processing may also be performed by the headset terminal 314.

[0594] [Fourth Embodiment]

[0595] Figure 7 shows an example of the configuration of the data processing system 410 according to the fourth embodiment.

[0596] As shown in Figure 7, the data processing system 410 includes a data processing device 12 and a robot 414. An example of the data processing device 12 is a server.

[0597] The data processing device 12 comprises a computer 22, a database 24, and a communication interface 26. The computer 22 is an example of a "computer" related to the technology of this disclosure. The computer 22 comprises a processor 28, RAM 30, and storage 32. The processor 28, RAM 30, and storage 32 are connected to a bus 34. The database 24 and the communication interface 26 are also connected to the bus 34. The communication interface 26 is connected to a network 54. An example of the network 54 is a WAN (Wide Area Network) and / or a LAN (Local Area Network).

[0598] The robot 414 includes a computer 36, a microphone 238, a speaker 240, a camera 42, a communication interface 44, and a controlled object 443. The computer 36 includes a processor 46, RAM 48, and storage 50. The processor 46, RAM 48, and storage 50 are connected to a bus 52. The microphone 238, speaker 240, camera 42, and controlled object 443 are also connected to the bus 52.

[0599] The microphone 238 receives voice signals from the user 20 and receives instructions from the user 20. The microphone 238 captures the voice signals from the user 20, converts the captured voice into audio data, and outputs it to the processor 46. The speaker 240 outputs audio according to the instructions from the processor 46.

[0600] Camera 42 is a small digital camera equipped with an optical system including a lens, aperture, and shutter, and an image sensor such as a CMOS (Complementary Metal-Oxide-Semiconductor) image sensor or a CCD (Charge Coupled Device) image sensor, and captures images of the area around the user 20 (for example, an imaging range defined by a field of view equivalent to the width of a typical healthy person's field of vision).

[0601] Communication interface 44 is connected to network 54. Communication interfaces 44 and 26 are responsible for the exchange of various information between processor 46 and processor 28 via network 54. The exchange of various information between processor 46 and processor 28 using communication interfaces 44 and 26 is performed in a secure manner.

[0602] The controlled object 443 includes a display device, LEDs in the eyes, and motors that drive the arms, hands, and feet. The posture and gestures of the robot 414 are controlled by controlling the motors of the arms, hands, and feet. Some of the robot 414's emotions can be expressed by controlling these motors. Furthermore, the robot 414's facial expressions can also be expressed by controlling the illumination state of the LEDs in its eyes.

[0603] Figure 8 shows an example of the main functions of the data processing device 12 and the robot 414. As shown in Figure 8, the data processing device 12 performs specific processing using the processor 28. The storage 32 stores the specific processing program 56.

[0604] The specific processing program 56 is an example of a "program" relating to the technology of this disclosure. The processor 28 reads the specific processing program 56 from the storage 32 and executes the read specific processing program 56 on the RAM 30. The specific processing is realized by the processor 28 operating as a specific processing unit 290 in accordance with the specific processing program 56 executed on the RAM 30.

[0605] The storage 32 stores the data generation model 58 and the emotion identification model 59. The data generation model 58 and the emotion identification model 59 are used by the identification processing unit 290.

[0606] In robot 414, the processor 46 performs the reception output processing. The storage 50 stores the reception output program 60. The processor 46 reads the reception output program 60 from the storage 50 and executes the read reception output program 60 on the RAM 48. The reception output processing is realized by the processor 46 operating as a control unit 46A according to the reception output program 60 executed on the RAM 48.

[0607] Next, the specific processing performed by the specific processing unit 290 of the data processing device 12 will be described. In the following description, the data processing device 12 will be referred to as the "server" and the robot 414 as the "terminal".

[0608] The AR / VR teaching system of the present invention utilizes real-time motion capture and feedback functions to support users in acquiring skills. The specific implementation of this system is described below.

[0609] First, the device uses a motion capture system to capture the user's movements. The motion capture system acquires the positional information and movement of each segment of the user's body in real time. This makes it possible to collect accurate data at the moment the user is performing actions such as playing music or participating in sports. The collected data is transmitted to the device through sensors and camera devices worn by the user.

[0610] Next, the terminal sends the captured user motion data to the server. The server runs a dedicated analysis algorithm to analyze the received data and compare it with the training motion data. The algorithm used here is designed to analyze elements such as the timing, speed, and angle of each motion in detail. The server identifies which parts of the user's motion can be improved and generates the necessary feedback information.

[0611] After the analysis is complete, the server generates feedback on the user's actions and sends it to the terminal. The terminal then presents this feedback information to the user through an AR / VR device. Specifically, the user's actions, overlaid on the teacher's ideal actions, are displayed in real time on the head-mounted display or AR glasses worn by the user. This allows the user to intuitively understand which parts of the actions are different visually.

[0612] Another distinctive feature of the system is that the device can provide feedback information to the user in the form of audio and video. For example, if a user is learning to play the violin, the device can overlay the teacher's hand movements, showing in detail what each finger is doing. With the addition of audio assistance, users can receive feedback in a format that suits their individual learning style.

[0613] This entire process aims to dramatically improve user learning efficiency and provide a personalized, interactive learning experience. In the future, this technology is expected to be applied beyond education to support motor-based learning in various fields.

[0614] The following describes the processing flow.

[0615] Step 1:

[0616] The terminal activates the motion capture device and prepares to collect user movement data in real time. The motion capture device begins acquiring location information for each segment.

[0617] Step 2:

[0618] The terminal packets the user's activity data it has acquired and sends that data to the server using a low-latency communication protocol. The data is streamed continuously and in real time.

[0619] Step 3:

[0620] The server receives the user's behavior data and performs a comparative analysis with the teacher's behavior data that has already been registered. Using artificial intelligence, it meticulously analyzes the differences between the user's and the teacher's behavior to identify specific areas for improvement and discrepancies.

[0621] Step 4:

[0622] Based on the improvements identified in the analysis and the ideal operating model, the server generates visual data that overlays the user's movements onto the virtual space. This visual data includes guidelines on how the user should adjust their movements.

[0623] Step 5:

[0624] The server sends the generated visual data to the terminal, which then displays it on the AR / VR device. Users can adjust their actions while visually confirming their own actions and the teacher's actions.

[0625] Step 6:

[0626] In addition to visual data, the device provides users with audio and video feedback through a feedback system. This allows users to understand areas for improvement from multiple perspectives and learn more efficiently.

[0627] Step 7:

[0628] The user modifies their movements based on the feedback provided and records their movements again using the motion capture device. This creates a cycle of continuous improvement through feedback.

[0629] (Example 1)

[0630] Next, we will describe Example 1. In the following description, the data processing device 12 will be referred to as the "server" and the robot 414 as the "terminal".

[0631] Conventional motion instruction systems have a problem where users cannot intuitively understand the difference between their own movements and ideal movements, leading to decreased instruction efficiency. Furthermore, the technical hurdles to ensuring real-time analysis of motion data are high, resulting in delays in feedback, which is a challenge.

[0632] The identification process performed by the identification processing unit 290 of the data processing device 12 in Example 1 is realized by the following means.

[0633] In this invention, the server includes an information processing device that compares and analyzes the user's movement information with that of the instructor, a presentation device that overlays the user's movements onto an abstract space based on the analysis results, and a reporting device that provides the user with areas for improvement in their movements. This allows the user to intuitively and in real time understand areas for improvement in their own movements, enabling effective learning.

[0634] A "motion detection device" is a device that measures the user's body movements and posture in real time and collects that information.

[0635] An "information processing device" is a device used to analyze collected behavioral information, and in particular, it has the function of comparing the actions of the user and the instructor.

[0636] A "presentation device" is a device that visually displays operational information to the user based on analysis results, and is a device that enables the superimposition and display of operations in an abstract space.

[0637] A "reporting device" is a device that presents the user with instructions for improvements and operational modifications based on the analysis results.

[0638] An "intelligent algorithm" is an automated analysis method that analyzes motion information and effectively compares the actions of the user and the instructor.

[0639] "Communication means" refers to technical means for sending and receiving information between a terminal and an information management device with low latency and efficiency.

[0640] This invention is an AR / VR teaching system for assisting users in acquiring skills. The system includes a motion detection device, an information processing device, a presentation device, and a reporting device, and provides real-time analysis and feedback of motion.

[0641] The device detects the user's movements using a motion capture device attached to the user's body. This device digitizes the movement and position of each part of the user's body in real time, providing precise motion information. The motion capture device includes multiple sensors and high-resolution cameras.

[0642] The server receives motion information transmitted from the terminal and processes it using intelligent algorithms. Software on the server compares the user's motion data with the instructor's ideal motion data and performs analysis. This analysis includes elements such as motion speed, angle, and timing. The analysis results are then processed into feedback for presentation to the user.

[0643] The terminal visually presents the feedback received from the server to the user using an AR / VR device. The display device overlays the analyzed movements with the ideal movements through the head-mounted display or AR glasses worn by the user. This allows the user to intuitively understand their own movements and make necessary adjustments.

[0644] Users can improve their performance based on feedback from the device. The reporting device provides users with detailed instructions, via audio and video, on which parts of their performance should be improved and how. For example, if a user is learning to play the violin, improvements would be suggested via audio guidance, along with a digital overlay demonstrating the correct finger movements.

[0645] An example of a prompt for a generative AI model is, "Design a system to help a user learn the correct way to play the violin." Based on this prompt, the model will compare the actions of the user and the instructor in real time, select how to generate effective feedback, and provide specific suggestions.

[0646] The flow of the specific processing in Example 1 will be explained using Figure 11.

[0647] Step 1:

[0648] The user prepares to begin the process. They attach the motion capture device to their body and verify its connection to the system. The device prepares to detect the user's body movements in real time. The user's action (e.g., playing the violin) is set as input.

[0649] Step 2:

[0650] The terminal acquires user movement information from the motion capture device. The device records the user's movements and position, converts this information into digital data format, and transmits it to the terminal. As output, the user's movement data is acquired and temporarily stored on the terminal.

[0651] Step 3:

[0652] The terminal sends acquired motion data to the server. The data is transmitted over the network, allowing the server to perform real-time analysis. The input is the user's real-time motion data, and this data is sent to the server as output.

[0653] Step 4:

[0654] The server analyzes the received user motion data using intelligent algorithms. It compares the user's motion information with the instructor's ideal motion and performs data calculations to evaluate the accuracy, speed, angle, etc. of the motion. The input to the analysis is the user's motion information, and the output is the analysis result.

[0655] Step 5:

[0656] The server generates feedback information based on the analysis results. It prepares suggestions for improving the user's actions in text, audio, and video formats. This process clarifies to the user which actions they should correct and how. Detailed feedback information is generated as output.

[0657] Step 6:

[0658] The server sends the generated feedback information to the terminal. The feedback is formatted so that it is presented appropriately on the AR / VR device used by the user. The input is the feedback information, and the output is the transmission of data to the terminal.

[0659] Step 7:

[0660] The terminal displays the received feedback on the user's AR / VR device. The display device allows the user to overlay and re-examine the ideal action against their own action. This enables the user to receive direct visual feedback. Visual feedback is provided to the user as output.

[0661] Step 8:

[0662] The user modifies their movements based on the feedback provided. Specifically, they adjust their body movements by referring to the ideal movements displayed on the device. The improved movements are recorded again by the motion capture device, and the process is repeated. In this step, the input is the feedback information, and the output is the user's improved movements.

[0663] (Application Example 1)

[0664] Next, we will explain Application Example 1. In the following explanation, the data processing device 12 will be referred to as the "server" and the robot 414 as the "terminal".

[0665] In manufacturing environments, there is a problem with inefficient training methods when workers learn complex robot operation procedures. Furthermore, a lack of flexibility to accommodate individual learning styles and the absence of real-time feedback can delay the learning process. This can potentially impact the accuracy and productivity of operations.

[0666] The specific processing performed by the specific processing unit 290 of the data processing device 12 in Application Example 1 is realized by the following means.

[0667] In this invention, the server includes a motion measurement device for measuring movements in real time, an information processing device for comparing and analyzing the user's motion information and the instructor's motion information, and a display means for overlaying the user's movements onto a virtual environment based on the analysis results. This allows the worker to receive immediate visual and auditory feedback, enabling efficient acquisition of operating procedures based on their individual learning style.

[0668] A "motion measurement device" is a device that measures a user's physical movements in real time and acquires motion information.

[0669] An "information processing device" is a device used to compare and analyze acquired motion information with the motion information of an instructor.

[0670] "Display means" refers to a means for visually displaying the user's actions overlaid on a virtual environment based on the analysis results.

[0671] A "notification method" is a means of notifying the user of areas where functionality has been improved as feedback.

[0672] An "interface means" is a means of providing instructions to a worker visually and aurally, and supporting efficient learning.

[0673] "Machine learning technology" is an artificial intelligence technology used to analyze motion information and identify differences between the user's actions and the instructor's actions.

[0674] "Exchange means" refers to a means for exchanging information between a terminal and an information processing device with low latency and efficiency.

[0675] The server manages a device that uses motion measurement equipment to measure the user's physical movements in real time and acquire that motion information. This device is equipped with multiple sensors and has the function of accurately capturing data on the user's limbs and posture. The acquired motion data is immediately transmitted to the server via communication means.

[0676] The server analyzes this operational information through an information processing device and compares it with standard operational data from pre-registered instructors. The analysis algorithm uses machine learning techniques to identify differences in operation and determine which areas need improvement. Based on the results, the server generates feedback data specifically indicating the areas for improvement.

[0677] The terminal receives feedback data sent from the server and notifies the user through a display device. A head-mounted display or smart glasses are used as the display device, allowing the user to visually understand areas for improvement by overlaying their own actions with ideal actions within a virtual environment. Furthermore, the terminal can enhance learning efficiency by utilizing interface devices to present a combination of visual and auditory feedback.

[0678] A typical prompt for the entire system would be: "We have installed a new robotic arm in the factory. Please tell us how to build an AR / VR system that provides real-time feedback on the correct operating procedures." This allows for a specific inquiry about improving motion management.

[0679] This system enables workers in the factory work environment to efficiently acquire advanced skills and improve the precision and safety of their operations.

[0680] The flow of a specific process in Application Example 1 will be explained using Figure 12.

[0681] Step 1:

[0682] A motion measurement device captures the user's physical movements as they operate the robot using sensors. The acquired motion information is transmitted to a terminal as detailed data such as the user's position and speed. This allows for real-time input of motion data.

[0683] Step 2:

[0684] The terminal transmits the received operational data to the server via a communication means. The server converts this data into the format required for analysis and generates a dataset for analysis. The operational data obtained as input is output to the server's analysis system.

[0685] Step 3:

[0686] The server uses machine learning algorithms to compare user movement data with instructor movement data. Data processing analyzes movement timing, speed, angle, etc., identifying areas for improvement. Analysis results are generated and output as specific areas for improvement.

[0687] Step 4:

[0688] The server generates feedback data on user actions based on the analysis results. This data includes visual improvement instructions and procedures. The generated feedback data is output to a display device.

[0689] Step 5:

[0690] The terminal receives feedback data sent from the server and presents it to the user through a display device. Using a head-mounted display or smart glasses, the user is shown a virtual overlay of the ideal behavior. This allows the user to intuitively understand which parts need improvement.

[0691] Step 6:

[0692] Based on the feedback provided, the user modifies their actions and operates the robot again. The interface provides additional visual and auditory guidance, further enhancing the user's learning. This allows the user to input new actions in response to the feedback.

[0693] Furthermore, an emotion engine that estimates the user's emotions may be incorporated. That is, the identification processing unit 290 may use the emotion identification model 59 to estimate the user's emotions and perform identification processing using the user's emotions.

[0694] The present invention is a system that enables interactive feedback that takes into account the learner's emotional state by combining an emotion engine with an AR / VR teaching system for supporting motion learning. This system includes a motion capture device, a processing device, a display device, a feedback device, a communication means, and an emotion engine.

[0695] The device uses a motion capture system to record the user's body movements in detail in order to collect user motion data in real time. In addition, the emotion engine acquires the user's facial expression data with a camera and analyzes this data to recognize the user's emotions. The emotion engine uses a facial expression analysis algorithm to identify multiple emotions in real time, such as joy, sadness, surprise, anger, fear, disgust, and neutral expression.

[0696] Next, the device sends the collected motion data and emotion data to the server. The server uses the received motion data to perform a comparative analysis with the teacher's actions. Artificial intelligence is used in the analysis to perform a precise comparison of actions and identify specific areas for improvement. Emotion data is also analyzed simultaneously, and feedback is adjusted according to the user's emotions.

[0697] The server then generates feedback on the user's actions based on the analysis results and sends it to the terminal. For visual feedback, an AR / VR device overlays the teacher's ideal actions with the user's actions in the user's field of view. Additionally, voice feedback optimized for the user's emotions is provided. For example, if motivation is low, positive voice assistance is enhanced to increase confidence in improving actions.

[0698] For example, if a user is learning dance moves and the emotion engine detects a decrease in the user's concentration, the device can provide visual feedback along with voice messages of encouragement and relaxation, urging the user to take their time.

[0699] In this way, by considering the user's actions and emotions as an integrated whole, this system provides a more personalized learning experience and greatly supports the user's skill acquisition process.

[0700] The following describes the processing flow.

[0701] Step 1:

[0702] The device acquires user movement data in real time using a motion capture system. This includes collecting positional information of joints and skeleton using sensors or cameras worn by the user.

[0703] Step 2:

[0704] The device acquires user facial expression data through its camera and analyzes it using an emotion engine. It also determines the user's emotional state in real time.

[0705] Step 3:

[0706] The device sends user behavioral data and emotional data to the server. This transmission is performed with low latency, ensuring that data is transferred without compromising real-time capabilities.

[0707] Step 4:

[0708] The server uses the received behavioral data to begin an analysis that compares it to the teacher's behavioral data. This analysis uses artificial intelligence to identify differences and areas for improvement in the behavior in detail.

[0709] Step 5:

[0710] In parallel, the server analyzes data from the emotion engine and determines the tone and content of feedback based on the user's emotion index. If the user's emotions are unstable, it uses a reinforcement learning algorithm to select encouraging or adapted guidance.

[0711] Step 6:

[0712] The server generates feedback information based on the analysis results and sends it to the terminal. The feedback is provided visually and audibly and is customized to enhance user motivation.

[0713] Step 7:

[0714] The device provides feedback to the user via an AR / VR device. The user can view a video in which the teacher's movement model and their own movements are overlaid, and make corrections as needed. Voice assistance is also provided in parallel, allowing the user to receive guidance on specific ways to improve.

[0715] (Example 2)

[0716] Next, we will describe Example 2. In the following description, the data processing device 12 will be referred to as the "server" and the robot 414 as the "terminal".

[0717] When learners acquire a skill, simply comparing the accuracy of the skill makes it difficult to provide appropriate instruction tailored to each individual's learning progress and emotional state. Furthermore, standardized feedback can make it difficult for learners to maintain motivation.

[0718] The identification process performed by the identification processing unit 290 of the data processing device 12 in Example 2 is realized by the following means.

[0719] In this invention, the server includes an action acquisition device, a processing device, and an emotion recognition device. This allows for more effective action learning by not only comparing the user's action information with standard actions, but also providing feedback that takes into account the user's emotional state.

[0720] A "motion acquisition device" is a device that collects the user's body movements in real time and plays a role in acquiring specific motion information of the user.

[0721] A "calculation unit" is a device that compares and analyzes user operation information with standard operation information, performing calculations to identify deviations and areas for improvement in operation.

[0722] A "display means" is a means of providing information visually and displaying the user's actions overlaid on a virtual environment based on the analysis results.

[0723] An "information provision device" is a device that provides users with feedback on areas for improvement in their operation, issuing specific instructions and information to support their learning.

[0724] An "emotion recognition device" is a device that identifies a user's emotional state from their facial expressions, and it acquires emotional information using an emotion analysis algorithm.

[0725] A "feedback adjustment device" is a device that optimizes the content and format of the feedback provided based on the user's emotional state.

[0726] A "communication device" is a device that enables low-latency information communication between a terminal and an information processing device, allowing for high-speed and stable data exchange.

[0727] This system is designed to effectively support the user's motion learning and is configured as follows: First, the terminal collects the user's motion data in real time using a motion acquisition device. This motion acquisition device includes cameras and sensors to accurately capture the user's body movements. Furthermore, the terminal utilizes an emotion recognition device to acquire the user's facial expression data via the camera and analyzes their emotional state from it.

[0728] The collected data is transmitted to the server via a low-latency communication device. The server uses a computing unit to perform calculations to compare the user's behavioral data with standard behavioral data. Machine learning techniques are applied to these calculations to identify deviations and areas for improvement. Simultaneously, the emotional data obtained by the emotion recognition device is analyzed by a feedback adjustment device to generate feedback tailored to the user's emotional state.

[0729] The generated feedback is transmitted to the terminal via an information provider and provided to the user. Visual feedback utilizes display means, overlaying the user's actions with ideal actions in a virtual environment. Emotion-based voice feedback is also provided, playing a role in increasing user motivation.

[0730] For example, if a user practicing dance wants to improve the timing of their movements, the system can collect motion data for each step, analyze the differences, and provide feedback. At the same time, if the user's concentration is waning, it can provide a voice message encouraging them to relax.

[0731] Examples of prompts for a generative AI model:

[0732] "Please describe a system that generates appropriate feedback when a user is losing focus while practicing dance moves."

[0733] In this way, this system not only promotes the improvement of users' operational skills but also provides emotional support, resulting in a more fulfilling learning experience.

[0734] The flow of the specific processing in Example 2 will be explained using Figure 13.

[0735] Step 1:

[0736] The terminal uses a motion acquisition device to collect user motion data in real time. The input is the user's physical movements, which are captured by cameras and sensors and converted into numerical data such as the position and angle of the movements. The output is detailed motion data for analysis. During this process, the distance the user moves and the angles of their joints are specifically recorded.

[0737] Step 2:

[0738] The device uses an emotion recognition device to acquire and analyze the user's facial expression data. The input is image data of the user's face, which is applied to an analysis algorithm to extract emotional states such as joy and sadness. The output is an emotion label and data indicating the intensity of that emotion. Specifically, the degree of the user's smile and eyebrow movements are analyzed.

[0739] Step 3:

[0740] The terminal sends the acquired behavioral and emotional data to the server. The input consists of the previously collected behavioral and emotional data, and the output is the generation of data packets to be sent to the server. In this step, the communication device plays the role of transmitting the data with low latency.

[0741] Step 4:

[0742] The server uses a computing device to compare and analyze operational data with standard operational data. The input consists of user operational data and pre-prepared standard operational data. Machine learning is used to calculate the differences in operation and identify areas for improvement. The output is an analysis showing the deviations in operation. Specifically, timing delays and movement precision are quantified.

[0743] Step 5:

[0744] The server generates feedback using a feedback adjustment device based on emotional data. The input is an emotional label and its intensity, which is used to construct a feedback message optimized for the user's emotional state. The output is the adjusted feedback message, which includes positive content and, where necessary, words of encouragement.

[0745] Step 6:

[0746] The server sends the generated feedback to the terminal. The terminal provides this feedback to the user using visual and auditory means. The input is feedback data from the server, and the output is a visual overlay display and audio instructions. Specifically, this involves visualization of actions using AR / VR devices and audio guidance using speakers.

[0747] Step 7:

[0748] The user receives feedback and attempts to improve the operation. The input is feedback from the device, and the output is the improved operation. In this process, the user can learn by following specific instructions and correcting the operation.

[0749] (Application Example 2)

[0750] Next, we will explain application example 2. In the following explanation, the data processing device 12 will be referred to as the "server" and the robot 414 as the "terminal".

[0751] In the industrial sector, efficient and safe operation of workers and machinery requires both advanced skills acquisition and emotional support. However, conventional technologies struggle to provide personalized feedback in real time that considers the emotional state of individual workers, in addition to evaluating their movements. Therefore, there is a need for more effective and motivating methods.

[0752] The specific processing performed by the specific processing unit 290 of the data processing device 12 in Application Example 2 is realized by the following means.

[0753] In this invention, the server includes processing means for comparing and examining user behavior information and instructor behavior information, emotion analysis means for analyzing user emotional information, and adjustment means for optimizing output means considering the results of the emotional information analysis. This enables feedback that combines highly accurate evaluation of behavior with emotional support.

[0754] A "detection device" is a device used to record operations and changes in real time.

[0755] The "processing means" refers to a function that analyzes acquired motion information and compares it with a standard motion.

[0756] A "display means" is a device that visually superimposes actions in a virtual space based on the results of the study.

[0757] An "output means" is a device that provides the user with suggestions for improving the operation and feedback.

[0758] An "emotional analysis tool" is a function that analyzes the user's emotions and identifies their state.

[0759] A "modification mechanism" is a system for optimizing the content of feedback based on analyzed emotional information.

[0760] A "communication method" is a system for sending and receiving information with low latency between a terminal device and an information processing device.

[0761] The system in this invention primarily analyzes the actions and emotions of workers in real time and provides personalized feedback. In this system, the server, terminal, and user function as follows:

[0762] First, the terminal acquires the worker's movements in real time using a motion capture device. Examples of motion capture devices include Arduino and Kinect. The terminal also acquires facial expression data via a camera, collecting basic data for analyzing the user's emotions.

[0763] Next, the terminal sends the acquired motion data and emotion data to the server. On the server side, the motion data is analyzed using processing equipment. Specifically, a machine learning algorithm is used to compare the user's motion information with the instructor's motion information to identify accuracy and areas for improvement. In addition, emotion analysis equipment is used to analyze facial expression data and evaluate the user's emotional state.

[0764] Based on these analysis results, the server generates optimized feedback for the output method. This feedback is provided visually on the AR / VR device through the display method. Furthermore, the adjustment method generates audio feedback tailored to the user's emotional state, effectively supporting their motivation.

[0765] To give a concrete example, imagine a scenario in a training session for assembling a new automobile model, where workers are learning precise movements. In this case, it would be possible to capture the workers' movements in detail, evaluate the accuracy of those movements, and then provide visual and audio guidance for the actual movements.

[0766] When using a generative AI model, examples of prompt statements include the following:

[0767] "Propose a program for training new car assembly that provides optimal movement guidance and generates feedback tailored to the worker's emotions. Consider worker movement and emotional data to specifically identify feedback that enhances motivation."

[0768] The flow of a specific process in Application Example 2 will be explained using Figure 14.

[0769] Step 1:

[0770] The device uses a motion capture system to acquire user movement data in real time. The input is raw data from the motion capture sensor, and the output is a stream of movement data. The movement involves recording the user's body movements in detail and converting them into a format that can be transferred as data.

[0771] Step 2:

[0772] The device uses a camera to acquire facial expression data and performs analysis using an emotion analysis tool. The input is image data acquired from the camera, and the output is the analyzed emotional state. Specifically, the device passes the image data through a facial expression analysis algorithm to identify emotions such as joy, anger, sadness, and happiness in real time.

[0773] Step 3:

[0774] The terminal transmits acquired behavioral and emotional data to the server. The input is a data stream related to behavior and emotion, and the output is a successful transmission of data to the server. The terminal efficiently compresses the data and transfers it to the server using a low-latency communication protocol.

[0775] Step 4:

[0776] The server uses processing tools to compare and analyze the user's motion data with the instructor's motion data. The input consists of the user's motion data and existing baseline data, and the output is the extraction of motion similarities and areas for improvement. The server applies machine learning algorithms to perform motion evaluation.

[0777] Step 5:

[0778] The server evaluates facial expression data using emotion analysis tools and determines the emotional state. The input is processed facial expression data, and the output is the user's emotional state and its intensity. The server uses an emotion recognition algorithm to identify emotional patterns.

[0779] Step 6:

[0780] The server optimizes output methods based on analysis results to generate feedback. Inputs are the results of behavioral comparisons and sentiment evaluations, while output is detailed guidance in the form of visual and auditory feedback. The server customizes the content to ensure the feedback is delivered to the user most effectively.

[0781] Step 7:

[0782] The terminal provides feedback received from the server to the user using display means. Input is feedback data, and output is visual guidance via AR / VR devices and audio instructions via speakers. The terminal overlays the instructor's actions with the user's actions to highlight areas for improvement.

[0783] The specific processing unit 290 transmits the result of the specific processing to the robot 414. In the robot 414, the control unit 46A causes the speaker 240 and the controlled object 443 to output the result of the specific processing. The microphone 238 acquires audio indicating user input for the result of the specific processing. The control unit 46A transmits the audio data indicating user input acquired by the microphone 238 to the data processing unit 12. In the data processing unit 12, the specific processing unit 290 acquires the audio data.

[0784] Data generation model 58 is a type of so-called generative AI (Artificial Intelligence). One example of data generation model 58 is ChatGPT (Internet Search).<URL: https: / / openai.com / blog / chatgpt> ), Gemini (Internet search) <url: https: gemini.google.com ?hl="ja">Examples of generative AI include the following. The data generation model 58 is obtained by performing deep learning on a neural network. The data generation model 58 is input with prompts containing instructions, and with inference data such as audio data representing speech, text data representing text, and image data representing images. The data generation model 58 infers from the input inference data according to the instructions indicated by the prompts, and outputs the inference results in data formats such as audio data and text data. Here, inference refers to, for example, analysis, classification, prediction, and / or summarization.

[0785] In the above embodiment, an example was given in which specific processing is performed by the data processing device 12, but the technology of this disclosure is not limited thereto, and the specific processing may also be performed by the robot 414.

[0786] Furthermore, the emotion identification model 59, acting as an emotion engine, may determine the user's emotion according to a specific mapping. Specifically, the emotion identification model 59 may determine the user's emotion according to a specific mapping, which is an emotion map (see Figure 9). Similarly, the emotion identification model 59 may also determine the robot's emotion, and the identification processing unit 290 may perform identification processing using the robot's emotion.

[0787] Figure 9 shows an emotion map 400 in which multiple emotions are mapped. In the emotion map 400, emotions are arranged in concentric circles radiating from the center. The closer to the center of the concentric circles, the more primitive the emotions are located. Further out of the concentric circles, emotions representing states and actions arising from mental states are located. Emotion is a concept that includes feelings and mental states. On the left side of the concentric circles, emotions that are generally generated from reactions occurring in the brain are located. On the right side of the concentric circles, emotions that are generally induced by situational judgment are located. Above and below the concentric circles, emotions that are generally generated from reactions occurring in the brain and induced by situational judgment are located. In addition, the emotion of "pleasure" is located on the upper side of the concentric circles, and the emotion of "displeasure" is located on the lower side. Thus, in the emotion map 400, multiple emotions are mapped based on the structure in which emotions arise, and emotions that are likely to occur simultaneously are mapped close together.

[0788] These emotions are distributed at the 3 o'clock position on the Emotion Map 400, and usually fluctuate between feelings of security and anxiety. In the right half of the Emotion Map 400, situational awareness takes precedence over internal feelings, resulting in a calm impression.

[0789] The inside of the Emotion Map 400 represents inner thoughts, while the outside represents actions. Therefore, the further you go from the outside of the Emotion Map 400, the more visible (expressed in actions) your emotions become.

[0790] Here, human emotions are based on various balances, such as posture and blood sugar levels. When these balances deviate from the ideal, it results in discomfort, and when they approach the ideal, it results in pleasure. Similarly, in robots, cars, motorcycles, etc., emotions can be created based on various balances, such as posture and battery level. When these balances deviate from the ideal, it results in discomfort, and when they approach the ideal, it results in pleasure. The emotion map can be generated, for example, based on Dr. Mitsuyoshi's emotion map (Research on a system for analyzing brain physiological signals of speech emotion recognition and emotion, Tokushima University, doctoral dissertation: https: / / ci.nii.ac.jp / naid / 500000375379). The left half of the emotion map contains emotions belonging to a region called "response," where sensation is dominant. The right half of the emotion map contains emotions belonging to a region called "situation," where situational awareness is dominant.

[0791] The emotion map defines two emotions that promote learning. One is the emotion around the middle of the negative "repentance" and "reflection" on the situation side. In other words, it is when the robot experiences negative emotions such as "I never want to feel this way again" or "I don't want to be scolded again." The other is the emotion around the positive "desire" on the reaction side. In other words, it is when the robot has positive feelings such as "I want more" or "I want to know more."

[0792] The emotion identification model 59 inputs user input into a pre-trained neural network, obtains emotion values ​​representing each emotion shown in the emotion map 400, and determines the user's emotion. This neural network is pre-trained based on multiple training data sets, which are combinations of user input and emotion values ​​representing each emotion shown in the emotion map 400. Furthermore, this neural network is trained so that emotions located close together have similar values, as shown in the emotion map 900 in Figure 10. Figure 10 shows an example where multiple emotions such as "reassured," "calm," and "confident" have similar emotion values.

[0793] The above description primarily focuses on the functions of the data processing device 12 in relation to this disclosure. However, the system related to this disclosure is not necessarily implemented on a server. The system related to this disclosure may be implemented as a general information processing system. This disclosure may be implemented, for example, as a software program that runs on a personal computer or as an application that runs on a smartphone. The method related to this disclosure may be provided to users in SaaS (Software as a Service) format.

[0794] In the above embodiment, an example was given in which a specific process is performed by a single computer 22. However, the technology of this disclosure is not limited thereto, and a distributed processing of the specific process may be performed by multiple computers, including computer 22. For example, a data generation model 58 may be provided in an external device of the data processing device 12, and the external device may generate data according to the input data.

[0795] In the above embodiment, an example was given in which the specific processing program 56 is stored in the storage 32, but the technology of this disclosure is not limited thereto. For example, the specific processing program 56 may be stored in a portable, computer-readable, non-temporary storage medium such as a USB (Universal Serial Bus) memory. The specific processing program 56 stored in the non-temporary storage medium is installed in the computer 22 of the data processing device 12. The processor 28 executes specific processing according to the specific processing program 56.

[0796] Alternatively, the specific processing program 56 may be stored in a storage device such as a server connected to the data processing device 12 via the network 54, and the specific processing program 56 may be downloaded and installed on the computer 22 in response to a request from the data processing device 12.

[0797] Furthermore, it is not necessary to store the entirety of the specific processing program 56 in a storage device such as a server connected to the data processing device 12 via the network 54, or to store the entirety of the specific processing program 56 in the storage 32; it is acceptable to store only a portion of the specific processing program 56.

[0798] The following types of processors can be used as hardware resources to perform specific processing. Examples of processors include a CPU, a general-purpose processor that functions as a hardware resource to perform specific processing by executing software, i.e., a program. Other examples of processors include dedicated electrical circuits, such as FPGAs (Field-Programmable Gate Arrays), PLDs (Programmable Logic Devices), or ASICs (Application Specific Integrated Circuits), which have circuit configurations specifically designed to perform specific processing. All of these processors have built-in or connected memory, and all of them perform specific processing by using memory.

[0799] The hardware resource that performs a specific process may consist of one of these various processors, or it may consist of a combination of two or more processors of the same or different types (for example, a combination of multiple FPGAs, or a combination of a CPU and an FPGA). Alternatively, the hardware resource that performs a specific process may consist of a single processor.

[0800] Examples of configurations using a single processor include, firstly, a configuration in which one or more CPUs and software are combined to form a single processor, and this processor functions as a hardware resource that performs a specific process. Secondly, there is a configuration using a processor that realizes the functions of the entire system, including multiple hardware resources that perform a specific process, on a single IC chip, as exemplified by SoCs (System-on-a-chip). In this way, a specific process is realized using one or more of the above types of processors as hardware resources.

[0801] Furthermore, the hardware structure of these various processors can more specifically utilize electrical circuits that combine circuit elements such as semiconductor devices. Also, the specific processing described above is merely an example. Therefore, it goes without saying that unnecessary steps can be deleted, new steps added, or the processing order rearranged, as long as it does not deviate from the main purpose.

[0802] The descriptions and illustrations presented above are detailed explanations of the technical aspects of this disclosure and are merely examples of the technical aspects. For example, the above descriptions of the structure, function, operation, and effect are examples of the structure, function, operation, and effect of the technical aspects of this disclosure. Therefore, it goes without saying that you may delete unnecessary parts, add new elements, or replace elements in the descriptions and illustrations presented above, as long as you do not deviate from the essence of the technical aspects of this disclosure. Furthermore, in order to avoid confusion and facilitate understanding of the technical aspects of this disclosure, explanations of common technical knowledge and the like that do not require special explanation to enable the implementation of the technical aspects of this disclosure have been omitted from the descriptions and illustrations presented above.

[0803] All documents, patent applications, and technical standards described herein are incorporated by reference to the same extent as if each individual document, patent application, and technical standard were specifically and individually noted as being incorporated by reference.

[0804] The following is further disclosed regarding the embodiments described above.

[0805] (Claim 1)

[0806] A motion capture device for capturing movements in real time,

[0807] A processing unit that compares and analyzes user behavior data and teacher behavior data,

[0808] A display device that overlays and displays the user's actions in a virtual space based on the analysis results,

[0809] A feedback device that provides users with feedback on areas for improvement in operation,

[0810] A system that includes this.

[0811] (Claim 2)

[0812] The system according to claim 1, characterized in that it uses artificial intelligence for comparative analysis of actions.

[0813] (Claim 3)

[0814] The system according to claim 1, characterized by including communication means for performing low-latency data communication between a terminal and a server.

[0815] "Example 1"

[0816] (Claim 1)

[0817] A motion detection device for detecting motion in real time,

[0818] An information processing device that compares and analyzes the user's movement information and the instructor's movement information,

[0819] A presentation device that superimposes the user's actions onto an abstract space based on the analysis results,

[0820] A reporting device that provides users with information on areas for improvement in operation,

[0821] A system that includes this.

[0822] (Claim 2)

[0823] The system according to claim 1, characterized in that it employs an intelligent algorithm for comparative analysis of actions.

[0824] (Claim 3)

[0825] The system according to claim 1, characterized by including communication means for performing low-latency information communication between a terminal and an information management device.

[0826] "Application Example 1"

[0827] (Claim 1)

[0828] A motion measurement device for measuring motion in real time,

[0829] An information processing device that compares and analyzes the user's movement information and the instructor's movement information,

[0830] A display means that overlays the user's actions onto a virtual environment based on the analysis results,

[0831] A notification means to inform the user of improvements to the operation,

[0832] An interface means that provides instructions to workers visually and aurally,

[0833] A system that includes this.

[0834] (Claim 2)

[0835] The system according to claim 1, characterized in that machine learning technology is used for comparative analysis of actions.

[0836] (Claim 3)

[0837] The system according to claim 1, characterized by including an exchange means for performing low-latency information exchange between a terminal and an information processing device.

[0838] "Example 2 of combining an emotion engine"

[0839] (Claim 1)

[0840] A motion acquisition device for acquiring motion in real time,

[0841] A computing device that compares and analyzes user operation information and standard operation information,

[0842] A display means that overlays the user's actions onto a virtual environment based on the comparison results,

[0843] An information provision device that provides users with suggestions for improving the operation,

[0844] An emotion recognition device that recognizes the emotional state from the user's facial expressions,

[0845] A feedback adjustment device that adjusts feedback based on the user's emotional state,

[0846] A system that includes this.

[0847] (Claim 2)

[0848] The system according to claim 1, characterized in that machine learning is applied to the comparative analysis of actions.

[0849] (Claim 3)

[0850] The system according to claim 1, characterized by including a communication device that performs low-latency information communication between a terminal and an information processing device.

[0851] "Application example 2 of combining emotional engines"

[0852] (Claim 1)

[0853] A detection device for recording motion in real time,

[0854] A processing means for comparing and examining the user's action information and the instructor's action information,

[0855] A display means that overlays the user's actions onto a virtual space based on the results of the examination,

[0856] An output means for providing feedback to the user on areas for improvement in operation,

[0857] A means for analyzing the emotional information of users,

[0858] An adjustment means for optimizing the output means considering the results of the analysis of emotional information,

[0859] A system that includes this.

[0860] (Claim 2)

[0861] The system according to claim 1, characterized in that machine learning is used for comparative analysis of operations.

[0862] (Claim 3)

[0863] The system according to claim 1, characterized by including communication means for performing low-latency information communication between a terminal device and an information processing device. [Explanation of Symbols]

[0864] 10, 210, 310, 410 Data Processing Systems 12 Data Processing Devices 14 Smart Devices 214 Smart Glasses 314 Headset-type terminal 414 Robots< / url:> < / url:> < / url:> < / url:>

Claims

1. A motion capture device for capturing movements in real time, A processing unit that compares and analyzes user behavior data and teacher behavior data, A display device that overlays and displays the user's actions in a virtual space based on the analysis results, A feedback device that provides users with feedback on areas for improvement in operation, A system that includes this.

2. The system according to claim 1, characterized in that it uses artificial intelligence for comparative analysis of actions.

3. The system according to claim 1, characterized by including communication means for performing low-latency data communication between a terminal and a server.

Citation Information

Patent Citations

  • Persona chatbot control method and system

    JP2022180282A