system
A system analyzes user exercise videos to provide accurate feedback by comparing feature points with professional data, improving movement accuracy and reducing injury risk.
Patent Information
- Authority / Receiving Office
- JP · JP
- Patent Type
- Applications
- Current Assignee / Owner
- SOFTBANK GROUP CORP
- Filing Date
- 2024-10-18
- Publication Date
- 2026-05-01
AI Technical Summary
Users often practice exercises incorrectly due to limited methods for confirming the accuracy of their movements, leading to potential injuries or inefficiencies.
A system that analyzes user video data by extracting feature points and comparing them with professional data, providing specific feedback through multimodal technology and machine learning to improve movement accuracy.
Enables users to receive precise feedback for improving their movements, enhancing skill acquisition and reducing the risk of injuries through detailed analysis.
Smart Images

Figure 2026073414000001_ABST
Abstract
Description
Technical Field
[0001] The technology of the present disclosure relates to a system.
Background Art
[0002] Patent Document 1 discloses a method for controlling a persona chatbot, which is performed by at least one processor, and includes steps of receiving a user utterance, adding the user utterance to a prompt including an instruction sentence related to an explanation of a character of the chatbot, encoding the prompt, and inputting the encoded prompt into a language model to generate a chatbot utterance in response to the user utterance.
Prior Art Documents
Patent Documents
[0003]
Patent Document 1
Summary of the Invention
Problems to be Solved by the Invention
[0004] Currently, many users are doing exercises and training in their own ways while referring to online videos, but there are limited ways to confirm whether their movements are accurate. For this reason, there is a high possibility of repeating incorrect forms and inefficient practice methods, which may cause injuries or ineffective results. The purpose of the present invention is to provide a method for users to easily analyze their own practice videos and obtain professional advice.
Means for Solving the Problems
[0005] This invention provides a system that receives video data captured by a user, extracts feature points from it, and analyzes differences in movement by comparing it with other specialized video data. Furthermore, based on this comparison, it generates and provides specific advice to the user for improving their movements. By using multimodal technology, it achieves detailed analysis utilizing audio and text data, and machine learning models enable highly accurate feedback.
[0006] A "user" refers to an individual who uses this system to upload video data of their own movement and receives analysis results and advice.
[0007] "Video data" refers to digital data in video format filmed by the user, which records the details of their physical movements.
[0008] A "feature point" is an identifiable point within video data that indicates a specific position, angle, or part of a movement, and refers to numerical information that forms the basis of motion analysis.
[0009] "Comparison" refers to the process of comparing feature points extracted from the user's video data with feature points extracted from other specialized video data, and identifying the differences between the two.
[0010] "Advice" refers to information that includes specific guidance and recommendations for improving user behavior, generated based on the results of video data comparison.
[0011] "Multimodal technology" is a technical method that integrates and analyzes multiple different types of data, such as video, audio, and text, to provide more precise information.
[0012] A "machine learning model" is an algorithm used for data analysis. It is a statistical model that learns patterns from past data and makes predictions or judgments about new data. [Brief explanation of the drawing]
[0013] [Figure 1] This is a conceptual diagram showing an example of the configuration of a data processing system according to the first embodiment. [Figure 2] This is a conceptual diagram showing an example of the essential functions of a data processing device and a smart device according to the first embodiment. [Figure 3] This is a conceptual diagram showing an example of the configuration of a data processing system according to the second embodiment. [Figure 4] This is a conceptual diagram showing an example of the main functions of a data processing device and smart glasses according to the second embodiment. [Figure 5] This is a conceptual diagram showing an example of the configuration of a data processing system according to the third embodiment. [Figure 6] This is a conceptual diagram showing an example of the main functions of a data processing device and a headset-type terminal according to the third embodiment. [Figure 7] This is a conceptual diagram showing an example of the configuration of a data processing system according to the fourth embodiment. [Figure 8] This is a conceptual diagram showing an example of the main functions of a data processing device and a robot according to the fourth embodiment. [Figure 9] This shows an emotion map where multiple emotions are mapped. [Figure 10] This shows an emotion map where multiple emotions are mapped. [Figure 11] This is a sequence diagram showing the processing flow of the data processing system in Example 1. [Figure 12] This is a sequence diagram showing the processing flow of the data processing system in Application Example 1. [Figure 13] This is a sequence diagram showing the processing flow of the data processing system in Example 2, which incorporates an emotion engine. [Figure 14] This is a sequence diagram showing the processing flow of the data processing system in Application Example 2, which combines an emotion engine. [Modes for carrying out the invention]
[0014] An example of an embodiment of a system according to the technology of the present disclosure will be described below with reference to the accompanying drawings.
[0015] First, the terms used in the following description will be explained.
[0016] In the following embodiments, a numbered processor (hereinafter simply referred to as "processor") may be a single arithmetic unit or a combination of multiple arithmetic units. Also, the processor may be a single type of arithmetic unit or a combination of multiple types of arithmetic units. Examples of arithmetic units include a CPU (Central Processing Unit), a GPU (Graphics Processing Unit), a GPGPU (General-Purpose computing on Graphics Processing Units), an APU (Accelerated Processing Unit), and the like.
[0017] In the following embodiments, a numbered RAM (Random Access Memory) is a memory in which information is temporarily stored and is used as a work memory by the processor.
[0018] In the following embodiments, a numbered storage is one or more non-volatile storage devices that store various programs, various parameters, and the like. Examples of non-volatile storage devices include flash memory (SSD (Solid State Drive)), magnetic disks (e.g., hard disks), or magnetic tapes, and the like.
[0019] In the following embodiments, the signed communication interface (I / F) is an interface that includes a communication processor and an antenna, etc. The communication interface manages communication between multiple computers. Examples of communication standards applicable to the communication interface include wireless communication standards such as 5G (5th Generation Mobile Communication System), Wi-Fi (registered trademark), or Bluetooth (registered trademark).
[0020] In the following embodiments, "A and / or B" is synonymous with "at least one of A and B." That is, "A and / or B" means that it may be A alone, or B alone, or a combination of A and B. Furthermore, in this specification, the same concept as "A and / or B" applies when expressing three or more things linked by "and / or."
[0021] [First Embodiment]
[0022] Figure 1 shows an example of the configuration of the data processing system 10 according to the first embodiment.
[0023] As shown in Figure 1, the data processing system 10 includes a data processing device 12 and a smart device 14. An example of the data processing device 12 is a server.
[0024] The data processing device 12 comprises a computer 22, a database 24, and a communication interface 26. The computer 22 is an example of a "computer" related to the technology of this disclosure. The computer 22 comprises a processor 28, RAM 30, and storage 32. The processor 28, RAM 30, and storage 32 are connected to a bus 34. The database 24 and the communication interface 26 are also connected to the bus 34. The communication interface 26 is connected to a network 54. An example of the network 54 is a WAN (Wide Area Network) and / or a LAN (Local Area Network).
[0025] The smart device 14 comprises a computer 36, a reception device 38, an output device 40, a camera 42, and a communication interface 44. The computer 36 comprises a processor 46, RAM 48, and storage 50. The processor 46, RAM 48, and storage 50 are connected to a bus 52. The reception device 38, output device 40, and camera 42 are also connected to the bus 52.
[0026] The reception device 38 is equipped with a touch panel 38A and a microphone 38B, etc., and receives user input. The touch panel 38A receives user input by detecting contact with an object (e.g., a pen or finger). The microphone 38B receives user input by detecting the user's voice. The control unit 46A transmits data indicating the user input received by the touch panel 38A and microphone 38B to the data processing device 12. In the data processing device 12, the specific processing unit 290 acquires the data indicating the user input.
[0027] The output device 40 includes a display 40A and a speaker 40B, and presents data to the user 20 by outputting the data in a form perceptible to the user 20 (e.g., audio and / or text). The display 40A displays visible information such as text and images according to instructions from the processor 46. The speaker 40B outputs audio according to instructions from the processor 46. The camera 42 is a small digital camera equipped with an optical system such as a lens, aperture, and shutter, and an image sensor such as a CMOS (Complementary Metal-Oxide-Semiconductor) image sensor or a CCD (Charge Coupled Device) image sensor.
[0028] Communication interface 44 is connected to network 54. Communication interfaces 44 and 26 are responsible for the exchange of various types of information between processor 46 and processor 28 via network 54.
[0029] Figure 2 shows an example of the main functions of the data processing device 12 and the smart device 14.
[0030] As shown in Figure 2, in the data processing device 12, a specific processing is performed by the processor 28. A specific processing program 56 is stored in the storage 32. The specific processing program 56 is an example of a "program" related to the technology of this disclosure. The processor 28 reads the specific processing program 56 from the storage 32 and executes the read specific processing program 56 on the RAM 30. The specific processing is realized by the processor 28 operating as a specific processing unit 290 according to the specific processing program 56 executed on the RAM 30.
[0031] The storage 32 stores the data generation model 58 and the emotion identification model 59. The data generation model 58 and the emotion identification model 59 are used by the identification processing unit 290.
[0032] In the smart device 14, the processor 46 performs the reception output processing. The storage 50 stores the reception output program 60. The reception output program 60 is used in conjunction with a specific processing program 56 by the data processing system 10. The processor 46 reads the reception output program 60 from the storage 50 and executes the read reception output program 60 on the RAM 48. The reception output processing is realized by the processor 46 operating as a control unit 46A according to the reception output program 60 executed on the RAM 48.
[0033] Next, the specific processing performed by the specific processing unit 290 of the data processing device 12 will be described. In the following description, the data processing device 12 will be referred to as the "server" and the smart device 14 as the "terminal".
[0034] This invention is a system for analyzing and providing feedback on motor movements. The process begins with the user uploading video data of their own movements to a server via a terminal. The server extracts feature points to analyze the received video data. Feature points are data such as coordinates, angles, and velocities that characterize elements of movement. Next, the server references reference video data from professionals within the same category and compares it to the user's feature points.
[0035] The server utilizes AI and multimodal technologies to assess the quality of the user's actions based on these comparison results and generates specific improvement advice. This advice can be in text format or, if necessary, supplementary visual guides. For example, it might generate feedback such as, "Your arm position at the top of your swing is insufficient. Try to keep it more horizontal." The server sends this feedback to the user's device, allowing them to review their actions and identify areas for improvement.
[0036] In particular, by utilizing multimodal technology, it is possible to perform a more comprehensive movement evaluation by incorporating audio and subtitle information in addition to video analysis. Based on the feedback, users can learn specifically which parts they need to correct in their next training session, enabling more efficient improvement. This system can be applied not only to sports but also to learning dance and other physical movements, supporting users in acquiring appropriate skills even through self-study.
[0037] The following describes the processing flow.
[0038] Step 1:
[0039] The user prepares video data of their exercise movements using their device and launches the system's application. The user uploads the recorded video to the server via the application. During the upload, the user selects the target exercise category.
[0040] Step 2:
[0041] The server securely stores the received video data and prepares it for analysis. The server uses image processing techniques to identify keypoints for each frame in order to extract the motion features contained in this data.
[0042] Step 3:
[0043] The server compares the extracted feature points with feature points from professional reference video data stored within the same category. Here, AI is used to quantify the degree of similarity and differences between the two sets of movements.
[0044] Step 4:
[0045] The server analyzes the comparison results and generates specific feedback on the user's actions. The generated advice includes points for improvement and guidelines on correct form. The feedback is provided not only in text format but also with visual aids as needed.
[0046] Step 5:
[0047] The server sends the generated feedback to the user's device and notifies the user in response. The user receives the feedback on their device and can use it to improve future training.
[0048] (Example 1)
[0049] Next, we will describe Example 1. In the following description, the data processing device 12 will be referred to as the "server," and the smart device 14 will be referred to as the "terminal."
[0050] Conventional exercise analysis systems have the challenge of making it difficult for users to receive specific feedback that accurately helps them improve their exercise. Furthermore, they often fail to comprehensively utilize diverse information in exercise evaluation, making it difficult to provide useful advice to users.
[0051] The identification process performed by the identification processing unit 290 of the data processing device 12 in Example 1 is realized by the following means.
[0052] In this invention, the server includes a medium for receiving video data including motion capture by the user, a medium for extracting feature points such as body coordinates and angles from the video data, a medium for comparing the feature points with other reference video data, a medium for generating suggestions for improving the user's motion based on the comparison results using deep learning technology, and a medium for providing the suggestions to the user. This allows the user to obtain more comprehensive and specific information for improving their motion, enabling them to efficiently improve their skills.
[0053] The term "user" refers to a person who accesses the system and receives analysis and feedback on their motor movements.
[0054] "Video data" refers to digital data that includes visual information of movement captured by the user.
[0055] "Feature points" are sets of data extracted from video data to characterize body movements, such as coordinates and angles.
[0056] "Reference video data" refers to reference data provided by professionals or other entities for use as a point of comparison.
[0057] "Deep learning technology" is a technique used in the field of artificial intelligence that utilizes multi-layer neural networks to analyze and predict data.
[0058] A "suggestion" refers to specific instructions or advice generated based on comparison results and other factors, aimed at improving user behavior.
[0059] "Multimodal technology" refers to technologies for integrating and analyzing data in different formats (e.g., audio, video, and text information).
[0060] This invention realizes a motion analysis system, which primarily relies on the cooperation of a server and a terminal. First, the user uses a terminal such as a smartphone or camera to record their movements. This terminal is responsible for saving the video data of the movements as a file and uploading it to the server via the internet.
[0061] The server is responsible for processing the received video data. This analysis utilizes techniques such as extracting feature points from the video, including body coordinates and angle information related to movement, using libraries like OpenPose. The extracted feature points are then compared with reference video data. This reference data includes professional movements stored in a database.
[0062] The server uses deep learning technology during the comparison process. This allows it to evaluate user behavior through the comparison of feature points, and the generative AI model generates specific suggestions for improving user behavior. These suggestions are output as text and, in some cases, as visual guides, and are ultimately sent from the server to the terminal.
[0063] Users can receive suggestions provided on their device and use the feedback to improve their exercise. For example, if a user wants to improve their golf swing, they can enter "How can I improve this golf swing form?" as a prompt and receive specific advice based on the result.
[0064] This system allows users to efficiently improve their skills while receiving feedback on their movements. Furthermore, multimodal technology complements the analysis, enabling comprehensive movement analysis that incorporates audio and text information in addition to video data. This makes it easier for users to learn various forms of movement, from sports to dance.
[0065] The flow of the specific processing in Example 1 will be explained using Figure 11.
[0066] Step 1:
[0067] The user saves video data of their exercise to their device. Specifically, they use the video recording function of their smartphone or camera to film themselves exercising and save the file to the device's storage. This video data becomes the input for the next process.
[0068] Step 2:
[0069] The terminal uploads the stored video data to the server via the internet. Specifically, the user uses a dedicated application for the system to send video files to the server. This operation provides the video data as input for server processing.
[0070] Step 3:
[0071] The server extracts movement-related physical feature points from the received video data. Specifically, it uses video analysis techniques to extract data about body joints and posture. This involves digital image processing using software such as OpenPose. The output of this processing is feature point data for analysis.
[0072] Step 4:
[0073] The server compares the extracted feature points with reference video data. It refers to a reference database containing examples of actions performed by professionals and calculates the distance and positional differences between feature points. This allows it to evaluate how well the user's actions match the reference. Deep learning models are used to improve the accuracy of the evaluation. This evaluation data then serves as input for the next step.
[0074] Step 5:
[0075] The server uses a generated AI model based on the evaluation results to generate specific suggestions for improving the user's performance. Specifically, it inputs prompt sentences into the AI model and generates improvement advice such as, "Your arm position at the top of your swing is insufficient. Try to keep it more horizontal." The generated suggestions are output as text and visual guides.
[0076] Step 6:
[0077] The server sends the generated suggestions to the device. Users can review the feedback displayed on their device and use it to improve their exercise. Specifically, they read the feedback on the device screen and consciously try to correct it during their next workout. This feedback supports the user's skill improvement.
[0078] (Application Example 1)
[0079] Next, we will explain Application Example 1. In the following explanation, the data processing device 12 will be referred to as the "server," and the smart device 14 will be referred to as the "terminal."
[0080] In modern factories, maximizing robot operational efficiency and minimizing errors is a crucial challenge for improving productivity. However, current monitoring and feedback systems are inadequate, making it difficult to pinpoint detailed areas for improvement in robotic movement. Furthermore, real-time capabilities are often lacking in the process of comparing and analyzing movements. This results in delays in flexible responses and immediate improvements in the manufacturing process.
[0081] The specific processing performed by the specific processing unit 290 of the data processing device 12 in Application Example 1 is realized by the following means.
[0082] In this invention, the server includes a device for receiving video data containing motion capture by a user, a function for extracting feature points from the video data, and a function for comparing the feature points with reference data. This enables real-time analysis of the movements of robots used in factories, allowing for specific guidance to improve efficiency and reduce errors.
[0083] A "user" is an individual or organization that uses a system to analyze and improve their own behavior and actions.
[0084] "Motor activity" refers to all bodily movements and mechanical operations performed with a specific purpose.
[0085] "Video data" refers to data containing visual information recorded by cameras, sensors, and other devices.
[0086] A "device" refers to hardware or software designed to perform a specific function.
[0087] A "feature point" is data containing specific coordinates, angles, velocities, and other information extracted to represent the characteristics of movement or action.
[0088] "Reference data" refers to data that records professional or standard behavior and is used for comparison and evaluation.
[0089] "Comparison" is the act of comparing two or more different sets of data or pieces of information to identify their similarities and differences.
[0090] "Efficiency" refers to the degree of ability to achieve a goal while minimizing resources and time.
[0091] "Error" is a term that refers to an unexpected malfunction or discrepancy in a system or its operation.
[0092] "Guidance" refers to specific advice and suggestions for improvement and optimization provided by the system.
[0093] This invention is a system that uses video data of movement captured by a user to provide feedback for improving the efficiency and precision of that movement. The server uses the following technologies for this purpose.
[0094] First, the user uses their device to capture video data from a camera or sensor and uploads it to the server. A dedicated application is installed on the device, and the data is reliably and efficiently transmitted to the server through this application. Specifically, video processing libraries such as OpenCV are used to capture the video and extract its features.
[0095] Next, the server extracts feature points from the received video data and performs a comparison. Here, learning algorithms, particularly deep learning and generative AI models, are used to evaluate the extracted feature points against reference data. This process is performed in real time, enabling immediate feedback.
[0096] Feedback is generated as numerical data and visual guides and sent to the terminal. In this process, a generating AI model is used as an aid, providing appropriate advice to the user based on the prompts it generates. A specific example is providing instructions to optimize the movement of a robotic arm on a factory production line.
[0097] Examples of prompts for a generative AI model include the following:
[0098] "Factory Robot Motion Analysis: Analyze the motion for efficient bottle picking and provide feedback. Include specific areas for improvement regarding the arm's movement path and speed."
[0099] This system allows users to receive real-time feedback, enabling them to quickly improve the efficiency and precision of their exercise.
[0100] The flow of a specific process in Application Example 1 will be explained using Figure 12.
[0101] Step 1:
[0102] The user uses a device to record their movements. The recorded video data is captured in real time using video processing libraries such as OpenCV. The input is the video acquired from the camera, and this data is sent directly to the server. The output is the video data uploaded to the server.
[0103] Step 2:
[0104] The server extracts feature points from the received video data. Given video data as input, image processing techniques are used to calculate feature points such as coordinates, angles, and velocities that represent the characteristics of motion in each frame. This results in a dataset of the analyzed feature points as output.
[0105] Step 3:
[0106] The server compares the extracted feature points with the reference data. The reference data and feature point datasets serve as input, and a learning algorithm is used to analyze the similarities and differences between them. The output is feedback data indicating the quality of the operation as an evaluation result.
[0107] Step 4:
[0108] The server generates feedback based on the evaluation results. Using an AI model, particularly a generative AI model, it generates specific areas for improvement based on prompts. The input is the evaluation results, and the output is feedback provided to the user in the form of text or visual guides.
[0109] Step 5:
[0110] The terminal receives feedback from the server and provides it to the user. Feedback data is sent to the terminal as input, and the information is presented to the user visually or audibly through a dedicated application. The output is a presentation of areas for improvement and advice that the user can review.
[0111] Furthermore, an emotion engine that estimates the user's emotions may be incorporated. That is, the identification processing unit 290 may use the emotion identification model 59 to estimate the user's emotions and perform identification processing using the user's emotions.
[0112] The system based on the present invention receives the user's movement as video data and generates advice based on it, and in addition, combines this with an emotion engine that recognizes the user's emotions to provide more personalized feedback.
[0113] The user first uses a device to record their movement and uploads the video data to the server. The server receives the video data and stores it securely. The server extracts characteristic points of the movement from the video data and compares them with other reference movement data. Machine learning models are used for the comparison to achieve highly accurate movement analysis.
[0114] The server further uses an emotion engine to understand the user's emotional state. This engine has algorithms that analyze the user's facial expressions from video or recognize emotions from audio data. The recognized emotional state of the user is reflected in the adjustment of feedback to the user. For example, if the server detects that the user is feeling stressed, it can provide positive feedback to boost their motivation.
[0115] The generated advice includes specific points for improvement and is sensitive to the user's feelings. The server sends this to the terminal, allowing the user to review the feedback. The user then uses this feedback to plan how to improve their performance in the next exercise session.
[0116] As an example, let's assume a user is practicing their golf swing. This system evaluates the consistency of the swing from video footage and can also recognize whether the user is feeling confused or anxious from their facial expressions and voice in the video. As a result, the server provides feedback such as, "Your swing is improving well. Next, try to relax and focus more on your rhythm."
[0117] This invention allows users to receive feedback tailored to their individual emotional state, going beyond simple motion analysis, and further enhance the effectiveness of their training.
[0118] The following describes the processing flow.
[0119] Step 1:
[0120] The user uses a device to film their exercise movements and uploads the video data to a server via a dedicated app. At this time, the user selects the appropriate category according to the type of movement.
[0121] Step 2:
[0122] The server stores the received video data and extracts feature points of motion. These feature points are identified from each frame using image processing techniques. This allows the specific elements of the motion to be quantified.
[0123] Step 3:
[0124] The server compares feature points extracted from the user's video with feature points from professional video data registered as a reference. This comparison is performed by a machine learning model, which calculates the degree of similarity and differences for each action.
[0125] Step 4:
[0126] The server uses an emotion engine to analyze the user's facial expressions and voice data to recognize their emotional state. The emotions the user displays in the video (e.g., concentration, confusion, joy) are detected in this step.
[0127] Step 5:
[0128] The server generates customized advice based on the comparison results of the actions and the user's emotional state. For example, it might create feedback such as, "Relaxing will make your arm movements smoother."
[0129] Step 6:
[0130] The server generates feedback which is then provided to the user via the terminal. The user can review the feedback on the terminal and use it to improve their next training session.
[0131] (Example 2)
[0132] Next, we will describe Example 2. In the following description, the data processing device 12 will be referred to as the "server" and the smart device 14 as the "terminal".
[0133] When attempting to improve motor skills, there is a challenge in providing personalized feedback that takes into account each user's emotional state. Furthermore, there is a need to integrate and analyze diverse data formats to achieve more accurate motion analysis and advice generation.
[0134] The identification process performed by the identification processing unit 290 of the data processing device 12 in Example 2 is realized by the following means.
[0135] In this invention, the server includes means for receiving visual information, including motor movements acquired by the user; means for extracting characteristic information from the visual information; and means for comparing standard video data with the characteristic information. This makes it possible to provide optimal feedback tailored to the individual user's characteristics and emotional state.
[0136] A "user" refers to a person who uses the system to analyze their own movement patterns and receive advice for improvement.
[0137] "Visual information" refers to image or video data that includes recordings of movement, and includes details of the movements captured by the user.
[0138] "Characteristic information" refers to the characteristic points and movement patterns of motion extracted from visual information, and this constitutes the basic data for analyzing motion.
[0139] "Standard video data" refers to video data of past motions used as a reference, and is used for comparison with characteristic information.
[0140] A "machine learning algorithm" refers to a group of mathematical methods used to learn patterns from large amounts of data and analyze their behavioral characteristics.
[0141] "Emotion recognition means" encompasses methods and technologies for analyzing a user's emotional state and are used to recognize a user's emotions from visual and auditory information.
[0142] "Advice" refers to information generated after considering motion analysis and emotional state, which includes specific instructions and suggestions for the user to improve their motor skills.
[0143] This invention is a system that aims to allow users to analyze their own motor movements in detail and receive feedback based on their individual emotional state. Specific embodiments of this system are described below.
[0144] The user first films their own movements using a device. This device can be any recording device, including a smartphone or camera. The captured visual information is uploaded from the device to a server via the internet.
[0145] The server extracts characteristic information from the received visual information. This involves using image processing libraries such as OpenCV to detect feature points for detailed analysis of user movements. Next, machine learning algorithms are used to compare the extracted characteristic information with pre-prepared standard video data. In this comparison process, artificial intelligence frameworks such as TENSORFLOW® are used to analyze and evaluate user movements.
[0146] Furthermore, the server uses video and audio data to identify the user's emotional state. This incorporates a speech recognition engine and facial expression analysis algorithm as means of emotion recognition. This analysis provides emotional information, such as whether the user is feeling tense or anxious, which influences the feedback.
[0147] The generated feedback includes specific advice for improving the user's movements, as well as emotionally responsive psychological support. For example, it might say, "Analysis of your golf swing video shows that your swing is improving well. Next time, try to relax and focus more on your rhythm." This allows the user to make precise improvements to their movements.
[0148] An example of a prompt message might be: "Consider the analysis results of the user's golf swing video and their emotional state, and generate improvement advice. The user is performing a swing and appears confused. The feedback should be positive and motivating."
[0149] As described above, this invention can more effectively enhance athletic ability and mental stability by analyzing the individual user's motor movements and providing personalized feedback based on their emotional state.
[0150] The flow of the specific processing in Example 2 will be explained using Figure 13.
[0151] Step 1:
[0152] The user uses a device to record their movement. The input at this time is visual information recording the user's movement. The device temporarily stores this information and prepares to upload it to the server. The output is the visual information ready to be sent to the server.
[0153] Step 2:
[0154] The server receives visual information sent from the terminal and stores it in a database. The input is visual information from the user, which the server receives and stores securely. The output is visual information ready for analysis.
[0155] Step 3:
[0156] The server extracts characteristic information from the received visual information. At this stage, the input is stored visual information, and the server uses an image processing library such as OpenCV to detect feature points of motion from the image data. The output is characteristic information including the feature points.
[0157] Step 4:
[0158] The server compares extracted characteristic information with standard video data. The input consists of characteristic information and standard video data. The server uses machine learning algorithms such as TensorFlow to compare these and analyze the consistency of the operation and areas requiring improvement. The output is performance evaluation data based on the analysis results.
[0159] Step 5:
[0160] The server uses emotion recognition to determine the user's emotional state. Visual and auditory information are used as input, and the server utilizes an emotion recognition engine to analyze the user's emotions. The output is data on the user's psychological state obtained from their facial expressions and voice.
[0161] Step 6:
[0162] The server generates feedback using a generative AI model based on performance evaluation data and emotional state. The input consists of analysis results and psychological state data. Prompt sentences are input to the generative AI model, which generates specific, emotionally sensitive advice tailored to the user. The output is the generated feedback message.
[0163] Step 7:
[0164] The server sends the generated feedback to the terminal. The input is the feedback message. The output is the feedback message that the user receives through the terminal. This allows the user to review the feedback and use it to improve their next exercise session.
[0165] (Application Example 2)
[0166] Next, we will explain application example 2. In the following explanation, the data processing device 12 will be referred to as a "server" and the smart device 14 as a "terminal".
[0167] In modern fitness and training environments, there is a demand for rapid and appropriate personalized feedback to individual users. However, conventional systems struggle to understand not only a user's movement but also their emotional state and provide feedback that takes those emotions into account. Furthermore, there is a need for systems that can individually adjust advice on areas for improvement in user movement in real time.
[0168] The specific processing performed by the specific processing unit 290 of the data processing device 12 in Application Example 2 is realized by the following means.
[0169] In this invention, the server includes means for receiving video information including motion capture by the user, means for extracting motion characteristics from the video information, means for comparing the motion characteristics with other video information, and means for analyzing the user's emotional state and adjusting the advice based on the results. This makes it possible to effectively provide individual users with advice for motion improvement while taking their emotional state into consideration.
[0170] "Video information" refers to visual data that records movement actions captured by the user, and is used for motion analysis.
[0171] "Movement characteristics" refer to various data points related to the body's posture and movement during athletic activity, and these are used as the basis for analyzing the movement.
[0172] A "means of comparison" refers to a method for evaluating the quality of a user's actions by comparing the extracted characteristics of their actions with other reference data.
[0173] "Advice" refers to guidance information provided to improve the user's motor skills, including specific areas for improvement and content designed to boost motivation.
[0174] "Emotional state" refers to a psychological state inferred from the user's facial expressions, voice, etc., and reflects emotions such as stress, anxiety, and enjoyment during exercise.
[0175] "Means of adjustment" refer to methods for modifying the content and method of providing advice according to the user's emotional state, thereby achieving more effective feedback.
[0176] The system for implementing this invention begins with the user recording their movements with a smartphone or other device while exercising and uploading the video information to a server. The server receives the video information and extracts the characteristics of the movement from it. Specifically, it uses a machine learning model to calculate data points such as posture, angle, and speed of the movement. The software used mainly consists of machine learning libraries such as TensorFlow and PyTorch.
[0177] Furthermore, the server analyzes the user's emotional state from the video information. This analysis uses the Emotion API from Microsoft® Azure® Cognitive Services, based on the user's facial expressions and voice data. This analysis assesses the user's psychological state during exercise, such as their stress levels and relaxation levels.
[0178] Based on this data, the server generates advice for the user. This advice specifically points out areas for improvement and uses considerate language to enhance motivation. Finally, the generated advice is sent to the user's device, allowing them to refer to the feedback during their next exercise session.
[0179] As a concrete example, let's say there's a user practicing their golf swing. This user films their swing and sends the video information to a server. The server analyzes the swing form, and if it detects through emotion analysis that the user is confused, it provides advice such as, "Your swing form is good. Next time, try to relax and continue practicing while focusing on your rhythm."
[0180] An example of a prompt message is: "Analyze your exercise form and emotional state, and provide specific areas for improvement and advice on how to relax."
[0181] The flow of a specific process in Application Example 2 will be explained using Figure 14.
[0182] Step 1:
[0183] Users record their exercise using their smartphones. During this process, video information is captured on the device. The device's camera function is used to thoroughly record every detail of the movement.
[0184] Step 2:
[0185] The terminal uploads the captured video information to the server. The terminal transfers the video data to the server via the network connection. The input is the video data, and the output is the server confirming receipt.
[0186] Step 3:
[0187] The server extracts motion characteristics from the received video information. Using motion analysis software, it analyzes important posture and movement patterns in the video using machine learning algorithms (TensorFlow or PyTorch). The input is video data, and the output is feature data.
[0188] Step 4:
[0189] The server compares feature data with reference data. It uses past motion data as a reference to analyze how well it matches current motion. The input is feature data and reference data, and the output is the comparison result.
[0190] Step 5:
[0191] The server analyzes emotional states from video information. It uses the Microsoft Azure Cognitive Services Emotion API to infer user emotions from video and audio. Input is video and audio data, and output is emotional state data.
[0192] Step 6:
[0193] The server generates advice using comparison results and emotional state data. It creates advice that combines points for performance improvement with positive feedback tailored to the emotional state. The input is the comparison results and emotional state data, and the output is the generated advice.
[0194] Step 7:
[0195] The server sends the generated advice to the user's device. This advice is forwarded to the device as feedback, allowing the user to utilize it during their next exercise session. The input is the generated advice, and the output is the display of the advice on the device.
[0196] The specific processing unit 290 transmits the result of the specific processing to the smart device 14. In the smart device 14, the control unit 46A causes the output device 40 to output the result of the specific processing. The microphone 38B acquires audio indicating user input for the result of the specific processing. The control unit 46A transmits the audio data indicating user input acquired by the microphone 38B to the data processing device 12. In the data processing device 12, the specific processing unit 290 acquires the audio data.
[0197] Data generation model 58 is a so-called generative AI (Artificial Intelligence). An example of data generation model 58 is ChatGPT (registered trademark) (Internet search).<URL: https: / / openai.com / blog / chatgpt> ), Gemini (registered trademark) (Internet search) <url: https: gemini.google.com ?hl="ja">Examples of generative AI include the following. The data generation model 58 is obtained by performing deep learning on a neural network. The data generation model 58 is input with prompts containing instructions, and with inference data such as audio data representing speech, text data representing text, and image data representing images. The data generation model 58 infers from the input inference data according to the instructions indicated by the prompts, and outputs the inference results in data formats such as audio data and text data. Here, inference refers to, for example, analysis, classification, prediction, and / or summarization.
[0198] In the above embodiment, an example was given in which specific processing is performed by the data processing device 12, but the technology of this disclosure is not limited thereto, and the specific processing may also be performed by the smart device 14.
[0199] [Second Embodiment]
[0200] Figure 3 shows an example of the configuration of the data processing system 210 according to the second embodiment.
[0201] As shown in Figure 3, the data processing system 210 includes a data processing device 12 and smart glasses 214. An example of the data processing device 12 is a server.
[0202] The data processing device 12 comprises a computer 22, a database 24, and a communication interface 26. The computer 22 is an example of a "computer" related to the technology of this disclosure. The computer 22 comprises a processor 28, RAM 30, and storage 32. The processor 28, RAM 30, and storage 32 are connected to a bus 34. The database 24 and the communication interface 26 are also connected to the bus 34. The communication interface 26 is connected to a network 54. An example of the network 54 is a WAN (Wide Area Network) and / or a LAN (Local Area Network).
[0203] The smart glasses 214 include a computer 36, a microphone 238, a speaker 240, a camera 42, and a communication interface 44. The computer 36 includes a processor 46, RAM 48, and storage 50. The processor 46, RAM 48, and storage 50 are connected to a bus 52. The microphone 238, speaker 240, and camera 42 are also connected to the bus 52.
[0204] The microphone 238 receives voice signals from the user 20 and receives instructions from the user 20. The microphone 238 captures the voice signals from the user 20, converts the captured voice into audio data, and outputs it to the processor 46. The speaker 240 outputs audio according to the instructions from the processor 46.
[0205] Camera 42 is a small digital camera equipped with an optical system including a lens, aperture, and shutter, and an image sensor such as a CMOS (Complementary Metal-Oxide-Semiconductor) image sensor or a CCD (Charge Coupled Device) image sensor, and captures images of the area around the user 20 (for example, an imaging range defined by a field of view equivalent to the width of a typical healthy person's field of vision).
[0206] Communication interface 44 is connected to network 54. Communication interfaces 44 and 26 are responsible for the exchange of various information between processor 46 and processor 28 via network 54. The exchange of various information between processor 46 and processor 28 using communication interfaces 44 and 26 is performed in a secure manner.
[0207] Figure 4 shows an example of the main functions of the data processing device 12 and the smart glasses 214. As shown in Figure 4, the data processing device 12 performs specific processing using the processor 28. The storage 32 stores the specific processing program 56.
[0208] The specific processing program 56 is an example of a "program" relating to the technology of this disclosure. The processor 28 reads the specific processing program 56 from the storage 32 and executes the read specific processing program 56 on the RAM 30. The specific processing is realized by the processor 28 operating as a specific processing unit 290 in accordance with the specific processing program 56 executed on the RAM 30.
[0209] The storage 32 stores the data generation model 58 and the emotion identification model 59. The data generation model 58 and the emotion identification model 59 are used by the identification processing unit 290.
[0210] In the smart glasses 214, the processor 46 performs the reception output processing. The storage 50 stores the reception output program 60. The processor 46 reads the reception output program 60 from the storage 50 and executes the read reception output program 60 on the RAM 48. The reception output processing is realized by the processor 46 operating as a control unit 46A according to the reception output program 60 executed on the RAM 48.
[0211] Next, the identification processing performed by the identification processing unit 290 of the data processing device 12 will be described. In the following description, the data processing device 12 will be referred to as the "server" and the smart glasses 214 will be referred to as the "terminal".
[0212] This invention is a system for analyzing and providing feedback on motor movements. The process begins with the user uploading video data of their own movements to a server via a terminal. The server extracts feature points to analyze the received video data. Feature points are data such as coordinates, angles, and velocities that characterize elements of movement. Next, the server references reference video data from professionals within the same category and compares it to the user's feature points.
[0213] The server utilizes AI and multimodal technologies to assess the quality of the user's actions based on these comparison results and generates specific improvement advice. This advice can be in text format or, if necessary, supplementary visual guides. For example, it might generate feedback such as, "Your arm position at the top of your swing is insufficient. Try to keep it more horizontal." The server sends this feedback to the user's device, allowing them to review their actions and identify areas for improvement.
[0214] In particular, by utilizing multimodal technology, it is possible to perform a more comprehensive movement evaluation by incorporating audio and subtitle information in addition to video analysis. Based on the feedback, users can learn specifically which parts they need to correct in their next training session, enabling more efficient improvement. This system can be applied not only to sports but also to learning dance and other physical movements, supporting users in acquiring appropriate skills even through self-study.
[0215] The following describes the processing flow.
[0216] Step 1:
[0217] The user prepares video data of their exercise movements using their device and launches the system's application. The user uploads the recorded video to the server via the application. During the upload, the user selects the target exercise category.
[0218] Step 2:
[0219] The server securely stores the received video data and prepares it for analysis. The server uses image processing techniques to identify keypoints for each frame in order to extract the motion features contained in this data.
[0220] Step 3:
[0221] The server compares the extracted feature points with feature points from professional reference video data stored within the same category. Here, AI is used to quantify the degree of similarity and differences between the two sets of movements.
[0222] Step 4:
[0223] The server analyzes the comparison results and generates specific feedback on the user's actions. The generated advice includes points for improvement and guidelines on correct form. The feedback is provided not only in text format but also with visual aids as needed.
[0224] Step 5:
[0225] The server sends the generated feedback to the user's device and notifies the user in response. The user receives the feedback on their device and can use it to improve future training.
[0226] (Example 1)
[0227] Next, we will describe Example 1. In the following description, the data processing device 12 will be referred to as the "server," and the smart glasses 214 will be referred to as the "terminal."
[0228] Conventional exercise analysis systems have the challenge of making it difficult for users to receive specific feedback that accurately helps them improve their exercise. Furthermore, they often fail to comprehensively utilize diverse information in exercise evaluation, making it difficult to provide useful advice to users.
[0229] The identification process performed by the identification processing unit 290 of the data processing device 12 in Example 1 is realized by the following means.
[0230] In this invention, the server includes a medium for receiving video data including motion capture by the user, a medium for extracting feature points such as body coordinates and angles from the video data, a medium for comparing the feature points with other reference video data, a medium for generating suggestions for improving the user's motion based on the comparison results using deep learning technology, and a medium for providing the suggestions to the user. This allows the user to obtain more comprehensive and specific information for improving their motion, enabling them to efficiently improve their skills.
[0231] The term "user" refers to a person who accesses the system and receives analysis and feedback on their motor movements.
[0232] "Video data" refers to digital data that includes visual information of movement captured by the user.
[0233] "Feature points" are sets of data extracted from video data to characterize body movements, such as coordinates and angles.
[0234] "Reference video data" refers to reference data provided by professionals or other entities for use as a point of comparison.
[0235] "Deep learning technology" is a technique used in the field of artificial intelligence that utilizes multi-layer neural networks to analyze and predict data.
[0236] A "suggestion" refers to specific instructions or advice generated based on comparison results and other factors, aimed at improving user behavior.
[0237] "Multimodal technology" refers to technologies for integrating and analyzing data in different formats (e.g., audio, video, and text information).
[0238] This invention realizes a motion analysis system, which primarily relies on the cooperation of a server and a terminal. First, the user uses a terminal such as a smartphone or camera to record their movements. This terminal is responsible for saving the video data of the movements as a file and uploading it to the server via the internet.
[0239] The server is responsible for processing the received video data. This analysis utilizes techniques such as extracting feature points from the video, including body coordinates and angle information related to movement, using libraries like OpenPose. The extracted feature points are then compared with reference video data. This reference data includes professional movements stored in a database.
[0240] The server uses deep learning technology during the comparison process. This allows it to evaluate user behavior through the comparison of feature points, and the generative AI model generates specific suggestions for improving user behavior. These suggestions are output as text and, in some cases, as visual guides, and are ultimately sent from the server to the terminal.
[0241] Users can receive suggestions provided on their device and use the feedback to improve their exercise. For example, if a user wants to improve their golf swing, they can enter "How can I improve this golf swing form?" as a prompt and receive specific advice based on the result.
[0242] This system allows users to efficiently improve their skills while receiving feedback on their movements. Furthermore, multimodal technology complements the analysis, enabling comprehensive movement analysis that incorporates audio and text information in addition to video data. This makes it easier for users to learn various forms of movement, from sports to dance.
[0243] The flow of the specific processing in Example 1 will be explained using Figure 11.
[0244] Step 1:
[0245] The user saves video data of their exercise to their device. Specifically, they use the video recording function of their smartphone or camera to film themselves exercising and save the file to the device's storage. This video data becomes the input for the next process.
[0246] Step 2:
[0247] The terminal uploads the stored video data to the server via the internet. Specifically, the user uses a dedicated application for the system to send video files to the server. This operation provides the video data as input for server processing.
[0248] Step 3:
[0249] The server extracts movement-related physical feature points from the received video data. Specifically, it uses video analysis techniques to extract data about body joints and posture. This involves digital image processing using software such as OpenPose. The output of this processing is feature point data for analysis.
[0250] Step 4:
[0251] The server compares the extracted feature points with reference video data. It refers to a reference database containing examples of actions performed by professionals and calculates the distance and positional differences between feature points. This allows it to evaluate how well the user's actions match the reference. Deep learning models are used to improve the accuracy of the evaluation. This evaluation data then serves as input for the next step.
[0252] Step 5:
[0253] The server uses a generated AI model based on the evaluation results to generate specific suggestions for improving the user's performance. Specifically, it inputs prompt sentences into the AI model and generates improvement advice such as, "Your arm position at the top of your swing is insufficient. Try to keep it more horizontal." The generated suggestions are output as text and visual guides.
[0254] Step 6:
[0255] The server sends the generated suggestions to the device. Users can review the feedback displayed on their device and use it to improve their exercise. Specifically, they read the feedback on the device screen and consciously try to correct it during their next workout. This feedback supports the user's skill improvement.
[0256] (Application Example 1)
[0257] Next, we will explain Application Example 1. In the following explanation, the data processing device 12 will be referred to as the "server," and the smart glasses 214 will be referred to as the "terminal."
[0258] In modern factories, maximizing robot operational efficiency and minimizing errors is a crucial challenge for improving productivity. However, current monitoring and feedback systems are inadequate, making it difficult to pinpoint detailed areas for improvement in robotic movement. Furthermore, real-time capabilities are often lacking in the process of comparing and analyzing movements. This results in delays in flexible responses and immediate improvements in the manufacturing process.
[0259] The specific processing performed by the specific processing unit 290 of the data processing device 12 in Application Example 1 is realized by the following means.
[0260] In this invention, the server includes a device for receiving video data containing motion capture by a user, a function for extracting feature points from the video data, and a function for comparing the feature points with reference data. This enables real-time analysis of the movements of robots used in factories, allowing for specific guidance to improve efficiency and reduce errors.
[0261] A "user" is an individual or organization that uses a system to analyze and improve their own behavior and actions.
[0262] "Motor activity" refers to all bodily movements and mechanical operations performed with a specific purpose.
[0263] "Video data" refers to data containing visual information recorded by cameras, sensors, and other devices.
[0264] A "device" refers to hardware or software designed to perform a specific function.
[0265] A "feature point" is data containing specific coordinates, angles, velocities, and other information extracted to represent the characteristics of movement or action.
[0266] "Reference data" refers to data that records professional or standard behavior and is used for comparison and evaluation.
[0267] "Comparison" is the act of comparing two or more different sets of data or pieces of information to identify their similarities and differences.
[0268] "Efficiency" refers to the degree of ability to achieve a goal while minimizing resources and time.
[0269] "Error" is a term that refers to an unexpected malfunction or discrepancy in a system or its operation.
[0270] "Guidance" refers to specific advice and suggestions for improvement and optimization provided by the system.
[0271] This invention is a system that uses video data of movement captured by a user to provide feedback for improving the efficiency and precision of that movement. The server uses the following technologies for this purpose.
[0272] First, the user uses their device to capture video data from a camera or sensor and uploads it to the server. A dedicated application is installed on the device, and the data is reliably and efficiently transmitted to the server through this application. Specifically, video processing libraries such as OpenCV are used to capture the video and extract its features.
[0273] Next, the server extracts feature points from the received video data and performs a comparison. Here, learning algorithms, particularly deep learning and generative AI models, are used to evaluate the extracted feature points against reference data. This process is performed in real time, enabling immediate feedback.
[0274] Feedback is generated as numerical data and visual guides and sent to the terminal. In this process, a generating AI model is used as an aid, providing appropriate advice to the user based on the prompts it generates. A specific example is providing instructions to optimize the movement of a robotic arm on a factory production line.
[0275] Examples of prompts for a generative AI model include the following:
[0276] "Factory Robot Motion Analysis: Analyze the motion for efficient bottle picking and provide feedback. Include specific areas for improvement regarding the arm's movement path and speed."
[0277] This system allows users to receive real-time feedback, enabling them to quickly improve the efficiency and precision of their exercise.
[0278] The process of the specific processing in Application Example 1 will be described using FIG. 12.
[0279] Step 1:
[0280] The user uses the terminal to capture a motion action. The captured video data is captured in real time by utilizing a video processing library such as OpenCV. The input is the video obtained from the camera, and this data is directly sent to the server. The output is the video data uploaded to the server.
[0281] Step 2:
[0282] The server extracts feature points from the received video data. The video data is given as the input, and using image processing technology, feature points such as coordinates, angles, and speeds indicating the characteristics of the movement are calculated from each frame. As a result, a dataset of the analyzed feature points is obtained as the output.
[0283] Step 3:
[0284] The server compares the extracted feature points with the reference data. The reference data and the dataset of the feature points are used as the input, and the similarity and differences between the two are analyzed using a learning algorithm. The output is feedback data indicating the quality of the movement as the evaluation result.
[0285] Step 4:
[0286] The server generates feedback based on the evaluation result. Using an AI model, especially a generative AI model, specific improvement points are generated based on the prompt text. The input is the evaluation result, and the output is the feedback in the form of text or visual guidance provided to the user.
[0287] Step 5:
[0288] The terminal receives feedback from the server and provides it to the user. Feedback data is sent to the terminal as input, and the information is presented to the user visually or audibly through a dedicated application. The output is a presentation of areas for improvement and advice that the user can review.
[0289] Furthermore, an emotion engine that estimates the user's emotions may be incorporated. That is, the identification processing unit 290 may use the emotion identification model 59 to estimate the user's emotions and perform identification processing using the user's emotions.
[0290] The system based on the present invention receives the user's movement as video data and generates advice based on it, and in addition, combines this with an emotion engine that recognizes the user's emotions to provide more personalized feedback.
[0291] The user first uses a device to record their movement and uploads the video data to the server. The server receives the video data and stores it securely. The server extracts characteristic points of the movement from the video data and compares them with other reference movement data. Machine learning models are used for the comparison to achieve highly accurate movement analysis.
[0292] The server further uses an emotion engine to understand the user's emotional state. This engine has algorithms that analyze the user's facial expressions from video or recognize emotions from audio data. The recognized emotional state of the user is reflected in the adjustment of feedback to the user. For example, if the server detects that the user is feeling stressed, it can provide positive feedback to boost their motivation.
[0293] The generated advice includes specific points for improvement and is sensitive to the user's feelings. The server sends this to the terminal, allowing the user to review the feedback. The user then uses this feedback to plan how to improve their performance in the next exercise session.
[0294] As an example, let's assume a user is practicing their golf swing. This system evaluates the consistency of the swing from video footage and can also recognize whether the user is feeling confused or anxious from their facial expressions and voice in the video. As a result, the server provides feedback such as, "Your swing is improving well. Next, try to relax and focus more on your rhythm."
[0295] This invention allows users to receive feedback tailored to their individual emotional state, going beyond simple motion analysis, and further enhance the effectiveness of their training.
[0296] The following describes the processing flow.
[0297] Step 1:
[0298] The user uses a device to film their exercise movements and uploads the video data to a server via a dedicated app. At this time, the user selects the appropriate category according to the type of movement.
[0299] Step 2:
[0300] The server stores the received video data and extracts feature points of motion. These feature points are identified from each frame using image processing techniques. This allows the specific elements of the motion to be quantified.
[0301] Step 3:
[0302] The server compares feature points extracted from the user's video with feature points from professional video data registered as a reference. This comparison is performed by a machine learning model, which calculates the degree of similarity and differences for each action.
[0303] Step 4:
[0304] The server analyzes the user's facial expressions and voice data using an emotion engine to recognize the emotional state. The emotions shown by the user in the video (e.g., concentration, confusion, joy) are detected at this step.
[0305] Step 5:
[0306] Based on the comparison result of the actions and the user's emotional state, the server generates customized advice. For example, feedback such as "Your arm movements will become smoother by relaxing" is created.
[0307] Step 6:
[0308] The server provides the feedback generated to the user through the terminal. The user can check the feedback on the terminal and apply it to the next training.
[0309] (Example 2)
[0310] Next, Example 2 will be described. In the following description, the data processing device 12 is referred to as the "server", and the smart glasses 214 are referred to as the "terminal".
[0311] When attempting to improve movement actions, there is a problem that it is difficult to provide personalized feedback considering the individual emotional states of users. Also, it is required to integratively analyze various data formats to achieve more accurate movement analysis and advice generation.
[0312] The specific processing by the specific processing unit 290 of the data processing device 12 in Example 2 is realized by the following means.
[0313] In this invention, the server includes means for receiving visual information including the movement actions acquired by the user, means for extracting characteristic information from the visual information, and means for comparing based on the standard video data and the characteristic information. Thereby, it becomes possible to provide optimal feedback according to the characteristics and emotional states of individual users.
[0314] A "user" refers to a person who uses the system to analyze their own movement patterns and receive advice for improvement.
[0315] "Visual information" refers to image or video data that includes recordings of movement, and includes details of the movements captured by the user.
[0316] "Characteristic information" refers to the characteristic points and movement patterns of motion extracted from visual information, and this constitutes the basic data for analyzing motion.
[0317] "Standard video data" refers to video data of past motions used as a reference, and is used for comparison with characteristic information.
[0318] A "machine learning algorithm" refers to a group of mathematical methods used to learn patterns from large amounts of data and analyze their behavioral characteristics.
[0319] "Emotion recognition means" encompasses methods and technologies for analyzing a user's emotional state and are used to recognize a user's emotions from visual and auditory information.
[0320] "Advice" refers to information generated after considering motion analysis and emotional state, which includes specific instructions and suggestions for the user to improve their motor skills.
[0321] This invention is a system that aims to allow users to analyze their own motor movements in detail and receive feedback based on their individual emotional state. Specific embodiments of this system are described below.
[0322] The user first films their own movements using a device. This device can be any recording device, including a smartphone or camera. The captured visual information is uploaded from the device to a server via the internet.
[0323] The server extracts characteristic information from the received visual information. This involves using image processing libraries such as OpenCV to detect feature points for detailed analysis of user movements. Next, machine learning algorithms are used to compare the extracted characteristic information with pre-prepared standard video data. In this comparison process, artificial intelligence frameworks such as TensorFlow are used to analyze and evaluate the user's movements.
[0324] Furthermore, the server uses video and audio data to identify the user's emotional state. This incorporates a speech recognition engine and facial expression analysis algorithm as means of emotion recognition. This analysis provides emotional information, such as whether the user is feeling tense or anxious, which influences the feedback.
[0325] The generated feedback includes specific advice for improving the user's movements, as well as emotionally responsive psychological support. For example, it might say, "Analysis of your golf swing video shows that your swing is improving well. Next time, try to relax and focus more on your rhythm." This allows the user to make precise improvements to their movements.
[0326] An example of a prompt message might be: "Consider the analysis results of the user's golf swing video and their emotional state, and generate improvement advice. The user is performing a swing and appears confused. The feedback should be positive and motivating."
[0327] As described above, this invention can more effectively enhance athletic ability and mental stability by analyzing the individual user's motor movements and providing personalized feedback based on their emotional state.
[0328] The flow of the specific processing in Example 2 will be explained using Figure 13.
[0329] Step 1:
[0330] The user uses a device to record their movement. The input at this time is visual information recording the user's movement. The device temporarily stores this information and prepares to upload it to the server. The output is the visual information ready to be sent to the server.
[0331] Step 2:
[0332] The server receives visual information sent from the terminal and stores it in a database. The input is visual information from the user, which the server receives and stores securely. The output is visual information ready for analysis.
[0333] Step 3:
[0334] The server extracts characteristic information from the received visual information. At this stage, the input is stored visual information, and the server uses an image processing library such as OpenCV to detect feature points of motion from the image data. The output is characteristic information including the feature points.
[0335] Step 4:
[0336] The server compares extracted characteristic information with standard video data. The input consists of characteristic information and standard video data. The server uses machine learning algorithms such as TensorFlow to compare these and analyze the consistency of the operation and areas requiring improvement. The output is performance evaluation data based on the analysis results.
[0337] Step 5:
[0338] The server uses emotion recognition to determine the user's emotional state. Visual and auditory information are used as input, and the server utilizes an emotion recognition engine to analyze the user's emotions. The output is data on the user's psychological state obtained from their facial expressions and voice.
[0339] Step 6:
[0340] The server generates feedback using a generative AI model based on performance evaluation data and emotional state. The input consists of analysis results and psychological state data. Prompt sentences are input to the generative AI model, which generates specific, emotionally sensitive advice tailored to the user. The output is the generated feedback message.
[0341] Step 7:
[0342] The server sends the generated feedback to the terminal. The input is the feedback message. The output is the feedback message that the user receives through the terminal. This allows the user to review the feedback and use it to improve their next exercise session.
[0343] (Application Example 2)
[0344] Next, we will explain application example 2. In the following explanation, the data processing device 12 will be referred to as the "server," and the smart glasses 214 will be referred to as the "terminal."
[0345] In modern fitness and training environments, there is a demand for rapid and appropriate personalized feedback to individual users. However, conventional systems struggle to understand not only a user's movement but also their emotional state and provide feedback that takes those emotions into account. Furthermore, there is a need for systems that can individually adjust advice on areas for improvement in user movement in real time.
[0346] The specific processing performed by the specific processing unit 290 of the data processing device 12 in Application Example 2 is realized by the following means.
[0347] In this invention, the server includes means for receiving video information including motion capture by the user, means for extracting motion characteristics from the video information, means for comparing the motion characteristics with other video information, and means for analyzing the user's emotional state and adjusting the advice based on the results. This makes it possible to effectively provide individual users with advice for motion improvement while taking their emotional state into consideration.
[0348] "Video information" refers to visual data that records movement actions captured by the user, and is used for motion analysis.
[0349] "Movement characteristics" refer to various data points related to the body's posture and movement during athletic activity, and these are used as the basis for analyzing the movement.
[0350] A "means of comparison" refers to a method for evaluating the quality of a user's actions by comparing the extracted characteristics of their actions with other reference data.
[0351] "Advice" refers to guidance information provided to improve the user's motor skills, including specific areas for improvement and content designed to boost motivation.
[0352] "Emotional state" refers to a psychological state inferred from the user's facial expressions, voice, etc., and reflects emotions such as stress, anxiety, and enjoyment during exercise.
[0353] "Means of adjustment" refer to methods for modifying the content and method of providing advice according to the user's emotional state, thereby achieving more effective feedback.
[0354] The system for implementing this invention begins with the user recording their movements with a smartphone or other device while exercising and uploading the video information to a server. The server receives the video information and extracts the characteristics of the movement from it. Specifically, it uses a machine learning model to calculate data points such as posture, angle, and speed of the movement. The software used mainly consists of machine learning libraries such as TensorFlow and PyTorch.
[0355] Furthermore, the server analyzes the user's emotional state from the video information. This analysis uses the Microsoft Azure Cognitive Services Emotion API based on the user's facial expressions and voice data. This analysis assesses the user's psychological state during exercise, such as their stress levels and relaxation levels.
[0356] Based on this data, the server generates advice for the user. This advice specifically points out areas for improvement and uses considerate language to enhance motivation. Finally, the generated advice is sent to the user's device, allowing them to refer to the feedback during their next exercise session.
[0357] As a concrete example, let's say there's a user practicing their golf swing. This user films their swing and sends the video information to a server. The server analyzes the swing form, and if it detects through emotion analysis that the user is confused, it provides advice such as, "Your swing form is good. Next time, try to relax and continue practicing while focusing on your rhythm."
[0358] An example of a prompt message is: "Analyze your exercise form and emotional state, and provide specific areas for improvement and advice on how to relax."
[0359] The flow of a specific process in Application Example 2 will be explained using Figure 14.
[0360] Step 1:
[0361] Users record their exercise using their smartphones. During this process, video information is captured on the device. The device's camera function is used to thoroughly record every detail of the movement.
[0362] Step 2:
[0363] The terminal uploads the captured video information to the server. The terminal transfers the video data to the server via the network connection. The input is the video data, and the output is the server confirming receipt.
[0364] Step 3:
[0365] The server extracts motion characteristics from the received video information. Using motion analysis software, it analyzes important posture and movement patterns in the video using machine learning algorithms (TensorFlow or PyTorch). The input is video data, and the output is feature data.
[0366] Step 4:
[0367] The server compares feature data with reference data. It uses past motion data as a reference to analyze how well it matches current motion. The input is feature data and reference data, and the output is the comparison result.
[0368] Step 5:
[0369] The server analyzes emotional states from video information. It uses the Microsoft Azure Cognitive Services Emotion API to infer user emotions from video and audio. Input is video and audio data, and output is emotional state data.
[0370] Step 6:
[0371] The server generates advice using comparison results and emotional state data. It creates advice that combines points for performance improvement with positive feedback tailored to the emotional state. The input is the comparison results and emotional state data, and the output is the generated advice.
[0372] Step 7:
[0373] The server sends the generated advice to the user's device. This advice is forwarded to the device as feedback, allowing the user to utilize it during their next exercise session. The input is the generated advice, and the output is the display of the advice on the device.
[0374] The specific processing unit 290 transmits the result of the specific processing to the smart glasses 214. In the smart glasses 214, the control unit 46A causes the speaker 240 to output the result of the specific processing. The microphone 238 acquires audio indicating user input for the result of the specific processing. The control unit 46A transmits the audio data indicating user input acquired by the microphone 238 to the data processing unit 12. In the data processing unit 12, the specific processing unit 290 acquires the audio data.
[0375] Data generation model 58 is a type of so-called generative AI (Artificial Intelligence). One example of data generation model 58 is ChatGPT (Internet search<URL: https: / / openai.com / blog / chatgpt> ), Gemini (Internet search) <url: https: gemini.google.com ?hl="ja">Examples of generative AI include the following. The data generation model 58 is obtained by performing deep learning on a neural network. The data generation model 58 is input with prompts containing instructions, and with inference data such as audio data representing speech, text data representing text, and image data representing images. The data generation model 58 infers from the input inference data according to the instructions indicated by the prompts, and outputs the inference results in data formats such as audio data and text data. Here, inference refers to, for example, analysis, classification, prediction, and / or summarization.
[0376] In the above embodiment, an example was given in which specific processing is performed by the data processing device 12, but the technology of this disclosure is not limited thereto, and the specific processing may also be performed by the smart glasses 214.
[0377] [Third Embodiment]
[0378] Figure 5 shows an example of the configuration of the data processing system 310 according to the third embodiment.
[0379] As shown in Figure 5, the data processing system 310 includes a data processing device 12 and a headset terminal 314. An example of the data processing device 12 is a server.
[0380] The data processing device 12 comprises a computer 22, a database 24, and a communication interface 26. The computer 22 is an example of a "computer" related to the technology of this disclosure. The computer 22 comprises a processor 28, RAM 30, and storage 32. The processor 28, RAM 30, and storage 32 are connected to a bus 34. The database 24 and the communication interface 26 are also connected to the bus 34. The communication interface 26 is connected to a network 54. An example of the network 54 is a WAN (Wide Area Network) and / or a LAN (Local Area Network).
[0381] The headset terminal 314 includes a computer 36, a microphone 238, a speaker 240, a camera 42, a communication interface 44, and a display 343. The computer 36 includes a processor 46, RAM 48, and storage 50. The processor 46, RAM 48, and storage 50 are connected to a bus 52. The microphone 238, speaker 240, camera 42, and display 343 are also connected to the bus 52.
[0382] The microphone 238 receives voice signals from the user 20 and receives instructions from the user 20. The microphone 238 captures the voice signals from the user 20, converts the captured voice into audio data, and outputs it to the processor 46. The speaker 240 outputs audio according to the instructions from the processor 46.
[0383] Camera 42 is a small digital camera equipped with an optical system including a lens, aperture, and shutter, and an image sensor such as a CMOS (Complementary Metal-Oxide-Semiconductor) image sensor or a CCD (Charge Coupled Device) image sensor, and captures images of the area around the user 20 (for example, an imaging range defined by a field of view equivalent to the width of a typical healthy person's field of vision).
[0384] Communication interface 44 is connected to network 54. Communication interfaces 44 and 26 are responsible for the exchange of various information between processor 46 and processor 28 via network 54. The exchange of various information between processor 46 and processor 28 using communication interfaces 44 and 26 is performed in a secure manner.
[0385] Figure 6 shows an example of the main functions of the data processing device 12 and the headset terminal 314. As shown in Figure 6, the data processing device 12 performs specific processing using the processor 28. The storage 32 stores the specific processing program 56.
[0386] The specific processing program 56 is an example of a "program" relating to the technology of this disclosure. The processor 28 reads the specific processing program 56 from the storage 32 and executes the read specific processing program 56 on the RAM 30. The specific processing is realized by the processor 28 operating as a specific processing unit 290 in accordance with the specific processing program 56 executed on the RAM 30.
[0387] The storage 32 stores the data generation model 58 and the emotion identification model 59. The data generation model 58 and the emotion identification model 59 are used by the identification processing unit 290.
[0388] In the headset terminal 314, the processor 46 performs the reception output processing. The storage 50 stores the reception output program 60. The processor 46 reads the reception output program 60 from the storage 50 and executes the read reception output program 60 on the RAM 48. The reception output processing is realized by the processor 46 operating as a control unit 46A according to the reception output program 60 executed on the RAM 48.
[0389] Next, the specific processing performed by the specific processing unit 290 of the data processing device 12 will be described. In the following description, the data processing device 12 will be referred to as the "server" and the headset terminal 314 will be referred to as the "terminal".
[0390] This invention is a system for analyzing and providing feedback on motor movements. The process begins with the user uploading video data of their own movements to a server via a terminal. The server extracts feature points to analyze the received video data. Feature points are data such as coordinates, angles, and velocities that characterize elements of movement. Next, the server references reference video data from professionals within the same category and compares it to the user's feature points.
[0391] The server utilizes AI and multimodal technologies to assess the quality of the user's actions based on these comparison results and generates specific improvement advice. This advice can be in text format or, if necessary, supplementary visual guides. For example, it might generate feedback such as, "Your arm position at the top of your swing is insufficient. Try to keep it more horizontal." The server sends this feedback to the user's device, allowing them to review their actions and identify areas for improvement.
[0392] In particular, by utilizing multimodal technology, it is possible to perform a more comprehensive movement evaluation by incorporating audio and subtitle information in addition to video analysis. Based on the feedback, users can learn specifically which parts they need to correct in their next training session, enabling more efficient improvement. This system can be applied not only to sports but also to learning dance and other physical movements, supporting users in acquiring appropriate skills even through self-study.
[0393] The following describes the processing flow.
[0394] Step 1:
[0395] The user prepares video data of their exercise movements using their device and launches the system's application. The user uploads the recorded video to the server via the application. During the upload, the user selects the target exercise category.
[0396] Step 2:
[0397] The server securely stores the received video data and prepares it for analysis. The server uses image processing techniques to identify keypoints for each frame in order to extract the motion features contained in this data.
[0398] Step 3:
[0399] The server compares the extracted feature points with feature points from professional reference video data stored within the same category. Here, AI is used to quantify the degree of similarity and differences between the two sets of movements.
[0400] Step 4:
[0401] The server analyzes the comparison results and generates specific feedback on the user's actions. The generated advice includes points for improvement and guidelines on correct form. The feedback is provided not only in text format but also with visual aids as needed.
[0402] Step 5:
[0403] The server sends the generated feedback to the user's device and notifies the user in response. The user receives the feedback on their device and can use it to improve future training.
[0404] (Example 1)
[0405] Next, we will describe Example 1. In the following description, the data processing device 12 will be referred to as the "server," and the headset-type terminal 314 will be referred to as the "terminal."
[0406] Conventional exercise analysis systems have the challenge of making it difficult for users to receive specific feedback that accurately helps them improve their exercise. Furthermore, they often fail to comprehensively utilize diverse information in exercise evaluation, making it difficult to provide useful advice to users.
[0407] The identification process performed by the identification processing unit 290 of the data processing device 12 in Example 1 is realized by the following means.
[0408] In this invention, the server includes a medium for receiving video data including motion capture by the user, a medium for extracting feature points such as body coordinates and angles from the video data, a medium for comparing the feature points with other reference video data, a medium for generating suggestions for improving the user's motion based on the comparison results using deep learning technology, and a medium for providing the suggestions to the user. This allows the user to obtain more comprehensive and specific information for improving their motion, enabling them to efficiently improve their skills.
[0409] The term "user" refers to a person who accesses the system and receives analysis and feedback on their motor movements.
[0410] "Video data" refers to digital data that includes visual information of movement captured by the user.
[0411] "Feature points" are sets of data extracted from video data to characterize body movements, such as coordinates and angles.
[0412] "Reference video data" refers to reference data provided by professionals or other entities for use as a point of comparison.
[0413] "Deep learning technology" is a technique used in the field of artificial intelligence that utilizes multi-layer neural networks to analyze and predict data.
[0414] A "suggestion" refers to specific instructions or advice generated based on comparison results and other factors, aimed at improving user behavior.
[0415] "Multimodal technology" refers to technologies for integrating and analyzing data in different formats (e.g., audio, video, and text information).
[0416] This invention realizes a motion analysis system, which primarily relies on the cooperation of a server and a terminal. First, the user uses a terminal such as a smartphone or camera to record their movements. This terminal is responsible for saving the video data of the movements as a file and uploading it to the server via the internet.
[0417] The server is responsible for processing the received video data. This analysis utilizes techniques such as extracting feature points from the video, including body coordinates and angle information related to movement, using libraries like OpenPose. The extracted feature points are then compared with reference video data. This reference data includes professional movements stored in a database.
[0418] The server uses deep learning technology during the comparison process. This allows it to evaluate user behavior through the comparison of feature points, and the generative AI model generates specific suggestions for improving user behavior. These suggestions are output as text and, in some cases, as visual guides, and are ultimately sent from the server to the terminal.
[0419] Users can receive suggestions provided on their device and use the feedback to improve their exercise. For example, if a user wants to improve their golf swing, they can enter "How can I improve this golf swing form?" as a prompt and receive specific advice based on the result.
[0420] This system allows users to efficiently improve their skills while receiving feedback on their movements. Furthermore, multimodal technology complements the analysis, enabling comprehensive movement analysis that incorporates audio and text information in addition to video data. This makes it easier for users to learn various forms of movement, from sports to dance.
[0421] The flow of the specific processing in Example 1 will be explained using Figure 11.
[0422] Step 1:
[0423] The user saves video data of their exercise to their device. Specifically, they use the video recording function of their smartphone or camera to film themselves exercising and save the file to the device's storage. This video data becomes the input for the next process.
[0424] Step 2:
[0425] The terminal uploads the stored video data to the server via the internet. Specifically, the user uses a dedicated application for the system to send video files to the server. This operation provides the video data as input for server processing.
[0426] Step 3:
[0427] The server extracts movement-related physical feature points from the received video data. Specifically, it uses video analysis techniques to extract data about body joints and posture. This involves digital image processing using software such as OpenPose. The output of this processing is feature point data for analysis.
[0428] Step 4:
[0429] The server compares the extracted feature points with reference video data. It refers to a reference database containing examples of actions performed by professionals and calculates the distance and positional differences between feature points. This allows it to evaluate how well the user's actions match the reference. Deep learning models are used to improve the accuracy of the evaluation. This evaluation data then serves as input for the next step.
[0430] Step 5:
[0431] The server uses a generated AI model based on the evaluation results to generate specific suggestions for improving the user's performance. Specifically, it inputs prompt sentences into the AI model and generates improvement advice such as, "Your arm position at the top of your swing is insufficient. Try to keep it more horizontal." The generated suggestions are output as text and visual guides.
[0432] Step 6:
[0433] The server sends the generated suggestions to the device. Users can review the feedback displayed on their device and use it to improve their exercise. Specifically, they read the feedback on the device screen and consciously try to correct it during their next workout. This feedback supports the user's skill improvement.
[0434] (Application Example 1)
[0435] Next, we will explain Application Example 1. In the following explanation, the data processing device 12 will be referred to as the "server," and the headset-type terminal 314 will be referred to as the "terminal."
[0436] In modern factories, maximizing robot operational efficiency and minimizing errors is a crucial challenge for improving productivity. However, current monitoring and feedback systems are inadequate, making it difficult to pinpoint detailed areas for improvement in robotic movement. Furthermore, real-time capabilities are often lacking in the process of comparing and analyzing movements. This results in delays in flexible responses and immediate improvements in the manufacturing process.
[0437] The specific processing performed by the specific processing unit 290 of the data processing device 12 in Application Example 1 is realized by the following means.
[0438] In this invention, the server includes a device for receiving video data containing motion capture by a user, a function for extracting feature points from the video data, and a function for comparing the feature points with reference data. This enables real-time analysis of the movements of robots used in factories, allowing for specific guidance to improve efficiency and reduce errors.
[0439] A "user" is an individual or organization that uses a system to analyze and improve their own behavior and actions.
[0440] "Motor activity" refers to all bodily movements and mechanical operations performed with a specific purpose.
[0441] "Video data" refers to data containing visual information recorded by cameras, sensors, and other devices.
[0442] A "device" refers to hardware or software designed to perform a specific function.
[0443] A "feature point" is data containing specific coordinates, angles, velocities, and other information extracted to represent the characteristics of movement or action.
[0444] "Reference data" refers to data that records professional or standard behavior and is used for comparison and evaluation.
[0445] "Comparison" is the act of comparing two or more different sets of data or pieces of information to identify their similarities and differences.
[0446] "Efficiency" refers to the degree of ability to achieve a goal while minimizing resources and time.
[0447] "Error" is a term that refers to an unexpected malfunction or discrepancy in a system or its operation.
[0448] "Guidance" refers to specific advice and suggestions for improvement and optimization provided by the system.
[0449] This invention is a system that uses video data of movement captured by a user to provide feedback for improving the efficiency and precision of that movement. The server uses the following technologies for this purpose.
[0450] First, the user uses their device to capture video data from a camera or sensor and uploads it to the server. A dedicated application is installed on the device, and the data is reliably and efficiently transmitted to the server through this application. Specifically, video processing libraries such as OpenCV are used to capture the video and extract its features.
[0451] Next, the server extracts feature points from the received video data and performs a comparison. Here, learning algorithms, particularly deep learning and generative AI models, are used to evaluate the extracted feature points against reference data. This process is performed in real time, enabling immediate feedback.
[0452] Feedback is generated as numerical data and visual guides and sent to the terminal. In this process, a generating AI model is used as an aid, providing appropriate advice to the user based on the prompts it generates. A specific example is providing instructions to optimize the movement of a robotic arm on a factory production line.
[0453] Examples of prompts for a generative AI model include the following:
[0454] "Factory Robot Motion Analysis: Analyze the motion for efficient bottle picking and provide feedback. Include specific areas for improvement regarding the arm's movement path and speed."
[0455] This system allows users to receive real-time feedback, enabling them to quickly improve the efficiency and precision of their exercise.
[0456] The flow of a specific process in Application Example 1 will be explained using Figure 12.
[0457] Step 1:
[0458] The user uses a device to record their movements. The recorded video data is captured in real time using video processing libraries such as OpenCV. The input is the video acquired from the camera, and this data is sent directly to the server. The output is the video data uploaded to the server.
[0459] Step 2:
[0460] The server extracts feature points from the received video data. Given video data as input, image processing techniques are used to calculate feature points such as coordinates, angles, and velocities that represent the characteristics of motion in each frame. This results in a dataset of the analyzed feature points as output.
[0461] Step 3:
[0462] The server compares the extracted feature points with the reference data. The reference data and feature point datasets serve as input, and a learning algorithm is used to analyze the similarities and differences between them. The output is feedback data indicating the quality of the operation as an evaluation result.
[0463] Step 4:
[0464] The server generates feedback based on the evaluation results. Using an AI model, particularly a generative AI model, it generates specific areas for improvement based on prompts. The input is the evaluation results, and the output is feedback provided to the user in the form of text or visual guides.
[0465] Step 5:
[0466] The terminal receives feedback from the server and provides it to the user. Feedback data is sent to the terminal as input, and the information is presented to the user visually or audibly through a dedicated application. The output is a presentation of areas for improvement and advice that the user can review.
[0467] Furthermore, an emotion engine that estimates the user's emotions may be incorporated. That is, the identification processing unit 290 may use the emotion identification model 59 to estimate the user's emotions and perform identification processing using the user's emotions.
[0468] The system based on the present invention receives the user's movement as video data and generates advice based on it, and in addition, combines this with an emotion engine that recognizes the user's emotions to provide more personalized feedback.
[0469] The user first uses a device to record their movement and uploads the video data to the server. The server receives the video data and stores it securely. The server extracts characteristic points of the movement from the video data and compares them with other reference movement data. Machine learning models are used for the comparison to achieve highly accurate movement analysis.
[0470] The server further uses an emotion engine to understand the user's emotional state. This engine has algorithms that analyze the user's facial expressions from video or recognize emotions from audio data. The recognized emotional state of the user is reflected in the adjustment of feedback to the user. For example, if the server detects that the user is feeling stressed, it can provide positive feedback to boost their motivation.
[0471] The generated advice includes specific points for improvement and is sensitive to the user's feelings. The server sends this to the terminal, allowing the user to review the feedback. The user then uses this feedback to plan how to improve their performance in the next exercise session.
[0472] As an example, let's assume a user is practicing their golf swing. This system evaluates the consistency of the swing from video footage and can also recognize whether the user is feeling confused or anxious from their facial expressions and voice in the video. As a result, the server provides feedback such as, "Your swing is improving well. Next, try to relax and focus more on your rhythm."
[0473] This invention allows users to receive feedback tailored to their individual emotional state, going beyond simple motion analysis, and further enhance the effectiveness of their training.
[0474] The following describes the processing flow.
[0475] Step 1:
[0476] The user uses a device to film their exercise movements and uploads the video data to a server via a dedicated app. At this time, the user selects the appropriate category according to the type of movement.
[0477] Step 2:
[0478] The server stores the received video data and extracts feature points of motion. These feature points are identified from each frame using image processing techniques. This allows the specific elements of the motion to be quantified.
[0479] Step 3:
[0480] The server compares feature points extracted from the user's video with feature points from professional video data registered as a reference. This comparison is performed by a machine learning model, which calculates the degree of similarity and differences for each action.
[0481] Step 4:
[0482] The server uses an emotion engine to analyze the user's facial expressions and voice data to recognize their emotional state. The emotions the user displays in the video (e.g., concentration, confusion, joy) are detected in this step.
[0483] Step 5:
[0484] The server generates customized advice based on the comparison results of the actions and the user's emotional state. For example, it might create feedback such as, "Relaxing will make your arm movements smoother."
[0485] Step 6:
[0486] The server generates feedback which is then provided to the user via the terminal. The user can review the feedback on the terminal and use it to improve their next training session.
[0487] (Example 2)
[0488] Next, we will describe Example 2. In the following description, the data processing device 12 will be referred to as the "server," and the headset-type terminal 314 will be referred to as the "terminal."
[0489] When attempting to improve motor skills, there is a challenge in providing personalized feedback that takes into account each user's emotional state. Furthermore, there is a need to integrate and analyze diverse data formats to achieve more accurate motion analysis and advice generation.
[0490] The identification process performed by the identification processing unit 290 of the data processing device 12 in Example 2 is realized by the following means.
[0491] In this invention, the server includes means for receiving visual information, including motor movements acquired by the user; means for extracting characteristic information from the visual information; and means for comparing standard video data with the characteristic information. This makes it possible to provide optimal feedback tailored to the individual user's characteristics and emotional state.
[0492] A "user" refers to a person who uses the system to analyze their own movement patterns and receive advice for improvement.
[0493] "Visual information" refers to image or video data that includes recordings of movement, and includes details of the movements captured by the user.
[0494] "Characteristic information" refers to the characteristic points and movement patterns of motion extracted from visual information, and this constitutes the basic data for analyzing motion.
[0495] "Standard video data" refers to video data of past motions used as a reference, and is used for comparison with characteristic information.
[0496] A "machine learning algorithm" refers to a group of mathematical methods used to learn patterns from large amounts of data and analyze their behavioral characteristics.
[0497] "Emotion recognition means" encompasses methods and technologies for analyzing a user's emotional state and are used to recognize a user's emotions from visual and auditory information.
[0498] "Advice" refers to information generated after considering motion analysis and emotional state, which includes specific instructions and suggestions for the user to improve their motor skills.
[0499] This invention is a system that aims to allow users to analyze their own motor movements in detail and receive feedback based on their individual emotional state. Specific embodiments of this system are described below.
[0500] The user first films their own movements using a device. This device can be any recording device, including a smartphone or camera. The captured visual information is uploaded from the device to a server via the internet.
[0501] The server extracts characteristic information from the received visual information. This involves using image processing libraries such as OpenCV to detect feature points for detailed analysis of user movements. Next, machine learning algorithms are used to compare the extracted characteristic information with pre-prepared standard video data. In this comparison process, artificial intelligence frameworks such as TensorFlow are used to analyze and evaluate the user's movements.
[0502] Furthermore, the server uses video and audio data to identify the user's emotional state. This incorporates a speech recognition engine and facial expression analysis algorithm as means of emotion recognition. This analysis provides emotional information, such as whether the user is feeling tense or anxious, which influences the feedback.
[0503] The generated feedback includes specific advice for improving the user's movements, as well as emotionally responsive psychological support. For example, it might say, "Analysis of your golf swing video shows that your swing is improving well. Next time, try to relax and focus more on your rhythm." This allows the user to make precise improvements to their movements.
[0504] An example of a prompt message might be: "Consider the analysis results of the user's golf swing video and their emotional state, and generate improvement advice. The user is performing a swing and appears confused. The feedback should be positive and motivating."
[0505] As described above, this invention can more effectively enhance athletic ability and mental stability by analyzing the individual user's motor movements and providing personalized feedback based on their emotional state.
[0506] The flow of the specific processing in Example 2 will be explained using Figure 13.
[0507] Step 1:
[0508] The user uses a device to record their movement. The input at this time is visual information recording the user's movement. The device temporarily stores this information and prepares to upload it to the server. The output is the visual information ready to be sent to the server.
[0509] Step 2:
[0510] The server receives visual information sent from the terminal and stores it in a database. The input is visual information from the user, which the server receives and stores securely. The output is visual information ready for analysis.
[0511] Step 3:
[0512] The server extracts characteristic information from the received visual information. At this stage, the input is stored visual information, and the server uses an image processing library such as OpenCV to detect feature points of motion from the image data. The output is characteristic information including the feature points.
[0513] Step 4:
[0514] The server compares extracted characteristic information with standard video data. The input consists of characteristic information and standard video data. The server uses machine learning algorithms such as TensorFlow to compare these and analyze the consistency of the operation and areas requiring improvement. The output is performance evaluation data based on the analysis results.
[0515] Step 5:
[0516] The server uses emotion recognition to determine the user's emotional state. Visual and auditory information are used as input, and the server utilizes an emotion recognition engine to analyze the user's emotions. The output is data on the user's psychological state obtained from their facial expressions and voice.
[0517] Step 6:
[0518] The server generates feedback using a generative AI model based on performance evaluation data and emotional state. The input consists of analysis results and psychological state data. Prompt sentences are input to the generative AI model, which generates specific, emotionally sensitive advice tailored to the user. The output is the generated feedback message.
[0519] Step 7:
[0520] The server sends the generated feedback to the terminal. The input is the feedback message. The output is the feedback message that the user receives through the terminal. This allows the user to review the feedback and use it to improve their next exercise session.
[0521] (Application Example 2)
[0522] Next, we will explain application example 2. In the following explanation, the data processing device 12 will be referred to as the "server," and the headset-type terminal 314 will be referred to as the "terminal."
[0523] In modern fitness and training environments, there is a demand for rapid and appropriate personalized feedback to individual users. However, conventional systems struggle to understand not only a user's movement but also their emotional state and provide feedback that takes those emotions into account. Furthermore, there is a need for systems that can individually adjust advice on areas for improvement in user movement in real time.
[0524] The specific processing performed by the specific processing unit 290 of the data processing device 12 in Application Example 2 is realized by the following means.
[0525] In this invention, the server includes means for receiving video information including motion capture by the user, means for extracting motion characteristics from the video information, means for comparing the motion characteristics with other video information, and means for analyzing the user's emotional state and adjusting the advice based on the results. This makes it possible to effectively provide individual users with advice for motion improvement while taking their emotional state into consideration.
[0526] "Video information" refers to visual data that records movement actions captured by the user, and is used for motion analysis.
[0527] "Movement characteristics" refer to various data points related to the body's posture and movement during athletic activity, and these are used as the basis for analyzing the movement.
[0528] A "means of comparison" refers to a method for evaluating the quality of a user's actions by comparing the extracted characteristics of their actions with other reference data.
[0529] "Advice" refers to guidance information provided to improve the user's motor skills, including specific areas for improvement and content designed to boost motivation.
[0530] "Emotional state" refers to a psychological state inferred from the user's facial expressions, voice, etc., and reflects emotions such as stress, anxiety, and enjoyment during exercise.
[0531] "Means of adjustment" refer to methods for modifying the content and method of providing advice according to the user's emotional state, thereby achieving more effective feedback.
[0532] The system for implementing this invention begins with the user recording their movements with a smartphone or other device while exercising and uploading the video information to a server. The server receives the video information and extracts the characteristics of the movement from it. Specifically, it uses a machine learning model to calculate data points such as posture, angle, and speed of the movement. The software used mainly consists of machine learning libraries such as TensorFlow and PyTorch.
[0533] Furthermore, the server analyzes the user's emotional state from the video information. This analysis uses the Microsoft Azure Cognitive Services Emotion API based on the user's facial expressions and voice data. This analysis assesses the user's psychological state during exercise, such as their stress levels and relaxation levels.
[0534] Based on this data, the server generates advice for the user. This advice specifically points out areas for improvement and uses considerate language to enhance motivation. Finally, the generated advice is sent to the user's device, allowing them to refer to the feedback during their next exercise session.
[0535] As a concrete example, let's say there's a user practicing their golf swing. This user films their swing and sends the video information to a server. The server analyzes the swing form, and if it detects through emotion analysis that the user is confused, it provides advice such as, "Your swing form is good. Next time, try to relax and continue practicing while focusing on your rhythm."
[0536] An example of a prompt message is: "Analyze your exercise form and emotional state, and provide specific areas for improvement and advice on how to relax."
[0537] The flow of a specific process in Application Example 2 will be explained using Figure 14.
[0538] Step 1:
[0539] Users record their exercise using their smartphones. During this process, video information is captured on the device. The device's camera function is used to thoroughly record every detail of the movement.
[0540] Step 2:
[0541] The terminal uploads the captured video information to the server. The terminal transfers the video data to the server via the network connection. The input is the video data, and the output is the server confirming receipt.
[0542] Step 3:
[0543] The server extracts motion characteristics from the received video information. Using motion analysis software, it analyzes important posture and movement patterns in the video using machine learning algorithms (TensorFlow or PyTorch). The input is video data, and the output is feature data.
[0544] Step 4:
[0545] The server compares feature data with reference data. It uses past motion data as a reference to analyze how well it matches current motion. The input is feature data and reference data, and the output is the comparison result.
[0546] Step 5:
[0547] The server analyzes emotional states from video information. It uses the Microsoft Azure Cognitive Services Emotion API to infer user emotions from video and audio. Input is video and audio data, and output is emotional state data.
[0548] Step 6:
[0549] The server generates advice using comparison results and emotional state data. It creates advice that combines points for performance improvement with positive feedback tailored to the emotional state. The input is the comparison results and emotional state data, and the output is the generated advice.
[0550] Step 7:
[0551] The server sends the generated advice to the user's device. This advice is forwarded to the device as feedback, allowing the user to utilize it during their next exercise session. The input is the generated advice, and the output is the display of the advice on the device.
[0552] The specific processing unit 290 transmits the result of the specific processing to the headset terminal 314. In the headset terminal 314, the control unit 46A causes the speaker 240 and display 343 to output the result of the specific processing. The microphone 238 acquires audio indicating user input for the result of the specific processing. The control unit 46A transmits the audio data indicating user input acquired by the microphone 238 to the data processing unit 12. In the data processing unit 12, the specific processing unit 290 acquires the audio data.
[0553] Data generation model 58 is a type of so-called generative AI (Artificial Intelligence). One example of data generation model 58 is ChatGPT (Internet search<URL: https: / / openai.com / blog / chatgpt> ), Gemini (Internet search) <url: https: gemini.google.com ?hl="ja">Examples of generative AI include the following. The data generation model 58 is obtained by performing deep learning on a neural network. The data generation model 58 is input with prompts containing instructions, and with inference data such as audio data representing speech, text data representing text, and image data representing images. The data generation model 58 infers from the input inference data according to the instructions indicated by the prompts, and outputs the inference results in data formats such as audio data and text data. Here, inference refers to, for example, analysis, classification, prediction, and / or summarization.
[0554] In the above embodiment, an example was given in which specific processing is performed by the data processing device 12, but the technology of this disclosure is not limited thereto, and specific processing may also be performed by the headset terminal 314.
[0555] [Fourth Embodiment]
[0556] Figure 7 shows an example of the configuration of the data processing system 410 according to the fourth embodiment.
[0557] As shown in Figure 7, the data processing system 410 includes a data processing device 12 and a robot 414. An example of the data processing device 12 is a server.
[0558] The data processing device 12 comprises a computer 22, a database 24, and a communication interface 26. The computer 22 is an example of a "computer" related to the technology of this disclosure. The computer 22 comprises a processor 28, RAM 30, and storage 32. The processor 28, RAM 30, and storage 32 are connected to a bus 34. The database 24 and the communication interface 26 are also connected to the bus 34. The communication interface 26 is connected to a network 54. An example of the network 54 is a WAN (Wide Area Network) and / or a LAN (Local Area Network).
[0559] The robot 414 includes a computer 36, a microphone 238, a speaker 240, a camera 42, a communication interface 44, and a controlled object 443. The computer 36 includes a processor 46, RAM 48, and storage 50. The processor 46, RAM 48, and storage 50 are connected to a bus 52. The microphone 238, speaker 240, camera 42, and controlled object 443 are also connected to the bus 52.
[0560] The microphone 238 receives voice signals from the user 20 and receives instructions from the user 20. The microphone 238 captures the voice signals from the user 20, converts the captured voice into audio data, and outputs it to the processor 46. The speaker 240 outputs audio according to the instructions from the processor 46.
[0561] Camera 42 is a small digital camera equipped with an optical system including a lens, aperture, and shutter, and an image sensor such as a CMOS (Complementary Metal-Oxide-Semiconductor) image sensor or a CCD (Charge Coupled Device) image sensor, and captures images of the area around the user 20 (for example, an imaging range defined by a field of view equivalent to the width of a typical healthy person's field of vision).
[0562] Communication interface 44 is connected to network 54. Communication interfaces 44 and 26 are responsible for the exchange of various information between processor 46 and processor 28 via network 54. The exchange of various information between processor 46 and processor 28 using communication interfaces 44 and 26 is performed in a secure manner.
[0563] The controlled object 443 includes a display device, LEDs in the eyes, and motors that drive the arms, hands, and feet. The posture and gestures of the robot 414 are controlled by controlling the motors of the arms, hands, and feet. Some of the robot 414's emotions can be expressed by controlling these motors. Furthermore, the robot 414's facial expressions can also be expressed by controlling the illumination state of the LEDs in its eyes.
[0564] Figure 8 shows an example of the main functions of the data processing device 12 and the robot 414. As shown in Figure 8, the data processing device 12 performs specific processing using the processor 28. The storage 32 stores the specific processing program 56.
[0565] The specific processing program 56 is an example of a "program" relating to the technology of this disclosure. The processor 28 reads the specific processing program 56 from the storage 32 and executes the read specific processing program 56 on the RAM 30. The specific processing is realized by the processor 28 operating as a specific processing unit 290 in accordance with the specific processing program 56 executed on the RAM 30.
[0566] The storage 32 stores the data generation model 58 and the emotion identification model 59. The data generation model 58 and the emotion identification model 59 are used by the identification processing unit 290.
[0567] In robot 414, the processor 46 performs the reception output processing. The storage 50 stores the reception output program 60. The processor 46 reads the reception output program 60 from the storage 50 and executes the read reception output program 60 on the RAM 48. The reception output processing is realized by the processor 46 operating as a control unit 46A according to the reception output program 60 executed on the RAM 48.
[0568] Next, the specific processing performed by the specific processing unit 290 of the data processing device 12 will be described. In the following description, the data processing device 12 will be referred to as the "server" and the robot 414 as the "terminal".
[0569] This invention is a system for analyzing and providing feedback on motor movements. The process begins with the user uploading video data of their own movements to a server via a terminal. The server extracts feature points to analyze the received video data. Feature points are data such as coordinates, angles, and velocities that characterize elements of movement. Next, the server references reference video data from professionals within the same category and compares it to the user's feature points.
[0570] The server utilizes AI and multimodal technologies to assess the quality of the user's actions based on these comparison results and generates specific improvement advice. This advice can be in text format or, if necessary, supplementary visual guides. For example, it might generate feedback such as, "Your arm position at the top of your swing is insufficient. Try to keep it more horizontal." The server sends this feedback to the user's device, allowing them to review their actions and identify areas for improvement.
[0571] In particular, by utilizing multimodal technology, it is possible to perform a more comprehensive movement evaluation by incorporating audio and subtitle information in addition to video analysis. Based on the feedback, users can learn specifically which parts they need to correct in their next training session, enabling more efficient improvement. This system can be applied not only to sports but also to learning dance and other physical movements, supporting users in acquiring appropriate skills even through self-study.
[0572] The following describes the processing flow.
[0573] Step 1:
[0574] The user prepares video data of their exercise movements using their device and launches the system's application. The user uploads the recorded video to the server via the application. During the upload, the user selects the target exercise category.
[0575] Step 2:
[0576] The server securely stores the received video data and prepares it for analysis. The server uses image processing techniques to identify keypoints for each frame in order to extract the motion features contained in this data.
[0577] Step 3:
[0578] The server compares the extracted feature points with feature points from professional reference video data stored within the same category. Here, AI is used to quantify the degree of similarity and differences between the two sets of movements.
[0579] Step 4:
[0580] The server analyzes the comparison results and generates specific feedback on the user's actions. The generated advice includes points for improvement and guidelines on correct form. The feedback is provided not only in text format but also with visual aids as needed.
[0581] Step 5:
[0582] The server sends the generated feedback to the user's device and notifies the user in response. The user receives the feedback on their device and can use it to improve future training.
[0583] (Example 1)
[0584] Next, we will describe Example 1. In the following description, the data processing device 12 will be referred to as the "server" and the robot 414 as the "terminal".
[0585] Conventional exercise analysis systems have the challenge of making it difficult for users to receive specific feedback that accurately helps them improve their exercise. Furthermore, they often fail to comprehensively utilize diverse information in exercise evaluation, making it difficult to provide useful advice to users.
[0586] The identification process performed by the identification processing unit 290 of the data processing device 12 in Example 1 is realized by the following means.
[0587] In this invention, the server includes a medium for receiving video data including motion capture by the user, a medium for extracting feature points such as body coordinates and angles from the video data, a medium for comparing the feature points with other reference video data, a medium for generating suggestions for improving the user's motion based on the comparison results using deep learning technology, and a medium for providing the suggestions to the user. This allows the user to obtain more comprehensive and specific information for improving their motion, enabling them to efficiently improve their skills.
[0588] The term "user" refers to a person who accesses the system and receives analysis and feedback on their motor movements.
[0589] "Video data" refers to digital data that includes visual information of movement captured by the user.
[0590] "Feature points" are sets of data extracted from video data to characterize body movements, such as coordinates and angles.
[0591] "Reference video data" refers to reference data provided by professionals or other entities for use as a point of comparison.
[0592] "Deep learning technology" is a technique used in the field of artificial intelligence that utilizes multi-layer neural networks to analyze and predict data.
[0593] A "suggestion" refers to specific instructions or advice generated based on comparison results and other factors, aimed at improving user behavior.
[0594] "Multimodal technology" refers to technologies for integrating and analyzing data in different formats (e.g., audio, video, and text information).
[0595] This invention realizes a motion analysis system, which primarily relies on the cooperation of a server and a terminal. First, the user uses a terminal such as a smartphone or camera to record their movements. This terminal is responsible for saving the video data of the movements as a file and uploading it to the server via the internet.
[0596] The server is responsible for processing the received video data. This analysis utilizes techniques such as extracting feature points from the video, including body coordinates and angle information related to movement, using libraries like OpenPose. The extracted feature points are then compared with reference video data. This reference data includes professional movements stored in a database.
[0597] The server uses deep learning technology during the comparison process. This allows it to evaluate user behavior through the comparison of feature points, and the generative AI model generates specific suggestions for improving user behavior. These suggestions are output as text and, in some cases, as visual guides, and are ultimately sent from the server to the terminal.
[0598] Users can receive suggestions provided on their device and use the feedback to improve their exercise. For example, if a user wants to improve their golf swing, they can enter "How can I improve this golf swing form?" as a prompt and receive specific advice based on the result.
[0599] This system allows users to efficiently improve their skills while receiving feedback on their movements. Furthermore, multimodal technology complements the analysis, enabling comprehensive movement analysis that incorporates audio and text information in addition to video data. This makes it easier for users to learn various forms of movement, from sports to dance.
[0600] The flow of the specific processing in Example 1 will be explained using Figure 11.
[0601] Step 1:
[0602] The user saves video data of their exercise to their device. Specifically, they use the video recording function of their smartphone or camera to film themselves exercising and save the file to the device's storage. This video data becomes the input for the next process.
[0603] Step 2:
[0604] The terminal uploads the stored video data to the server via the internet. Specifically, the user uses a dedicated application for the system to send video files to the server. This operation provides the video data as input for server processing.
[0605] Step 3:
[0606] The server extracts movement-related physical feature points from the received video data. Specifically, it uses video analysis techniques to extract data about body joints and posture. This involves digital image processing using software such as OpenPose. The output of this processing is feature point data for analysis.
[0607] Step 4:
[0608] The server compares the extracted feature points with reference video data. It refers to a reference database containing examples of actions performed by professionals and calculates the distance and positional differences between feature points. This allows it to evaluate how well the user's actions match the reference. Deep learning models are used to improve the accuracy of the evaluation. This evaluation data then serves as input for the next step.
[0609] Step 5:
[0610] The server uses a generated AI model based on the evaluation results to generate specific suggestions for improving the user's performance. Specifically, it inputs prompt sentences into the AI model and generates improvement advice such as, "Your arm position at the top of your swing is insufficient. Try to keep it more horizontal." The generated suggestions are output as text and visual guides.
[0611] Step 6:
[0612] The server sends the generated suggestions to the device. Users can review the feedback displayed on their device and use it to improve their exercise. Specifically, they read the feedback on the device screen and consciously try to correct it during their next workout. This feedback supports the user's skill improvement.
[0613] (Application Example 1)
[0614] Next, we will explain Application Example 1. In the following explanation, the data processing device 12 will be referred to as the "server" and the robot 414 as the "terminal".
[0615] In modern factories, maximizing robot operational efficiency and minimizing errors is a crucial challenge for improving productivity. However, current monitoring and feedback systems are inadequate, making it difficult to pinpoint detailed areas for improvement in robotic movement. Furthermore, real-time capabilities are often lacking in the process of comparing and analyzing movements. This results in delays in flexible responses and immediate improvements in the manufacturing process.
[0616] The specific processing performed by the specific processing unit 290 of the data processing device 12 in Application Example 1 is realized by the following means.
[0617] In this invention, the server includes a device for receiving video data containing motion capture by a user, a function for extracting feature points from the video data, and a function for comparing the feature points with reference data. This enables real-time analysis of the movements of robots used in factories, allowing for specific guidance to improve efficiency and reduce errors.
[0618] A "user" is an individual or organization that uses a system to analyze and improve their own behavior and actions.
[0619] "Motor activity" refers to all bodily movements and mechanical operations performed with a specific purpose.
[0620] "Video data" refers to data containing visual information recorded by cameras, sensors, and other devices.
[0621] A "device" refers to hardware or software designed to perform a specific function.
[0622] A "feature point" is data containing specific coordinates, angles, velocities, and other information extracted to represent the characteristics of movement or action.
[0623] "Reference data" refers to data that records professional or standard behavior and is used for comparison and evaluation.
[0624] "Comparison" is the act of comparing two or more different sets of data or pieces of information to identify their similarities and differences.
[0625] "Efficiency" refers to the degree of ability to achieve a goal while minimizing resources and time.
[0626] "Error" is a term that refers to an unexpected malfunction or discrepancy in a system or its operation.
[0627] "Guidance" refers to specific advice and suggestions for improvement and optimization provided by the system.
[0628] This invention is a system that uses video data of movement captured by a user to provide feedback for improving the efficiency and precision of that movement. The server uses the following technologies for this purpose.
[0629] First, the user uses their device to capture video data from a camera or sensor and uploads it to the server. A dedicated application is installed on the device, and the data is reliably and efficiently transmitted to the server through this application. Specifically, video processing libraries such as OpenCV are used to capture the video and extract its features.
[0630] Next, the server extracts feature points from the received video data and performs a comparison. Here, learning algorithms, particularly deep learning and generative AI models, are used to evaluate the extracted feature points against reference data. This process is performed in real time, enabling immediate feedback.
[0631] Feedback is generated as numerical data and visual guides and sent to the terminal. In this process, a generating AI model is used as an aid, providing appropriate advice to the user based on the prompts it generates. A specific example is providing instructions to optimize the movement of a robotic arm on a factory production line.
[0632] Examples of prompts for a generative AI model include the following:
[0633] "Factory Robot Motion Analysis: Analyze the motion for efficient bottle picking and provide feedback. Include specific areas for improvement regarding the arm's movement path and speed."
[0634] This system allows users to receive real-time feedback, enabling them to quickly improve the efficiency and precision of their exercise.
[0635] The flow of a specific process in Application Example 1 will be explained using Figure 12.
[0636] Step 1:
[0637] The user uses a device to record their movements. The recorded video data is captured in real time using video processing libraries such as OpenCV. The input is the video acquired from the camera, and this data is sent directly to the server. The output is the video data uploaded to the server.
[0638] Step 2:
[0639] The server extracts feature points from the received video data. Given video data as input, image processing techniques are used to calculate feature points such as coordinates, angles, and velocities that represent the characteristics of motion in each frame. This results in a dataset of the analyzed feature points as output.
[0640] Step 3:
[0641] The server compares the extracted feature points with the reference data. The reference data and feature point datasets serve as input, and a learning algorithm is used to analyze the similarities and differences between them. The output is feedback data indicating the quality of the operation as an evaluation result.
[0642] Step 4:
[0643] The server generates feedback based on the evaluation results. Using an AI model, particularly a generative AI model, it generates specific areas for improvement based on prompts. The input is the evaluation results, and the output is feedback provided to the user in the form of text or visual guides.
[0644] Step 5:
[0645] The terminal receives feedback from the server and provides it to the user. Feedback data is sent to the terminal as input, and the information is presented to the user visually or audibly through a dedicated application. The output is a presentation of areas for improvement and advice that the user can review.
[0646] Furthermore, an emotion engine that estimates the user's emotions may be incorporated. That is, the identification processing unit 290 may use the emotion identification model 59 to estimate the user's emotions and perform identification processing using the user's emotions.
[0647] The system based on the present invention receives the user's movement as video data and generates advice based on it, and in addition, combines this with an emotion engine that recognizes the user's emotions to provide more personalized feedback.
[0648] The user first uses a device to record their movement and uploads the video data to the server. The server receives the video data and stores it securely. The server extracts characteristic points of the movement from the video data and compares them with other reference movement data. Machine learning models are used for the comparison to achieve highly accurate movement analysis.
[0649] The server further uses an emotion engine to understand the user's emotional state. This engine has algorithms that analyze the user's facial expressions from video or recognize emotions from audio data. The recognized emotional state of the user is reflected in the adjustment of feedback to the user. For example, if the server detects that the user is feeling stressed, it can provide positive feedback to boost their motivation.
[0650] The generated advice includes specific points for improvement and is sensitive to the user's feelings. The server sends this to the terminal, allowing the user to review the feedback. The user then uses this feedback to plan how to improve their performance in the next exercise session.
[0651] As an example, let's assume a user is practicing their golf swing. This system evaluates the consistency of the swing from video footage and can also recognize whether the user is feeling confused or anxious from their facial expressions and voice in the video. As a result, the server provides feedback such as, "Your swing is improving well. Next, try to relax and focus more on your rhythm."
[0652] This invention allows users to receive feedback tailored to their individual emotional state, going beyond simple motion analysis, and further enhance the effectiveness of their training.
[0653] The following describes the processing flow.
[0654] Step 1:
[0655] The user uses a device to film their exercise movements and uploads the video data to a server via a dedicated app. At this time, the user selects the appropriate category according to the type of movement.
[0656] Step 2:
[0657] The server stores the received video data and extracts feature points of motion. These feature points are identified from each frame using image processing techniques. This allows the specific elements of the motion to be quantified.
[0658] Step 3:
[0659] The server compares feature points extracted from the user's video with feature points from professional video data registered as a reference. This comparison is performed by a machine learning model, which calculates the degree of similarity and differences for each action.
[0660] Step 4:
[0661] The server uses an emotion engine to analyze the user's facial expressions and voice data to recognize their emotional state. The emotions the user displays in the video (e.g., concentration, confusion, joy) are detected in this step.
[0662] Step 5:
[0663] The server generates customized advice based on the comparison results of the actions and the user's emotional state. For example, it might create feedback such as, "Relaxing will make your arm movements smoother."
[0664] Step 6:
[0665] The server generates feedback which is then provided to the user via the terminal. The user can review the feedback on the terminal and use it to improve their next training session.
[0666] (Example 2)
[0667] Next, we will describe Example 2. In the following description, the data processing device 12 will be referred to as the "server" and the robot 414 as the "terminal".
[0668] When attempting to improve motor skills, there is a challenge in providing personalized feedback that takes into account each user's emotional state. Furthermore, there is a need to integrate and analyze diverse data formats to achieve more accurate motion analysis and advice generation.
[0669] The identification process performed by the identification processing unit 290 of the data processing device 12 in Example 2 is realized by the following means.
[0670] In this invention, the server includes means for receiving visual information, including motor movements acquired by the user; means for extracting characteristic information from the visual information; and means for comparing standard video data with the characteristic information. This makes it possible to provide optimal feedback tailored to the individual user's characteristics and emotional state.
[0671] A "user" refers to a person who uses the system to analyze their own movement patterns and receive advice for improvement.
[0672] "Visual information" refers to image or video data that includes recordings of movement, and includes details of the movements captured by the user.
[0673] "Characteristic information" refers to the characteristic points and movement patterns of motion extracted from visual information, and this constitutes the basic data for analyzing motion.
[0674] "Standard video data" refers to video data of past motions used as a reference, and is used for comparison with characteristic information.
[0675] A "machine learning algorithm" refers to a group of mathematical methods used to learn patterns from large amounts of data and analyze their behavioral characteristics.
[0676] "Emotion recognition means" encompasses methods and technologies for analyzing a user's emotional state and are used to recognize a user's emotions from visual and auditory information.
[0677] "Advice" refers to information generated after considering motion analysis and emotional state, which includes specific instructions and suggestions for the user to improve their motor skills.
[0678] This invention is a system that aims to allow users to analyze their own motor movements in detail and receive feedback based on their individual emotional state. Specific embodiments of this system are described below.
[0679] The user first films their own movements using a device. This device can be any recording device, including a smartphone or camera. The captured visual information is uploaded from the device to a server via the internet.
[0680] The server extracts characteristic information from the received visual information. This involves using image processing libraries such as OpenCV to detect feature points for detailed analysis of user movements. Next, machine learning algorithms are used to compare the extracted characteristic information with pre-prepared standard video data. In this comparison process, artificial intelligence frameworks such as TensorFlow are used to analyze and evaluate the user's movements.
[0681] Furthermore, the server uses video and audio data to identify the user's emotional state. This incorporates a speech recognition engine and facial expression analysis algorithm as means of emotion recognition. This analysis provides emotional information, such as whether the user is feeling tense or anxious, which influences the feedback.
[0682] The generated feedback includes specific advice for improving the user's movements, as well as emotionally responsive psychological support. For example, it might say, "Analysis of your golf swing video shows that your swing is improving well. Next time, try to relax and focus more on your rhythm." This allows the user to make precise improvements to their movements.
[0683] An example of a prompt message might be: "Consider the analysis results of the user's golf swing video and their emotional state, and generate improvement advice. The user is performing a swing and appears confused. The feedback should be positive and motivating."
[0684] As described above, this invention can more effectively enhance athletic ability and mental stability by analyzing the individual user's motor movements and providing personalized feedback based on their emotional state.
[0685] The flow of the specific processing in Example 2 will be explained using Figure 13.
[0686] Step 1:
[0687] The user uses a device to record their movement. The input at this time is visual information recording the user's movement. The device temporarily stores this information and prepares to upload it to the server. The output is the visual information ready to be sent to the server.
[0688] Step 2:
[0689] The server receives visual information sent from the terminal and stores it in a database. The input is visual information from the user, which the server receives and stores securely. The output is visual information ready for analysis.
[0690] Step 3:
[0691] The server extracts characteristic information from the received visual information. At this stage, the input is stored visual information, and the server uses an image processing library such as OpenCV to detect feature points of motion from the image data. The output is characteristic information including the feature points.
[0692] Step 4:
[0693] The server compares extracted characteristic information with standard video data. The input consists of characteristic information and standard video data. The server uses machine learning algorithms such as TensorFlow to compare these and analyze the consistency of the operation and areas requiring improvement. The output is performance evaluation data based on the analysis results.
[0694] Step 5:
[0695] The server uses emotion recognition to determine the user's emotional state. Visual and auditory information are used as input, and the server utilizes an emotion recognition engine to analyze the user's emotions. The output is data on the user's psychological state obtained from their facial expressions and voice.
[0696] Step 6:
[0697] The server generates feedback using a generative AI model based on performance evaluation data and emotional state. The input consists of analysis results and psychological state data. Prompt sentences are input to the generative AI model, which generates specific, emotionally sensitive advice tailored to the user. The output is the generated feedback message.
[0698] Step 7:
[0699] The server sends the generated feedback to the terminal. The input is the feedback message. The output is the feedback message that the user receives through the terminal. This allows the user to review the feedback and use it to improve their next exercise session.
[0700] (Application Example 2)
[0701] Next, we will explain application example 2. In the following explanation, the data processing device 12 will be referred to as the "server" and the robot 414 as the "terminal".
[0702] In modern fitness and training environments, there is a demand for rapid and appropriate personalized feedback to individual users. However, conventional systems struggle to understand not only a user's movement but also their emotional state and provide feedback that takes those emotions into account. Furthermore, there is a need for systems that can individually adjust advice on areas for improvement in user movement in real time.
[0703] The specific processing performed by the specific processing unit 290 of the data processing device 12 in Application Example 2 is realized by the following means.
[0704] In this invention, the server includes means for receiving video information including motion capture by the user, means for extracting motion characteristics from the video information, means for comparing the motion characteristics with other video information, and means for analyzing the user's emotional state and adjusting the advice based on the results. This makes it possible to effectively provide individual users with advice for motion improvement while taking their emotional state into consideration.
[0705] "Video information" refers to visual data that records movement actions captured by the user, and is used for motion analysis.
[0706] "Movement characteristics" refer to various data points related to the body's posture and movement during athletic activity, and these are used as the basis for analyzing the movement.
[0707] A "means of comparison" refers to a method for evaluating the quality of a user's actions by comparing the extracted characteristics of their actions with other reference data.
[0708] "Advice" refers to guidance information provided to improve the user's motor skills, including specific areas for improvement and content designed to boost motivation.
[0709] "Emotional state" refers to a psychological state inferred from the user's facial expressions, voice, etc., and reflects emotions such as stress, anxiety, and enjoyment during exercise.
[0710] "Means of adjustment" refer to methods for modifying the content and method of providing advice according to the user's emotional state, thereby achieving more effective feedback.
[0711] The system for implementing this invention begins with the user recording their movements with a smartphone or other device while exercising and uploading the video information to a server. The server receives the video information and extracts the characteristics of the movement from it. Specifically, it uses a machine learning model to calculate data points such as posture, angle, and speed of the movement. The software used mainly consists of machine learning libraries such as TensorFlow and PyTorch.
[0712] Furthermore, the server analyzes the user's emotional state from the video information. This analysis uses the Microsoft Azure Cognitive Services Emotion API based on the user's facial expressions and voice data. This analysis assesses the user's psychological state during exercise, such as their stress levels and relaxation levels.
[0713] Based on this data, the server generates advice for the user. This advice specifically points out areas for improvement and uses considerate language to enhance motivation. Finally, the generated advice is sent to the user's device, allowing them to refer to the feedback during their next exercise session.
[0714] As a concrete example, let's say there's a user practicing their golf swing. This user films their swing and sends the video information to a server. The server analyzes the swing form, and if it detects through emotion analysis that the user is confused, it provides advice such as, "Your swing form is good. Next time, try to relax and continue practicing while focusing on your rhythm."
[0715] An example of a prompt message is: "Analyze your exercise form and emotional state, and provide specific areas for improvement and advice on how to relax."
[0716] The flow of a specific process in Application Example 2 will be explained using Figure 14.
[0717] Step 1:
[0718] Users record their exercise using their smartphones. During this process, video information is captured on the device. The device's camera function is used to thoroughly record every detail of the movement.
[0719] Step 2:
[0720] The terminal uploads the captured video information to the server. The terminal transfers the video data to the server via the network connection. The input is the video data, and the output is the server confirming receipt.
[0721] Step 3:
[0722] The server extracts motion characteristics from the received video information. Using motion analysis software, it analyzes important posture and movement patterns in the video using machine learning algorithms (TensorFlow or PyTorch). The input is video data, and the output is feature data.
[0723] Step 4:
[0724] The server compares feature data with reference data. It uses past motion data as a reference to analyze how well it matches current motion. The input is feature data and reference data, and the output is the comparison result.
[0725] Step 5:
[0726] The server analyzes emotional states from video information. It uses the Microsoft Azure Cognitive Services Emotion API to infer user emotions from video and audio. Input is video and audio data, and output is emotional state data.
[0727] Step 6:
[0728] The server generates advice using comparison results and emotional state data. It creates advice that combines points for performance improvement with positive feedback tailored to the emotional state. The input is the comparison results and emotional state data, and the output is the generated advice.
[0729] Step 7:
[0730] The server sends the generated advice to the user's device. This advice is forwarded to the device as feedback, allowing the user to utilize it during their next exercise session. The input is the generated advice, and the output is the display of the advice on the device.
[0731] The specific processing unit 290 transmits the result of the specific processing to the robot 414. In the robot 414, the control unit 46A causes the speaker 240 and the controlled object 443 to output the result of the specific processing. The microphone 238 acquires audio indicating user input for the result of the specific processing. The control unit 46A transmits the audio data indicating user input acquired by the microphone 238 to the data processing unit 12. In the data processing unit 12, the specific processing unit 290 acquires the audio data.
[0732] Data generation model 58 is a type of so-called generative AI (Artificial Intelligence). One example of data generation model 58 is ChatGPT (Internet search<URL: https: / / openai.com / blog / chatgpt> ), Gemini (Internet search) <url: https: gemini.google.com ?hl="ja">Examples of generative AI include the following. The data generation model 58 is obtained by performing deep learning on a neural network. The data generation model 58 is input with prompts containing instructions, and with inference data such as audio data representing speech, text data representing text, and image data representing images. The data generation model 58 infers from the input inference data according to the instructions indicated by the prompts, and outputs the inference results in data formats such as audio data and text data. Here, inference refers to, for example, analysis, classification, prediction, and / or summarization.
[0733] In the above embodiment, an example was given in which the specific processing is performed by the data processing device 12, but the technology of this disclosure is not limited thereto, and the specific processing may also be performed by the robot 414.
[0734] Furthermore, the emotion identification model 59, acting as an emotion engine, may determine the user's emotion according to a specific mapping. Specifically, the emotion identification model 59 may determine the user's emotion according to a specific mapping, which is an emotion map (see Figure 9). Similarly, the emotion identification model 59 may also determine the robot's emotion, and the identification processing unit 290 may perform identification processing using the robot's emotion.
[0735] Figure 9 shows an emotion map 400 in which multiple emotions are mapped. In the emotion map 400, emotions are arranged in concentric circles radiating from the center. The closer to the center of the concentric circles, the more primitive the emotions are located. Further out of the concentric circles, emotions representing states and actions arising from mental states are located. Emotion is a concept that includes feelings and mental states. On the left side of the concentric circles, emotions that are generally generated from reactions occurring in the brain are located. On the right side of the concentric circles, emotions that are generally induced by situational judgment are located. Above and below the concentric circles, emotions that are generally generated from reactions occurring in the brain and induced by situational judgment are located. In addition, the emotion of "pleasure" is located on the upper side of the concentric circles, and the emotion of "displeasure" is located on the lower side. Thus, in the emotion map 400, multiple emotions are mapped based on the structure in which emotions arise, and emotions that are likely to occur simultaneously are mapped close together.
[0736] These emotions are distributed at the 3 o'clock position on the Emotion Map 400, and usually fluctuate between feelings of security and anxiety. In the right half of the Emotion Map 400, situational awareness takes precedence over internal feelings, resulting in a calm impression.
[0737] The inside of the Emotion Map 400 represents inner thoughts, while the outside represents actions. Therefore, the further you go from the outside of the Emotion Map 400, the more visible (expressed in actions) your emotions become.
[0738] Here, human emotions are based on various balances, such as posture and blood sugar levels. When these balances deviate from the ideal, it results in discomfort, and when they approach the ideal, it results in pleasure. Similarly, in robots, cars, motorcycles, etc., emotions can be created based on various balances, such as posture and battery level. When these balances deviate from the ideal, it results in discomfort, and when they approach the ideal, it results in pleasure. The emotion map can be generated, for example, based on Dr. Mitsuyoshi's emotion map (Research on a system for analyzing brain physiological signals of speech emotion recognition and emotion, Tokushima University, doctoral dissertation: https: / / ci.nii.ac.jp / naid / 500000375379). The left half of the emotion map contains emotions belonging to a region called "response," where sensation is dominant. The right half of the emotion map contains emotions belonging to a region called "situation," where situational awareness is dominant.
[0739] The emotion map defines two emotions that promote learning. One is the emotion around the middle of the negative "repentance" and "reflection" on the situation side. In other words, it is when the robot experiences negative emotions such as "I never want to feel this way again" or "I don't want to be scolded again." The other is the emotion around the positive "desire" on the reaction side. In other words, it is when the robot has positive feelings such as "I want more" or "I want to know more."
[0740] The emotion identification model 59 inputs user input into a pre-trained neural network, obtains emotion values representing each emotion shown in the emotion map 400, and determines the user's emotion. This neural network is pre-trained based on multiple training data sets, which are combinations of user input and emotion values representing each emotion shown in the emotion map 400. Furthermore, this neural network is trained so that emotions located close together have similar values, as shown in the emotion map 900 in Figure 10. Figure 10 shows an example where multiple emotions such as "reassured," "calm," and "confident" have similar emotion values.
[0741] The above description primarily focuses on the functions of the data processing device 12 in relation to this disclosure. However, the system related to this disclosure is not necessarily implemented on a server. The system related to this disclosure may be implemented as a general information processing system. This disclosure may be implemented, for example, as a software program that runs on a personal computer or as an application that runs on a smartphone. The method related to this disclosure may be provided to users in SaaS (Software as a Service) format.
[0742] In the above embodiment, an example was given in which a specific process is performed by a single computer 22. However, the technology of this disclosure is not limited thereto, and a distributed processing of the specific process may be performed by multiple computers, including computer 22. For example, a data generation model 58 may be provided in an external device of the data processing device 12, and the external device may generate data according to the input data.
[0743] In the above embodiment, an example was given in which the specific processing program 56 is stored in the storage 32, but the technology of this disclosure is not limited thereto. For example, the specific processing program 56 may be stored in a portable, computer-readable, non-temporary storage medium such as a USB (Universal Serial Bus) memory. The specific processing program 56 stored in the non-temporary storage medium is installed in the computer 22 of the data processing device 12. The processor 28 executes specific processing according to the specific processing program 56.
[0744] Alternatively, the specific processing program 56 may be stored in a storage device such as a server connected to the data processing device 12 via the network 54, and the specific processing program 56 may be downloaded and installed on the computer 22 in response to a request from the data processing device 12.
[0745] Furthermore, it is not necessary to store the entirety of the specific processing program 56 in a storage device such as a server connected to the data processing device 12 via the network 54, or to store the entirety of the specific processing program 56 in the storage 32; it is acceptable to store only a portion of the specific processing program 56.
[0746] The following types of processors can be used as hardware resources to perform specific processing. Examples of processors include a CPU, a general-purpose processor that functions as a hardware resource to perform specific processing by executing software, i.e., a program. Other examples of processors include dedicated electrical circuits, such as FPGAs (Field-Programmable Gate Arrays), PLDs (Programmable Logic Devices), or ASICs (Application Specific Integrated Circuits), which have circuit configurations specifically designed to perform specific processing. All of these processors have built-in or connected memory, and all of them perform specific processing by using memory.
[0747] The hardware resource that performs a specific process may consist of one of these various processors, or it may consist of a combination of two or more processors of the same or different types (for example, a combination of multiple FPGAs, or a combination of a CPU and an FPGA). Alternatively, the hardware resource that performs a specific process may consist of a single processor.
[0748] Examples of configurations using a single processor include, firstly, a configuration in which one or more CPUs and software are combined to form a single processor, and this processor functions as a hardware resource that performs a specific process. Secondly, there is a configuration using a processor that realizes the functions of the entire system, including multiple hardware resources that perform a specific process, on a single IC chip, as exemplified by SoCs (System-on-a-chip). In this way, a specific process is realized using one or more of the above types of processors as hardware resources.
[0749] Furthermore, the hardware structure of these various processors can more specifically utilize electrical circuits that combine circuit elements such as semiconductor devices. Also, the specific processing described above is merely an example. Therefore, it goes without saying that unnecessary steps can be deleted, new steps added, or the processing order rearranged, as long as it does not deviate from the main purpose.
[0750] The descriptions and illustrations presented above are detailed explanations of the technical aspects of this disclosure and are merely examples of the technical aspects. For example, the above descriptions of the structure, function, operation, and effect are examples of the structure, function, operation, and effect of the technical aspects of this disclosure. Therefore, it goes without saying that you may delete unnecessary parts, add new elements, or replace elements in the descriptions and illustrations presented above, as long as you do not deviate from the essence of the technical aspects of this disclosure. Furthermore, in order to avoid confusion and facilitate understanding of the technical aspects of this disclosure, explanations of common technical knowledge and the like that do not require special explanation to enable the implementation of the technical aspects of this disclosure have been omitted from the descriptions and illustrations presented above.
[0751] All documents, patent applications, and technical standards described herein are incorporated by reference to the same extent as if each individual document, patent application, and technical standard were specifically and individually noted to be incorporated by reference.
[0752] The following is further disclosed regarding the embodiments described above.
[0753] (Claim 1)
[0754] A means for receiving video data including motion capture by the user,
[0755] A means for extracting feature points from the aforementioned video data,
[0756] A means for comparing other video data with the aforementioned feature points,
[0757] A means for generating advice for improving user behavior based on comparison results,
[0758] Means for providing the aforementioned advice to the user,
[0759] A system that includes this.
[0760] (Claim 2)
[0761] The system according to claim 1, which uses multimodal technology to analyze audio and text data from the other video data and generates the advice.
[0762] (Claim 3)
[0763] The system according to claim 1, which uses a machine learning model for extracting and comparing the aforementioned feature points.
[0764] "Example 1"
[0765] (Claim 1)
[0766] A medium for receiving video data including motion capture by the user,
[0767] A medium for extracting feature points such as body coordinates and angles from the aforementioned video data,
[0768] A medium used for comparison with other reference video data based on the aforementioned feature points,
[0769] A medium that uses deep learning technology to generate suggestions for improving user behavior based on comparison results,
[0770] The medium through which the aforementioned proposal is provided to the user,
[0771] A system that includes this.
[0772] (Claim 2)
[0773] The system according to claim 1, which uses multimodal technology to analyze various types of information, analyzes audio and text information from the other video data, and generates the proposal.
[0774] (Claim 3)
[0775] The system according to claim 1, which uses a generative AI model for extracting and comparing the aforementioned feature points.
[0776] "Application Example 1"
[0777] (Claim 1)
[0778] A device that receives video data including motion capture by the user,
[0779] A function to extract feature points from the aforementioned video data,
[0780] A function to compare based on reference data and the aforementioned feature points,
[0781] A device that generates advice for improving user behavior based on comparison results,
[0782] A device that provides the aforementioned advice to the user,
[0783] A device that films the operation of the device and analyzes its feature points in real time,
[0784] A device that provides specific guidance to improve the efficiency of operations based on the difference from the standard operation,
[0785] A system that includes this.
[0786] (Claim 2)
[0787] The system according to claim 1, which uses multimodal technology to analyze acoustic and textual information from the reference data and generates the advice.
[0788] (Claim 3)
[0789] The system according to claim 1, wherein a learning algorithm is used for extracting and comparing the aforementioned feature points.
[0790] "Example 2 of combining an emotion engine"
[0791] (Claim 1)
[0792] A means for receiving visual information, including motor movements acquired by the user,
[0793] Means for extracting characteristic information from the aforementioned visual information,
[0794] A means for comparing standard video data with the aforementioned characteristic information,
[0795] A means for performing the above comparison using a machine learning algorithm and obtaining the analysis results,
[0796] An emotion recognition means that analyzes video or audio to recognize the user's emotional state,
[0797] A means for generating advice that considers user behavior improvement and psychological impact based on the aforementioned analysis results and emotional state,
[0798] Means for providing the aforementioned advice to the user,
[0799] A system that includes this.
[0800] (Claim 2)
[0801] The system according to claim 1, which analyzes audio and text information from the standard video data using various data formats and generates the advice.
[0802] (Claim 3)
[0803] The system according to claim 1, which uses an artificial intelligence model for extracting and comparing the aforementioned characteristic information.
[0804] "Application example 2 when combining with an emotional engine"
[0805] (Claim 1)
[0806] A means for receiving video information including motion capture by the user,
[0807] A means for extracting the characteristics of the motion from the aforementioned video information,
[0808] A means for comparing other video information with the characteristics of the aforementioned operation,
[0809] A means for generating advice for improving user behavior based on comparison results,
[0810] A means of analyzing the user's emotional state and adjusting the advice based on the results,
[0811] Means for providing the aforementioned advice to the user,
[0812] A system that includes this.
[0813] (Claim 2)
[0814] The system according to claim 1, which uses multimodal technology to analyze audio and document information from the other video information and generates the advice.
[0815] (Claim 3)
[0816] The system according to claim 1, which uses a machine learning model to extract and compare the characteristics of the aforementioned operations. [Explanation of Symbols]
[0817] 10, 210, 310, 410 Data Processing Systems 12 Data Processing Devices 14 Smart Devices 214 Smart Glasses 314 Headset-type terminal 414 Robots< / url:> < / url:> < / url:> < / url:>
Claims
1. A means for receiving video data including motion capture by the user, A means for extracting feature points from the aforementioned video data, A means for comparing other video data with the aforementioned feature points, A means for generating advice for improving user behavior based on comparison results, Means for providing the aforementioned advice to the user, A system that includes this.
2. The system according to claim 1, which uses multimodal technology to analyze audio and text data from the other video data and generates the advice.
3. The system according to claim 1, wherein a machine learning model is used for extracting and comparing the aforementioned feature points.
Citation Information
Patent Citations
Persona chatbot control method and system
JP2022180282A