System
The system provides objective and detailed AI-driven feedback to athletes, addressing the challenge of subjective form analysis and enabling efficient, customized training improvements.
Patent Information
- Application Number
- JP2024123956
- Authority / Receiving Office
- JP · JP
- Patent Type
- Applications
- Current Assignee / Owner
- Filing Date
- 2024-07-30
- Publication Date
- 2026-02-12
AI Technical Summary
Athletes struggle to objectively analyze their form and receive specific, customized training guidance, as conventional methods rely on subjective evaluations and are time-consuming and expensive.
A system that includes a preprocessing unit, a motion analysis model, and a feedback generation unit to analyze exercisers' form using AI, providing detailed and efficient feedback.
Enables athletes to identify specific form issues and receive customized training advice efficiently, improving their performance through continuous feedback loops.
Smart Images

Figure 2026022439000001_ABST
Abstract
Description
[Technical Field]
[0001] The technology of the present disclosure relates to a system. [Background technology]
[0002] Patent document 1 discloses a persona chatbot control method performed by at least one processor, the method including the steps of receiving a user utterance, adding the user utterance to a prompt including an instruction sentence related to a description of the chatbot character, encoding the prompt, and inputting the encoded prompt into a language model to generate a chatbot utterance in response to the user utterance. [Prior art documents] [Patent documents]
[0003] [Patent Document 1] Japanese Patent Publication No. 2022-180282 Summary of the Invention [Problem to be solved by the invention]
[0004] In sports, the quality of an athlete's form has a significant impact on performance. However, many athletes find it difficult to accurately identify problems with their own form and know how to improve. Conventional training methods rely on subjective evaluations and advice from limited perspectives, often failing to provide objective and detailed feedback. Furthermore, receiving individually customized training instruction requires significant time and expense. Therefore, there is a need for a means to objectively analyze problems with athletes' form and provide specific improvement measures. [Means for solving the problem]
[0005] The present invention relates to a system that includes a preprocessing unit for receiving video data captured by an exerciser and processing each frame contained in the video data, a motion analysis model for analyzing the preprocessed frames as input, a motion analysis model for generating feedback on the exerciser's form based on the analysis results, and a motion analysis model for providing the generated feedback to the exerciser. This system allows exercisers to analyze their own exercise form in detail and clearly identify specific issues and areas for improvement. Furthermore, the use of an artificial intelligence model enables efficient and accurate form analysis, allowing exercisers to easily receive individually customized training guidance.
[0006] An "athlete" is an individual who engages in sports or fitness.
[0007] "Video data" refers to moving images recorded in digital format, and is a video file that includes the movements and form of an athlete.
[0008] "Means for receiving" refers to communication equipment or software for receiving, storing, or processing data transmitted from an external source.
[0009] A "preprocessing means" is a device or algorithm that performs operations to convert each frame in the video data into a format suitable for the analytical model.
[0010] A "motion analysis model" is an artificial intelligence model trained to analyze an athlete's form and evaluate and provide feedback on their movements.
[0011] The "analyzing means" is a device or software that inputs the preprocessed data into a motion analysis model and obtains the results.
[0012] The "means for generating feedback" is a device or algorithm that generates an evaluation and comments on areas for improvement for the athlete based on the analysis results.
[0013] The "means for providing" refers to a communication means or a display device for presenting the generated feedback to the exerciser. [Brief explanation of the drawings]
[0014] [Figure 1] 1 is a conceptual diagram showing an example of the configuration of a data processing system according to a first embodiment. [Figure 2] 1 is a conceptual diagram showing an example of main functions of a data processing device and a smart device according to a first embodiment. [Figure 3] FIG. 10 is a conceptual diagram showing an example of the configuration of a data processing system according to a second embodiment. [Figure 4] FIG. 10 is a conceptual diagram showing an example of main functions of a data processing device and smart glasses according to a second embodiment. [Figure 5] FIG. 10 is a conceptual diagram showing an example of the configuration of a data processing system according to a third embodiment. [Figure 6] FIG. 11 is a conceptual diagram showing an example of main functions of a data processing device and a headset-type terminal according to a third embodiment. [Figure 7] FIG. 10 is a conceptual diagram showing an example of the configuration of a data processing system according to a fourth embodiment. [Figure 8] FIG. 10 is a conceptual diagram showing an example of main functions of a data processing device and a robot according to a fourth embodiment. [Figure 9] 1 shows an emotion map onto which multiple emotions are mapped. [Figure 10] 1 shows an emotion map onto which multiple emotions are mapped. [Figure 11] FIG. 3 is a sequence diagram showing a processing flow of the data processing system according to the first embodiment. [Figure 12] FIG. 10 is a sequence diagram showing the flow of processing in the data processing system in Application Example 1. [Figure 13] FIG. 10 is a sequence diagram showing the flow of processing in the data processing system according to the second embodiment when an emotion engine is combined. [Figure 14]FIG. 10 is a sequence diagram showing the flow of processing in the data processing system in Application Example 2 when an emotion engine is combined. DETAILED DESCRIPTION OF THE INVENTION
[0015] An example of an embodiment of a system according to the technology of the present disclosure will be described below with reference to the accompanying drawings.
[0016] First, the terms used in the following description will be explained.
[0017] In the following embodiments, a coded processor (hereinafter simply referred to as a "processor") may be a single arithmetic device or a combination of multiple arithmetic devices. Furthermore, a processor may be a single type of arithmetic device or a combination of multiple types of arithmetic devices. Examples of arithmetic devices include a CPU (Central Processing Unit), a GPU (Graphics Processing Unit), a GPGPU (General-Purpose computing on Graphics Processing Units), and an APU (Accelerated Processing Unit).
[0018] In the following embodiments, a coded RAM (Random Access Memory) is a memory in which information is temporarily stored and is used as a working memory by a processor.
[0019] In the following embodiments, the coded storage is one or more non-volatile storage devices that store various programs, various parameters, etc. Examples of non-volatile storage devices include flash memory (SSD (Solid State Drive)), magnetic disks (e.g., hard disks), and magnetic tapes.
[0020] In the following embodiments, a communication I / F (Interface) with a symbol is an interface including a communication processor, an antenna, etc. The communication I / F controls communication between multiple computers. Examples of communication standards applied to the communication I / F include wireless communication standards including 5G (5th Generation Mobile Communication System), Wi-Fi (registered trademark), Bluetooth (registered trademark), etc.
[0021] In the following embodiments, "A and / or B" is synonymous with "at least one of A and B." In other words, "A and / or B" means that it may be only A, only B, or a combination of A and B. Furthermore, in this specification, the same concept as "A and / or B" is also applied when three or more things are expressed connected by "and / or."
[0022] [First embodiment]
[0023] FIG. 1 shows an example of the configuration of a data processing system 10 according to the first embodiment.
[0024] 1, a data processing system 10 includes a data processing device 12 and a smart device 14. An example of the data processing device 12 is a server.
[0025] The data processing device 12 includes a computer 22, a database 24, and a communication I / F 26. The computer 22 is an example of a "computer" according to the technology of the present disclosure. The computer 22 includes a processor 28, a RAM 30, and a storage 32. The processor 28, the RAM 30, and the storage 32 are connected to a bus 34. The database 24 and the communication I / F 26 are also connected to the bus 34. The communication I / F 26 is connected to a network 54. Examples of the network 54 include a WAN (Wide Area Network) and / or a LAN (Local Area Network).
[0026] The smart device 14 includes a computer 36, a reception device 38, an output device 40, a camera 42, and a communication I / F 44. The computer 36 includes a processor 46, a RAM 48, and a storage 50. The processor 46, the RAM 48, and the storage 50 are connected to a bus 52. The reception device 38, the output device 40, and the camera 42 are also connected to the bus 52.
[0027] The reception device 38 includes a touch panel 38A, a microphone 38B, and the like, and receives user input. The touch panel 38A detects contact with an indicator (for example, a pen or a finger) to receive user input by the touch of the indicator. The microphone 38B detects the user's voice to receive user input by voice. The control unit 46A transmits data indicating the user input received by the touch panel 38A and the microphone 38B to the data processing device 12. In the data processing device 12, the specific processing unit 290 acquires the data indicating the user input.
[0028] The output device 40 includes a display 40A and a speaker 40B, and presents data to the user 20 by outputting the data in a form of expression that the user 20 can perceive (for example, audio and / or text). The display 40A displays visible information such as text and images in accordance with instructions from the processor 46. The speaker 40B outputs audio in accordance with instructions from the processor 46. The camera 42 is a compact digital camera equipped with an optical system including a lens, aperture, and shutter, and an imaging element such as a CMOS (Complementary Metal-Oxide-Semiconductor) image sensor or a CCD (Charge Coupled Device) image sensor.
[0029] The communication I / F 44 is connected to a network 54. The communication I / Fs 44 and 26 control the exchange of various information between the processor 46 and the processor 28 via the network 54.
[0030] FIG. 2 shows an example of the main functions of the data processing device 12 and the smart device 14.
[0031] 2, in the data processing device 12, a specific process is performed by the processor 28. A specific processing program 56 is stored in the storage 32. The specific processing program 56 is an example of a "program" according to the technology of the present disclosure. The processor 28 reads the specific processing program 56 from the storage 32 and executes the read specific processing program 56 on the RAM 30. The specific process is realized by the processor 28 operating as a specific processing unit 290 in accordance with the specific processing program 56 executed on the RAM 30.
[0032] The storage 32 stores a data generation model 58 and an emotion identification model 59. The data generation model 58 and the emotion identification model 59 are used by the identification processing unit 290.
[0033] In the smart device 14, the processor 46 performs the reception output process. The storage 50 stores a reception output program 60. The reception output program 60 is used in conjunction with the specific processing program 56 by the data processing system 10. The processor 46 reads the reception output program 60 from the storage 50 and executes the read reception output program 60 on the RAM 48. The reception output process is realized by the processor 46 operating as the control unit 46A in accordance with the reception output program 60 executed on the RAM 48.
[0034] Next, a description will be given of the specific processing performed by the specific processing unit 290 of the data processing device 12. In the following description, the data processing device 12 will be referred to as a "server" and the smart device 14 will be referred to as a "terminal."
[0035] This invention relates to a system that allows athletes to analyze their own form and receive specific advice for improving their technique. This system allows athletes to take video data using devices such as smartphones or computers, upload it to a server, and then analyzes the video data using an AI model to provide detailed feedback.
[0036] Specific Embodiments of the System
[0037] 1. Upload your video
[0038] Users upload videos of their exercise form from their devices to the server, where the uploaded videos are temporarily stored.
[0039] 2. Receiving and storing video data
[0040] The server receives the video data sent by the user and saves it under a file name such as "uploaded_video.mp4." This saved video data will then be analyzed.
[0041] 3. Preparing and Loading the AI Model
[0042] The server loads a pre-trained motion analysis model (e.g., a deep learning model) that analyzes an athlete's form and provides an assessment of specific movements.
[0043] 4. Loading the video and splitting it into frames
[0044] The server extracts the stored video file frame by frame. Since video is usually 30 frames per second, it is processed as 30 images per second.
[0045] 5. Frame Preprocessing
[0046] The server performs preprocessing such as frame resizing and normalization to convert each frame into a format suitable for the analytical model. Specifically, the frame is resized to 224x224 pixels and the pixel values are normalized to the range 0 to 1.
[0047] 6. Form Analysis
[0048] The server then feeds the preprocessed frames into a motion analysis model to evaluate form, measuring specific movement parameters such as knee angle and body tilt, and returns the results.
[0049] 7. Generate feedback
[0050] The server generates feedback for each frame based on the analysis results, and provides the user with specific advice such as "Good form. Keep going like this" or "Review the way you bend your knees."
[0051] 8. Providing Feedback
[0052] The server compiles the generated feedback and sends it back to the user's device, which then displays it to the user in a visually understandable format.
[0053] 9. Review feedback and incorporate it into your training plan
[0054] Users can view the feedback displayed on their device to understand where their form is lacking, and use that feedback to restructure their training plan and address specific areas for improvement.
[0055] 10. Re-recording and analyzing the video
[0056] Users can re-record a video with a new form and re-upload it into the system, repeating this process to continually improve their form.
[0057] As described above, the present invention allows athletes to improve the quality of their training by analyzing their own form in detail and receiving specific, individually customized feedback. This system is expected to make sports skill improvement even more efficient.
[0058] The processing flow will be explained below.
[0059] Step 1:
[0060] User: Record a video of their exercise form using a smartphone or computer. Review the video to ensure the necessary parts are properly recorded.
[0061] Step 2:
[0062] User: Access the dedicated upload page from their device, select the video file they shot on that page, and click the upload button.
[0063] Step 3:
[0064] Terminal: The selected video file is sent to the server as binary data. Once the sending is complete, a message indicating success is displayed to the user.
[0065] Step 4:
[0066] Server: Receives the binary data sent from the device. Temporarily saves the received data as "uploaded_video.mp4".
[0067] Step 5:
[0068] Server: Uses the OpenCV library to read the saved video file. Opens the video file and prepares to extract data frame by frame.
[0069] Step 6:
[0070] Server: Loads a pre-trained motion analysis model (e.g., "pose_analysis_model.h5") that analyzes an athlete's form and evaluates specific movements.
[0071] Step 7:
[0072] Server: Divide the video into frames. For example, if the video is shot at 30 frames per second, extract 30 frames per second.
[0073] Step 8:
[0074] Server: Preprocess each frame by resizing it to 224x224 pixels and normalizing pixel values to the range 0 to 1.
[0075] Step 9:
[0076] Server: Inputs the preprocessed frames into the motion analysis model, which outputs the motion estimation results for each frame.
[0077] Step 10:
[0078] Server: Generates feedback based on the analysis results. For example, if the model determines that the knee angle is inappropriate, it generates specific advice such as "Reconsider how you bend your knees."
[0079] Step 11:
[0080] Server: Compiles the generated feedback into a list and prepares it for sending back to the user.
[0081] Step 12:
[0082] Server: Sends the feedback list to the user's terminal.
[0083] Step 13:
[0084] Terminal: Receives feedback sent from the server and displays it to the user. It presents comments in a format that is easy for the user to understand.
[0085] Step 14:
[0086] User: Check the feedback displayed on the device to understand problems and areas for improvement in their exercise form, and restructure their training plan based on this.
[0087] Step 15:
[0088] User: Record a video with improved form, upload it to the system in the same way, and request analysis. By repeating this process, you can continuously improve your form.
[0089] Example 1
[0090] Next, a description will be given of Example 1. In the following description, the data processing device 12 will be referred to as a "server" and the smart device 14 will be referred to as a "terminal."
[0091] Conventional exercise form analysis systems have the drawback of requiring users to film their own form, send the video data to a dedicated device, and obtain analysis results, all of which require a complex and time-consuming process. Furthermore, the accuracy of the analysis results is low, making it difficult for users to identify specific areas for improvement. Furthermore, many systems do not provide instant feedback, making it difficult to incorporate the results into effective training plans.
[0092] The specific processing by the specific processing unit 290 of the data processing device 12 in the first embodiment is realized by the following means.
[0093] In this invention, the server includes means for receiving video data taken by an exerciser, means for saving the video data, means for dividing the video data into frames, means for preprocessing each of the frames, means for analyzing the preprocessed frames using a generative AI model, means for generating feedback on the exerciser's form based on the analysis results, means for providing the generated feedback to the exerciser, and means for displaying the provided feedback on the exerciser's device. This provides high-precision feedback in real time, allowing the user to quickly identify areas for improvement in their form and implement an effective training plan.
[0094] An "athlete" refers to an individual who exercises, photographs their own exercise form, and uses the data for analysis.
[0095] "Video data" refers to video files taken by the athlete, and contains information to be analyzed.
[0096] "Server" refers to a computer system that receives, stores, processes, and provides analysis results of video data.
[0097] "Each frame" refers to the individual images that make up the video data, and the information per frame is used for analysis.
[0098] "Preprocessing" refers to data manipulation such as resizing and normalization to convert each frame into a format suitable for the analytical model.
[0099] A "generative AI model" is a pre-trained AI (artificial intelligence) model used to analyze and evaluate exercise form.
[0100] "Analysis results" refers to the evaluation data of exercise form obtained by the generative AI model, and includes specific movement parameters (such as knee angle and body tilt).
[0101] "Feedback" refers to advice and evaluations generated based on the analysis results, and is information that athletes can use to improve their form.
[0102] "Terminal" refers to the device used by the exerciser, such as a smartphone or computer, on which feedback is displayed.
[0103] This invention relates to a system that allows athletes to analyze their own form and receive specific advice for improving their technique. The system allows athletes to take video data using devices such as smartphones or computers, upload it to a server, and then analyzes the video data using a generative AI model to provide detailed feedback.
[0104] Hardware and software used
[0105] Hardware: smartphones, computers, servers
[0106] Software: Video recording app, video uploading tool, motion analysis AI model (deep learning model)
[0107] Data processing and calculation methods
[0108] Users can upload videos of their exercise form to a server using a smartphone or computer, for example, via a dedicated application or browser.
[0109] The server receives the video data sent by the user and temporarily stores it in a specific directory (e.g., "uploaded_videos / ") on the server with the file name "uploaded_video.mp4". This data is used in the subsequent analysis steps.
[0110] The server loads pre-trained AI models for motion analysis, which use deep learning techniques to analyze athletic form, including models trained using TensorFlow and PyTorch.
[0111] The server loads the saved video file and uses tools such as ffmpeg to extract the video frame by frame. Since video is typically 30 frames per second, this translates to 30 images per second.
[0112] Each frame is converted into a format suitable for the analytical model. In this step, the frame is resized to 224x224 pixels and the pixel values are normalized to the range 0 to 1. This process is often performed using OpenCV or PIL (Python Imaging Library).
[0113] The preprocessed frames are fed into a generative AI model to evaluate the form. The model measures movement parameters such as knee angle and body tilt and returns the analysis results. For example, the model outputs data such as "knee angle is 45 degrees" for each frame.
[0114] The server generates feedback based on the analysis results, including specific advice such as "Try to widen your knee angle by about 5 degrees." This feedback is created for each frame and is customized for each user.
[0115] The generated feedback is sent back from the server to the user's device, which then displays it to the user in a visually understandable format, for example, highlighting key points in addition to displaying them as text messages within the application.
[0116] Users can view the feedback displayed on their device, understand the issues with their form, and use that feedback to restructure their training plan and address specific areas for improvement, such as paying special attention to knee angle during their next workout.
[0117] Examples of concrete examples and prompts
[0118] For example, suppose a user wants to improve their running form. The user films their running form with their smartphone and uploads the video to the system. The server analyzes the video and generates feedback to the user, such as, "The angle of your right knee is not appropriate. Try bending your right knee a bit more while running." The user can then revise their training plan based on this feedback and review their form again.
[0119] Prompt Sentence Examples
[0120] "I uploaded a video of my running form. Please help me analyze where my form needs improvement."
[0121] "Please analyze the video of my form and tell me specific areas for improvement."
[0122] The system allows exercisers to receive detailed and specific feedback to continually improve their form.
[0123] The flow of the identification process in the first embodiment will be described with reference to FIG.
[0124] Step 1:
[0125] The user takes a video of their exercise form on the device.
[0126] Input: Video of the exercise form (e.g. "run.mp4")
[0127] Output: Recorded video file
[0128] Specific movements: Using a smartphone or computer, participants will be videotaped while performing various exercises, such as running, jumping, and squatting.
[0129] Step 2:
[0130] The user uploads the video they have taken from their terminal to the server.
[0131] Input: Recorded video file
[0132] Output: Video data uploaded to the server
[0133] Specific operation: Use a dedicated application or web browser to send the video file to the server. For example, click the button, select "run.mp4", and start uploading.
[0134] Step 3:
[0135] The server receives and stores the video data sent by the user.
[0136] Input: Video data received from the user
[0137] Output: Video file saved on the server (e.g. "uploaded_video.mp4")
[0138] What happens: The server stores the uploaded data in a specific directory (e.g. "uploaded_videos / ").
[0139] Step 4:
[0140] The server loads a pre-trained generative AI model.
[0141] Input: A trained generative AI model
[0142] Output: Generative AI model loaded into memory
[0143] What it does: Loads a motion analysis model trained using TensorFlow or PyTorch, and the system is ready for analysis.
[0144] Step 5:
[0145] The server reads the saved video file and splits it into frames.
[0146] Input: Video file stored on the server ("uploaded_video.mp4")
[0147] Output: Images separated by frames (e.g. frame 1, frame 2, ...)
[0148] Specific operation: Using a video processing tool such as ffmpeg, the video is divided into 30 frames per second and each frame is obtained as an image file.
[0149] Step 6:
[0150] The server performs pre-processing on each frame.
[0151] Input: Images split into frames
[0152] Output: Preprocessed frame (resized and normalized)
[0153] Specific operation: Using OpenCV and PIL (Python Imaging Library), each frame is resized to 224x224 pixels and the value of each pixel is normalized to the range 0 to 1.
[0154] Step 7:
[0155] The server feeds the preprocessed frames into a generative AI model to analyze the form.
[0156] Input: Preprocessed frame
[0157] Output: Analysis results (motion parameters, e.g. knee angle, body tilt)
[0158] Specific movement: The preprocessed frame is input into the generative AI model, movement parameters are measured, and an evaluation result is output. For example, the result may be "Frame 10: Knee angle 45 degrees."
[0159] Step 8:
[0160] The server generates feedback based on the analysis results.
[0161] Input: Analysis results
[0162] Output: Generated feedback (e.g., "Try to open your knees by about 5 degrees.")
[0163] Specific actions: Based on the analysis results for each frame, specific advice on how to improve your athletic form is generated. For example, "Your right knee angle is not appropriate, so try bending it a little more while running."
[0164] Step 9:
[0165] The server transmits the generated feedback to the user's terminal.
[0166] Input: Generated feedback
[0167] Output: Feedback sent to the user's device
[0168] Specific operation: The server compiles the feedback data and sends it to the user's device, for example, by sending the feedback data back using an HTTP request.
[0169] Step 10:
[0170] The terminal displays the received feedback to the user.
[0171] Input: Feedback sent by the server
[0172] Output: The feedback screen presented to the user.
[0173] What it does: The device visualizes the feedback it receives and presents it to the user in a user-friendly format, for example by using text messages or highlighting to emphasize key points.
[0174] Step 11:
[0175] The user checks the feedback displayed on the device and reflects it in their training plan.
[0176] Input: Feedback displayed on the device
[0177] Output: Modified training plan
[0178] What it does: Users can use the feedback to understand where their form needs improvement and focus on those areas during their next workout.
[0179] Step 12:
[0180] The user then shoots the video again with the improved form and uploads it to the system again.
[0181] Input: Improved video
[0182] Output: New video data uploaded to the system
[0183] What it does: The user takes a photo of a new form and repeats the steps above to provide data to the system, which then continually improves the form.
[0184] (Application example 1)
[0185] Next, a description will be given of Application Example 1. In the following description, the data processing device 12 will be referred to as a "server" and the smart device 14 will be referred to as a "terminal."
[0186] Conventional motion analysis systems have been used only to provide feedback to athletes and general exercisers to improve their exercise form, but the concept has not been applied to the motion analysis of exercise machines and equipment in factories. In factories, accurate analysis of motion and appropriate feedback are essential to maintain and improve the efficiency and accuracy of highly automated equipment and robots. However, current systems have difficulty meeting these needs, and have the problem of being unable to respond immediately when a decline in motion efficiency or accuracy occurs.
[0187] The specific processing by the specific processing unit 290 of the data processing device 12 in the application example 1 is realized by the following means.
[0188] In this invention, the server includes: means for receiving video data captured by an exerciser; preprocessing means for processing each frame included in the video data; means for analyzing the preprocessed frames using a motion analysis model; means for generating feedback on the exerciser's form based on the analysis results; means for providing the generated feedback to the exerciser; a device for capturing the movements of exercise equipment or devices in a factory; means for receiving and preprocessing the captured movement data; means for evaluating movements using an AI model for analyzing the preprocessed movement data; means for generating feedback on improving the efficiency and accuracy of movements based on the evaluation results; and means for providing the generated feedback to personnel involved in operating the factory equipment. This enables the movement analysis of exercise equipment or devices in a factory, and enables quick and appropriate feedback to be provided to improve efficiency and accuracy.
[0189] "Exercise participant" refers to the person or equipment whose exercise form is being analyzed and feedback is being provided.
[0190] "Video data" refers to data containing visual information that records the movements of an exerciser or the movements of exercise equipment.
[0191] "Means for receiving" refers to the technical methods and equipment for accepting video data from an external device.
[0192] "Preprocessing means" refers to a method or device that performs processing to convert video data into a format suitable for the analysis model.
[0193] "Movement analysis model" refers to an artificial intelligence model used to analyze an athlete's form and movement data.
[0194] "Means for analysis" refers to the techniques and devices used to input preprocessed data into a motion analysis model and obtain results.
[0195] "Means for generating feedback" refers to techniques and devices for generating information that indicates specific areas for improvement and direction based on the analysis results.
[0196] "Means of providing" refers to the techniques and devices used to communicate the generated feedback to the exerciser or person in charge.
[0197] "Motion equipment or devices" refers to machines, robots, and other equipment that perform specific movements in a factory.
[0198] "Motion Data" refers to data, including visual information, that records the movements of exercise equipment or devices.
[0199] "Means for evaluating" refers to techniques or devices for evaluating the efficiency and accuracy of a movement using a motion analysis model based on preprocessed movement data.
[0200] This invention is a system based on the "athlete's form analysis system" that is applied to the analysis of the movements of exercise machines or equipment in factories and the provision of feedback. This system can monitor the efficiency and accuracy of the movements of robots and machines in factories and provide necessary advice for improvement.
[0201] First, cameras are placed to capture the operation of the motion machines and equipment used in the factory. This video data is sent to a server. The server then acquires this video data using a receiving means. The hardware used in this case can be a high-performance camera.
[0202] The server then processes the received video data using preprocessing tools. Specifically, it splits the video data into frames and converts them into a format suitable for the AI model. This preprocessing uses OpenCV and PIL (Pillow) to resize and normalize the images, allowing for detailed analysis of the motion.
[0203] The preprocessed data is then input into an AI model, the motion analysis model. This model uses deep learning to analyze the motion frame by frame and measure specific motion parameters. PyTorch is used for the analysis, and the model's motion is evaluated.
[0204] The server generates feedback regarding improvements to the efficiency and accuracy of the operation based on the analysis results. The feedback generation means creates information indicating specific advice and areas for improvement based on the analysis results. For example, feedback such as "The accuracy of the operation is low. Adjustments are required" may be generated.
[0205] Finally, the server provides the generated feedback to personnel involved in the operation of the factory equipment using a means for providing the feedback, which can then be used by the personnel to adjust the robots and equipment to improve the efficiency and accuracy of their operations.
[0206] As a specific example, a robot arm in a factory is filmed as it assembles parts, and the video data is uploaded to the system. The system analyzes the video data and generates specific feedback such as, "The robot arm's movement is delayed. Please adjust the angle at which the part is attached." Based on this, the worker can adjust the robot arm's movement, improving the efficiency of the assembly work.
[0207] An example of a prompt sentence to be input to the generative AI model is as follows:
[0208] Consider a scenario where a system analyzes the behavior of a factory robot and provides feedback on its efficiency and accuracy. Specifically, imagine a robotic arm takes a video of itself assembling a part, and an AI model analyzes its behavior and provides advice on how to improve it.
[0209] The flow of the specific processing in the application example 1 will be described with reference to FIG.
[0210] Step 1:
[0211] Users use cameras to capture the operation of exercise equipment and devices in the factory and upload the video data to a server.
[0212] Input: Video data containing user-recorded actions.
[0213] Output: Video data uploaded to the server.
[0214] Specific operation: The movement is filmed using a high-performance camera, and the user uploads the captured video data to a server via a dedicated application.
[0215] Step 2:
[0216] The server obtains the uploaded video data using a receiving means and temporarily stores it.
[0217] Input: User uploaded video data.
[0218] Output: Video data temporarily stored on the server.
[0219] Specific operation: The server saves the video data in a dedicated folder called "uploaded_video.mp4".
[0220] Step 3:
[0221] The server uses a pre-processing means to divide the received video data into frames.
[0222] Input: Video data stored on the server.
[0223] Output: Image data separated by frames.
[0224] Specific operation: Using OpenCV, video data is extracted frame by frame and divided at a rate of 30 frames per second.
[0225] Step 4:
[0226] The server uses pre-processing means to convert each frame into a format suitable for image analysis.
[0227] Input: Each split frame.
[0228] Output: Image data in a format suitable for analytical models.
[0229] What it does: Uses PIL (Pillow) to resize the image (to 224x224 pixels) and normalize it (convert pixel values to the range 0 to 1).
[0230] Step 5:
[0231] The server inputs the preprocessed frames into the AI model and analyzes the behavior.
[0232] Input: Preprocessed image data.
[0233] Output: Motion analysis result data.
[0234] Specific operation: Using PyTorch, preprocessed data is input to a deep learning model and analysis results are output.
[0235] Step 6:
[0236] The server generates feedback on the efficiency and accuracy of the operation based on the analysis results.
[0237] Input: Motion analysis result data.
[0238] Output: Feedback with suggestions for improvement and advice.
[0239] Specific actions: Evaluate the analysis results and use feedback generation methods to generate specific advice such as "The accuracy of the action is low. Adjustments are required."
[0240] Step 7:
[0241] The server provides the generated feedback to personnel involved in operating the factory equipment.
[0242] Input: Feedback data.
[0243] Output: Feedback provided to the agent.
[0244] What it does: Sends feedback to the device and presents it to the agent in a visual format via the application.
[0245] Step 8:
[0246] Based on the feedback, the staff will adjust the operation of the exercise equipment and devices and take further photographs to confirm the effectiveness.
[0247] Input: Adjustment information based on feedback.
[0248] Output: New video data recording the adjusted behavior.
[0249] Specific Action: Interpret the feedback provided and change the settings of factory equipment or devices, then photograph the new action and upload it back into the system.
[0250] Furthermore, an emotion engine that estimates the user's emotion may be combined. That is, the identification processing unit 290 may estimate the user's emotion using the emotion identification model 59 and perform identification processing using the user's emotion.
[0251] The present invention relates to a system that allows athletes to analyze their own form and receive feedback to improve their technique. The system includes a means for receiving and analyzing video data taken by the athlete, and a means for recognizing the athlete's emotions using an emotion engine. This allows for further optimization of the feedback provided to the athlete, supporting effective training.
[0252] Specific Embodiments of the System
[0253] 1. Upload your video
[0254] Users can record their own exercise form on their device, check the recorded video data, and send it to the server via a dedicated upload page.
[0255] 2. Receiving and storing video data
[0256] The server receives the video data sent by the user and temporarily stores it, which is then analyzed.
[0257] 3. Preparing the AI model
[0258] The server loads a pre-trained motion analysis model, which is used to individually analyze the athlete's form.
[0259] 4. Loading the video
[0260] The server reads the stored video file and splits it into frames, which typically consist of 30 frames per second.
[0261] 5. Frame Preprocessing
[0262] The server resizes and normalizes each frame into a format suitable for the analytical model, making the frame ready to be input into the AI model.
[0263] 6. Form Analysis
[0264] The server runs the pre-processed frames through a motion analysis model to evaluate the athlete's form, resulting in specific movement parameters such as knee angle and body posture.
[0265] 7. Emotion Recognition with Emotion Engine
[0266] The server extracts the facial expressions and voice data of the athlete from the frames and inputs them into the emotion engine, which then recognizes the athlete's emotional state from these data.
[0267] 8. Generate feedback
[0268] The server combines the results of the motion analysis with the output of the emotion engine to generate feedback for the exerciser. For example, if the exerciser's form is good, the server may comment, "Good form. Keep it up." If the exerciser's emotions are low, the server may add encouraging comments such as, "Great progress. Keep going with confidence."
[0269] 9. Providing Feedback
[0270] The server transmits the generated feedback to the user, and the terminal displays the feedback to the user in a visually easy-to-understand format.
[0271] 10. Review feedback and incorporate it into your training plan
[0272] Users can review the displayed feedback, understand areas for improvement in their form, and receive emotional advice, and use this information to restructure their training plan and make specific improvements.
[0273] 11. Re-recording and analyzing the video
[0274] Users can then re-record the video with improved form and re-upload it to the system. By repeating this process, users can expect to continually improve their form and increase the effectiveness of their training.
[0275] In this way, the system of the present invention can improve the quality of training by providing feedback that takes into account not only the athlete's form but also their emotional state. By receiving accurate and specific advice, athletes can achieve effective and sustained improvement in their skills.
[0276] The processing flow will be explained below.
[0277] Step 1:
[0278] User: Record a video of their exercise form using a smartphone or computer. Review the video to ensure the necessary parts are properly recorded.
[0279] Step 2:
[0280] User: Access the dedicated upload page, select the video file you have taken, and click the upload button to send the video data to the server.
[0281] Step 3:
[0282] Terminal: The selected video file is sent to the server as binary data. Once the sending is complete, a message indicating success is displayed to the user.
[0283] Step 4:
[0284] Server: Receives the binary data sent from the device. Temporarily saves the received data as "uploaded_video.mp4".
[0285] Step 5:
[0286] Server: Uses OpenCV library to read the saved video file. Opens the video file and splits it into frames.
[0287] Step 6:
[0288] Server: Loads a pre-trained motion analysis model (e.g., a deep learning model) that analyzes an athlete's form and provides an assessment of specific movements.
[0289] Step 7:
[0290] Server: Divide the video into frames. For example, if the video has 30 frames per second, extract 30 frames per second.
[0291] Step 8:
[0292] Server: Preprocess each frame by resizing it to 224x224 pixels and normalizing pixel values to the range 0 to 1.
[0293] Step 9:
[0294] Server: The preprocessed frames are input into the motion analysis model to obtain form evaluation results.
[0295] Step 10:
[0296] Server: Before generating feedback, emotion recognition is performed using the emotion engine. The facial expressions of the athlete are extracted and input into the emotion engine.
[0297] Step 11:
[0298] Server: The emotion engine recognizes the emotion of the exerciser and acquires the emotion information, which includes happiness, stress, concentration, etc.
[0299] Step 12:
[0300] Server: Generates feedback by combining motion analysis results and emotional information. For example, even if the form is good, if the emotion is unstable, an encouraging comment is added.
[0301] Step 13:
[0302] Server: Sends the generated feedback list to the user's terminal.
[0303] Step 14:
[0304] Terminal: Receives feedback sent from the server and displays it to the user. It presents comments in a format that is easy for the user to understand.
[0305] Step 15:
[0306] User: Check the feedback displayed on the device, understand the problems with their exercise form and emotional advice, and restructure their training plan based on this.
[0307] Step 16:
[0308] User: Record a video with improved form, upload it to the system in the same way, and request analysis. By repeating this process, you can continuously improve your form.
[0309] Example 2
[0310] Next, a description will be given of Example 2. In the following description, the data processing device 12 will be referred to as a "server" and the smart device 14 will be referred to as a "terminal."
[0311] Conventional exercise form analysis systems do not provide feedback that takes into account the athlete's emotional state, which can lead to a decline in the quality of training. A means is needed to improve mental motivation and the quality of feedback while allowing athletes to focus on improving their own technique.
[0312] The specific processing by the specific processing unit 290 of the data processing device 12 in the second embodiment is realized by the following means.
[0313] In this invention, the server includes a means for receiving video data captured by an exerciser, a preprocessing means for processing each frame included in the video data, and a means for analyzing the preprocessed frames using a motion analysis model as input, thereby making it possible to generate optimal feedback by combining the analysis results and emotion recognition results.
[0314] An "athlete" is someone who uses their body to exercise or train.
[0315] "Video data" refers to video files taken by the athlete.
[0316] "Preprocessing means" refers to the process performed to convert each frame contained in the video data into a format suitable for the AI model.
[0317] A "motion analysis model" refers to an algorithm or system that uses artificial intelligence to analyze an athlete's form.
[0318] "Analysis results" refers to the evaluations and parameters obtained about the athlete's form by the motion analysis model.
[0319] "Feedback" refers to advice or comments provided to athletes based on the analysis results.
[0320] "Providing means" refers to a method or system for transmitting the generated feedback to the exerciser.
[0321] The "face area" refers to the part of the video data in which the athlete's face is shown.
[0322] "Audio data" refers to the audio track included in the video data.
[0323] "Emotion recognition" refers to the process of identifying an athlete's emotional state from video and audio data.
[0324] "Emotion recognition result" refers to information about the emotional state of the exerciser obtained by the emotion recognition means.
[0325] "Optimal feedback" refers to the advice and comments that are deemed most beneficial to the athlete, generated by combining analysis results and emotion recognition results.
[0326] This invention is a system that allows athletes to improve their technique by analyzing their own form and receiving feedback that takes emotions into account. The system includes a means for receiving and analyzing video data taken by the athlete, and a means for recognizing the athlete's emotions using an emotion engine. This allows for further optimization of feedback to the athlete, supporting effective training.
[0327] Hardware and software used
[0328] The system's hardware uses high-performance servers, specifically those equipped with NVIDIA GPUs. Furthermore, the deep learning framework TensorFlow is used for motion analysis, and Microsoft Azure Cognitive Services is used for emotion recognition. The combination of these technologies enables highly accurate, real-time analysis and feedback.
[0329] Specific operation of the system
[0330] 1. Upload your video
[0331] Users record their own exercise form on their device and send the video data to the server via a dedicated upload page.
[0332] As a specific example, access the specified URL from the device's browser, click the file selection button to select the video data, and then click the upload button.
[0333] 2. Receiving and storing video data
[0334] The server receives the video data sent by the user as an HTTP request and temporarily stores it in storage. When the data is saved, a unique ID is assigned to it.
[0335] 3. Preparing the AI model
[0336] The server loads a pre-trained motion analysis model into memory using the TensorFlow library, which is done only once when the server starts.
[0337] 4. Loading the video and splitting it into frames
[0338] The server reads the saved video file using the OpenCV library and splits the video into frames at a rate of 30 frames per second, each of which is stored in a list.
[0339] 5. Frame Preprocessing
[0340] The server resizes the frames and normalizes pixel values to the range 0 to 1. This preprocessing prepares the images for accurate analysis by the deep learning model.
[0341] 6. Analysis of exercise form
[0342] The server inputs the preprocessed frames into a motion analysis model to extract data such as knee angles and body posture. The analysis results are stored in a list.
[0343] 7. Emotion recognition
[0344] The server extracts the athlete's facial area and voice data from the frames and sends them to Microsoft Azure Cognitive Services for emotion recognition, which identifies the athlete's emotional state.
[0345] 8. Generate feedback
[0346] The server combines the results of motion analysis and emotion recognition to generate optimal feedback for the exerciser. For example, if the exerciser's form is good, the server may comment, "Good form. Keep it up." If the exerciser's emotions are low, the server may add encouraging comments such as, "Great progress. Keep going with confidence."
[0347] 9. Providing Feedback
[0348] The server sends the generated feedback to the user, who then visually displays it on the device. The user can then review the feedback and work on restructuring their training plan or making specific improvements.
[0349] 10. Re-recording and analyzing the video
[0350] Users can then re-record the video with improved form and re-upload it into the system, repeating this process to continually improve their form and increase their training effectiveness.
[0351] Prompt Sentence Examples
[0352] markdown
[0353] Write Python code to split a video into frames at 30fps, resize and normalize each frame, and feed it into a deep learning model to analyze athletic form. Then, use Microsoft Azure Cognitive Services to recognize emotions and generate feedback that integrates the analysis results and emotional information.
[0354] In this way, the system of the present invention can improve the quality of training by providing feedback that takes into account not only the athlete's form but also their emotional state. By receiving precise and specific advice, athletes can achieve effective and sustained improvement in their technique.
[0355] The flow of the identification process in the second embodiment will be described with reference to FIG.
[0356] Step 1:
[0357] Recording and uploading videos
[0358] Users use their device to record their own exercise form and send the video data to the server from a dedicated upload page. They access the specified URL from their device's browser, click the file selection button to select the video data, and then click the upload button to send it.
[0359] Input: Video file of the athlete's exercise form
[0360] Output: Video data sent to the server
[0361] Step 2:
[0362] Receiving and storing video data
[0363] The server receives the video data sent by the user as an HTTP request and temporarily stores it in storage. The server assigns a unique ID to the received file and stores it in a folder structure. This saving operation is performed asynchronously, returning a prompt response to the user.
[0364] Input: Video data sent by the user
[0365] Output: Video file saved in storage
[0366] Step 3:
[0367] Preparing the AI model
[0368] The server loads a pre-trained motion analysis model into memory using the TensorFlow library, a process that is performed only once when the server starts.
[0369] Input: The trained motion analysis model file
[0370] Output: Motion analysis model loaded into memory
[0371] Step 4:
[0372] Video loading and frame splitting
[0373] The server finds the stored video file and splits the video into frames using the OpenCV library by opening the video file, capturing frames at a rate of 30 frames per second, and storing them in a list.
[0374] Input: Video file saved in storage
[0375] Output: Each frame stored in a list
[0376] Step 5:
[0377] Frame Preprocessing
[0378] The server resizes each frame and normalizes pixel values to a range between 0 and 1. This preprocessing converts the frames into a format suitable for AI models. Specifically, it resizes each frame to a uniform size and normalizes the values per pixel.
[0379] Input: Each frame stored in a list
[0380] Output: Preprocessed frames
[0381] Step 6:
[0382] Analysis of athletic form
[0383] The server inputs the preprocessed frames into the motion analysis model to extract data such as knee angle and body posture, performs the analysis using the model's predict method, and saves the results in a list.
[0384] Input: Preprocessed frame
[0385] Output: Exercise form data as analysis results (e.g. knee angle, body posture, etc.)
[0386] Step 7:
[0387] emotion recognition
[0388] The server extracts the athlete's facial area and voice data from the frames and inputs them into the emotion engine. Specifically, it sends an HTTP request to the API endpoint of Microsoft Azure Cognitive Services to recognize emotions from facial expressions and voice.
[0389] Input: Face area and audio data in the frame
[0390] Output: Emotion recognition results (e.g., joy, sadness, anger, etc.)
[0391] Step 8:
[0392] Generate feedback
[0393] The server combines the results of motion analysis and emotion recognition to generate feedback for the exerciser. Specifically, it embeds comments based on the analysis results and encouraging messages based on the exerciser's emotional state into templates.
[0394] Input: Motion analysis results, emotion recognition results
[0395] Output: The generated feedback message
[0396] Step 9:
[0397] Providing Feedback
[0398] The server sends the generated feedback to the user, and the device displays it visually, either by returning a feedback message as an HTTP response or by delivering it to the user via email or push notification.
[0399] Input: The generated feedback message
[0400] Output: Feedback displayed on the terminal
[0401] Step 10:
[0402] Review feedback and incorporate it into your training plan
[0403] Users can view feedback on their device, understand areas for improvement in their form, and receive emotional advice, which they can use to restructure their training plan and make specific improvements.
[0404] Input: Feedback displayed on the device
[0405] Output: Improved training plan
[0406] Step 11:
[0407] Re-shooting and analyzing the video
[0408] Users can then re-record the video with improved form and re-upload it to the system, which will enable continuous improvement of form and increased training effectiveness.
[0409] Input: New video file taken with improved form
[0410] Output: Newly uploaded video data to the server
[0411] (Application example 2)
[0412] Next, a description will be given of Application Example 2. In the following description, the data processing device 12 will be referred to as a "server" and the smart device 14 will be referred to as a "terminal."
[0413] Conventional motion analysis systems only provide feedback on the athlete's form, but are unable to provide appropriate feedback that takes into account the athlete's emotional state. This can result in training motivation and effectiveness not being maximized. Furthermore, in operational environments such as factory robots, not only is improving work accuracy and efficiency important, but also managing the operator's stress is also an important issue.
[0414] The specific processing by the specific processing unit 290 of the data processing device 12 in Application Example 2 is realized by the following means. In this invention, the server includes means for receiving video data taken by an exerciser, preprocessing means for processing each frame included in the video data, means for analyzing the preprocessed frames using a motion analysis model as input, means for generating feedback on the exerciser's form based on the analysis results, means for recognizing emotions based on the exerciser's facial expressions and voice data, means for optimizing and generating feedback taking into account the emotion recognition results, and means for providing the generated feedback to the exerciser. This makes it possible to provide feedback that takes into account not only the form of the exerciser and operator but also their overall state, including emotions.
[0415] "Video data" refers to video footage of the movements of an athlete or robot.
[0416] "Preprocessing" refers to processes such as resizing and normalization that are performed to convert frames of video data into a format suitable for the analysis model.
[0417] "Movement analysis model" refers to an artificial intelligence model used to analyze the movements of an athlete or robot and evaluate their form and accuracy.
[0418] "Feedback" refers to advice or comments provided to an exerciser or operator based on the analysis results or emotion recognition results.
[0419] "Emotion recognition" refers to analyzing the facial expressions and voices of athletes and operators contained in video data to identify their emotional state.
[0420] "Optimization" refers to adjusting feedback based on acquired data to provide more effective advice.
[0421] A "frame" refers to an individual still image that makes up video data.
[0422] The present invention relates to a system that analyzes the operation of a factory robot and supports its efficient and highly accurate movement.
[0423] The system uses video data captured by the user to analyze the accuracy and efficiency of factory robot operation and provides feedback. It also recognizes the emotional state of the user and workers and optimizes feedback as needed.
[0424] Hardware:
[0425] Camera: Used to capture video data.
[0426] GPU: Used to run AI models at high speed.
[0427] software:
[0428] OpenCV: Used to load videos and split them into frames.
[0429] PyTorch: Used to run AI models for motion analysis.
[0430] EmotionEngine: A custom library for emotion recognition.
[0431] MovementAnalyzer: A custom library for performing movement analysis.
[0432] The server performs the following process:
[0433] First, the user uses a camera to record the movements of an athlete or a factory robot. The captured video data is then sent to the server, which receives it and temporarily stores it. The stored video data is then divided into frames, resized and normalized as preprocessing, and converted into a format suitable for the AI model.
[0434] The preprocessed frames are input into a motion analysis model to evaluate the form and movement accuracy of the athlete or robot. The evaluation results are obtained as specific movement parameters. In addition, the server extracts facial expressions and voice data of the user or worker from the video data and inputs them into the Emotion Engine. The Emotion Engine analyzes this data and identifies their emotional state.
[0435] The server combines the results of motion analysis and emotion recognition to generate feedback. For example, if the motion accuracy is high, it will comment, "Your operation is accurate. Please keep going," and if the emotion is stressed, it will generate advice such as, "Your operation is good, but you seem to be stressed. Please take a short break."
[0436] The server sends the generated feedback to the user's device, which then visually presents it to the user. The user can then review the feedback, understand the areas for improvement in their own operations and movements, and receive emotional advice, which they can then incorporate into their training plan.
[0437] As a concrete example, consider a robot operator at a factory who is currently learning a new operating procedure. The operator films his or her operation with a camera and uploads it to the system. The system analyzes the video data and provides specific feedback such as, "Your vehicle handling is good, but some improvement is needed in conveyor line operation." It also recognizes fatigue from the operator's facial expression and adds advice such as, "Take a short break and refresh yourself."
[0438] Example prompt sentence:
[0439] Upload a video of a factory robot being operated. Analyze the robot's movements and the operator's facial expressions to generate specific feedback to improve accuracy. Include comments based on the operator's level of fatigue and stress.
[0440] The flow of the specific processing in the application example 2 will be described with reference to FIG.
[0441] Program processing flow of the application example system
[0442] Step 1:
[0443] The server receives video data captured by the user of an athlete or a factory robot. Specifically, the user sends the video captured with a camera to the server from the upload page. The input is the captured video data, and the output is a video file stored on the server.
[0444] Step 2:
[0445] The server divides the received video data into frames. Specifically, it uses OpenCV to read the video data and extract each frame. The input is a saved video file, and the output is multiple frame images.
[0446] Step 3:
[0447] The server preprocesses each frame. Specifically, it resizes and normalizes the frame image to a format suitable for the AI model. The input is the extracted frame image, and the output is the preprocessed frame image.
[0448] Step 4:
[0449] The server inputs the preprocessed frames into a motion analysis model. Specifically, the preprocessed frame images are input into a motion analysis model running on PyTorch to obtain motion parameters. The input is the preprocessed frame images, and the output is the motion parameters.
[0450] Step 5:
[0451] The server extracts the user's facial expression and voice data from the frames. Specifically, it uses OpenCV and voice analysis libraries to extract the necessary information from the frame images and voice data. The input is the preprocessed frame images and voice data, and the output is the user's facial expression and voice data.
[0452] Step 6:
[0453] The server inputs the extracted facial and voice data into an emotion recognition engine. Specifically, it uses the EmotionEngine to identify the emotional state. The input is the user's facial and voice data, and the output is the emotion recognition result.
[0454] Step 7:
[0455] The server generates feedback based on the results of motion analysis and emotion recognition. Specifically, it combines the analysis data to generate appropriate comments and advice. The input is the motion parameters of the motion analysis and the emotion recognition results, and the output is the generated feedback.
[0456] Step 8:
[0457] The server transmits the generated feedback to the user's terminal. Specifically, the server converts the generated feedback into a format for visual display and transmits it to the user's terminal. The input is the generated feedback, and the output is the feedback displayed on the user's terminal.
[0458] Step 9:
[0459] The user checks the feedback displayed on the device and understands the areas for improvement in their own operations and behaviors, as well as emotional advice. Specifically, they restructure their training plan based on the feedback and work on specific improvements. The input is the feedback displayed on the device, and the output is the improved operations and behaviors.
[0460] summary
[0461] This step enables the system to analyze the movements of athletes and factory robots and provide appropriate feedback that takes emotions into account.
[0462] The specific processing unit 290 transmits the result of the specific processing to the smart device 14. In the smart device 14, the control unit 46A causes the output device 40 to output the result of the specific processing. The microphone 38B acquires audio indicating a user input regarding the result of the specific processing. The control unit 46A transmits audio data indicating the user input acquired by the microphone 38B to the data processing device 12. In the data processing device 12, the specific processing unit 290 acquires the audio data.
[0463] The data generation model 58 is a so-called generative AI (Artificial Intelligence). An example of the data generation model 58 is ChatGPT (Internet Search<URL: https: / / openai.com / blog / chatgpt> ), Gemini (Internet search <url: https: gemini.google.com ?hl="ja">) and other generation AIs. The data generation model 58 is obtained by performing deep learning on a neural network. A prompt including an instruction is input to the data generation model 58, and inference data such as voice data indicating voice, text data indicating text, and image data indicating an image is also input. The data generation model 58 performs inference on the input inference data in accordance with the instruction indicated by the prompt, and outputs the inference result in a data format such as voice data and text data. Here, inference refers to, for example, analysis, classification, prediction, and / or summarization.
[0464] In the above embodiment, an example in which the specific process is performed by the data processing device 12 has been given, but the technology of the present disclosure is not limited to this, and the specific process may be performed by the smart device 14.
[0465] [Second embodiment]
[0466] FIG. 3 shows an example of the configuration of a data processing system 210 according to the second embodiment.
[0467] 3, the data processing system 210 includes the data processing device 12 and smart glasses 214. An example of the data processing device 12 is a server.
[0468] The data processing device 12 includes a computer 22, a database 24, and a communication I / F 26. The computer 22 is an example of a "computer" according to the technology of the present disclosure. The computer 22 includes a processor 28, a RAM 30, and a storage 32. The processor 28, the RAM 30, and the storage 32 are connected to a bus 34. The database 24 and the communication I / F 26 are also connected to the bus 34. The communication I / F 26 is connected to a network 54. Examples of the network 54 include a WAN (Wide Area Network) and / or a LAN (Local Area Network).
[0469] The smart glasses 214 include a computer 36, a microphone 238, a speaker 240, a camera 42, and a communication I / F 44. The computer 36 includes a processor 46, a RAM 48, and a storage 50. The processor 46, the RAM 48, and the storage 50 are connected to a bus 52. The microphone 238, the speaker 240, and the camera 42 are also connected to the bus 52.
[0470] The microphone 238 receives instructions and the like from the user 20 by receiving voice uttered by the user 20. The microphone 238 captures the voice uttered by the user 20, converts the captured voice into audio data, and outputs it to the processor 46. The speaker 240 outputs audio in accordance with instructions from the processor 46.
[0471] Camera 42 is a small digital camera equipped with an optical system including a lens, aperture, and shutter, and an imaging element such as a CMOS (Complementary Metal-Oxide-Semiconductor) image sensor or a CCD (Charge Coupled Device) image sensor, and captures images of the surroundings of user 20 (for example, an imaging range defined by an angle of view equivalent to the field of vision of a typical healthy person).
[0472] The communication I / F 44 is connected to a network 54. The communication I / Fs 44 and 26 are responsible for the exchange of various information between the processor 46 and the processor 28 via the network 54. The exchange of various information between the processor 46 and the processor 28 using the communication I / Fs 44 and 26 is carried out in a secure state.
[0473] Fig. 4 shows an example of the main functions of the data processing device 12 and the smart glasses 214. As shown in Fig. 4, in the data processing device 12, a specific process is performed by the processor 28. A specific process program 56 is stored in the storage 32.
[0474] The specific processing program 56 is an example of a "program" according to the technology of the present disclosure. The processor 28 reads the specific processing program 56 from the storage 32 and executes the read specific processing program 56 on the RAM 30. The specific processing is realized by the processor 28 operating as a specific processing unit 290 in accordance with the specific processing program 56 executed on the RAM 30.
[0475] The storage 32 stores a data generation model 58 and an emotion identification model 59. The data generation model 58 and the emotion identification model 59 are used by the identification processing unit 290.
[0476] In the smart glasses 214, the reception output process is performed by the processor 46. A reception output program 60 is stored in the storage 50. The processor 46 reads the reception output program 60 from the storage 50 and executes the read reception output program 60 on the RAM 48. The reception output process is realized by the processor 46 operating as the control unit 46A in accordance with the reception output program 60 executed on the RAM 48.
[0477] Next, a description will be given of the identification process performed by the identification processing unit 290 of the data processing device 12. In the following description, the data processing device 12 will be referred to as the "server" and the smart glasses 214 will be referred to as the "terminal."
[0478] This invention relates to a system that allows athletes to analyze their own form and receive specific advice for improving their technique. This system allows athletes to take video data using devices such as smartphones or computers, upload it to a server, and then analyzes the video data using an AI model to provide detailed feedback.
[0479] Specific Embodiments of the System
[0480] 1. Upload your video
[0481] Users upload videos of their exercise form from their devices to the server, where the uploaded videos are temporarily stored.
[0482] 2. Receiving and storing video data
[0483] The server receives the video data sent by the user and saves it under a file name such as "uploaded_video.mp4." This saved video data will then be analyzed.
[0484] 3. Preparing and Loading the AI Model
[0485] The server loads a pre-trained motion analysis model (e.g., a deep learning model) that analyzes an athlete's form and provides an assessment of specific movements.
[0486] 4. Loading the video and splitting it into frames
[0487] The server extracts the stored video file frame by frame. Since video is usually 30 frames per second, it is processed as 30 images per second.
[0488] 5. Frame Preprocessing
[0489] The server performs preprocessing such as frame resizing and normalization to convert each frame into a format suitable for the analytical model. Specifically, the frame is resized to 224x224 pixels and the pixel values are normalized to the range 0 to 1.
[0490] 6. Form Analysis
[0491] The server then feeds the preprocessed frames into a motion analysis model to evaluate form, measuring specific movement parameters such as knee angle and body tilt, and returns the results.
[0492] 7. Generate feedback
[0493] The server generates feedback for each frame based on the analysis results, and provides the user with specific advice such as "Good form. Keep going like this" or "Review the way you bend your knees."
[0494] 8. Providing Feedback
[0495] The server compiles the generated feedback and sends it back to the user's device, which then displays it to the user in a visually understandable format.
[0496] 9. Review feedback and incorporate it into your training plan
[0497] Users can view the feedback displayed on their device to understand where their form is lacking, and use that feedback to restructure their training plan and address specific areas for improvement.
[0498] 10. Re-recording and analyzing the video
[0499] Users can re-record a video with a new form and re-upload it into the system, repeating this process to continually improve their form.
[0500] As described above, the present invention allows athletes to improve the quality of their training by analyzing their own form in detail and receiving specific, individually customized feedback. This system is expected to make sports skill improvement even more efficient.
[0501] The processing flow will be explained below.
[0502] Step 1:
[0503] User: Record a video of their exercise form using a smartphone or computer. Review the video to ensure the necessary parts are properly recorded.
[0504] Step 2:
[0505] User: Access the dedicated upload page from their device, select the video file they shot on that page, and click the upload button.
[0506] Step 3:
[0507] Terminal: The selected video file is sent to the server as binary data. Once the sending is complete, a message indicating success is displayed to the user.
[0508] Step 4:
[0509] Server: Receives the binary data sent from the device. Temporarily saves the received data as "uploaded_video.mp4".
[0510] Step 5:
[0511] Server: Uses the OpenCV library to read the saved video file. Opens the video file and prepares to extract data frame by frame.
[0512] Step 6:
[0513] Server: Loads a pre-trained motion analysis model (e.g., "pose_analysis_model.h5") that analyzes an athlete's form and evaluates specific movements.
[0514] Step 7:
[0515] Server: Divide the video into frames. For example, if the video is shot at 30 frames per second, extract 30 frames per second.
[0516] Step 8:
[0517] Server: Preprocess each frame by resizing it to 224x224 pixels and normalizing pixel values to the range 0 to 1.
[0518] Step 9:
[0519] Server: Inputs the preprocessed frames into the motion analysis model, which outputs the motion estimation results for each frame.
[0520] Step 10:
[0521] Server: Generates feedback based on the analysis results. For example, if the model determines that the knee angle is inappropriate, it generates specific advice such as "Reconsider how you bend your knees."
[0522] Step 11:
[0523] Server: Compiles the generated feedback into a list and prepares it for sending back to the user.
[0524] Step 12:
[0525] Server: Sends the feedback list to the user's terminal.
[0526] Step 13:
[0527] Terminal: Receives feedback sent from the server and displays it to the user. It presents comments in a format that is easy for the user to understand.
[0528] Step 14:
[0529] User: Check the feedback displayed on the device to understand problems and areas for improvement in their exercise form, and restructure their training plan based on this.
[0530] Step 15:
[0531] User: Record a video with improved form, upload it to the system in the same way, and request analysis. By repeating this process, you can continuously improve your form.
[0532] Example 1
[0533] Next, a description will be given of Example 1. In the following description, the data processing device 12 will be referred to as a "server" and the smart glasses 214 will be referred to as a "terminal."
[0534] Conventional exercise form analysis systems have the drawback of requiring users to film their own form, send the video data to a dedicated device, and obtain analysis results, all of which require a complex and time-consuming process. Furthermore, the accuracy of the analysis results is low, making it difficult for users to identify specific areas for improvement. Furthermore, many systems do not provide instant feedback, making it difficult to incorporate the results into effective training plans.
[0535] The specific processing by the specific processing unit 290 of the data processing device 12 in the first embodiment is realized by the following means.
[0536] In this invention, the server includes means for receiving video data taken by an exerciser, means for saving the video data, means for dividing the video data into frames, means for preprocessing each of the frames, means for analyzing the preprocessed frames using a generative AI model, means for generating feedback on the exerciser's form based on the analysis results, means for providing the generated feedback to the exerciser, and means for displaying the provided feedback on the exerciser's device. This provides high-precision feedback in real time, allowing the user to quickly identify areas for improvement in their form and implement an effective training plan.
[0537] An "athlete" refers to an individual who exercises, photographs their own exercise form, and uses the data for analysis.
[0538] "Video data" refers to video files taken by the athlete, and contains information to be analyzed.
[0539] "Server" refers to a computer system that receives, stores, processes, and provides analysis results of video data.
[0540] "Each frame" refers to the individual images that make up the video data, and the information per frame is used for analysis.
[0541] "Preprocessing" refers to data manipulation such as resizing and normalization to convert each frame into a format suitable for the analytical model.
[0542] A "generative AI model" is a pre-trained AI (artificial intelligence) model used to analyze and evaluate exercise form.
[0543] "Analysis results" refers to the evaluation data of exercise form obtained by the generative AI model, and includes specific movement parameters (such as knee angle and body tilt).
[0544] "Feedback" refers to advice and evaluations generated based on the analysis results, and is information that athletes can use to improve their form.
[0545] "Terminal" refers to the device used by the exerciser, such as a smartphone or computer, that displays the feedback.
[0546] This invention relates to a system that allows athletes to analyze their own form and receive specific advice for improving their technique. The system allows athletes to take video data using devices such as smartphones or computers, upload it to a server, and then analyzes the video data using a generative AI model to provide detailed feedback.
[0547] Hardware and software used
[0548] Hardware: smartphones, computers, servers
[0549] Software: Video recording app, video uploading tool, motion analysis AI model (deep learning model)
[0550] Data processing and calculation methods
[0551] Users can upload videos of their exercise form to a server using a smartphone or computer, for example, via a dedicated application or browser.
[0552] The server receives the video data sent by the user and temporarily stores it in a specific directory (e.g., "uploaded_videos / ") on the server with the file name "uploaded_video.mp4". This data is used in the subsequent analysis steps.
[0553] The server loads pre-trained AI models for motion analysis, which use deep learning techniques to analyze athletic form, including models trained using TensorFlow and PyTorch.
[0554] The server loads the saved video file and uses tools such as ffmpeg to extract the video frame by frame. Since video is typically 30 frames per second, this translates to 30 images per second.
[0555] Each frame is converted into a format suitable for the analytical model. In this step, the frame is resized to 224x224 pixels and the pixel values are normalized to the range 0 to 1. This process is often performed using OpenCV or PIL (Python Imaging Library).
[0556] The preprocessed frames are fed into a generative AI model to evaluate the form. The model measures movement parameters such as knee angle and body tilt and returns the analysis results. For example, the model outputs data such as "knee angle is 45 degrees" for each frame.
[0557] The server generates feedback based on the analysis results, including specific advice such as "Try to widen your knee angle by about 5 degrees." This feedback is created for each frame and is customized for each user.
[0558] The generated feedback is sent back from the server to the user's device, which then displays it to the user in a visually understandable format, for example, highlighting key points in addition to displaying them as text messages within the application.
[0559] Users can view the feedback displayed on their device, understand the issues with their form, and use that feedback to restructure their training plan and address specific areas for improvement, such as paying special attention to knee angle during their next workout.
[0560] Examples of specific examples and prompts
[0561] For example, suppose a user wants to improve their running form. The user films their running form with their smartphone and uploads the video to the system. The server analyzes the video and generates feedback to the user, such as, "The angle of your right knee is not appropriate. Try bending your right knee a bit more while running." The user can then revise their training plan based on this feedback and review their form again.
[0562] Prompt Sentence Examples
[0563] "I uploaded a video of my running form. Please help me analyze where my form needs improvement."
[0564] "Please analyze the video of my form and tell me specific areas for improvement."
[0565] The system allows exercisers to receive detailed and specific feedback to continually improve their form.
[0566] The flow of the identification process in the first embodiment will be described with reference to FIG.
[0567] Step 1:
[0568] The user takes a video of their exercise form on the device.
[0569] Input: A video of your exercise form (e.g. "run.mp4")
[0570] Output: Recorded video file
[0571] Specific movements: Using a smartphone or computer, participants will be videotaped while performing various exercises, such as running, jumping, and squatting.
[0572] Step 2:
[0573] The user uploads the video they have taken from their terminal to the server.
[0574] Input: Recorded video file
[0575] Output: Video data uploaded to the server
[0576] Specific operation: Use a dedicated application or web browser to send the video file to the server. For example, click the button, select "run.mp4", and start uploading.
[0577] Step 3:
[0578] The server receives and stores the video data sent by the user.
[0579] Input: Video data received from the user
[0580] Output: Video file saved on the server (e.g. "uploaded_video.mp4")
[0581] What happens: The server stores the uploaded data in a specific directory (e.g. "uploaded_videos / ").
[0582] Step 4:
[0583] The server loads a pre-trained generative AI model.
[0584] Input: A trained generative AI model
[0585] Output: Generative AI model loaded into memory
[0586] What it does: Loads a motion analysis model trained using TensorFlow or PyTorch, and the system is ready for analysis.
[0587] Step 5:
[0588] The server reads the saved video file and splits it into frames.
[0589] Input: Video file stored on the server ("uploaded_video.mp4")
[0590] Output: Images separated by frames (e.g. frame 1, frame 2, ...)
[0591] Specific operation: Using a video processing tool such as ffmpeg, the video is divided into 30 frames per second and each frame is obtained as an image file.
[0592] Step 6:
[0593] The server performs pre-processing on each frame.
[0594] Input: Images split into frames
[0595] Output: Preprocessed frame (resized and normalized)
[0596] Specific operation: Using OpenCV and PIL (Python Imaging Library), each frame is resized to 224x224 pixels and the value of each pixel is normalized to the range 0 to 1.
[0597] Step 7:
[0598] The server feeds the preprocessed frames into a generative AI model to analyze the form.
[0599] Input: Preprocessed frame
[0600] Output: Analysis results (motion parameters, e.g. knee angle, body tilt)
[0601] Specific movement: The preprocessed frame is input into the generative AI model, movement parameters are measured, and an evaluation result is output. For example, the result may be "Frame 10: Knee angle 45 degrees."
[0602] Step 8:
[0603] The server generates feedback based on the analysis results.
[0604] Input: Analysis results
[0605] Output: Generated feedback (e.g. "Try to open your knees by about 5 degrees.")
[0606] Specific actions: Based on the analysis results for each frame, specific advice on how to improve your athletic form is generated. For example, "Your right knee angle is not appropriate, so try bending it a little more while running."
[0607] Step 9:
[0608] The server transmits the generated feedback to the user's terminal.
[0609] Input: Generated feedback
[0610] Output: Feedback sent to the user's device
[0611] Specific operation: The server compiles the feedback data and sends it to the user's device, for example, by sending the feedback data back using an HTTP request.
[0612] Step 10:
[0613] The terminal displays the received feedback to the user.
[0614] Input: Feedback sent by the server
[0615] Output: The feedback screen presented to the user.
[0616] What it does: The device visualizes the feedback it receives and presents it to the user in a user-friendly format, for example by using text messages or highlighting to emphasize key points.
[0617] Step 11:
[0618] The user checks the feedback displayed on the device and reflects it in their training plan.
[0619] Input: Feedback displayed on the device
[0620] Output: Modified training plan
[0621] What it does: Users can use the feedback to understand where their form needs improvement and focus on those areas during their next workout.
[0622] Step 12:
[0623] The user then shoots the video again with the improved form and uploads it to the system again.
[0624] Input: Improved video
[0625] Output: New video data uploaded to the system
[0626] What it does: The user takes a photo of a new form and repeats the steps above to provide data to the system, which then continually improves the form.
[0627] (Application example 1)
[0628] Next, a description will be given of Application Example 1. In the following description, the data processing device 12 will be referred to as a "server" and the smart glasses 214 will be referred to as a "terminal."
[0629] Conventional motion analysis systems have been used only to provide feedback to athletes and general exercisers to improve their exercise form, but the concept has not been applied to the motion analysis of exercise machines and equipment in factories. In factories, accurate analysis of motion and appropriate feedback are essential to maintain and improve the efficiency and accuracy of highly automated equipment and robots. However, current systems have difficulty meeting these needs, and have the problem of being unable to respond immediately when a decline in motion efficiency or accuracy occurs.
[0630] The specific processing by the specific processing unit 290 of the data processing device 12 in the application example 1 is realized by the following means.
[0631] In this invention, the server includes: means for receiving video data captured by an exerciser; preprocessing means for processing each frame included in the video data; means for analyzing the preprocessed frames using a motion analysis model; means for generating feedback on the exerciser's form based on the analysis results; means for providing the generated feedback to the exerciser; a device for capturing the movements of exercise equipment or devices in a factory; means for receiving and preprocessing the captured movement data; means for evaluating movements using an AI model for analyzing the preprocessed movement data; means for generating feedback on improving the efficiency and accuracy of movements based on the evaluation results; and means for providing the generated feedback to personnel involved in operating the factory equipment. This enables the movement analysis of exercise equipment or devices in a factory, and enables quick and appropriate feedback to be provided to improve efficiency and accuracy.
[0632] "Exercise participant" refers to the person or equipment whose exercise form is being analyzed and feedback is being provided.
[0633] "Video data" refers to data containing visual information that records the movements of an exerciser or the movements of exercise equipment.
[0634] "Means for receiving" refers to the technical methods and equipment for accepting video data from an external device.
[0635] "Preprocessing means" refers to a method or device that performs processing to convert video data into a format suitable for the analysis model.
[0636] "Movement analysis model" refers to an artificial intelligence model used to analyze an athlete's form and movement data.
[0637] "Means for analysis" refers to the techniques and devices used to input preprocessed data into a motion analysis model and obtain results.
[0638] "Means for generating feedback" refers to techniques and devices for generating information that indicates specific areas for improvement and direction based on the analysis results.
[0639] "Means of providing" refers to the techniques and devices used to communicate the generated feedback to the exerciser or person in charge.
[0640] "Motion equipment or devices" refers to machines, robots, and other equipment that perform specific movements in a factory.
[0641] "Motion Data" refers to data, including visual information, that records the movements of exercise equipment or devices.
[0642] "Means for evaluating" refers to techniques or devices for evaluating the efficiency and accuracy of a movement using a motion analysis model based on preprocessed movement data.
[0643] This invention is a system based on the "athlete's form analysis system" that is applied to the analysis of the movements of exercise machines or equipment in factories and the provision of feedback. This system can monitor the efficiency and accuracy of the movements of robots and machines in factories and provide necessary advice for improvement.
[0644] First, cameras are placed to capture the operation of the motion machines and equipment used in the factory. This video data is sent to a server. The server then acquires this video data using a receiving means. The hardware used in this case can be a high-performance camera.
[0645] The server then processes the received video data using preprocessing tools. Specifically, it divides the video data into frames and converts them into a format suitable for the AI model. This preprocessing uses OpenCV and PIL (Pillow) to resize and normalize the images, allowing for detailed analysis of the motion.
[0646] The preprocessed data is then input into an AI model, the motion analysis model. This model uses deep learning to analyze the motion frame by frame and measure specific motion parameters. PyTorch is used for the analysis, and the model's motion is evaluated.
[0647] The server generates feedback regarding improvements to the efficiency and accuracy of the operation based on the analysis results. The feedback generation means creates information indicating specific advice and areas for improvement based on the analysis results. For example, feedback such as "The accuracy of the operation is low. Adjustments are required" may be generated.
[0648] Finally, the server provides the generated feedback to personnel involved in the operation of the factory equipment using a means for providing the feedback, which can then be used by the personnel to adjust the robots and equipment to improve the efficiency and accuracy of their operations.
[0649] As a specific example, a robot arm in a factory is filmed as it assembles parts, and the video data is uploaded to the system. The system analyzes the video data and generates specific feedback such as, "The robot arm's movement is delayed. Please adjust the angle at which the part is attached." Based on this, the worker can adjust the robot arm's movement, improving the efficiency of the assembly work.
[0650] An example of a prompt sentence to be input to the generative AI model is as follows:
[0651] Consider a scenario where a system analyzes the behavior of a factory robot and provides feedback on its efficiency and accuracy. Specifically, imagine a robotic arm takes a video of itself assembling a part, and an AI model analyzes its behavior and provides advice on how to improve it.
[0652] The flow of the specific processing in the application example 1 will be described with reference to FIG.
[0653] Step 1:
[0654] Users use cameras to capture the operation of exercise equipment and devices in the factory and upload the video data to a server.
[0655] Input: Video data containing user-recorded actions.
[0656] Output: Video data uploaded to the server.
[0657] Specific operation: The movement is filmed using a high-performance camera, and the user uploads the captured video data to a server via a dedicated application.
[0658] Step 2:
[0659] The server obtains the uploaded video data using a receiving means and temporarily stores it.
[0660] Input: User uploaded video data.
[0661] Output: Video data temporarily stored on the server.
[0662] Specific operation: The server saves the video data in a dedicated folder called "uploaded_video.mp4".
[0663] Step 3:
[0664] The server uses a pre-processing means to divide the received video data into frames.
[0665] Input: Video data stored on the server.
[0666] Output: Image data separated by frames.
[0667] Specific operation: Using OpenCV, video data is extracted frame by frame and divided at a rate of 30 frames per second.
[0668] Step 4:
[0669] The server uses pre-processing means to convert each frame into a format suitable for image analysis.
[0670] Input: Each split frame.
[0671] Output: Image data in a format suitable for analytical models.
[0672] What it does: Uses PIL (Pillow) to resize the image (to 224x224 pixels) and normalize it (convert pixel values to the range 0 to 1).
[0673] Step 5:
[0674] The server inputs the preprocessed frames into the AI model and analyzes the behavior.
[0675] Input: Preprocessed image data.
[0676] Output: Motion analysis result data.
[0677] Specific operation: Using PyTorch, preprocessed data is input to a deep learning model and analysis results are output.
[0678] Step 6:
[0679] The server generates feedback on the efficiency and accuracy of the operation based on the analysis results.
[0680] Input: Motion analysis result data.
[0681] Output: Feedback with suggestions for improvement and advice.
[0682] Specific actions: Evaluate the analysis results and use feedback generation methods to generate specific advice such as "The accuracy of the action is low. Adjustments are required."
[0683] Step 7:
[0684] The server provides the generated feedback to personnel involved in operating the factory equipment.
[0685] Input: Feedback data.
[0686] Output: Feedback provided to the agent.
[0687] What it does: Sends feedback to the device and presents it to the agent in a visual format via the application.
[0688] Step 8:
[0689] Based on the feedback, the staff will adjust the operation of the exercise equipment and devices and take further photographs to confirm the effectiveness.
[0690] Input: Adjustment information based on feedback.
[0691] Output: New video data recording the adjusted behavior.
[0692] Specific Action: Interpret the feedback provided and change the settings of factory equipment or devices, then photograph the new action and upload it back into the system.
[0693] Furthermore, an emotion engine that estimates the user's emotion may be further combined. That is, the identification processing unit 290 may estimate the user's emotion using the emotion identification model 59, and perform identification processing using the user's emotion.
[0694] The present invention relates to a system that allows athletes to analyze their own form and receive feedback to improve their technique. The system includes a means for receiving and analyzing video data taken by the athlete, and a means for recognizing the athlete's emotions using an emotion engine. This allows for further optimization of the feedback provided to the athlete, supporting effective training.
[0695] Specific Embodiments of the System
[0696] 1. Upload your video
[0697] Users can record their own exercise form on their device, check the recorded video data, and send it to the server via a dedicated upload page.
[0698] 2. Receiving and storing video data
[0699] The server receives the video data sent by the user and temporarily stores it, which is then analyzed.
[0700] 3. Preparing the AI model
[0701] The server loads a pre-trained motion analysis model, which is used to individually analyze the athlete's form.
[0702] 4. Loading the video
[0703] The server reads the stored video file and splits it into frames, which typically consist of 30 frames per second.
[0704] 5. Frame Preprocessing
[0705] The server resizes and normalizes each frame into a format suitable for the analytical model, making the frame ready to be input into the AI model.
[0706] 6. Form Analysis
[0707] The server runs the pre-processed frames through a motion analysis model to evaluate the athlete's form, resulting in specific movement parameters such as knee angle and body posture.
[0708] 7. Emotion Recognition with Emotion Engine
[0709] The server extracts the facial expressions and voice data of the athlete from the frames and inputs them into the emotion engine, which then recognizes the athlete's emotional state from these data.
[0710] 8. Generate feedback
[0711] The server combines the results of the motion analysis with the output of the emotion engine to generate feedback for the exerciser. For example, if the exerciser's form is good, the server may comment, "Good form. Keep it up." If the exerciser's emotions are low, the server may add encouraging comments such as, "Great progress. Keep going with confidence."
[0712] 9. Providing Feedback
[0713] The server transmits the generated feedback to the user, and the terminal displays the feedback to the user in a visually easy-to-understand format.
[0714] 10. Review feedback and incorporate it into your training plan
[0715] Users can review the displayed feedback, understand areas for improvement in their form, and receive emotional advice, and use this information to restructure their training plan and make specific improvements.
[0716] 11. Re-recording and analyzing the video
[0717] Users can then re-record the video with improved form and re-upload it to the system. By repeating this process, users can expect to continually improve their form and increase the effectiveness of their training.
[0718] In this way, the system of the present invention can improve the quality of training by providing feedback that takes into account not only the athlete's form but also their emotional state. By receiving accurate and specific advice, athletes can achieve effective and sustained improvement in their skills.
[0719] The processing flow will be explained below.
[0720] Step 1:
[0721] User: Record a video of their exercise form using a smartphone or computer. Review the video to ensure the necessary parts are properly recorded.
[0722] Step 2:
[0723] User: Access the dedicated upload page, select the video file you have taken, and click the upload button to send the video data to the server.
[0724] Step 3:
[0725] Terminal: The selected video file is sent to the server as binary data. Once the sending is complete, a message indicating success is displayed to the user.
[0726] Step 4:
[0727] Server: Receives the binary data sent from the device. Temporarily saves the received data as "uploaded_video.mp4".
[0728] Step 5:
[0729] Server: Uses OpenCV library to read the saved video file. Opens the video file and splits it into frames.
[0730] Step 6:
[0731] Server: Loads a pre-trained motion analysis model (e.g., a deep learning model) that analyzes an athlete's form and provides an assessment of specific movements.
[0732] Step 7:
[0733] Server: Divide the video into frames. For example, if the video has 30 frames per second, extract 30 frames per second.
[0734] Step 8:
[0735] Server: Preprocess each frame by resizing it to 224x224 pixels and normalizing pixel values to the range 0 to 1.
[0736] Step 9:
[0737] Server: The preprocessed frames are input into the motion analysis model to obtain form evaluation results.
[0738] Step 10:
[0739] Server: Before generating feedback, emotion recognition is performed using the emotion engine. The facial expressions of the athlete are extracted and input into the emotion engine.
[0740] Step 11:
[0741] Server: The emotion engine recognizes the emotion of the exerciser and acquires the emotion information, which includes happiness, stress, concentration, etc.
[0742] Step 12:
[0743] Server: Generates feedback by combining motion analysis results and emotional information. For example, even if the form is good, if the emotion is unstable, an encouraging comment is added.
[0744] Step 13:
[0745] Server: Sends the generated feedback list to the user's terminal.
[0746] Step 14:
[0747] Terminal: Receives feedback sent from the server and displays it to the user. It presents comments in a format that is easy for the user to understand.
[0748] Step 15:
[0749] User: Check the feedback displayed on the device, understand the problems with their exercise form and emotional advice, and restructure their training plan based on this.
[0750] Step 16:
[0751] User: Record a video with improved form, upload it to the system in the same way, and request analysis. By repeating this process, you can continuously improve your form.
[0752] Example 2
[0753] Next, a description will be given of Example 2. In the following description, the data processing device 12 will be referred to as a "server" and the smart glasses 214 will be referred to as a "terminal."
[0754] Conventional exercise form analysis systems do not provide feedback that takes into account the athlete's emotional state, which can lead to a decline in the quality of training. A means is needed to improve mental motivation and the quality of feedback while allowing athletes to focus on improving their own technique.
[0755] The specific processing by the specific processing unit 290 of the data processing device 12 in the second embodiment is realized by the following means.
[0756] In this invention, the server includes a means for receiving video data captured by an exerciser, a preprocessing means for processing each frame included in the video data, and a means for analyzing the preprocessed frames using a motion analysis model as input, thereby making it possible to generate optimal feedback by combining the analysis results and emotion recognition results.
[0757] An "athlete" is someone who uses their body to exercise or train.
[0758] "Video data" refers to video files taken by the athlete.
[0759] "Preprocessing means" refers to the process performed to convert each frame contained in the video data into a format suitable for the AI model.
[0760] A "motion analysis model" refers to an algorithm or system that uses artificial intelligence to analyze an athlete's form.
[0761] "Analysis results" refers to the evaluations and parameters obtained about the athlete's form by the motion analysis model.
[0762] "Feedback" refers to advice or comments provided to athletes based on the analysis results.
[0763] "Providing means" refers to a method or system for transmitting the generated feedback to the exerciser.
[0764] The "face area" refers to the part of the video data in which the athlete's face is shown.
[0765] "Audio data" refers to the audio track included in the video data.
[0766] "Emotion recognition" refers to the process of identifying an athlete's emotional state from video and audio data.
[0767] "Emotion recognition result" refers to information about the emotional state of the exerciser obtained by the emotion recognition means.
[0768] "Optimal feedback" refers to the advice and comments that are deemed most beneficial to the athlete, generated by combining analysis results and emotion recognition results.
[0769] This invention is a system that allows athletes to improve their technique by analyzing their own form and receiving feedback that takes emotions into account. The system includes a means for receiving and analyzing video data taken by the athlete, and a means for recognizing the athlete's emotions using an emotion engine. This allows for further optimization of feedback to the athlete, supporting effective training.
[0770] Hardware and software used
[0771] The system's hardware uses high-performance servers, specifically those equipped with NVIDIA GPUs. Furthermore, the deep learning framework TensorFlow is used for motion analysis, and Microsoft Azure Cognitive Services is used for emotion recognition. The combination of these technologies enables highly accurate, real-time analysis and feedback.
[0772] Specific operation of the system
[0773] 1. Upload your video
[0774] Users record their own exercise form on their device and send the video data to the server via a dedicated upload page.
[0775] As a specific example, access the specified URL from the device's browser, click the file selection button to select the video data, and then click the upload button.
[0776] 2. Receiving and storing video data
[0777] The server receives the video data sent by the user as an HTTP request and temporarily stores it in storage. When the data is saved, a unique ID is assigned to it.
[0778] 3. Preparing the AI model
[0779] The server loads a pre-trained motion analysis model into memory using the TensorFlow library, which is done only once when the server starts.
[0780] 4. Loading the video and splitting it into frames
[0781] The server reads the saved video file using the OpenCV library and splits the video into frames at a rate of 30 frames per second, each of which is stored in a list.
[0782] 5. Frame Preprocessing
[0783] The server resizes the frames and normalizes pixel values to the range 0 to 1. This preprocessing prepares the images for accurate analysis by the deep learning model.
[0784] 6. Analysis of exercise form
[0785] The server inputs the preprocessed frames into a motion analysis model to extract data such as knee angles and body posture. The analysis results are stored in a list.
[0786] 7. Emotion recognition
[0787] The server extracts the athlete's facial area and voice data from the frames and sends them to Microsoft Azure Cognitive Services for emotion recognition, which identifies the athlete's emotional state.
[0788] 8. Generate feedback
[0789] The server combines the results of motion analysis and emotion recognition to generate optimal feedback for the exerciser. For example, if the exerciser's form is good, the server may comment, "Good form. Keep it up." If the exerciser's emotions are low, the server may add encouraging comments such as, "Great progress. Keep going with confidence."
[0790] 9. Providing Feedback
[0791] The server sends the generated feedback to the user, who then visually displays it on the device. The user can then review the feedback and work on restructuring their training plan or making specific improvements.
[0792] 10. Re-recording and analyzing the video
[0793] Users can then re-record the video with improved form and re-upload it into the system, repeating this process to continually improve their form and increase the effectiveness of their training.
[0794] Prompt Sentence Examples
[0795] markdown
[0796] Write Python code to split a video into frames at 30fps, resize and normalize each frame, and feed it into a deep learning model to analyze athletic form. Then, use Microsoft Azure Cognitive Services to recognize emotions and generate feedback that integrates the analysis results and emotional information.
[0797] In this way, the system of the present invention can improve the quality of training by providing feedback that takes into account not only the athlete's form but also their emotional state. By receiving precise and specific advice, athletes can achieve effective and sustained improvement in their technique.
[0798] The flow of the identification process in the second embodiment will be described with reference to FIG.
[0799] Step 1:
[0800] Recording and uploading videos
[0801] Users use their device to record their own exercise form and send the video data to the server from a dedicated upload page. They access the specified URL from their device's browser, click the file selection button to select the video data, and then click the upload button to send it.
[0802] Input: Video file of the athlete's exercise form
[0803] Output: Video data sent to the server
[0804] Step 2:
[0805] Receiving and storing video data
[0806] The server receives the video data sent by the user as an HTTP request and temporarily stores it in storage. The server assigns a unique ID to the received file and stores it in a folder structure. This saving operation is performed asynchronously, returning a prompt response to the user.
[0807] Input: Video data sent by the user
[0808] Output: Video file saved in storage
[0809] Step 3:
[0810] Preparing the AI model
[0811] The server loads a pre-trained motion analysis model into memory using the TensorFlow library, a process that is performed only once when the server starts.
[0812] Input: The trained motion analysis model file
[0813] Output: Motion analysis model loaded into memory
[0814] Step 4:
[0815] Video loading and frame splitting
[0816] The server finds the stored video file and splits the video into frames using the OpenCV library by opening the video file, capturing frames at a rate of 30 frames per second, and storing them in a list.
[0817] Input: Video file saved in storage
[0818] Output: Each frame stored in a list
[0819] Step 5:
[0820] Frame Preprocessing
[0821] The server resizes each frame and normalizes pixel values to a range between 0 and 1. This preprocessing converts the frames into a format suitable for AI models. Specifically, it resizes each frame to a uniform size and normalizes the values per pixel.
[0822] Input: Each frame stored in a list
[0823] Output: Preprocessed frames
[0824] Step 6:
[0825] Analysis of athletic form
[0826] The server inputs the preprocessed frames into the motion analysis model to extract data such as knee angle and body posture, performs the analysis using the model's predict method, and saves the results in a list.
[0827] Input: Preprocessed frame
[0828] Output: Exercise form data as analysis results (e.g. knee angle, body posture, etc.)
[0829] Step 7:
[0830] emotion recognition
[0831] The server extracts the athlete's facial area and voice data from the frames and inputs them into the emotion engine. Specifically, it sends an HTTP request to the API endpoint of Microsoft Azure Cognitive Services to recognize emotions from facial expressions and voice.
[0832] Input: Face area and audio data in the frame
[0833] Output: Emotion recognition results (e.g., joy, sadness, anger, etc.)
[0834] Step 8:
[0835] Generate feedback
[0836] The server combines the results of motion analysis and emotion recognition to generate feedback for the exerciser. Specifically, it embeds comments based on the analysis results and encouraging messages based on the exerciser's emotional state into templates.
[0837] Input: Motion analysis results, emotion recognition results
[0838] Output: The generated feedback message
[0839] Step 9:
[0840] Providing feedback
[0841] The server sends the generated feedback to the user, and the device displays it visually, either by returning a feedback message as an HTTP response or by delivering it to the user via email or push notification.
[0842] Input: The generated feedback message
[0843] Output: Feedback displayed on the terminal
[0844] Step 10:
[0845] Review feedback and incorporate it into your training plan
[0846] Users can view feedback on their device, understand areas for improvement in their form, and receive emotional advice, which they can use to restructure their training plan and make specific improvements.
[0847] Input: Feedback displayed on the device
[0848] Output: Improved training plan
[0849] Step 11:
[0850] Re-shooting and analyzing the video
[0851] Users can then re-record the video with improved form and re-upload it to the system, which will enable continuous improvement of form and increased training effectiveness.
[0852] Input: New video file taken with improved form
[0853] Output: Newly uploaded video data to the server
[0854] (Application example 2)
[0855] Next, a description will be given of Application Example 2. In the following description, the data processing device 12 will be referred to as a "server" and the smart glasses 214 will be referred to as a "terminal."
[0856] Conventional motion analysis systems only provide feedback on the athlete's form, but are unable to provide appropriate feedback that takes into account the athlete's emotional state. This can result in training motivation and effectiveness not being maximized. Furthermore, in operational environments such as factory robots, not only is improving work accuracy and efficiency important, but also managing the operator's stress is also an important issue.
[0857] The specific processing by the specific processing unit 290 of the data processing device 12 in Application Example 2 is realized by the following means. In this invention, the server includes means for receiving video data taken by an exerciser, preprocessing means for processing each frame included in the video data, means for analyzing the preprocessed frames using a motion analysis model as input, means for generating feedback on the exerciser's form based on the analysis results, means for recognizing emotions based on the exerciser's facial expressions and voice data, means for optimizing and generating feedback taking into account the emotion recognition results, and means for providing the generated feedback to the exerciser. This makes it possible to provide feedback that takes into account not only the form of the exerciser and operator but also their overall state, including emotions.
[0858] "Video data" refers to video footage of the movements of an athlete or robot.
[0859] "Preprocessing" refers to processes such as resizing and normalization that are performed to convert frames of video data into a format suitable for the analysis model.
[0860] "Movement analysis model" refers to an artificial intelligence model used to analyze the movements of an athlete or robot and evaluate their form and accuracy.
[0861] "Feedback" refers to advice or comments provided to an exerciser or operator based on the analysis results or emotion recognition results.
[0862] "Emotion recognition" refers to analyzing the facial expressions and voices of athletes and operators contained in video data to identify their emotional state.
[0863] "Optimization" refers to adjusting feedback based on acquired data to provide more effective advice.
[0864] A "frame" refers to an individual still image that makes up video data.
[0865] The present invention relates to a system that analyzes the operation of a factory robot and supports its efficient and highly accurate movement.
[0866] The system uses video data captured by the user to analyze the accuracy and efficiency of factory robot operation and provides feedback. It also recognizes the emotional state of the user and workers and optimizes feedback as needed.
[0867] Hardware:
[0868] Camera: Used to capture video data.
[0869] GPU: Used to run AI models at high speed.
[0870] software:
[0871] OpenCV: Used to load videos and split them into frames.
[0872] PyTorch: Used to run AI models for motion analysis.
[0873] EmotionEngine: A custom library for emotion recognition.
[0874] MovementAnalyzer: A custom library for performing movement analysis.
[0875] The server performs the following process:
[0876] First, the user uses a camera to record the movements of an athlete or a factory robot. The captured video data is then sent to the server, which receives it and temporarily stores it. The stored video data is then divided into frames, resized and normalized as preprocessing, and converted into a format suitable for the AI model.
[0877] The preprocessed frames are input into a motion analysis model to evaluate the form and movement accuracy of the athlete or robot. The evaluation results are obtained as specific movement parameters. In addition, the server extracts facial expressions and voice data of the user or worker from the video data and inputs them into the Emotion Engine. The Emotion Engine analyzes this data and identifies their emotional state.
[0878] The server combines the results of motion analysis and emotion recognition to generate feedback. For example, if the motion accuracy is high, it will comment, "Your operation is accurate. Please keep going," and if the emotion is stressed, it will generate advice such as, "Your operation is good, but you seem to be stressed. Please take a short break."
[0879] The server sends the generated feedback to the user's device, which then visually presents it to the user. The user can then review the feedback, understand the areas for improvement in their own operations and movements, and receive emotional advice, which they can then incorporate into their training plan.
[0880] As a concrete example, consider a robot operator at a factory who is currently learning a new operating procedure. The operator films his or her operation with a camera and uploads it to the system. The system analyzes the video data and provides specific feedback such as, "Your vehicle handling is good, but some improvement is needed in conveyor line operation." It also recognizes fatigue from the operator's facial expression and adds advice such as, "Take a short break and refresh yourself."
[0881] Example prompt sentence:
[0882] Upload a video of a factory robot operating. Analyze the robot's movements and the operator's facial expressions to generate specific feedback to improve accuracy. Include comments based on the operator's level of fatigue and stress.
[0883] The flow of the specific processing in the application example 2 will be described with reference to FIG.
[0884] Program processing flow of the application example system
[0885] Step 1:
[0886] The server receives video data captured by the user of an athlete or a factory robot. Specifically, the user sends the video captured with a camera to the server from the upload page. The input is the captured video data, and the output is a video file stored on the server.
[0887] Step 2:
[0888] The server divides the received video data into frames. Specifically, it uses OpenCV to read the video data and extract each frame. The input is a saved video file, and the output is multiple frame images.
[0889] Step 3:
[0890] The server preprocesses each frame. Specifically, it resizes and normalizes the frame image to a format suitable for the AI model. The input is the extracted frame image, and the output is the preprocessed frame image.
[0891] Step 4:
[0892] The server inputs the preprocessed frames into a motion analysis model. Specifically, the preprocessed frame images are input into a motion analysis model running on PyTorch to obtain motion parameters. The input is the preprocessed frame images, and the output is the motion parameters.
[0893] Step 5:
[0894] The server extracts the user's facial expression and voice data from the frames. Specifically, it uses OpenCV and voice analysis libraries to extract the necessary information from the frame images and voice data. The input is the preprocessed frame images and voice data, and the output is the user's facial expression and voice data.
[0895] Step 6:
[0896] The server inputs the extracted facial and voice data into an emotion recognition engine. Specifically, it uses the EmotionEngine to identify the emotional state. The input is the user's facial and voice data, and the output is the emotion recognition result.
[0897] Step 7:
[0898] The server generates feedback based on the results of motion analysis and emotion recognition. Specifically, it combines the analysis data to generate appropriate comments and advice. The input is the motion parameters of the motion analysis and the emotion recognition results, and the output is the generated feedback.
[0899] Step 8:
[0900] The server transmits the generated feedback to the user's terminal. Specifically, the server converts the generated feedback into a format for visual display and transmits it to the user's terminal. The input is the generated feedback, and the output is the feedback displayed on the user's terminal.
[0901] Step 9:
[0902] The user checks the feedback displayed on the device and understands the areas for improvement in their own operations and behaviors, as well as emotional advice. Specifically, they restructure their training plan based on the feedback and work on specific improvements. The input is the feedback displayed on the device, and the output is the improved operations and behaviors.
[0903] summary
[0904] This step enables the system to analyze the movements of athletes and factory robots and provide appropriate feedback that takes emotions into account.
[0905] The specific processing unit 290 transmits the result of the specific processing to the smart glasses 214. In the smart glasses 214, the control unit 46A causes the speaker 240 to output the result of the specific processing. The microphone 238 acquires audio indicating a user input regarding the result of the specific processing. The control unit 46A transmits audio data indicating the user input acquired by the microphone 238 to the data processing device 12. In the data processing device 12, the specific processing unit 290 acquires the audio data.
[0906] The data generation model 58 is a so-called generative AI (Artificial Intelligence). An example of the data generation model 58 is ChatGPT (Internet Search<URL: https: / / openai.com / blog / chatgpt> ), Gemini (Internet search <url: https: gemini.google.com ?hl="ja">) and other generation AIs. The data generation model 58 is obtained by performing deep learning on a neural network. A prompt including an instruction is input to the data generation model 58, and inference data such as voice data indicating voice, text data indicating text, and image data indicating an image is also input. The data generation model 58 performs inference on the input inference data in accordance with the instruction indicated by the prompt, and outputs the inference result in a data format such as voice data and text data. Here, inference refers to, for example, analysis, classification, prediction, and / or summarization.
[0907] In the above embodiment, an example in which the specific processing is performed by the data processing device 12 has been given, but the technology of the present disclosure is not limited to this, and the specific processing may be performed by the smart glasses 214.
[0908] [Third embodiment]
[0909] FIG. 5 shows an example of the configuration of a data processing system 310 according to the third embodiment.
[0910] 5, the data processing system 310 includes the data processing device 12 and a headset type terminal 314. An example of the data processing device 12 is a server.
[0911] The data processing device 12 includes a computer 22, a database 24, and a communication I / F 26. The computer 22 is an example of a "computer" according to the technology of the present disclosure. The computer 22 includes a processor 28, a RAM 30, and a storage 32. The processor 28, the RAM 30, and the storage 32 are connected to a bus 34. The database 24 and the communication I / F 26 are also connected to the bus 34. The communication I / F 26 is connected to a network 54. Examples of the network 54 include a WAN (Wide Area Network) and / or a LAN (Local Area Network).
[0912] The headset type terminal 314 includes a computer 36, a microphone 238, a speaker 240, a camera 42, a communication I / F 44, and a display 343. The computer 36 includes a processor 46, a RAM 48, and a storage 50. The processor 46, the RAM 48, and the storage 50 are connected to a bus 52. The microphone 238, the speaker 240, the camera 42, and the display 343 are also connected to the bus 52.
[0913] The microphone 238 receives instructions and the like from the user 20 by receiving voice uttered by the user 20. The microphone 238 captures the voice uttered by the user 20, converts the captured voice into audio data, and outputs it to the processor 46. The speaker 240 outputs audio in accordance with instructions from the processor 46.
[0914] Camera 42 is a small digital camera equipped with an optical system including a lens, aperture, and shutter, and an imaging element such as a CMOS (Complementary Metal-Oxide-Semiconductor) image sensor or a CCD (Charge Coupled Device) image sensor, and captures images of the surroundings of user 20 (for example, an imaging range defined by an angle of view equivalent to the field of vision of a typical healthy person).
[0915] The communication I / F 44 is connected to a network 54. The communication I / Fs 44 and 26 are responsible for the exchange of various information between the processor 46 and the processor 28 via the network 54. The exchange of various information between the processor 46 and the processor 28 using the communication I / Fs 44 and 26 is carried out in a secure state.
[0916] Fig. 6 shows an example of the main functions of the data processing device 12 and the headset type terminal 314. As shown in Fig. 6, in the data processing device 12, a specific process is performed by the processor 28. A specific process program 56 is stored in the storage 32.
[0917] The specific processing program 56 is an example of a "program" according to the technology of the present disclosure. The processor 28 reads the specific processing program 56 from the storage 32 and executes the read specific processing program 56 on the RAM 30. The specific processing is realized by the processor 28 operating as a specific processing unit 290 in accordance with the specific processing program 56 executed on the RAM 30.
[0918] The storage 32 stores a data generation model 58 and an emotion identification model 59. The data generation model 58 and the emotion identification model 59 are used by the identification processing unit 290.
[0919] In the headset type terminal 314, a reception output process is performed by the processor 46. A reception output program 60 is stored in the storage 50. The processor 46 reads the reception output program 60 from the storage 50 and executes the read reception output program 60 on the RAM 48. The reception output process is realized by the processor 46 operating as the control unit 46A in accordance with the reception output program 60 executed on the RAM 48.
[0920] Next, a description will be given of the identification process performed by the identification processing unit 290 of the data processing device 12. In the following description, the data processing device 12 will be referred to as the "server" and the headset type terminal 314 will be referred to as the "terminal."
[0921] This invention relates to a system that allows athletes to analyze their own form and receive specific advice for improving their technique. This system allows athletes to take video data using devices such as smartphones or computers, upload it to a server, and then analyzes the video data using an AI model to provide detailed feedback.
[0922] Specific Embodiments of the System
[0923] 1. Upload your video
[0924] Users upload videos of their exercise form from their devices to the server, where the uploaded videos are temporarily stored.
[0925] 2. Receiving and storing video data
[0926] The server receives the video data sent by the user and saves it under a file name such as "uploaded_video.mp4." This saved video data will then be analyzed.
[0927] 3. Preparing and Loading the AI Model
[0928] The server loads a pre-trained motion analysis model (e.g., a deep learning model) that analyzes an athlete's form and provides an assessment of specific movements.
[0929] 4. Loading the video and splitting it into frames
[0930] The server extracts the stored video file frame by frame. Since video is usually 30 frames per second, it is processed as 30 images per second.
[0931] 5. Frame Preprocessing
[0932] The server performs preprocessing such as frame resizing and normalization to convert each frame into a format suitable for the analytical model. Specifically, the frame is resized to 224x224 pixels and the pixel values are normalized to the range 0 to 1.
[0933] 6. Form Analysis
[0934] The server then feeds the preprocessed frames into a motion analysis model to evaluate form, measuring specific movement parameters such as knee angle and body tilt, and returns the results.
[0935] 7. Generate feedback
[0936] The server generates feedback for each frame based on the analysis results, and provides the user with specific advice such as "Good form. Keep going like this" or "Review the way you bend your knees."
[0937] 8. Providing Feedback
[0938] The server compiles the generated feedback and sends it back to the user's device, which then displays it to the user in a visually understandable format.
[0939] 9. Review feedback and incorporate it into your training plan
[0940] Users can view the feedback displayed on their device to understand where their form is lacking, and use that feedback to restructure their training plan and address specific areas for improvement.
[0941] 10. Re-recording and analyzing the video
[0942] Users can re-record a video with a new form and re-upload it into the system, repeating this process to continually improve their form.
[0943] As described above, the present invention allows athletes to improve the quality of their training by analyzing their own form in detail and receiving specific, individually customized feedback. This system is expected to make sports skill improvement even more efficient.
[0944] The processing flow will be explained below.
[0945] Step 1:
[0946] User: Record a video of their exercise form using a smartphone or computer. Review the video to ensure the necessary parts are properly recorded.
[0947] Step 2:
[0948] User: Access the dedicated upload page from their device, select the video file they shot on that page, and click the upload button.
[0949] Step 3:
[0950] Terminal: The selected video file is sent to the server as binary data. Once the sending is complete, a message indicating success is displayed to the user.
[0951] Step 4:
[0952] Server: Receives the binary data sent from the device. Temporarily saves the received data as "uploaded_video.mp4".
[0953] Step 5:
[0954] Server: Uses the OpenCV library to read the saved video file. Opens the video file and prepares to extract data frame by frame.
[0955] Step 6:
[0956] Server: Loads a pre-trained motion analysis model (e.g., "pose_analysis_model.h5") that analyzes an athlete's form and evaluates specific movements.
[0957] Step 7:
[0958] Server: Divide the video into frames. For example, if the video is shot at 30 frames per second, extract 30 frames per second.
[0959] Step 8:
[0960] Server: Preprocess each frame by resizing it to 224x224 pixels and normalizing pixel values to the range 0 to 1.
[0961] Step 9:
[0962] Server: Inputs the preprocessed frames into the motion analysis model, which outputs the motion estimation results for each frame.
[0963] Step 10:
[0964] Server: Generates feedback based on the analysis results. For example, if the model determines that the knee angle is inappropriate, it generates specific advice such as "Reconsider how you bend your knees."
[0965] Step 11:
[0966] Server: Compiles the generated feedback into a list and prepares it for sending back to the user.
[0967] Step 12:
[0968] Server: Sends the feedback list to the user's terminal.
[0969] Step 13:
[0970] Terminal: Receives feedback sent from the server and displays it to the user. It presents comments in a format that is easy for the user to understand.
[0971] Step 14:
[0972] User: Check the feedback displayed on the device to understand problems and areas for improvement in their exercise form, and restructure their training plan based on this.
[0973] Step 15:
[0974] User: Record a video with improved form, upload it to the system in the same way, and request analysis. By repeating this process, you can continuously improve your form.
[0975] Example 1
[0976] Next, a description will be given of Example 1. In the following description, the data processing device 12 will be referred to as a "server" and the headset type terminal 314 will be referred to as a "terminal."
[0977] Conventional exercise form analysis systems have the drawback of requiring users to film their own form, send the video data to a dedicated device, and obtain analysis results, all of which require a complex and time-consuming process. Furthermore, the accuracy of the analysis results is low, making it difficult for users to identify specific areas for improvement. Furthermore, many systems do not provide instant feedback, making it difficult to incorporate the results into effective training plans.
[0978] The specific processing by the specific processing unit 290 of the data processing device 12 in the first embodiment is realized by the following means.
[0979] In this invention, the server includes means for receiving video data taken by an exerciser, means for saving the video data, means for dividing the video data into frames, means for preprocessing each of the frames, means for analyzing the preprocessed frames using a generative AI model, means for generating feedback on the exerciser's form based on the analysis results, means for providing the generated feedback to the exerciser, and means for displaying the provided feedback on the exerciser's device. This provides high-precision feedback in real time, allowing the user to quickly identify areas for improvement in their form and implement an effective training plan.
[0980] An "athlete" refers to an individual who exercises, photographs their own exercise form, and uses the data for analysis.
[0981] "Video data" refers to video files taken by the athlete, and contains information to be analyzed.
[0982] "Server" refers to a computer system that receives, stores, processes, and provides analysis results of video data.
[0983] "Each frame" refers to the individual images that make up the video data, and the information per frame is used for analysis.
[0984] "Preprocessing" refers to data manipulation such as resizing and normalization to convert each frame into a format suitable for the analytical model.
[0985] A "generative AI model" is a pre-trained AI (artificial intelligence) model used to analyze and evaluate exercise form.
[0986] "Analysis results" refers to the evaluation data of exercise form obtained by the generative AI model, and includes specific movement parameters (such as knee angle and body tilt).
[0987] "Feedback" refers to advice and evaluations generated based on the analysis results, and is information that athletes can use to improve their form.
[0988] "Terminal" refers to the device used by the exerciser, such as a smartphone or computer, that displays the feedback.
[0989] This invention relates to a system that allows athletes to analyze their own form and receive specific advice for improving their technique. The system allows athletes to take video data using devices such as smartphones or computers, upload it to a server, and then analyzes the video data using a generative AI model to provide detailed feedback.
[0990] Hardware and software used
[0991] Hardware: smartphones, computers, servers
[0992] Software: Video recording app, video uploading tool, motion analysis AI model (deep learning model)
[0993] Data processing and calculation methods
[0994] Users can upload videos of their exercise form to a server using a smartphone or computer, for example, via a dedicated application or browser.
[0995] The server receives the video data sent by the user and temporarily stores it in a specific directory (e.g., "uploaded_videos / ") on the server with the file name "uploaded_video.mp4". This data is used in the subsequent analysis steps.
[0996] The server loads pre-trained AI models for motion analysis, which use deep learning techniques to analyze athletic form, including models trained using TensorFlow and PyTorch.
[0997] The server loads the saved video file and uses tools such as ffmpeg to extract the video frame by frame. Since video is typically 30 frames per second, this translates to 30 images per second.
[0998] Each frame is converted into a format suitable for the analytical model. In this step, the frame is resized to 224x224 pixels and the pixel values are normalized to the range 0 to 1. This process is often performed using OpenCV or PIL (Python Imaging Library).
[0999] The preprocessed frames are fed into a generative AI model to evaluate the form. The model measures movement parameters such as knee angle and body tilt and returns the analysis results. For example, the model outputs data such as "knee angle is 45 degrees" for each frame.
[1000] The server generates feedback based on the analysis results, including specific advice such as "Try to widen your knee angle by about 5 degrees." This feedback is created for each frame and is customized for each user.
[1001] The generated feedback is sent back from the server to the user's device, which then displays it to the user in a visually understandable format, for example, highlighting key points in addition to displaying them as text messages within the application.
[1002] Users can view the feedback displayed on their device, understand the issues with their form, and use that feedback to restructure their training plan and address specific areas for improvement, such as paying special attention to knee angle during their next workout.
[1003] Examples of specific examples and prompts
[1004] For example, suppose a user wants to improve their running form. The user films their running form with their smartphone and uploads the video to the system. The server analyzes the video and generates feedback to the user, such as, "The angle of your right knee is not appropriate. Try bending your right knee a bit more while running." The user can then revise their training plan based on this feedback and review their form again.
[1005] Prompt Sentence Examples
[1006] "I uploaded a video of my running form. Please help me analyze where my form needs improvement."
[1007] "Please analyze the video of my form and tell me specific areas for improvement."
[1008] The system allows exercisers to receive detailed and specific feedback to continually improve their form.
[1009] The flow of the identification process in the first embodiment will be described with reference to FIG.
[1010] Step 1:
[1011] The user takes a video of their exercise form on the device.
[1012] Input: A video of your exercise form (e.g. "run.mp4")
[1013] Output: Recorded video file
[1014] Specific movements: Using a smartphone or computer, participants will be videotaped while performing various exercises, such as running, jumping, and squatting.
[1015] Step 2:
[1016] The user uploads the video they have taken from their terminal to the server.
[1017] Input: Recorded video file
[1018] Output: Video data uploaded to the server
[1019] Specific operation: Use a dedicated application or web browser to send the video file to the server. For example, click the button, select "run.mp4", and start uploading.
[1020] Step 3:
[1021] The server receives and stores the video data sent by the user.
[1022] Input: Video data received from the user
[1023] Output: Video file saved on the server (e.g. "uploaded_video.mp4")
[1024] What happens: The server stores the uploaded data in a specific directory (e.g. "uploaded_videos / ").
[1025] Step 4:
[1026] The server loads a pre-trained generative AI model.
[1027] Input: A trained generative AI model
[1028] Output: Generative AI model loaded into memory
[1029] What it does: Loads a motion analysis model trained using TensorFlow or PyTorch, and the system is ready for analysis.
[1030] Step 5:
[1031] The server reads the saved video file and splits it into frames.
[1032] Input: Video file stored on the server ("uploaded_video.mp4")
[1033] Output: Images separated by frames (e.g. frame 1, frame 2, ...)
[1034] Specific operation: Using a video processing tool such as ffmpeg, the video is divided into 30 frames per second and each frame is obtained as an image file.
[1035] Step 6:
[1036] The server performs pre-processing on each frame.
[1037] Input: Images split into frames
[1038] Output: Preprocessed frame (resized and normalized)
[1039] Specific operation: Using OpenCV and PIL (Python Imaging Library), each frame is resized to 224x224 pixels and the value of each pixel is normalized to the range 0 to 1.
[1040] Step 7:
[1041] The server feeds the preprocessed frames into a generative AI model to analyze the form.
[1042] Input: Preprocessed frame
[1043] Output: Analysis results (motion parameters, e.g. knee angle, body tilt)
[1044] Specific movement: The preprocessed frame is input into the generative AI model, movement parameters are measured, and an evaluation result is output. For example, the result may be "Frame 10: Knee angle 45 degrees."
[1045] Step 8:
[1046] The server generates feedback based on the analysis results.
[1047] Input: Analysis results
[1048] Output: Generated feedback (e.g. "Try to open your knees by about 5 degrees.")
[1049] Specific actions: Based on the analysis results for each frame, specific advice on how to improve your athletic form is generated. For example, "Your right knee angle is not appropriate, so try bending it a little more while running."
[1050] Step 9:
[1051] The server transmits the generated feedback to the user's terminal.
[1052] Input: Generated feedback
[1053] Output: Feedback sent to the user's device
[1054] Specific operation: The server compiles the feedback data and sends it to the user's device, for example, by sending the feedback data back using an HTTP request.
[1055] Step 10:
[1056] The terminal displays the received feedback to the user.
[1057] Input: Feedback sent by the server
[1058] Output: The feedback screen presented to the user.
[1059] What it does: The device visualizes the feedback it receives and presents it to the user in a user-friendly format, for example by using text messages or highlighting to emphasize key points.
[1060] Step 11:
[1061] The user checks the feedback displayed on the device and reflects it in their training plan.
[1062] Input: Feedback displayed on the device
[1063] Output: Modified training plan
[1064] What it does: Users can use the feedback to understand where their form needs improvement and focus on those areas during their next workout.
[1065] Step 12:
[1066] The user then shoots the video again with the improved form and uploads it to the system again.
[1067] Input: Improved video
[1068] Output: New video data uploaded to the system
[1069] What it does: The user takes a photo of a new form and repeats the steps above to provide data to the system, which then continually improves the form.
[1070] (Application example 1)
[1071] Next, a description will be given of Application Example 1. In the following description, the data processing device 12 will be referred to as a "server" and the headset type terminal 314 will be referred to as a "terminal."
[1072] Conventional motion analysis systems have been used only to provide feedback to athletes and general exercisers to improve their exercise form, but the concept has not been applied to the motion analysis of exercise machines and equipment in factories. In factories, accurate analysis of motion and appropriate feedback are essential to maintain and improve the efficiency and accuracy of highly automated equipment and robots. However, current systems have difficulty meeting these needs, and have the problem of being unable to respond immediately when a decline in motion efficiency or accuracy occurs.
[1073] The specific processing by the specific processing unit 290 of the data processing device 12 in the application example 1 is realized by the following means.
[1074] In this invention, the server includes: means for receiving video data captured by an exerciser; preprocessing means for processing each frame included in the video data; means for analyzing the preprocessed frames using a motion analysis model; means for generating feedback on the exerciser's form based on the analysis results; means for providing the generated feedback to the exerciser; a device for capturing the movements of exercise equipment or devices in a factory; means for receiving and preprocessing the captured movement data; means for evaluating movements using an AI model for analyzing the preprocessed movement data; means for generating feedback on improving the efficiency and accuracy of movements based on the evaluation results; and means for providing the generated feedback to personnel involved in operating the factory equipment. This enables the movement analysis of exercise equipment or devices in a factory, and enables quick and appropriate feedback to be provided to improve efficiency and accuracy.
[1075] "Exercise participant" refers to the person or equipment whose exercise form is being analyzed and feedback is being provided.
[1076] "Video data" refers to data containing visual information that records the movements of an exerciser or the movements of exercise equipment.
[1077] "Means for receiving" refers to the technical methods and equipment for accepting video data from an external device.
[1078] "Preprocessing means" refers to a method or device that performs processing to convert video data into a format suitable for the analysis model.
[1079] "Movement analysis model" refers to an artificial intelligence model used to analyze an athlete's form and movement data.
[1080] "Means for analysis" refers to the techniques and devices used to input preprocessed data into a motion analysis model and obtain results.
[1081] "Means for generating feedback" refers to techniques and devices for generating information that indicates specific areas for improvement and direction based on the analysis results.
[1082] "Means of providing" refers to the techniques and devices used to communicate the generated feedback to the exerciser or person in charge.
[1083] "Motion equipment or devices" refers to machines, robots, and other equipment that perform specific movements in a factory.
[1084] "Motion Data" refers to data, including visual information, that records the movements of exercise equipment or devices.
[1085] "Means for evaluating" refers to techniques or devices for evaluating the efficiency and accuracy of a movement using a motion analysis model based on preprocessed movement data.
[1086] This invention is a system based on the "athlete's form analysis system" that is applied to the analysis of the movements of exercise machines or equipment in factories and the provision of feedback. This system can monitor the efficiency and accuracy of the movements of robots and machines in factories and provide necessary advice for improvement.
[1087] First, cameras are placed to capture the operation of the motion machines and equipment used in the factory. This video data is sent to a server. The server then acquires this video data using a receiving means. The hardware used in this case can be a high-performance camera.
[1088] The server then processes the received video data using preprocessing tools. Specifically, it divides the video data into frames and converts them into a format suitable for the AI model. This preprocessing uses OpenCV and PIL (Pillow) to resize and normalize the images, allowing for detailed analysis of the motion.
[1089] The preprocessed data is then input into an AI model, the motion analysis model. This model uses deep learning to analyze the motion frame by frame and measure specific motion parameters. PyTorch is used for the analysis, and the model's motion is evaluated.
[1090] The server generates feedback regarding improvements to the efficiency and accuracy of the operation based on the analysis results. The feedback generation means creates information indicating specific advice and areas for improvement based on the analysis results. For example, feedback such as "The accuracy of the operation is low. Adjustments are required" may be generated.
[1091] Finally, the server provides the generated feedback to personnel involved in the operation of the factory equipment using a means for providing the feedback, which can then be used by the personnel to adjust the robots and equipment to improve the efficiency and accuracy of their operations.
[1092] As a specific example, a robot arm in a factory is filmed as it assembles parts, and the video data is uploaded to the system. The system analyzes the video data and generates specific feedback such as, "The robot arm's movement is delayed. Please adjust the angle at which the part is attached." Based on this, the worker can adjust the robot arm's movement, improving the efficiency of the assembly work.
[1093] An example of a prompt sentence to be input to the generative AI model is as follows:
[1094] Consider a scenario where a system analyzes the behavior of a factory robot and provides feedback on its efficiency and accuracy. Specifically, imagine a robotic arm takes a video of itself assembling a part, and an AI model analyzes its behavior and provides advice on how to improve it.
[1095] The flow of the specific processing in the application example 1 will be described with reference to FIG.
[1096] Step 1:
[1097] Users use cameras to capture the operation of exercise equipment and devices in the factory and upload the video data to a server.
[1098] Input: Video data containing user-recorded actions.
[1099] Output: Video data uploaded to the server.
[1100] Specific operation: The movement is filmed using a high-performance camera, and the user uploads the captured video data to a server via a dedicated application.
[1101] Step 2:
[1102] The server obtains the uploaded video data using a receiving means and temporarily stores it.
[1103] Input: User uploaded video data.
[1104] Output: Video data temporarily stored on the server.
[1105] Specific operation: The server saves the video data in a dedicated folder called "uploaded_video.mp4".
[1106] Step 3:
[1107] The server uses a pre-processing means to divide the received video data into frames.
[1108] Input: Video data stored on the server.
[1109] Output: Image data separated by frames.
[1110] Specific operation: Using OpenCV, video data is extracted frame by frame and divided at a rate of 30 frames per second.
[1111] Step 4:
[1112] The server uses pre-processing means to convert each frame into a format suitable for image analysis.
[1113] Input: Each split frame.
[1114] Output: Image data in a format suitable for analytical models.
[1115] What it does: Uses PIL (Pillow) to resize the image (to 224x224 pixels) and normalize it (convert pixel values to the range 0 to 1).
[1116] Step 5:
[1117] The server inputs the preprocessed frames into the AI model and analyzes the behavior.
[1118] Input: Preprocessed image data.
[1119] Output: Motion analysis result data.
[1120] Specific operation: Using PyTorch, preprocessed data is input to a deep learning model and analysis results are output.
[1121] Step 6:
[1122] The server generates feedback on the efficiency and accuracy of the operation based on the analysis results.
[1123] Input: Motion analysis result data.
[1124] Output: Feedback with suggestions for improvement and advice.
[1125] Specific actions: Evaluate the analysis results and use feedback generation methods to generate specific advice such as "The accuracy of the action is low. Adjustments are required."
[1126] Step 7:
[1127] The server provides the generated feedback to personnel involved in operating the factory equipment.
[1128] Input: Feedback data.
[1129] Output: Feedback provided to the agent.
[1130] What it does: Sends feedback to the device and presents it to the agent in a visual format via the application.
[1131] Step 8:
[1132] Based on the feedback, the staff will adjust the operation of the exercise equipment and devices and take further photographs to confirm the effectiveness.
[1133] Input: Adjustment information based on feedback.
[1134] Output: New video data recording the adjusted behavior.
[1135] Specific Action: Interpret the feedback provided and change the settings of factory equipment or devices, then photograph the new action and upload it back into the system.
[1136] Furthermore, an emotion engine that estimates the user's emotion may be further combined. That is, the identification processing unit 290 may estimate the user's emotion using the emotion identification model 59, and perform identification processing using the user's emotion.
[1137] The present invention relates to a system that allows athletes to analyze their own form and receive feedback to improve their technique. The system includes a means for receiving and analyzing video data taken by the athlete, and a means for recognizing the athlete's emotions using an emotion engine. This allows for further optimization of the feedback provided to the athlete, supporting effective training.
[1138] Specific Embodiments of the System
[1139] 1. Upload your video
[1140] Users can record their own exercise form on their device, check the recorded video data, and send it to the server via a dedicated upload page.
[1141] 2. Receiving and storing video data
[1142] The server receives the video data sent by the user and temporarily stores it, which is then analyzed.
[1143] 3. Preparing the AI model
[1144] The server loads a pre-trained motion analysis model, which is used to individually analyze the athlete's form.
[1145] 4. Loading the video
[1146] The server reads the stored video file and splits it into frames, which typically consist of 30 frames per second.
[1147] 5. Frame Preprocessing
[1148] The server resizes and normalizes each frame into a format suitable for the analytical model, making the frame ready to be input into the AI model.
[1149] 6. Form Analysis
[1150] The server runs the pre-processed frames through a motion analysis model to evaluate the athlete's form, resulting in specific movement parameters such as knee angle and body posture.
[1151] 7. Emotion Recognition with Emotion Engine
[1152] The server extracts the facial expressions and voice data of the athlete from the frames and inputs them into the emotion engine, which then recognizes the athlete's emotional state from these data.
[1153] 8. Generate feedback
[1154] The server combines the results of the motion analysis with the output of the emotion engine to generate feedback for the exerciser. For example, if the exerciser's form is good, the server may comment, "Good form. Keep it up." If the exerciser's emotions are low, the server may add encouraging comments such as, "Great progress. Keep going with confidence."
[1155] 9. Providing Feedback
[1156] The server transmits the generated feedback to the user, and the terminal displays the feedback to the user in a visually easy-to-understand format.
[1157] 10. Review feedback and incorporate it into your training plan
[1158] Users can review the displayed feedback, understand areas for improvement in their form, and receive emotional advice, and use this information to restructure their training plan and make specific improvements.
[1159] 11. Re-recording and analyzing the video
[1160] Users can then re-record the video with improved form and re-upload it to the system. By repeating this process, users can expect to continually improve their form and increase the effectiveness of their training.
[1161] In this way, the system of the present invention can improve the quality of training by providing feedback that takes into account not only the athlete's form but also their emotional state. By receiving accurate and specific advice, athletes can achieve effective and sustained improvement in their skills.
[1162] The processing flow will be explained below.
[1163] Step 1:
[1164] User: Record a video of their exercise form using a smartphone or computer. Review the video to ensure the necessary parts are properly recorded.
[1165] Step 2:
[1166] User: Access the dedicated upload page, select the video file you have taken, and click the upload button to send the video data to the server.
[1167] Step 3:
[1168] Terminal: The selected video file is sent to the server as binary data. Once the sending is complete, a message indicating success is displayed to the user.
[1169] Step 4:
[1170] Server: Receives the binary data sent from the device. Temporarily saves the received data as "uploaded_video.mp4".
[1171] Step 5:
[1172] Server: Uses OpenCV library to read the saved video file. Opens the video file and splits it into frames.
[1173] Step 6:
[1174] Server: Loads a pre-trained motion analysis model (e.g., a deep learning model) that analyzes an athlete's form and provides an assessment of specific movements.
[1175] Step 7:
[1176] Server: Divide the video into frames. For example, if the video has 30 frames per second, extract 30 frames per second.
[1177] Step 8:
[1178] Server: Preprocess each frame by resizing it to 224x224 pixels and normalizing pixel values to the range 0 to 1.
[1179] Step 9:
[1180] Server: The preprocessed frames are input into the motion analysis model to obtain form evaluation results.
[1181] Step 10:
[1182] Server: Before generating feedback, emotion recognition is performed using the emotion engine. The facial expressions of the athlete are extracted and input into the emotion engine.
[1183] Step 11:
[1184] Server: The emotion engine recognizes the emotion of the exerciser and acquires the emotion information, which includes happiness, stress, concentration, etc.
[1185] Step 12:
[1186] Server: Generates feedback by combining motion analysis results and emotional information. For example, even if the form is good, if the emotion is unstable, an encouraging comment is added.
[1187] Step 13:
[1188] Server: Sends the generated feedback list to the user's terminal.
[1189] Step 14:
[1190] Terminal: Receives feedback sent from the server and displays it to the user. It presents comments in a format that is easy for the user to understand.
[1191] Step 15:
[1192] User: Check the feedback displayed on the device, understand the problems with their exercise form and emotional advice, and restructure their training plan based on this.
[1193] Step 16:
[1194] User: Record a video with improved form, upload it to the system in the same way, and request analysis. By repeating this process, you can continuously improve your form.
[1195] Example 2
[1196] Next, a description will be given of Example 2. In the following description, the data processing device 12 will be referred to as a "server" and the headset type terminal 314 will be referred to as a "terminal."
[1197] Conventional exercise form analysis systems do not provide feedback that takes into account the athlete's emotional state, which can lead to a decline in the quality of training. A means is needed to improve mental motivation and the quality of feedback while allowing athletes to focus on improving their own technique.
[1198] The specific processing by the specific processing unit 290 of the data processing device 12 in the second embodiment is realized by the following means.
[1199] In this invention, the server includes a means for receiving video data captured by an exerciser, a preprocessing means for processing each frame included in the video data, and a means for analyzing the preprocessed frames using a motion analysis model as input, thereby making it possible to generate optimal feedback by combining the analysis results and emotion recognition results.
[1200] An "athlete" is someone who uses their body to exercise or train.
[1201] "Video data" refers to video files taken by the athlete.
[1202] "Preprocessing means" refers to the process performed to convert each frame contained in the video data into a format suitable for the AI model.
[1203] A "motion analysis model" refers to an algorithm or system that uses artificial intelligence to analyze an athlete's form.
[1204] "Analysis results" refers to the evaluations and parameters obtained about the athlete's form by the motion analysis model.
[1205] "Feedback" refers to advice or comments provided to athletes based on the analysis results.
[1206] "Providing means" refers to a method or system for transmitting the generated feedback to the exerciser.
[1207] The "face area" refers to the part of the video data in which the athlete's face is shown.
[1208] "Audio data" refers to the audio track included in the video data.
[1209] "Emotion recognition" refers to the process of identifying an athlete's emotional state from video and audio data.
[1210] "Emotion recognition result" refers to information about the emotional state of the exerciser obtained by the emotion recognition means.
[1211] "Optimal feedback" refers to the advice and comments that are deemed most beneficial to the athlete, generated by combining analysis results and emotion recognition results.
[1212] This invention is a system that allows athletes to improve their technique by analyzing their own form and receiving feedback that takes emotions into account. The system includes a means for receiving and analyzing video data taken by the athlete, and a means for recognizing the athlete's emotions using an emotion engine. This allows for further optimization of feedback to the athlete, supporting effective training.
[1213] Hardware and software used
[1214] The system's hardware uses high-performance servers, specifically those equipped with NVIDIA GPUs. Furthermore, the deep learning framework TensorFlow is used for motion analysis, and Microsoft Azure Cognitive Services is used for emotion recognition. The combination of these technologies enables highly accurate, real-time analysis and feedback.
[1215] Specific operation of the system
[1216] 1. Upload your video
[1217] Users record their own exercise form on their device and send the video data to the server via a dedicated upload page.
[1218] As a specific example, access the specified URL from the device's browser, click the file selection button to select the video data, and then click the upload button.
[1219] 2. Receiving and storing video data
[1220] The server receives the video data sent by the user as an HTTP request and temporarily stores it in storage. When the data is saved, a unique ID is assigned to it.
[1221] 3. Preparing the AI model
[1222] The server loads a pre-trained motion analysis model into memory using the TensorFlow library, which is done only once when the server starts.
[1223] 4. Loading the video and splitting it into frames
[1224] The server reads the saved video file using the OpenCV library and splits the video into frames at a rate of 30 frames per second, each of which is stored in a list.
[1225] 5. Frame Preprocessing
[1226] The server resizes the frames and normalizes pixel values to the range 0 to 1. This preprocessing prepares the images for accurate analysis by the deep learning model.
[1227] 6. Analysis of exercise form
[1228] The server inputs the preprocessed frames into a motion analysis model to extract data such as knee angles and body posture. The analysis results are stored in a list.
[1229] 7. Emotion recognition
[1230] The server extracts the athlete's facial area and voice data from the frames and sends them to Microsoft Azure Cognitive Services for emotion recognition, which identifies the athlete's emotional state.
[1231] 8. Generate feedback
[1232] The server combines the results of motion analysis and emotion recognition to generate optimal feedback for the exerciser. For example, if the exerciser's form is good, the server may comment, "Good form. Keep it up." If the exerciser's emotions are low, the server may add encouraging comments such as, "Great progress. Keep going with confidence."
[1233] 9. Providing Feedback
[1234] The server sends the generated feedback to the user, who then visually displays it on the device. The user can then review the feedback and work on restructuring their training plan or making specific improvements.
[1235] 10. Re-recording and analyzing the video
[1236] Users can then re-record the video with improved form and re-upload it into the system, repeating this process to continually improve their form and increase the effectiveness of their training.
[1237] Prompt Sentence Examples
[1238] markdown
[1239] Write Python code to split a video into frames at 30fps, resize and normalize each frame, and feed it into a deep learning model to analyze athletic form. Then, use Microsoft Azure Cognitive Services to recognize emotions and generate feedback that integrates the analysis results and emotional information.
[1240] In this way, the system of the present invention can improve the quality of training by providing feedback that takes into account not only the athlete's form but also their emotional state. By receiving precise and specific advice, athletes can achieve effective and sustained improvement in their technique.
[1241] The flow of the identification process in the second embodiment will be described with reference to FIG.
[1242] Step 1:
[1243] Recording and uploading videos
[1244] Users use their device to record their own exercise form and send the video data to the server from a dedicated upload page. They access the specified URL from their device's browser, click the file selection button to select the video data, and then click the upload button to send it.
[1245] Input: Video file of the athlete's exercise form
[1246] Output: Video data sent to the server
[1247] Step 2:
[1248] Receiving and storing video data
[1249] The server receives the video data sent by the user as an HTTP request and temporarily stores it in storage. The server assigns a unique ID to the received file and stores it in a folder structure. This saving operation is performed asynchronously, returning a prompt response to the user.
[1250] Input: Video data sent by the user
[1251] Output: Video file saved in storage
[1252] Step 3:
[1253] Preparing the AI model
[1254] The server loads a pre-trained motion analysis model into memory using the TensorFlow library, a process that is performed only once when the server starts.
[1255] Input: The trained motion analysis model file
[1256] Output: Motion analysis model loaded into memory
[1257] Step 4:
[1258] Video loading and frame splitting
[1259] The server finds the stored video file and splits the video into frames using the OpenCV library by opening the video file, capturing frames at a rate of 30 frames per second, and storing them in a list.
[1260] Input: Video file saved in storage
[1261] Output: Each frame stored in a list
[1262] Step 5:
[1263] Frame Preprocessing
[1264] The server resizes each frame and normalizes pixel values to a range between 0 and 1. This preprocessing converts the frames into a format suitable for AI models. Specifically, it resizes each frame to a uniform size and normalizes the values per pixel.
[1265] Input: Each frame stored in a list
[1266] Output: Preprocessed frames
[1267] Step 6:
[1268] Analysis of athletic form
[1269] The server inputs the preprocessed frames into the motion analysis model to extract data such as knee angle and body posture, performs the analysis using the model's predict method, and saves the results in a list.
[1270] Input: Preprocessed frame
[1271] Output: Exercise form data as analysis results (e.g. knee angle, body posture, etc.)
[1272] Step 7:
[1273] emotion recognition
[1274] The server extracts the athlete's facial area and voice data from the frames and inputs them into the emotion engine. Specifically, it sends an HTTP request to the API endpoint of Microsoft Azure Cognitive Services to recognize emotions from facial expressions and voice.
[1275] Input: Face area and audio data in the frame
[1276] Output: Emotion recognition results (e.g., joy, sadness, anger, etc.)
[1277] Step 8:
[1278] Generate feedback
[1279] The server combines the results of motion analysis and emotion recognition to generate feedback for the exerciser. Specifically, it embeds comments based on the analysis results and encouraging messages based on the exerciser's emotional state into templates.
[1280] Input: Motion analysis results, emotion recognition results
[1281] Output: The generated feedback message
[1282] Step 9:
[1283] Providing feedback
[1284] The server sends the generated feedback to the user, and the device displays it visually, either by returning a feedback message as an HTTP response or by delivering it to the user via email or push notification.
[1285] Input: The generated feedback message
[1286] Output: Feedback displayed on the terminal
[1287] Step 10:
[1288] Review feedback and incorporate it into your training plan
[1289] Users can view feedback on their device, understand areas for improvement in their form, and receive emotional advice, which they can use to restructure their training plan and make specific improvements.
[1290] Input: Feedback displayed on the device
[1291] Output: Improved training plan
[1292] Step 11:
[1293] Re-shooting and analyzing the video
[1294] Users can then re-record the video with improved form and re-upload it to the system, which will enable continuous improvement of form and increased training effectiveness.
[1295] Input: New video file taken with improved form
[1296] Output: Newly uploaded video data to the server
[1297] (Application example 2)
[1298] Next, a description will be given of Application Example 2. In the following description, the data processing device 12 will be referred to as a "server" and the headset type terminal 314 will be referred to as a "terminal."
[1299] Conventional motion analysis systems only provide feedback on the athlete's form, but are unable to provide appropriate feedback that takes into account the athlete's emotional state. This can result in training motivation and effectiveness not being maximized. Furthermore, in operational environments such as factory robots, not only is improving work accuracy and efficiency important, but also managing the operator's stress is also an important issue.
[1300] The specific processing by the specific processing unit 290 of the data processing device 12 in Application Example 2 is realized by the following means. In this invention, the server includes means for receiving video data taken by an exerciser, preprocessing means for processing each frame included in the video data, means for analyzing the preprocessed frames using a motion analysis model as input, means for generating feedback on the exerciser's form based on the analysis results, means for recognizing emotions based on the exerciser's facial expressions and voice data, means for optimizing and generating feedback taking into account the emotion recognition results, and means for providing the generated feedback to the exerciser. This makes it possible to provide feedback that takes into account not only the form of the exerciser and operator but also their overall state, including emotions.
[1301] "Video data" refers to video footage of the movements of an athlete or robot.
[1302] "Preprocessing" refers to processes such as resizing and normalization that are performed to convert frames of video data into a format suitable for the analysis model.
[1303] "Movement analysis model" refers to an artificial intelligence model used to analyze the movements of an athlete or robot and evaluate their form and accuracy.
[1304] "Feedback" refers to advice or comments provided to an exerciser or operator based on the analysis results or emotion recognition results.
[1305] "Emotion recognition" refers to analyzing the facial expressions and voices of athletes and operators contained in video data to identify their emotional state.
[1306] "Optimization" refers to adjusting feedback based on acquired data to provide more effective advice.
[1307] A "frame" refers to an individual still image that makes up video data.
[1308] The present invention relates to a system that analyzes the operation of a factory robot and supports its efficient and highly accurate movement.
[1309] The system uses video data captured by the user to analyze the accuracy and efficiency of factory robot operation and provides feedback. It also recognizes the emotional state of the user and workers and optimizes feedback as needed.
[1310] Hardware:
[1311] Camera: Used to capture video data.
[1312] GPU: Used to run AI models at high speed.
[1313] software:
[1314] OpenCV: Used to load videos and split them into frames.
[1315] PyTorch: Used to run AI models for motion analysis.
[1316] EmotionEngine: A custom library for emotion recognition.
[1317] MovementAnalyzer: A custom library for performing movement analysis.
[1318] The server performs the following process:
[1319] First, the user uses a camera to record the movements of an athlete or a factory robot. The captured video data is then sent to the server, which receives it and temporarily stores it. The stored video data is then divided into frames, resized and normalized as preprocessing, and converted into a format suitable for the AI model.
[1320] The preprocessed frames are input into a motion analysis model to evaluate the form and movement accuracy of the athlete or robot. The evaluation results are obtained as specific movement parameters. In addition, the server extracts facial expressions and voice data of the user or worker from the video data and inputs them into the Emotion Engine. The Emotion Engine analyzes this data and identifies their emotional state.
[1321] The server combines the results of motion analysis and emotion recognition to generate feedback. For example, if the motion accuracy is high, it will comment, "Your operation is accurate. Please keep going," and if the emotion is stressed, it will generate advice such as, "Your operation is good, but you seem to be stressed. Please take a short break."
[1322] The server sends the generated feedback to the user's device, which then visually presents it to the user. The user can then review the feedback, understand the areas for improvement in their own operations and movements, and receive emotional advice, which they can then incorporate into their training plan.
[1323] As a concrete example, consider a robot operator at a factory who is currently learning a new operating procedure. The operator films his or her operation with a camera and uploads it to the system. The system analyzes the video data and provides specific feedback such as, "Your vehicle handling is good, but some improvement is needed in conveyor line operation." It also recognizes fatigue from the operator's facial expression and adds advice such as, "Take a short break and refresh yourself."
[1324] Example prompt sentence:
[1325] Upload a video of a factory robot operating. Analyze the robot's movements and the operator's facial expressions to generate specific feedback to improve accuracy. Include comments based on the operator's level of fatigue and stress.
[1326] The flow of the specific processing in the application example 2 will be described with reference to FIG.
[1327] Program processing flow of the application example system
[1328] Step 1:
[1329] The server receives video data captured by the user of an athlete or a factory robot. Specifically, the user sends the video captured with a camera to the server from the upload page. The input is the captured video data, and the output is a video file stored on the server.
[1330] Step 2:
[1331] The server divides the received video data into frames. Specifically, it uses OpenCV to read the video data and extract each frame. The input is a saved video file, and the output is multiple frame images.
[1332] Step 3:
[1333] The server preprocesses each frame. Specifically, it resizes and normalizes the frame image to a format suitable for the AI model. The input is the extracted frame image, and the output is the preprocessed frame image.
[1334] Step 4:
[1335] The server inputs the preprocessed frames into a motion analysis model. Specifically, the preprocessed frame images are input into a motion analysis model running on PyTorch to obtain motion parameters. The input is the preprocessed frame images, and the output is the motion parameters.
[1336] Step 5:
[1337] The server extracts the user's facial expression and voice data from the frames. Specifically, it uses OpenCV and voice analysis libraries to extract the necessary information from the frame images and voice data. The input is the preprocessed frame images and voice data, and the output is the user's facial expression and voice data.
[1338] Step 6:
[1339] The server inputs the extracted facial and voice data into an emotion recognition engine. Specifically, it uses the EmotionEngine to identify the emotional state. The input is the user's facial and voice data, and the output is the emotion recognition result.
[1340] Step 7:
[1341] The server generates feedback based on the results of motion analysis and emotion recognition. Specifically, it combines the analysis data to generate appropriate comments and advice. The input is the motion parameters of the motion analysis and the emotion recognition results, and the output is the generated feedback.
[1342] Step 8:
[1343] The server transmits the generated feedback to the user's terminal. Specifically, the server converts the generated feedback into a format for visual display and transmits it to the user's terminal. The input is the generated feedback, and the output is the feedback displayed on the user's terminal.
[1344] Step 9:
[1345] The user checks the feedback displayed on the device and understands the areas for improvement in their own operations and behaviors, as well as emotional advice. Specifically, they restructure their training plan based on the feedback and work on specific improvements. The input is the feedback displayed on the device, and the output is the improved operations and behaviors.
[1346] summary
[1347] This step enables the system to analyze the movements of athletes and factory robots and provide appropriate feedback that takes emotions into account.
[1348] The specific processing unit 290 transmits the result of the specific processing to the headset type terminal 314. In the headset type terminal 314, the control unit 46A causes the speaker 240 and the display 343 to output the result of the specific processing. The microphone 238 acquires audio indicating a user input regarding the result of the specific processing. The control unit 46A transmits audio data indicating the user input acquired by the microphone 238 to the data processing device 12. In the data processing device 12, the specific processing unit 290 acquires the audio data.
[1349] The data generation model 58 is a so-called generative AI (Artificial Intelligence). An example of the data generation model 58 is ChatGPT (Internet Search<URL: https: / / openai.com / blog / chatgpt> ), Gemini (Internet search <url: https: gemini.google.com ?hl="ja">) and other generation AIs. The data generation model 58 is obtained by performing deep learning on a neural network. A prompt including an instruction is input to the data generation model 58, and inference data such as voice data indicating voice, text data indicating text, and image data indicating an image is also input. The data generation model 58 performs inference on the input inference data in accordance with the instruction indicated by the prompt, and outputs the inference result in a data format such as voice data and text data. Here, inference refers to, for example, analysis, classification, prediction, and / or summarization.
[1350] In the above embodiment, an example was given in which the specific processing is performed by the data processing device 12, but the technology of the present disclosure is not limited to this, and the specific processing may be performed by the headset type terminal 314.
[1351] [Fourth embodiment]
[1352] FIG. 7 shows an example of the configuration of a data processing system 410 according to the fourth embodiment.
[1353] 7, a data processing system 410 includes a data processing device 12 and a robot 414. An example of the data processing device 12 is a server.
[1354] The data processing device 12 includes a computer 22, a database 24, and a communication I / F 26. The computer 22 is an example of a "computer" according to the technology of the present disclosure. The computer 22 includes a processor 28, a RAM 30, and a storage 32. The processor 28, the RAM 30, and the storage 32 are connected to a bus 34. The database 24 and the communication I / F 26 are also connected to the bus 34. The communication I / F 26 is connected to a network 54. Examples of the network 54 include a WAN (Wide Area Network) and / or a LAN (Local Area Network).
[1355] The robot 414 includes a computer 36, a microphone 238, a speaker 240, a camera 42, a communication I / F 44, and a control target 443. The computer 36 includes a processor 46, a RAM 48, and a storage 50. The processor 46, the RAM 48, and the storage 50 are connected to a bus 52. The microphone 238, the speaker 240, the camera 42, and the control target 443 are also connected to the bus 52.
[1356] The microphone 238 receives instructions and the like from the user 20 by receiving voice uttered by the user 20. The microphone 238 captures the voice uttered by the user 20, converts the captured voice into audio data, and outputs it to the processor 46. The speaker 240 outputs audio in accordance with instructions from the processor 46.
[1357] Camera 42 is a small digital camera equipped with an optical system including a lens, aperture, and shutter, and an imaging element such as a CMOS (Complementary Metal-Oxide-Semiconductor) image sensor or a CCD (Charge Coupled Device) image sensor, and captures images of the surroundings of user 20 (for example, an imaging range defined by an angle of view equivalent to the field of vision of a typical healthy person).
[1358] The communication I / F 44 is connected to a network 54. The communication I / Fs 44 and 26 are responsible for the exchange of various information between the processor 46 and the processor 28 via the network 54. The exchange of various information between the processor 46 and the processor 28 using the communication I / Fs 44 and 26 is carried out in a secure state.
[1359] The control object 443 includes a display device, LEDs in the eyes, and motors for driving the arms, hands, and feet. The posture and gestures of the robot 414 are controlled by controlling the motors of the arms, hands, and feet. Some of the emotions of the robot 414 can be expressed by controlling these motors. In addition, the facial expressions of the robot 414 can also be expressed by controlling the light emission state of the LEDs in the eyes of the robot 414.
[1360] Fig. 8 shows an example of the main functions of the data processing device 12 and the robot 414. As shown in Fig. 8, in the data processing device 12, a specific process is performed by the processor 28. A specific process program 56 is stored in the storage 32.
[1361] The specific processing program 56 is an example of a "program" according to the technology of the present disclosure. The processor 28 reads the specific processing program 56 from the storage 32 and executes the read specific processing program 56 on the RAM 30. The specific processing is realized by the processor 28 operating as a specific processing unit 290 in accordance with the specific processing program 56 executed on the RAM 30.
[1362] The storage 32 stores a data generation model 58 and an emotion identification model 59. The data generation model 58 and the emotion identification model 59 are used by the identification processing unit 290.
[1363] In the robot 414, the processor 46 performs the reception output process. A reception output program 60 is stored in the storage 50. The processor 46 reads the reception output program 60 from the storage 50 and executes the read reception output program 60 on the RAM 48. The reception output process is realized by the processor 46 operating as the control unit 46A in accordance with the reception output program 60 executed on the RAM 48.
[1364] Next, a description will be given of the specific processing performed by the specific processing unit 290 of the data processing device 12. In the following description, the data processing device 12 will be referred to as a "server" and the robot 414 will be referred to as a "terminal."
[1365] This invention relates to a system that allows athletes to analyze their own form and receive specific advice for improving their technique. This system allows athletes to take video data using devices such as smartphones or computers, upload it to a server, and then analyzes the video data using an AI model to provide detailed feedback.
[1366] Specific Embodiments of the System
[1367] 1. Upload your video
[1368] Users upload videos of their exercise form from their devices to the server, where the uploaded videos are temporarily stored.
[1369] 2. Receiving and storing video data
[1370] The server receives the video data sent by the user and saves it under a file name such as "uploaded_video.mp4." This saved video data will then be analyzed.
[1371] 3. Preparing and Loading the AI Model
[1372] The server loads a pre-trained motion analysis model (e.g., a deep learning model) that analyzes an athlete's form and provides an assessment of specific movements.
[1373] 4. Loading the video and splitting it into frames
[1374] The server extracts the stored video file frame by frame. Since video is usually 30 frames per second, it is processed as 30 images per second.
[1375] 5. Frame Preprocessing
[1376] The server performs preprocessing such as frame resizing and normalization to convert each frame into a format suitable for the analytical model. Specifically, the frame is resized to 224x224 pixels and the pixel values are normalized to the range 0 to 1.
[1377] 6. Form Analysis
[1378] The server then feeds the preprocessed frames into a motion analysis model to evaluate form, measuring specific movement parameters such as knee angle and body tilt, and returns the results.
[1379] 7. Generate feedback
[1380] The server generates feedback for each frame based on the analysis results, and provides the user with specific advice such as "Good form. Keep going like this" or "Review the way you bend your knees."
[1381] 8. Providing Feedback
[1382] The server compiles the generated feedback and sends it back to the user's device, which then displays it to the user in a visually understandable format.
[1383] 9. Review feedback and incorporate it into your training plan
[1384] Users can view the feedback displayed on their device to understand where their form is lacking, and use that feedback to restructure their training plan and address specific areas for improvement.
[1385] 10. Re-recording and analyzing the video
[1386] Users can re-record a video with a new form and re-upload it into the system, repeating this process to continually improve their form.
[1387] As described above, the present invention allows athletes to improve the quality of their training by analyzing their own form in detail and receiving specific, individually customized feedback. This system is expected to make sports skill improvement even more efficient.
[1388] The processing flow will be explained below.
[1389] Step 1:
[1390] User: Record a video of their exercise form using a smartphone or computer. Review the video to ensure the necessary parts are properly recorded.
[1391] Step 2:
[1392] User: Access the dedicated upload page from their device, select the video file they shot on that page, and click the upload button.
[1393] Step 3:
[1394] Terminal: The selected video file is sent to the server as binary data. Once the sending is complete, a message indicating success is displayed to the user.
[1395] Step 4:
[1396] Server: Receives the binary data sent from the device. Temporarily saves the received data as "uploaded_video.mp4".
[1397] Step 5:
[1398] Server: Uses the OpenCV library to read the saved video file. Opens the video file and prepares to extract data frame by frame.
[1399] Step 6:
[1400] Server: Loads a pre-trained motion analysis model (e.g., "pose_analysis_model.h5") that analyzes an athlete's form and evaluates specific movements.
[1401] Step 7:
[1402] Server: Divide the video into frames. For example, if the video is shot at 30 frames per second, extract 30 frames per second.
[1403] Step 8:
[1404] Server: Preprocess each frame by resizing it to 224x224 pixels and normalizing pixel values to the range 0 to 1.
[1405] Step 9:
[1406] Server: Inputs the preprocessed frames into the motion analysis model, which outputs the motion estimation results for each frame.
[1407] Step 10:
[1408] Server: Generates feedback based on the analysis results. For example, if the model determines that the knee angle is inappropriate, it generates specific advice such as "Reconsider how you bend your knees."
[1409] Step 11:
[1410] Server: Compiles the generated feedback into a list and prepares it for sending back to the user.
[1411] Step 12:
[1412] Server: Sends the feedback list to the user's terminal.
[1413] Step 13:
[1414] Terminal: Receives feedback sent from the server and displays it to the user. It presents comments in a format that is easy for the user to understand.
[1415] Step 14:
[1416] User: Check the feedback displayed on the device to understand problems and areas for improvement in their exercise form, and restructure their training plan based on this.
[1417] Step 15:
[1418] User: Record a video with improved form, upload it to the system in the same way, and request analysis. By repeating this process, you can continuously improve your form.
[1419] Example 1
[1420] Next, a description will be given of Example 1. In the following description, the data processing device 12 will be referred to as a "server" and the robot 414 will be referred to as a "terminal."
[1421] Conventional exercise form analysis systems have the drawback of requiring users to film their own form, send the video data to a dedicated device, and obtain analysis results, all of which require a complex and time-consuming process. Furthermore, the accuracy of the analysis results is low, making it difficult for users to identify specific areas for improvement. Furthermore, many systems do not provide instant feedback, making it difficult to incorporate the results into effective training plans.
[1422] The specific processing by the specific processing unit 290 of the data processing device 12 in the first embodiment is realized by the following means.
[1423] In this invention, the server includes means for receiving video data taken by an exerciser, means for saving the video data, means for dividing the video data into frames, means for preprocessing each of the frames, means for analyzing the preprocessed frames using a generative AI model, means for generating feedback on the exerciser's form based on the analysis results, means for providing the generated feedback to the exerciser, and means for displaying the provided feedback on the exerciser's device. This provides high-precision feedback in real time, allowing the user to quickly identify areas for improvement in their form and implement an effective training plan.
[1424] An "athlete" refers to an individual who exercises, photographs their own exercise form, and uses the data for analysis.
[1425] "Video data" refers to video files taken by the athlete, and contains information to be analyzed.
[1426] "Server" refers to a computer system that receives, stores, processes, and provides analysis results of video data.
[1427] "Each frame" refers to the individual images that make up the video data, and the information per frame is used for analysis.
[1428] "Preprocessing" refers to data manipulation such as resizing and normalization to convert each frame into a format suitable for the analytical model.
[1429] A "generative AI model" is a pre-trained AI (artificial intelligence) model used to analyze and evaluate exercise form.
[1430] "Analysis results" refers to the evaluation data of exercise form obtained by the generative AI model, and includes specific movement parameters (such as knee angle and body tilt).
[1431] "Feedback" refers to advice and evaluations generated based on the analysis results, and is information that athletes can use to improve their form.
[1432] "Terminal" refers to the device used by the exerciser, such as a smartphone or computer, that displays the feedback.
[1433] This invention relates to a system that allows athletes to analyze their own form and receive specific advice for improving their technique. The system allows athletes to take video data using devices such as smartphones or computers, upload it to a server, and then analyzes the video data using a generative AI model to provide detailed feedback.
[1434] Hardware and software used
[1435] Hardware: smartphones, computers, servers
[1436] Software: Video recording app, video uploading tool, motion analysis AI model (deep learning model)
[1437] Data processing and calculation methods
[1438] Users can upload videos of their exercise form to a server using a smartphone or computer, for example, via a dedicated application or browser.
[1439] The server receives the video data sent by the user and temporarily stores it in a specific directory (e.g., "uploaded_videos / ") on the server with the file name "uploaded_video.mp4". This data is used in the subsequent analysis steps.
[1440] The server loads pre-trained AI models for motion analysis, which use deep learning techniques to analyze athletic form, including models trained using TensorFlow and PyTorch.
[1441] The server loads the saved video file and uses tools such as ffmpeg to extract the video frame by frame. Since video is typically 30 frames per second, this translates to 30 images per second.
[1442] Each frame is converted into a format suitable for the analytical model. In this step, the frame is resized to 224x224 pixels and the pixel values are normalized to the range 0 to 1. This process is often performed using OpenCV or PIL (Python Imaging Library).
[1443] The preprocessed frames are fed into a generative AI model to evaluate the form. The model measures movement parameters such as knee angle and body tilt and returns the analysis results. For example, the model outputs data such as "knee angle is 45 degrees" for each frame.
[1444] The server generates feedback based on the analysis results, including specific advice such as "Try to widen your knee angle by about 5 degrees." This feedback is created for each frame and is customized for each user.
[1445] The generated feedback is sent back from the server to the user's device, which then displays it to the user in a visually understandable format, for example, highlighting key points in addition to displaying them as text messages within the application.
[1446] Users can view the feedback displayed on their device, understand the issues with their form, and use that feedback to restructure their training plan and address specific areas for improvement, such as paying special attention to knee angle during their next workout.
[1447] Examples of specific examples and prompts
[1448] For example, suppose a user wants to improve their running form. The user films their running form with their smartphone and uploads the video to the system. The server analyzes the video and generates feedback to the user, such as, "The angle of your right knee is not appropriate. Try bending your right knee a bit more while running." The user can then revise their training plan based on this feedback and review their form again.
[1449] Prompt Sentence Examples
[1450] "I uploaded a video of my running form. Please help me analyze where my form needs improvement."
[1451] "Please analyze the video of my form and tell me specific areas for improvement."
[1452] The system allows exercisers to receive detailed and specific feedback to continually improve their form.
[1453] The flow of the identification process in the first embodiment will be described with reference to FIG.
[1454] Step 1:
[1455] The user takes a video of their exercise form on the device.
[1456] Input: A video of your exercise form (e.g. "run.mp4")
[1457] Output: Recorded video file
[1458] Specific movements: Using a smartphone or computer, participants will be videotaped while performing various exercises, such as running, jumping, and squatting.
[1459] Step 2:
[1460] The user uploads the video they have taken from their terminal to the server.
[1461] Input: Recorded video file
[1462] Output: Video data uploaded to the server
[1463] Specific operation: Use a dedicated application or web browser to send the video file to the server. For example, click the button, select "run.mp4", and start uploading.
[1464] Step 3:
[1465] The server receives and stores the video data sent by the user.
[1466] Input: Video data received from the user
[1467] Output: Video file saved on the server (e.g. "uploaded_video.mp4")
[1468] What happens: The server stores the uploaded data in a specific directory (e.g. "uploaded_videos / ").
[1469] Step 4:
[1470] The server loads a pre-trained generative AI model.
[1471] Input: A trained generative AI model
[1472] Output: Generative AI model loaded into memory
[1473] What it does: Loads a motion analysis model trained using TensorFlow or PyTorch, and the system is ready for analysis.
[1474] Step 5:
[1475] The server reads the saved video file and splits it into frames.
[1476] Input: Video file stored on the server ("uploaded_video.mp4")
[1477] Output: Images separated by frames (e.g. frame 1, frame 2, ...)
[1478] Specific operation: Using a video processing tool such as ffmpeg, the video is divided into 30 frames per second and each frame is obtained as an image file.
[1479] Step 6:
[1480] The server performs pre-processing on each frame.
[1481] Input: Images split into frames
[1482] Output: Preprocessed frame (resized and normalized)
[1483] Specific operation: Using OpenCV and PIL (Python Imaging Library), each frame is resized to 224x224 pixels and the value of each pixel is normalized to the range 0 to 1.
[1484] Step 7:
[1485] The server feeds the preprocessed frames into a generative AI model to analyze the form.
[1486] Input: Preprocessed frame
[1487] Output: Analysis results (motion parameters, e.g. knee angle, body tilt)
[1488] Specific movement: The preprocessed frame is input into the generative AI model, movement parameters are measured, and an evaluation result is output. For example, the result may be "Frame 10: Knee angle 45 degrees."
[1489] Step 8:
[1490] The server generates feedback based on the analysis results.
[1491] Input: Analysis results
[1492] Output: Generated feedback (e.g. "Try to open your knees by about 5 degrees.")
[1493] Specific actions: Based on the analysis results for each frame, specific advice on how to improve your athletic form is generated. For example, "Your right knee angle is not appropriate, so try bending it a little more while running."
[1494] Step 9:
[1495] The server transmits the generated feedback to the user's terminal.
[1496] Input: Generated feedback
[1497] Output: Feedback sent to the user's device
[1498] Specific operation: The server compiles the feedback data and sends it to the user's device, for example, by sending the feedback data back using an HTTP request.
[1499] Step 10:
[1500] The terminal displays the received feedback to the user.
[1501] Input: Feedback sent by the server
[1502] Output: The feedback screen presented to the user.
[1503] What it does: The device visualizes the feedback it receives and presents it to the user in a user-friendly format, for example by using text messages or highlighting to emphasize key points.
[1504] Step 11:
[1505] The user checks the feedback displayed on the device and reflects it in their training plan.
[1506] Input: Feedback displayed on the device
[1507] Output: Modified training plan
[1508] What it does: Users can use the feedback to understand where their form needs improvement and focus on those areas during their next workout.
[1509] Step 12:
[1510] The user then shoots the video again with the improved form and uploads it to the system again.
[1511] Input: Improved video
[1512] Output: New video data uploaded to the system
[1513] What it does: The user takes a photo of a new form and repeats the steps above to provide data to the system, which then continually improves the form.
[1514] (Application example 1)
[1515] Next, a description will be given of Application Example 1. In the following description, the data processing device 12 will be referred to as a "server" and the robot 414 will be referred to as a "terminal."
[1516] Conventional motion analysis systems have been used only to provide feedback to athletes and general exercisers to improve their exercise form, but the concept has not been applied to the motion analysis of exercise machines and equipment in factories. In factories, accurate analysis of motion and appropriate feedback are essential to maintain and improve the efficiency and accuracy of highly automated equipment and robots. However, current systems have difficulty meeting these needs, and have the problem of being unable to respond immediately when a decline in motion efficiency or accuracy occurs.
[1517] The specific processing by the specific processing unit 290 of the data processing device 12 in the application example 1 is realized by the following means.
[1518] In this invention, the server includes: means for receiving video data captured by an exerciser; preprocessing means for processing each frame included in the video data; means for analyzing the preprocessed frames using a motion analysis model; means for generating feedback on the exerciser's form based on the analysis results; means for providing the generated feedback to the exerciser; a device for capturing the movements of exercise equipment or devices in a factory; means for receiving and preprocessing the captured movement data; means for evaluating movements using an AI model for analyzing the preprocessed movement data; means for generating feedback on improving the efficiency and accuracy of movements based on the evaluation results; and means for providing the generated feedback to personnel involved in operating the factory equipment. This enables the movement analysis of exercise equipment or devices in a factory, and enables quick and appropriate feedback to be provided to improve efficiency and accuracy.
[1519] "Exercise participant" refers to the person or equipment whose exercise form is being analyzed and feedback is being provided.
[1520] "Video data" refers to data containing visual information that records the movements of an exerciser or the movements of exercise equipment.
[1521] "Means for receiving" refers to the technical methods and equipment for accepting video data from an external device.
[1522] "Preprocessing means" refers to a method or device that performs processing to convert video data into a format suitable for the analysis model.
[1523] "Movement analysis model" refers to an artificial intelligence model used to analyze an athlete's form and movement data.
[1524] "Means for analysis" refers to the techniques and devices used to input preprocessed data into a motion analysis model and obtain results.
[1525] "Means for generating feedback" refers to techniques and devices for generating information that indicates specific areas for improvement and direction based on the analysis results.
[1526] "Means of providing" refers to the techniques and devices used to communicate the generated feedback to the exerciser or person in charge.
[1527] "Motion equipment or devices" refers to machines, robots, and other equipment that perform specific movements in a factory.
[1528] "Motion Data" refers to data, including visual information, that records the movements of exercise equipment or devices.
[1529] "Means for evaluating" refers to techniques or devices for evaluating the efficiency and accuracy of a movement using a motion analysis model based on preprocessed movement data.
[1530] This invention is a system based on the "athlete's form analysis system" that is applied to the analysis of the movements of exercise machines or equipment in factories and the provision of feedback. This system can monitor the efficiency and accuracy of the movements of robots and machines in factories and provide necessary advice for improvement.
[1531] First, cameras are placed to capture the operation of the motion machines and equipment used in the factory. This video data is sent to a server. The server then acquires this video data using a receiving means. The hardware used in this case can be a high-performance camera.
[1532] The server then processes the received video data using preprocessing tools. Specifically, it divides the video data into frames and converts them into a format suitable for the AI model. This preprocessing uses OpenCV and PIL (Pillow) to resize and normalize the images, allowing for detailed analysis of the motion.
[1533] The preprocessed data is then input into an AI model, the motion analysis model. This model uses deep learning to analyze the motion frame by frame and measure specific motion parameters. PyTorch is used for the analysis, and the model's motion is evaluated.
[1534] The server generates feedback regarding improvements to the efficiency and accuracy of the operation based on the analysis results. The feedback generation means creates information indicating specific advice and areas for improvement based on the analysis results. For example, feedback such as "The accuracy of the operation is low. Adjustments are required" may be generated.
[1535] Finally, the server provides the generated feedback to personnel involved in the operation of the factory equipment using a means for providing the feedback, which can then be used by the personnel to adjust the robots and equipment to improve the efficiency and accuracy of their operations.
[1536] As a specific example, a robot arm in a factory is filmed as it assembles parts, and the video data is uploaded to the system. The system analyzes the video data and generates specific feedback such as, "The robot arm's movement is delayed. Please adjust the angle at which the part is attached." Based on this, the worker can adjust the robot arm's movement, improving the efficiency of the assembly work.
[1537] An example of a prompt sentence to be input to the generative AI model is as follows:
[1538] Consider a scenario where a system analyzes the behavior of a factory robot and provides feedback on its efficiency and accuracy. Specifically, imagine a robotic arm takes a video of itself assembling a part, and an AI model analyzes its behavior and provides advice on how to improve it.
[1539] The flow of the specific processing in the application example 1 will be described with reference to FIG.
[1540] Step 1:
[1541] Users use cameras to capture the operation of exercise equipment and devices in the factory and upload the video data to a server.
[1542] Input: Video data containing user-recorded actions.
[1543] Output: Video data uploaded to the server.
[1544] Specific operation: The movement is filmed using a high-performance camera, and the user uploads the captured video data to a server via a dedicated application.
[1545] Step 2:
[1546] The server obtains the uploaded video data using a receiving means and temporarily stores it.
[1547] Input: User uploaded video data.
[1548] Output: Video data temporarily stored on the server.
[1549] Specific operation: The server saves the video data in a dedicated folder called "uploaded_video.mp4".
[1550] Step 3:
[1551] The server uses a pre-processing means to divide the received video data into frames.
[1552] Input: Video data stored on the server.
[1553] Output: Image data separated by frames.
[1554] Specific operation: Using OpenCV, video data is extracted frame by frame and divided at a rate of 30 frames per second.
[1555] Step 4:
[1556] The server uses pre-processing means to convert each frame into a format suitable for image analysis.
[1557] Input: Each split frame.
[1558] Output: Image data in a format suitable for analytical models.
[1559] What it does: Uses PIL (Pillow) to resize the image (to 224x224 pixels) and normalize it (convert pixel values to the range 0 to 1).
[1560] Step 5:
[1561] The server inputs the preprocessed frames into the AI model and analyzes the behavior.
[1562] Input: Preprocessed image data.
[1563] Output: Motion analysis result data.
[1564] Specific operation: Using PyTorch, preprocessed data is input to a deep learning model and analysis results are output.
[1565] Step 6:
[1566] The server generates feedback on the efficiency and accuracy of the operation based on the analysis results.
[1567] Input: Motion analysis result data.
[1568] Output: Feedback with suggestions for improvement and advice.
[1569] Specific actions: Evaluate the analysis results and use feedback generation methods to generate specific advice such as "The accuracy of the action is low. Adjustments are required."
[1570] Step 7:
[1571] The server provides the generated feedback to personnel involved in operating the factory equipment.
[1572] Input: Feedback data.
[1573] Output: Feedback provided to the agent.
[1574] What it does: Sends feedback to the device and presents it to the agent in a visual format via the application.
[1575] Step 8:
[1576] Based on the feedback, the staff will adjust the operation of the exercise equipment and devices and take further photographs to confirm the effectiveness.
[1577] Input: Adjustment information based on feedback.
[1578] Output: New video data recording the adjusted behavior.
[1579] Specific Action: Interpret the feedback provided and change the settings of factory equipment or devices, then photograph the new action and upload it back into the system.
[1580] Furthermore, an emotion engine that estimates the user's emotion may be further combined. That is, the identification processing unit 290 may estimate the user's emotion using the emotion identification model 59, and perform identification processing using the user's emotion.
[1581] The present invention relates to a system that allows athletes to analyze their own form and receive feedback to improve their technique. The system includes a means for receiving and analyzing video data taken by the athlete, and a means for recognizing the athlete's emotions using an emotion engine. This allows for further optimization of the feedback provided to the athlete, supporting effective training.
[1582] Specific Embodiments of the System
[1583] 1. Upload your video
[1584] Users can record their own exercise form on their device, check the recorded video data, and send it to the server via a dedicated upload page.
[1585] 2. Receiving and storing video data
[1586] The server receives the video data sent by the user and temporarily stores it, which is then analyzed.
[1587] 3. Preparing the AI model
[1588] The server loads a pre-trained motion analysis model, which is used to individually analyze the athlete's form.
[1589] 4. Loading the video
[1590] The server reads the stored video file and splits it into frames, which typically consist of 30 frames per second.
[1591] 5. Frame Preprocessing
[1592] The server resizes and normalizes each frame into a format suitable for the analytical model, making the frame ready to be input into the AI model.
[1593] 6. Form Analysis
[1594] The server runs the pre-processed frames through a motion analysis model to evaluate the athlete's form, resulting in specific movement parameters such as knee angle and body posture.
[1595] 7. Emotion Recognition with Emotion Engine
[1596] The server extracts the facial expressions and voice data of the athlete from the frames and inputs them into the emotion engine, which then recognizes the athlete's emotional state from these data.
[1597] 8. Generate feedback
[1598] The server combines the results of the motion analysis with the output of the emotion engine to generate feedback for the exerciser. For example, if the exerciser's form is good, the server may comment, "Good form. Keep it up." If the exerciser's emotions are low, the server may add encouraging comments such as, "Great progress. Keep going with confidence."
[1599] 9. Providing Feedback
[1600] The server transmits the generated feedback to the user, and the terminal displays the feedback to the user in a visually easy-to-understand format.
[1601] 10. Review feedback and incorporate it into your training plan
[1602] Users can review the displayed feedback, understand areas for improvement in their form, and receive emotional advice, and use this information to restructure their training plan and make specific improvements.
[1603] 11. Re-recording and analyzing the video
[1604] Users can then re-record the video with improved form and re-upload it to the system. By repeating this process, users can expect to continually improve their form and increase the effectiveness of their training.
[1605] In this way, the system of the present invention can improve the quality of training by providing feedback that takes into account not only the athlete's form but also their emotional state. By receiving accurate and specific advice, athletes can achieve effective and sustained improvement in their skills.
[1606] The processing flow will be explained below.
[1607] Step 1:
[1608] User: Record a video of their exercise form using a smartphone or computer. Review the video to ensure the necessary parts are properly recorded.
[1609] Step 2:
[1610] User: Access the dedicated upload page, select the video file you have taken, and click the upload button to send the video data to the server.
[1611] Step 3:
[1612] Terminal: The selected video file is sent to the server as binary data. Once the sending is complete, a message indicating success is displayed to the user.
[1613] Step 4:
[1614] Server: Receives the binary data sent from the device. Temporarily saves the received data as "uploaded_video.mp4".
[1615] Step 5:
[1616] Server: Uses OpenCV library to read the saved video file. Opens the video file and splits it into frames.
[1617] Step 6:
[1618] Server: Loads a pre-trained motion analysis model (e.g., a deep learning model) that analyzes an athlete's form and provides an assessment of specific movements.
[1619] Step 7:
[1620] Server: Divide the video into frames. For example, if the video has 30 frames per second, extract 30 frames per second.
[1621] Step 8:
[1622] Server: Preprocess each frame by resizing it to 224x224 pixels and normalizing pixel values to the range 0 to 1.
[1623] Step 9:
[1624] Server: The preprocessed frames are input into the motion analysis model to obtain form evaluation results.
[1625] Step 10:
[1626] Server: Before generating feedback, emotion recognition is performed using the emotion engine. The facial expressions of the athlete are extracted and input into the emotion engine.
[1627] Step 11:
[1628] Server: The emotion engine recognizes the emotion of the exerciser and acquires the emotion information, which includes happiness, stress, concentration, etc.
[1629] Step 12:
[1630] Server: Generates feedback by combining motion analysis results and emotional information. For example, even if the form is good, if the emotion is unstable, an encouraging comment is added.
[1631] Step 13:
[1632] Server: Sends the generated feedback list to the user's terminal.
[1633] Step 14:
[1634] Terminal: Receives feedback sent from the server and displays it to the user. It presents comments in a format that is easy for the user to understand.
[1635] Step 15:
[1636] User: Check the feedback displayed on the device, understand the problems with their exercise form and emotional advice, and restructure their training plan based on this.
[1637] Step 16:
[1638] User: Record a video with improved form, upload it to the system in the same way, and request analysis. By repeating this process, you can continuously improve your form.
[1639] Example 2
[1640] Next, a description will be given of Example 2. In the following description, the data processing device 12 will be referred to as a "server" and the robot 414 will be referred to as a "terminal."
[1641] Conventional exercise form analysis systems do not provide feedback that takes into account the athlete's emotional state, which can lead to a decline in the quality of training. A means is needed to improve mental motivation and the quality of feedback while allowing athletes to focus on improving their own technique.
[1642] The specific processing by the specific processing unit 290 of the data processing device 12 in the second embodiment is realized by the following means.
[1643] In this invention, the server includes a means for receiving video data captured by an exerciser, a preprocessing means for processing each frame included in the video data, and a means for analyzing the preprocessed frames using a motion analysis model as input, thereby making it possible to generate optimal feedback by combining the analysis results and emotion recognition results.
[1644] An "athlete" is someone who uses their body to exercise or train.
[1645] "Video data" refers to video files taken by the athlete.
[1646] "Preprocessing means" refers to the process performed to convert each frame contained in the video data into a format suitable for the AI model.
[1647] A "motion analysis model" refers to an algorithm or system that uses artificial intelligence to analyze an athlete's form.
[1648] "Analysis results" refers to the evaluations and parameters obtained about the athlete's form by the motion analysis model.
[1649] "Feedback" refers to advice or comments provided to athletes based on the analysis results.
[1650] "Providing means" refers to a method or system for transmitting the generated feedback to the exerciser.
[1651] The "face area" refers to the part of the video data in which the athlete's face is shown.
[1652] "Audio data" refers to the audio track included in the video data.
[1653] "Emotion recognition" refers to the process of identifying an athlete's emotional state from video and audio data.
[1654] "Emotion recognition result" refers to information about the emotional state of the exerciser obtained by the emotion recognition means.
[1655] "Optimal feedback" refers to the advice and comments that are deemed most beneficial to the athlete, generated by combining analysis results and emotion recognition results.
[1656] This invention is a system that allows athletes to improve their technique by analyzing their own form and receiving feedback that takes emotions into account. The system includes a means for receiving and analyzing video data taken by the athlete, and a means for recognizing the athlete's emotions using an emotion engine. This allows for further optimization of feedback to the athlete, supporting effective training.
[1657] Hardware and software used
[1658] The system's hardware uses high-performance servers, specifically those equipped with NVIDIA GPUs. Furthermore, the deep learning framework TensorFlow is used for motion analysis, and Microsoft Azure Cognitive Services is used for emotion recognition. The combination of these technologies enables highly accurate, real-time analysis and feedback.
[1659] Specific operation of the system
[1660] 1. Upload your video
[1661] Users record their own exercise form on their device and send the video data to the server via a dedicated upload page.
[1662] As a specific example, access the specified URL from the device's browser, click the file selection button to select the video data, and then click the upload button.
[1663] 2. Receiving and storing video data
[1664] The server receives the video data sent by the user as an HTTP request and temporarily stores it in storage. When the data is saved, a unique ID is assigned to it.
[1665] 3. Preparing the AI model
[1666] The server loads a pre-trained motion analysis model into memory using the TensorFlow library, which is done only once when the server starts.
[1667] 4. Loading the video and splitting it into frames
[1668] The server reads the saved video file using the OpenCV library and splits the video into frames at a rate of 30 frames per second, each of which is stored in a list.
[1669] 5. Frame Preprocessing
[1670] The server resizes the frames and normalizes pixel values to the range 0 to 1. This preprocessing prepares the images for accurate analysis by the deep learning model.
[1671] 6. Analysis of exercise form
[1672] The server inputs the preprocessed frames into a motion analysis model to extract data such as knee angles and body posture. The analysis results are stored in a list.
[1673] 7. Emotion recognition
[1674] The server extracts the athlete's facial area and voice data from the frames and sends them to Microsoft Azure Cognitive Services for emotion recognition, which identifies the athlete's emotional state.
[1675] 8. Generate feedback
[1676] The server combines the results of motion analysis and emotion recognition to generate optimal feedback for the exerciser. For example, if the exerciser's form is good, the server may comment, "Good form. Keep it up." If the exerciser's emotions are low, the server may add encouraging comments such as, "Great progress. Keep going with confidence."
[1677] 9. Providing Feedback
[1678] The server sends the generated feedback to the user, who then visually displays it on the device. The user can then review the feedback and work on restructuring their training plan or making specific improvements.
[1679] 10. Re-recording and analyzing the video
[1680] Users can then re-record the video with improved form and re-upload it into the system, repeating this process to continually improve their form and increase the effectiveness of their training.
[1681] Prompt Sentence Examples
[1682] markdown
[1683] Write Python code to split a video into frames at 30fps, resize and normalize each frame, and feed it into a deep learning model to analyze athletic form. Then, use Microsoft Azure Cognitive Services to recognize emotions and generate feedback that integrates the analysis results and emotional information.
[1684] In this way, the system of the present invention can improve the quality of training by providing feedback that takes into account not only the athlete's form but also their emotional state. By receiving precise and specific advice, athletes can achieve effective and sustained improvement in their technique.
[1685] The flow of the identification process in the second embodiment will be described with reference to FIG.
[1686] Step 1:
[1687] Recording and uploading videos
[1688] Users use their device to record their own exercise form and send the video data to the server from a dedicated upload page. They access the specified URL from their device's browser, click the file selection button to select the video data, and then click the upload button to send it.
[1689] Input: Video file of the athlete's exercise form
[1690] Output: Video data sent to the server
[1691] Step 2:
[1692] Receiving and storing video data
[1693] The server receives the video data sent by the user as an HTTP request and temporarily stores it in storage. The server assigns a unique ID to the received file and stores it in a folder structure. This saving operation is performed asynchronously, returning a prompt response to the user.
[1694] Input: Video data sent by the user
[1695] Output: Video file saved in storage
[1696] Step 3:
[1697] Preparing the AI model
[1698] The server loads a pre-trained motion analysis model into memory using the TensorFlow library, a process that is performed only once when the server starts.
[1699] Input: The trained motion analysis model file
[1700] Output: Motion analysis model loaded into memory
[1701] Step 4:
[1702] Video loading and frame splitting
[1703] The server finds the stored video file and splits the video into frames using the OpenCV library by opening the video file, capturing frames at a rate of 30 frames per second, and storing them in a list.
[1704] Input: Video file saved in storage
[1705] Output: Each frame stored in a list
[1706] Step 5:
[1707] Frame Preprocessing
[1708] The server resizes each frame and normalizes pixel values to a range between 0 and 1. This preprocessing converts the frames into a format suitable for AI models. Specifically, it resizes each frame to a uniform size and normalizes the values per pixel.
[1709] Input: Each frame stored in a list
[1710] Output: Preprocessed frames
[1711] Step 6:
[1712] Analysis of athletic form
[1713] The server inputs the preprocessed frames into the motion analysis model to extract data such as knee angle and body posture, performs the analysis using the model's predict method, and saves the results in a list.
[1714] Input: Preprocessed frame
[1715] Output: Exercise form data as analysis results (e.g. knee angle, body posture, etc.)
[1716] Step 7:
[1717] emotion recognition
[1718] The server extracts the athlete's facial area and voice data from the frames and inputs them into the emotion engine. Specifically, it sends an HTTP request to the API endpoint of Microsoft Azure Cognitive Services to recognize emotions from facial expressions and voice.
[1719] Input: Face area and audio data in the frame
[1720] Output: Emotion recognition results (e.g., joy, sadness, anger, etc.)
[1721] Step 8:
[1722] Generate feedback
[1723] The server combines the results of motion analysis and emotion recognition to generate feedback for the exerciser. Specifically, it embeds comments based on the analysis results and encouraging messages based on the exerciser's emotional state into templates.
[1724] Input: Motion analysis results, emotion recognition results
[1725] Output: The generated feedback message
[1726] Step 9:
[1727] Providing feedback
[1728] The server sends the generated feedback to the user, and the device displays it visually, either by returning a feedback message as an HTTP response or by delivering it to the user via email or push notification.
[1729] Input: The generated feedback message
[1730] Output: Feedback displayed on the terminal
[1731] Step 10:
[1732] Review feedback and incorporate it into your training plan
[1733] Users can view feedback on their device, understand areas for improvement in their form, and receive emotional advice, which they can use to restructure their training plan and make specific improvements.
[1734] Input: Feedback displayed on the device
[1735] Output: Improved training plan
[1736] Step 11:
[1737] Re-shooting and analyzing the video
[1738] Users can then re-record the video with improved form and re-upload it to the system, which will enable continuous improvement of form and increased training effectiveness.
[1739] Input: New video file taken with improved form
[1740] Output: Newly uploaded video data to the server
[1741] (Application example 2)
[1742] Next, a description will be given of Application Example 2. In the following description, the data processing device 12 will be referred to as a "server" and the robot 414 will be referred to as a "terminal."
[1743] Conventional motion analysis systems only provide feedback on the athlete's form, but are unable to provide appropriate feedback that takes into account the athlete's emotional state. This can result in training motivation and effectiveness not being maximized. Furthermore, in operational environments such as factory robots, not only is improving work accuracy and efficiency important, but also managing the operator's stress is also an important issue.
[1744] The specific processing by the specific processing unit 290 of the data processing device 12 in Application Example 2 is realized by the following means. In this invention, the server includes means for receiving video data taken by an exerciser, preprocessing means for processing each frame included in the video data, means for analyzing the preprocessed frames using a motion analysis model as input, means for generating feedback on the exerciser's form based on the analysis results, means for recognizing emotions based on the exerciser's facial expressions and voice data, means for optimizing and generating feedback taking into account the emotion recognition results, and means for providing the generated feedback to the exerciser. This makes it possible to provide feedback that takes into account not only the form of the exerciser and operator but also their overall state, including emotions.
[1745] "Video data" refers to video footage of the movements of an athlete or robot.
[1746] "Preprocessing" refers to processes such as resizing and normalization that are performed to convert frames of video data into a format suitable for the analysis model.
[1747] "Movement analysis model" refers to an artificial intelligence model used to analyze the movements of an athlete or robot and evaluate their form and accuracy.
[1748] "Feedback" refers to advice or comments provided to an exerciser or operator based on the analysis results or emotion recognition results.
[1749] "Emotion recognition" refers to analyzing the facial expressions and voices of athletes and operators contained in video data to identify their emotional state.
[1750] "Optimization" refers to adjusting feedback based on acquired data to provide more effective advice.
[1751] A "frame" refers to an individual still image that makes up video data.
[1752] The present invention relates to a system that analyzes the operation of a factory robot and supports its efficient and highly accurate movement.
[1753] The system uses video data captured by the user to analyze the accuracy and efficiency of factory robot operation and provides feedback. It also recognizes the emotional state of the user and workers and optimizes feedback as needed.
[1754] Hardware:
[1755] Camera: Used to capture video data.
[1756] GPU: Used to run AI models at high speed.
[1757] software:
[1758] OpenCV: Used to load videos and split them into frames.
[1759] PyTorch: Used to run AI models for motion analysis.
[1760] EmotionEngine: A custom library for emotion recognition.
[1761] MovementAnalyzer: A custom library for performing movement analysis.
[1762] The server performs the following process:
[1763] First, the user uses a camera to record the movements of an athlete or a factory robot. The captured video data is then sent to the server, which receives it and temporarily stores it. The stored video data is then divided into frames, resized and normalized as preprocessing, and converted into a format suitable for the AI model.
[1764] The preprocessed frames are input into a motion analysis model to evaluate the form and movement accuracy of the athlete or robot. The evaluation results are obtained as specific movement parameters. In addition, the server extracts facial expressions and voice data of the user or worker from the video data and inputs them into the Emotion Engine. The Emotion Engine analyzes this data and identifies their emotional state.
[1765] The server combines the results of motion analysis and emotion recognition to generate feedback. For example, if the motion accuracy is high, it will comment, "Your operation is accurate. Please keep going," and if the emotion is stressed, it will generate advice such as, "Your operation is good, but you seem to be stressed. Please take a short break."
[1766] The server sends the generated feedback to the user's device, which then visually presents it to the user. The user can then review the feedback, understand the areas for improvement in their own operations and movements, and receive emotional advice, which they can then incorporate into their training plan.
[1767] As a concrete example, consider a robot operator at a factory who is currently learning a new operating procedure. The operator films his or her operation with a camera and uploads it to the system. The system analyzes the video data and provides specific feedback such as, "Your vehicle handling is good, but some improvement is needed in conveyor line operation." It also recognizes fatigue from the operator's facial expression and adds advice such as, "Take a short break and refresh yourself."
[1768] Example prompt sentence:
[1769] Upload a video of a factory robot operating. Analyze the robot's movements and the operator's facial expressions to generate specific feedback to improve accuracy. Include comments based on the operator's level of fatigue and stress.
[1770] The flow of the specific processing in the application example 2 will be described with reference to FIG.
[1771] Program processing flow of the application example system
[1772] Step 1:
[1773] The server receives video data captured by the user of an athlete or a factory robot. Specifically, the user sends the video captured with a camera to the server from the upload page. The input is the captured video data, and the output is a video file stored on the server.
[1774] Step 2:
[1775] The server divides the received video data into frames. Specifically, it uses OpenCV to read the video data and extract each frame. The input is a saved video file, and the output is multiple frame images.
[1776] Step 3:
[1777] The server preprocesses each frame. Specifically, it resizes and normalizes the frame image to a format suitable for the AI model. The input is the extracted frame image, and the output is the preprocessed frame image.
[1778] Step 4:
[1779] The server inputs the preprocessed frames into a motion analysis model. Specifically, the preprocessed frame images are input into a motion analysis model running on PyTorch to obtain motion parameters. The input is the preprocessed frame images, and the output is the motion parameters.
[1780] Step 5:
[1781] The server extracts the user's facial expression and voice data from the frames. Specifically, it uses OpenCV and voice analysis libraries to extract the necessary information from the frame images and voice data. The input is the preprocessed frame images and voice data, and the output is the user's facial expression and voice data.
[1782] Step 6:
[1783] The server inputs the extracted facial and voice data into an emotion recognition engine. Specifically, it uses the EmotionEngine to identify the emotional state. The input is the user's facial and voice data, and the output is the emotion recognition result.
[1784] Step 7:
[1785] The server generates feedback based on the results of motion analysis and emotion recognition. Specifically, it combines the analysis data to generate appropriate comments and advice. The input is the motion parameters of the motion analysis and the emotion recognition results, and the output is the generated feedback.
[1786] Step 8:
[1787] The server transmits the generated feedback to the user's terminal. Specifically, the server converts the generated feedback into a format for visual display and transmits it to the user's terminal. The input is the generated feedback, and the output is the feedback displayed on the user's terminal.
[1788] Step 9:
[1789] The user checks the feedback displayed on the device and understands the areas for improvement in their own operations and behaviors, as well as emotional advice. Specifically, they restructure their training plan based on the feedback and work on specific improvements. The input is the feedback displayed on the device, and the output is the improved operations and behaviors.
[1790] summary
[1791] This step enables the system to analyze the movements of athletes and factory robots and provide appropriate feedback that takes emotions into account.
[1792] The specific processing unit 290 transmits the result of the specific processing to the robot 414. In the robot 414, the control unit 46A causes the speaker 240 and the control target 443 to output the result of the specific processing. The microphone 238 acquires voice indicating a user input regarding the result of the specific processing. The control unit 46A transmits voice data indicating the user input acquired by the microphone 238 to the data processing device 12. In the data processing device 12, the specific processing unit 290 acquires the voice data.
[1793] The data generation model 58 is a so-called generative AI (Artificial Intelligence). An example of the data generation model 58 is ChatGPT (Internet Search<URL: https: / / openai.com / blog / chatgpt> ), Gemini (Internet search <url: https: gemini.google.com ?hl="ja">) and other generation AIs. The data generation model 58 is obtained by performing deep learning on a neural network. A prompt including an instruction is input to the data generation model 58, and inference data such as voice data indicating voice, text data indicating text, and image data indicating an image is also input. The data generation model 58 performs inference on the input inference data in accordance with the instruction indicated by the prompt, and outputs the inference result in a data format such as voice data and text data. Here, inference refers to, for example, analysis, classification, prediction, and / or summarization.
[1794] In the above embodiment, an example was given in which the specific processing is performed by the data processing device 12, but the technology of the present disclosure is not limited to this, and the specific processing may be performed by the robot 414.
[1795] The emotion identification model 59 as an emotion engine may determine the user's emotion according to a specific mapping. Specifically, the emotion identification model 59 may determine the user's emotion according to an emotion map (see FIG. 9), which is a specific mapping. Similarly, the emotion identification model 59 may determine the robot's emotion, and the identification processing unit 290 may perform identification processing using the robot's emotion.
[1796] FIG. 9 is a diagram illustrating an emotion map 400 on which multiple emotions are mapped. In the emotion map 400, emotions are arranged in concentric circles radiating from the center. Emotions closer to the center of the concentric circles are more primitive. Emotions representing states and actions arising from a state of mind are arranged on the outer edges of the concentric circles. The concept of emotion includes both affect and mental states. Emotions generally generated from reactions occurring in the brain are arranged on the left side of the concentric circles. Emotions generally induced by situational judgment are arranged on the right side of the concentric circles. Emotions generally generated from reactions occurring in the brain and induced by situational judgment are arranged on the upper and lower sides of the concentric circles. Furthermore, the emotion of "pleasure" is arranged on the upper side of the concentric circles, and the emotion of "discomfort" is arranged on the lower side. In this way, in the emotion map 400, multiple emotions are mapped based on the structure by which emotions are generated, and emotions that tend to occur simultaneously are mapped close to each other.
[1797] These emotions are distributed in the 3 o'clock direction on emotion map 400, and typically fluctuate between relief and anxiety. In the right half of emotion map 400, situational awareness dominates over internal sensations, resulting in a sense of calm.
[1798] The inside of emotion map 400 represents what is going on in the mind, and the outside of emotion map 400 represents behavior, so the further you go outside emotion map 400, the more visible the emotions become (the more they are expressed in behavior).
[1799] Human emotions are based on various balances, such as posture and blood sugar levels. When these balances deviate from the ideal, a state of discomfort is indicated, and when they approach the ideal, a state of pleasure is indicated. Emotions can also be created for robots, automobiles, and motorcycles, based on various balances, such as posture and remaining battery life. When these balances deviate from the ideal, a state of discomfort is indicated, and when they approach the ideal, a state of pleasure is indicated. An emotion map can be generated, for example, based on Dr. Mitsuyoshi's emotion map (Research on Voice Emotion Recognition and Emotional Brain Physiological Signal Analysis Systems, Tokushima University, Doctoral Dissertation: https: / / ci.nii.ac.jp / naid / 500000375379). The left half of the emotion map lists emotions belonging to the "reaction" domain, where sensation is dominant. The right half of the emotion map lists emotions belonging to the "situation" domain, where situational awareness is dominant.
[1800] The emotion map defines two emotions that promote learning. One is a negative emotion on the situation side, around the middle of "repentance" or "reflection." In other words, this occurs when the robot experiences negative emotions such as "I never want to feel this way again" or "I don't want to be scolded again." The other is a positive emotion on the response side, around "desire." In other words, this occurs when the robot experiences positive feelings such as "I want more" or "I want to know more."
[1801] The emotion identification model 59 inputs user input into a pre-trained neural network, obtains emotion values indicating each emotion shown in the emotion map 400, and determines the user's emotion. This neural network is pre-trained based on multiple pieces of training data that are combinations of user input and emotion values indicating each emotion shown in the emotion map 400. Furthermore, this neural network is trained so that emotions that are located close to each other have similar values, as in the emotion map 900 shown in FIG. 10. FIG. 10 shows an example in which multiple emotions, "relieved," "calm," and "reassuring," have similar emotion values.
[1802] The system according to the present disclosure has been described above mainly with respect to the functions of the data processing device 12, but the system according to the present disclosure is not necessarily implemented on a server. The system according to the present disclosure may be implemented as a general information processing system. The present disclosure may be implemented, for example, as a software program running on a personal computer or an application running on a smartphone, etc. The method according to the present disclosure may be provided to users in the form of SaaS (Software as a Service).
[1803] In the above embodiment, an example was given in which the specific processing is performed by one computer 22, but the technology of the present disclosure is not limited to this, and the specific processing may be distributed and performed by a plurality of computers including the computer 22. For example, the data generation model 58 may be provided in an external device of the data processing device 12, and data may be generated in the external device in accordance with input data.
[1804] In the above embodiment, an example in which the specific processing program 56 is stored in the storage 32 has been described, but the technology of the present disclosure is not limited to this. For example, the specific processing program 56 may be stored in a portable, computer-readable, non-transitory storage medium such as a USB (Universal Serial Bus) memory. The specific processing program 56 stored in the non-transitory storage medium is installed in the computer 22 of the data processing device 12. The processor 28 executes the specific processing in accordance with the specific processing program 56.
[1805] Alternatively, the specific processing program 56 may be stored in a storage device such as a server connected to the data processing device 12 via the network 54, and the specific processing program 56 may be downloaded and installed on the computer 22 in response to a request from the data processing device 12.
[1806] It is not necessary to store all of the specific processing program 56 in a storage device such as a server connected to the data processing device 12 via the network 54, or to store all of the specific processing program 56 in the storage 32; only a portion of the specific processing program 56 may be stored.
[1807] The hardware resource for executing a specific process can be any of the following processors: An example of a processor is a CPU, which is a general-purpose processor that functions as a hardware resource for executing a specific process by executing software, i.e., a program. Another example of a processor is a dedicated electrical circuit, such as an FPGA (Field-Programmable Gate Array), a PLD (Programmable Logic Device), or an ASIC (Application Specific Integrated Circuit), which is a processor with a circuit configuration designed specifically for executing a specific process. Each processor has built-in or connected memory, and each processor uses the memory to execute the specific process.
[1808] The hardware resource that executes the specific processing may be configured with one of these various processors, or may be configured with a combination of two or more processors of the same or different types (for example, a combination of multiple FPGAs, or a combination of a CPU and an FPGA). Also, the hardware resource that executes the specific processing may be a single processor.
[1809] As an example of a system configured with a single processor, first, one processor is configured by combining one or more CPUs and software, and this processor functions as a hardware resource that executes a specific process. Second, there is a system that uses a processor that realizes the functions of an entire system including multiple hardware resources that execute a specific process on a single IC chip, as typified by SoC (System-on-a-chip). In this way, a specific process is realized using one or more of the above-mentioned various processors as hardware resources.
[1810] Furthermore, the hardware structure of these various processors can be, more specifically, an electric circuit that combines circuit elements such as semiconductor devices. The specific processing described above is merely an example. Therefore, it goes without saying that unnecessary steps may be deleted, new steps may be added, or the processing order may be rearranged, without departing from the spirit of the invention.
[1811] The above-described description and illustrations are a detailed explanation of the parts related to the technology of the present disclosure and are merely an example of the technology of the present disclosure. For example, the above description of the configuration, functions, actions, and effects is an explanation of an example of the configuration, functions, actions, and effects of the parts related to the technology of the present disclosure. Therefore, it goes without saying that unnecessary parts may be deleted, new elements may be added, or replacements may be made to the above-described description and illustrations within the scope of the gist of the technology of the present disclosure. Furthermore, to avoid confusion and facilitate understanding of the parts related to the technology of the present disclosure, the above-described description and illustrations omit explanations of common technical knowledge that do not require particular explanation to enable the implementation of the technology of the present disclosure.
[1812] All publications, patent applications, and technical standards mentioned in this specification are herein incorporated by reference to the same extent as if each individual publication, patent application, or technical standard was specifically and individually indicated to be incorporated by reference.
[1813] The following is further disclosed regarding the above embodiment.
[1814] (Claim 1)
[1815] A means for receiving video data captured by an exerciser;
[1816] a preprocessing means for processing each frame included in the video data;
[1817] means for analyzing the preprocessed frames using a motion analysis model as an input;
[1818] means for generating feedback regarding the athlete's form based on the analysis results;
[1819] means for providing the generated feedback to an exerciser;
[1820] A system including:
[1821] (Claim 2)
[1822] 10. The system of claim 1, further comprising means for storing video data captured by the exerciser.
[1823] (Claim 3)
[1824] 2. The system of claim 1, wherein the motion analysis model is an artificial intelligence model that performs analysis of an athlete's form.
[1825] "Example 1"
[1826] (Claim 1)
[1827] A means for receiving video data captured by an exerciser;
[1828] means for storing the video data;
[1829] means for dividing the video data into frames;
[1830] means for performing pre-processing on each of the frames;
[1831] means for analyzing the preprocessed frames using a generative AI model as input;
[1832] means for generating feedback regarding the athlete's form based on the analysis results;
[1833] means for providing the generated feedback to an exerciser;
[1834] means for displaying the provided feedback on an exerciser's terminal;
[1835] A system including:
[1836] (Claim 2)
[1837] 2. The system of claim 1, wherein the motion analysis model is an artificial intelligence model that performs analysis of an athlete's form.
[1838] (Claim 3)
[1839] The system of claim 1 , further comprising means for performing pre-processing of resizing and normalization on each of the frames.
[1840] "Application Example 1"
[1841] (Claim 1)
[1842] A means for receiving video data captured by an exerciser;
[1843] a preprocessing means for processing each frame included in the video data;
[1844] means for analyzing the preprocessed frames using a motion analysis model as an input;
[1845] means for generating feedback regarding the athlete's form based on the analysis results;
[1846] means for providing the generated feedback to an exerciser;
[1847] a device for photographing the operation of exercise equipment or devices in a factory;
[1848] means for receiving and preprocessing the captured motion data;
[1849] means for evaluating a motion using an AI model for analyzing the preprocessed motion data;
[1850] means for generating feedback regarding improvements in efficiency and accuracy of the operation based on the evaluation results;
[1851] means for providing said generated feedback to personnel involved in the operation of factory equipment;
[1852] A system including:
[1853] (Claim 2)
[1854] 10. The system of claim 1, further comprising means for storing video data captured by the exerciser.
[1855] (Claim 3)
[1856] 2. The system of claim 1, wherein the motion analysis model is an artificial intelligence model that performs analysis of an athlete's form.
[1857] "Example 2: Combining Emotion Engines"
[1858] (Claim 1)
[1859] A means for receiving video data captured by an exerciser;
[1860] a preprocessing means for processing each frame included in the video data;
[1861] means for analyzing the preprocessed frames using a motion analysis model as an input;
[1862] means for generating feedback regarding the athlete's form based on the analysis results;
[1863] means for providing the generated feedback to an exerciser;
[1864] a means for extracting facial areas and voice data of the athlete from the video data and recognizing emotions;
[1865] means for generating optimal feedback by combining the analysis result and the emotion recognition result;
[1866] A system including:
[1867] (Claim 2)
[1868] 10. The system of claim 1, further comprising means for storing video data captured by the exerciser.
[1869] (Claim 3)
[1870] 2. The system of claim 1, wherein the motion analysis model is an artificial intelligence model that performs analysis of an athlete's form.
[1871] "Application example 2 when combining emotion engines"
[1872] (Claim 1)
[1873] A means for receiving video data captured by an exerciser;
[1874] a preprocessing means for processing each frame included in the video data;
[1875] means for analyzing the preprocessed frames using a motion analysis model as an input;
[1876] means for generating feedback regarding the athlete's form based on the analysis results;
[1877] a means for recognizing emotions based on facial expressions and voice data of the exerciser;
[1878] means for optimizing and generating feedback in consideration of the emotion recognition results;
[1879] means for providing the generated feedback to an exerciser;
[1880] A system including:
[1881] (Claim 2)
[1882] 10. The system of claim 1, further comprising means for storing video data captured by the exerciser.
[1883] (Claim 3)
[1884] 2. The system of claim 1, wherein the motion analysis model is an artificial intelligence model that performs analysis of an athlete's form and emotions. [Explanation of symbols]
[1885] 10, 210, 310, 410 Data Processing Systems 12 Data Processing Device 14 Smart Devices 214 Smart Glasses 314 Headset-type terminal 414 Robot< / url:> < / url:> < / url:> < / url:>
Claims
1. A means for receiving video data captured by an exerciser; a preprocessing means for processing each frame included in the video data; means for analyzing the preprocessed frames using a motion analysis model as an input; means for generating feedback regarding the athlete's form based on the analysis results; means for providing the generated feedback to an exerciser; A system including:
2. The system of claim 1 , further comprising means for storing video data captured by the exerciser.
3. 2. The system according to claim 1, wherein the motion analysis model is an artificial intelligence model that performs analysis of an athlete's form.
Citation Information
Patent Citations
Persona chatbot control method and system
JP2022180282A