System
The system addresses the inefficiency in conventional sports training by using a device-server-terminal setup for real-time analysis and feedback, enabling athletes to intuitively understand and correct their movements for improved performance.
Patent Information
- Application Number
- JP2024116575
- Authority / Receiving Office
- JP · JP
- Patent Type
- Applications
- Current Assignee / Owner
- Filing Date
- 2024-07-19
- Publication Date
- 2026-01-29
AI Technical Summary
Conventional sports training systems struggle to efficiently identify subtle differences in an athlete's movement form and provide appropriate correction instructions, relying heavily on the trainer's skill and lacking intuitive and efficient methods for real-time feedback.
A system comprising a device for filming athletic performance, a server for analyzing video data, a comparison mechanism to identify discrepancies with optimal form data, and a terminal for displaying corrective instructions in an intuitive format, utilizing AI algorithms for real-time feedback.
Enables athletes to quickly and efficiently correct their form in real-time, improving overall performance by providing specific and easily understandable corrective instructions.
Smart Images

Figure 2026015101000001_ABST
Abstract
Description
[Technical Field]
[0001] The technology of the present disclosure relates to a system. [Background technology]
[0002] Patent document 1 discloses a persona chatbot control method performed by at least one processor, the method including the steps of receiving a user utterance, adding the user utterance to a prompt including an instruction sentence related to a description of the chatbot character, encoding the prompt, and inputting the encoded prompt into a language model to generate a chatbot utterance in response to the user utterance. [Prior art documents] [Patent documents]
[0003] [Patent Document 1] Japanese Patent Publication No. 2022-180282 Summary of the Invention [Problem to be solved by the invention]
[0004] In conventional sports training, it is extremely difficult to identify subtle differences in an athlete's movement form and provide appropriate correction instructions. As a result, improving athlete performance is inefficient and largely relies on the skill of the trainer. There is a need for a system that can solve these issues and efficiently improve athlete performance. [Means for solving the problem]
[0005] The present invention solves the above problems by providing a system including a device for filming an athlete's performance, a server that receives the filmed video, analysis means that analyzes the video data received by the server to extract the athlete's movement data, comparison means that compares the extracted movement data with optimal form data, instruction generation means that generates specific correction instructions based on the obtained differences, transmission means that transmits the generated correction instructions to a terminal, and display means that displays the received correction instructions to a user on the terminal.
[0006] "Athlete" refers to an individual who participates in a sport or athletic activity and seeks to improve their skills or performance.
[0007] "Device" refers to an electronic device or piece of equipment designed to perform a specific function, and in this case refers to a camera or smartphone used to film an athlete's performance.
[0008] A "server" is a computer system that provides data and services in response to requests from clients via a network, and in this case, its role is to receive, store, and analyze uploaded video data.
[0009] "Video Data" refers to data saved in the form of a video file that records an athlete's performance.
[0010] "Analysis means" refers to algorithms and software used to analyze video data and extract player movement data.
[0011] "Motion data" refers to data that shows the movement and position of each part of a player's body over time.
[0012] "Comparison means" refers to algorithms or software that compare extracted motion data with optimal form data and identify differences.
[0013] "Form data" refers to data that represents the ideal body position and movement for a specific movement.
[0014] "Instruction generation means" refers to algorithms or software for generating specific correction instructions based on the differences between the motion data and the form data.
[0015] "Corrective Instructions" refers to specific instructions to modify the movement or position of a specific body part in order to improve an athlete's performance.
[0016] "Transmission means" refers to a network connection or communication protocol for transmitting the generated correction instructions to the terminal.
[0017] "Terminal" refers to a device used by a user, such as a computer, smartphone, or tablet.
[0018] "Display means" refers to a display or software for visually displaying to a user the corrective instructions received on the terminal. [Brief explanation of the drawings]
[0019] [Figure 1] 1 is a conceptual diagram showing an example of the configuration of a data processing system according to a first embodiment. [Figure 2] 1 is a conceptual diagram showing an example of main functions of a data processing device and a smart device according to a first embodiment. [Figure 3] FIG. 10 is a conceptual diagram showing an example of the configuration of a data processing system according to a second embodiment. [Figure 4] FIG. 10 is a conceptual diagram showing an example of main functions of a data processing device and smart glasses according to a second embodiment. [Figure 5] FIG. 10 is a conceptual diagram showing an example of the configuration of a data processing system according to a third embodiment. [Figure 6] FIG. 11 is a conceptual diagram showing an example of main functions of a data processing device and a headset-type terminal according to a third embodiment. [Figure 7]FIG. 10 is a conceptual diagram showing an example of the configuration of a data processing system according to a fourth embodiment. [Figure 8] FIG. 10 is a conceptual diagram showing an example of main functions of a data processing device and a robot according to a fourth embodiment. [Figure 9] 1 shows an emotion map onto which multiple emotions are mapped. [Figure 10] 1 shows an emotion map onto which multiple emotions are mapped. [Figure 11] FIG. 3 is a sequence diagram showing a processing flow of the data processing system according to the first embodiment. [Figure 12] FIG. 10 is a sequence diagram showing the flow of processing in the data processing system in Application Example 1. [Figure 13] FIG. 10 is a sequence diagram showing the flow of processing in the data processing system according to the second embodiment when an emotion engine is combined. [Figure 14] FIG. 10 is a sequence diagram showing the flow of processing in the data processing system in Application Example 2 when an emotion engine is combined. DETAILED DESCRIPTION OF THE INVENTION
[0020] An example of an embodiment of a system according to the technology of the present disclosure will be described below with reference to the accompanying drawings.
[0021] First, the terms used in the following description will be explained.
[0022] In the following embodiments, a coded processor (hereinafter simply referred to as a "processor") may be a single arithmetic device or a combination of multiple arithmetic devices. Furthermore, a processor may be a single type of arithmetic device or a combination of multiple types of arithmetic devices. Examples of arithmetic devices include a CPU (Central Processing Unit), a GPU (Graphics Processing Unit), a GPGPU (General-Purpose computing on Graphics Processing Units), and an APU (Accelerated Processing Unit).
[0023] In the following embodiments, a coded RAM (Random Access Memory) is a memory in which information is temporarily stored and is used as a working memory by a processor.
[0024] In the following embodiments, the coded storage is one or more non-volatile storage devices that store various programs, various parameters, etc. Examples of non-volatile storage devices include flash memory (SSD (Solid State Drive)), magnetic disks (e.g., hard disks), and magnetic tapes.
[0025] In the following embodiments, a communication I / F (Interface) with a symbol is an interface including a communication processor, an antenna, etc. The communication I / F controls communication between multiple computers. Examples of communication standards applied to the communication I / F include wireless communication standards including 5G (5th Generation Mobile Communication System), Wi-Fi (registered trademark), Bluetooth (registered trademark), etc.
[0026] In the following embodiments, "A and / or B" is synonymous with "at least one of A and B." In other words, "A and / or B" means that it may be only A, only B, or a combination of A and B. Furthermore, in this specification, the same concept as "A and / or B" is also applied when three or more things are expressed connected by "and / or."
[0027] [First embodiment]
[0028] FIG. 1 shows an example of the configuration of a data processing system 10 according to the first embodiment.
[0029] 1, a data processing system 10 includes a data processing device 12 and a smart device 14. An example of the data processing device 12 is a server.
[0030] The data processing device 12 includes a computer 22, a database 24, and a communication I / F 26. The computer 22 is an example of a "computer" according to the technology of the present disclosure. The computer 22 includes a processor 28, a RAM 30, and a storage 32. The processor 28, the RAM 30, and the storage 32 are connected to a bus 34. The database 24 and the communication I / F 26 are also connected to the bus 34. The communication I / F 26 is connected to a network 54. Examples of the network 54 include a WAN (Wide Area Network) and / or a LAN (Local Area Network).
[0031] The smart device 14 includes a computer 36, a reception device 38, an output device 40, a camera 42, and a communication I / F 44. The computer 36 includes a processor 46, a RAM 48, and a storage 50. The processor 46, the RAM 48, and the storage 50 are connected to a bus 52. The reception device 38, the output device 40, and the camera 42 are also connected to the bus 52.
[0032] The reception device 38 includes a touch panel 38A, a microphone 38B, and the like, and receives user input. The touch panel 38A detects contact with an indicator (for example, a pen or a finger) to receive user input by the touch of the indicator. The microphone 38B detects the user's voice to receive user input by voice. The control unit 46A transmits data indicating the user input received by the touch panel 38A and the microphone 38B to the data processing device 12. In the data processing device 12, the specific processing unit 290 acquires the data indicating the user input.
[0033] The output device 40 includes a display 40A and a speaker 40B, and presents data to the user 20 by outputting the data in a form of expression that the user 20 can perceive (for example, audio and / or text). The display 40A displays visible information such as text and images in accordance with instructions from the processor 46. The speaker 40B outputs audio in accordance with instructions from the processor 46. The camera 42 is a compact digital camera equipped with an optical system including a lens, aperture, and shutter, and an imaging element such as a CMOS (Complementary Metal-Oxide-Semiconductor) image sensor or a CCD (Charge Coupled Device) image sensor.
[0034] The communication I / F 44 is connected to a network 54. The communication I / Fs 44 and 26 control the exchange of various information between the processor 46 and the processor 28 via the network 54.
[0035] FIG. 2 shows an example of the main functions of the data processing device 12 and the smart device 14.
[0036] 2, in the data processing device 12, a specific process is performed by the processor 28. A specific processing program 56 is stored in the storage 32. The specific processing program 56 is an example of a "program" according to the technology of the present disclosure. The processor 28 reads the specific processing program 56 from the storage 32 and executes the read specific processing program 56 on the RAM 30. The specific process is realized by the processor 28 operating as a specific processing unit 290 in accordance with the specific processing program 56 executed on the RAM 30.
[0037] The storage 32 stores a data generation model 58 and an emotion identification model 59. The data generation model 58 and the emotion identification model 59 are used by the identification processing unit 290.
[0038] In the smart device 14, the processor 46 performs the reception output process. The storage 50 stores a reception output program 60. The reception output program 60 is used in conjunction with the specific processing program 56 by the data processing system 10. The processor 46 reads the reception output program 60 from the storage 50 and executes the read reception output program 60 on the RAM 48. The reception output process is realized by the processor 46 operating as the control unit 46A in accordance with the reception output program 60 executed on the RAM 48.
[0039] Next, a description will be given of the specific processing performed by the specific processing unit 290 of the data processing device 12. In the following description, the data processing device 12 will be referred to as a "server" and the smart device 14 will be referred to as a "terminal."
[0040] The present invention relates to an AI sports trainer system for efficiently improving athletes' performance. The system includes a device for recording athletes' performance, a server for receiving and analyzing the video data, a means for generating and transmitting corrective instructions, and a terminal for displaying the instructions to the user.
[0041] System Overview
[0042] 1. Use of photography devices
[0043] Users can use their smartphones or dedicated cameras to film their pitching, batting, and other performances.
[0044] The video data is shot in high resolution, clearly recording the players' detailed movements.
[0045] 2. Uploading data
[0046] Users upload the videos they have taken to the server through a dedicated app. The app's intuitive operation makes it easy to select and upload videos.
[0047] 3. Data Receipt and Analysis
[0048] The server receives the video data and begins analysis. It uses powerful AI algorithms to model the player's movements in the video and extract each key element (e.g., shoulder position, elbow angle, etc.).
[0049] In this case, a deep learning model operates as an analytical tool to automatically generate motion data.
[0050] 4. Comparison with optimal form
[0051] The server compares the extracted motion data with optimal form data stored in a database, which is based on model athletes and their best past performances.
[0052] The comparison means relatively matches the motion data and the optimal form data for each frame and identifies specific differences.
[0053] 5. Generate corrective instructions
[0054] Based on the analysis of the discrepancies, the server generates specific corrective instructions, detailing what the player needs to correct and how.
[0055] For example, specific instructions for correcting the movement such as "pull your shoulders back a little more" are generated.
[0056] 6. Sending Corrective Instructions
[0057] The server generates and transmits corrective instructions to the terminal, which may include instructions in textual and visual formats.
[0058] In addition to text instructions, animations and illustrations may also be sent.
[0059] 7. Display of orthodontic instructions
[0060] The device receives correction instructions from the server and displays them to the user. The app has an easy-to-use interface and is designed to allow users to intuitively understand the instructions.
[0061] This allows users to learn how to correct their form in real time.
[0062] Specific examples
[0063] 1. Users film and upload pitching videos
[0064] Users record their pitching on their smartphones and upload the video to the server using a dedicated app.
[0065] 2. The server receives and analyzes the video
[0066] The server receives the video and uses a deep learning model to analyze and extract key points in the pitching motion.
[0067] 3. Comparison with optimal form
[0068] The server compares the analyzed movement data with optimal form data and detects differences such as "elbow position is too low" or "timing is too early."
[0069] 4. Generate corrective instructions
[0070] Based on the comparison results, the server generates specific corrective instructions, such as "keep your elbow 5 degrees higher on your next pitch."
[0071] 5. Sending and Displaying Instructions
[0072] The server generates and sends the correction instructions to the device, which then displays them to the user, allowing the user to receive effective feedback and instantly correct their form.
[0073] Such a system will enable players to efficiently correct their form in real time, which is expected to improve their overall performance.
[0074] The processing flow will be explained below.
[0075] Step 1:
[0076] Users use a smartphone or dedicated camera to film their own sports performance (pitching, batting, etc.) After finishing filming, they launch the app and select the video.
[0077] Step 2:
[0078] The user uploads the video they took using the app to the server. The app sends the user ID and the video data together, and confirms that the server has received the video.
[0079] Step 3:
[0080] The server receives the uploaded video data and stores it in a database. The video file is linked to a user ID and can be accessed later for analysis.
[0081] Step 4:
[0082] The server retrieves the stored video data and uses deep learning to analyze the player's movements in the video. Specifically, an AI algorithm analyzes each frame and extracts the positions of the player's joints and body parts.
[0083] Step 5:
[0084] The server compares the extracted motion data with the optimal form data, using ideal form data stored in a database and matching the current motion data point by point.
[0085] Step 6:
[0086] The server analyzes the comparison results and identifies specific differences, such as shoulders being too high or elbows being at an incorrect angle.
[0087] Step 7:
[0088] The server generates specific corrective instructions based on an analysis of the differences, which are specific and detailed instructions on how to modify the behavior.
[0089] Step 8:
[0090] The server sends the generated correction instructions to the terminal, which may be in text format or, if necessary, in visual format (images or animations).
[0091] Step 9:
[0092] The device receives the correction instructions from the server and displays them to the user. The app provides an interface that allows the user to easily check the instructions and intuitively understand specific improvement methods.
[0093] Step 10:
[0094] Users can correct their form based on the correction instructions they receive through the app, and then film and upload the video again to check the effectiveness of their corrections.
[0095] In this way, each step works in tandem to create a system that efficiently analyzes athletes' performance and provides specific methods for improvement.
[0096] Example 1
[0097] Next, a description will be given of Example 1. In the following description, the data processing device 12 will be referred to as a "server" and the smart device 14 will be referred to as a "terminal."
[0098] With conventional sports training systems, it was difficult to analyze athletes' movements and provide instructions for correcting their form quickly and efficiently. Furthermore, the process of uploading and analyzing video data was complicated, making it difficult for users to operate. Furthermore, conventional systems had a low level of expressiveness in providing correction instructions, making it difficult for users to intuitively understand them.
[0099] The specific processing by the specific processing unit 290 of the data processing device 12 in the first embodiment is realized by the following means.
[0100] In this invention, the server includes a device means for filming a player's performance, a server means for receiving the video filmed by the device, an analysis means for analyzing the video data received by the server and extracting the player's motion data, a comparison means for comparing the extracted motion data with optimal form data, an instruction generation means for generating specific correction instructions based on the differences obtained by the comparison means, a transmission means for transmitting the correction instructions generated by the instruction generation means to a terminal, a display means for displaying the correction instructions received by the terminal to a user, a filming means for filming video data at high resolution and recording the player's detailed movements, a means for easily uploading the video to the server via a dedicated app, a means for analyzing the player's movements in the video using a deep learning model and automatically generating movement data, and a means for sending correction instructions in text or visual format and displaying them in a format that the user can intuitively understand. This makes it possible to quickly and efficiently provide a player's motion analysis and form correction instructions, and to provide a training environment that is easy for users to operate and understand.
[0101] "Athlete" means a person engaged in sports activities and who participates in training and competition.
[0102] "Performance" refers to the overall sporting movements performed by athletes, including all physical movements related to the competition.
[0103] "Device" refers to equipment used to record athletes' performances, including smartphones and dedicated cameras.
[0104] A "server" is a computer system that receives video data sent from devices via a network and analyzes and processes the data.
[0105] "Analysis means" refers to the technology and algorithms used to process video data and extract and analyze player movement data.
[0106] "Movement data" is specific information about the player's movements extracted by the analysis means, and includes the positions and angles of the joints.
[0107] The "comparison means" refers to the technology and algorithm that compares the motion data with the pre-established optimal form data.
[0108] "Optimal form data" refers to ideal movement information set based on model athletes and their best past performances.
[0109] The "instruction generation means" refers to the technology and algorithms that generate corrective instructions for the player based on the analysis results of the movement data.
[0110] The "transmission means" is a communication technology for transmitting the generated correction instructions to the user's terminal.
[0111] A "terminal" is a device that receives correction instructions and displays them to the user, and includes a smartphone, tablet, etc.
[0112] "Display means" refers to the technology and interface for visually displaying to the user the corrective instructions received at the terminal.
[0113] "High resolution" refers to a level of image quality where the video data is highly detailed and the players' movements can be clearly recorded.
[0114] The "dedicated app" is application software that allows users to easily operate and upload video data.
[0115] A "deep learning model" is an artificial intelligence model that uses deep learning technology to analyze the movements of players in video and automatically generate movement data.
[0116] "Text format" refers to a format in which correction instructions are expressed as text.
[0117] "Visual format" refers to a format in which corrective instructions are conveyed through visual representations such as diagrams or animations.
[0118] This invention relates to an AI sports trainer system for efficiently improving athletes' performance. The system includes a device for recording athletes' performance, a server for receiving and analyzing the video data, a means for generating and transmitting corrective instructions, and a terminal for displaying the instructions to the user.
[0119] The system consists of the following main components:
[0120] 1. Device
[0121] The devices, which can include smartphones or dedicated cameras, are used to capture high-resolution footage of athletes' performances. For example, iPhones and dedicated sports cameras have high-resolution capabilities that allow for clear recording of athletes' detailed movements.
[0122] 2. Server
[0123] The server is a computer system that receives and analyzes video data, and often uses powerful cloud infrastructure such as Amazon Web Services (AWS) or Google Cloud Platform (GCP).
[0124] The server also uses deep learning frameworks such as TensorFlow and PyTorch to run deep learning models to analyze the player's movements in the received video.
[0125] As a specific example, a user can record their batting movements using an iPhone and then upload the video to a server on AWS via a dedicated app.
[0126] 3. Analysis method
[0127] The analysis method is integrated into the server and is a technology for extracting movement data from video data. It uses a deep learning model to analyze each key point of a player's movement.
[0128] For example, a TensorFlow model analyzes each frame in a video and extracts key data points such as the position of a player's shoulders and the angle of their elbows.
[0129] 4. Means of comparison
[0130] The server also includes a comparison mechanism that compares the extracted motion data with optimal form data, using the motion data of past best performances and model athletes.
[0131] For example, the server detects specific differences such as "elbows are positioned too low."
[0132] 5. Instruction generation means
[0133] The instruction generation means generates corrective instructions based on the differences obtained from the analysis and comparison means. The generated instructions are specific and show the player what parts to correct and how.
[0134] For example, specific instructions such as "Keep your elbow 5 degrees higher on your next pitch" are generated.
[0135] 6. Means of transmission
[0136] The transmission means is a technique for transmitting the server-generated correction instructions to the user's terminal using the secure HTTPS protocol.
[0137] 7. Terminal
[0138] The terminal is a device such as a smartphone or tablet that displays the orthodontic instructions received from the server to the user. A dedicated app is installed on the terminal, and it is designed to allow the user to intuitively understand the instructions.
[0139] For example, the smartphone app screen displays text instructions such as "Pull your shoulders back a little more," along with animations and illustrations.
[0140] These components enable the system to provide an environment for efficiently and quickly improving athlete performance.
[0141] Specific examples
[0142] 1. Users film and upload their performances
[0143] For example, a user can record their pitching motion using an iPhone and then upload the video to a server on AWS via a dedicated app.
[0144] Example prompt: "Upload a pitching video"
[0145] 2. The server receives and analyzes the video
[0146] The server receives the video data and uses TensorFlow to analyze and extract each key point of the player (e.g., shoulder position, elbow angle, etc.).
[0147] Example prompt: "Start behavior analysis using deep learning models."
[0148] 3. Comparison of motion data and optimal form data
[0149] The server compares the movement data with optimal form data in a database and detects specific differences, such as "elbow position too low."
[0150] Example prompt: "Your elbow position is 15 degrees off from optimal."
[0151] 4. Generate and send correction instructions
[0152] The server generates specific corrective instructions, such as "keep your elbow 5 degrees higher on your next pitch," and sends them to the user's device using the HTTPS protocol.
[0153] Example prompt: "Keep your elbow 5 degrees higher on your next pitch."
[0154] 5. The device will display correction instructions.
[0155] The user's smartphone app receives the correction instructions, displays "Keep your elbow 5 degrees higher on your next pitch," and plays an animation showing the movement.
[0156] Example prompt: "Push your shoulders back a little more."
[0157] In this way, the system provides a comprehensive solution for efficiently improving athletes' performance in sports.
[0158] The flow of the identification process in the first embodiment will be described with reference to FIG.
[0159] Step 1:
[0160] Users film their performances
[0161] Specifically, a user uses a smartphone or a dedicated camera to capture high-resolution footage of a sports performance (e.g., batting or pitching). The input is actual footage of the player's movements, and the output is a high-resolution video file.
[0162] Example: "User records batting action on iPhone"
[0163] Step 2:
[0164] User uploads video to server
[0165] Users upload the recorded video to the server through a dedicated app. They launch the app, select the recorded video file, and click the "Upload" button. The input is the recorded video file, and the output is the video data saved on the server.
[0166] Example: "Use the dedicated app to upload a video of your batting practice."
[0167] Step 3:
[0168] Server receives video
[0169] The server receives videos uploaded by users. The server saves the video files in cloud storage (e.g., AWS S3 bucket). The input is the uploaded video data, and the output is the video file saved in the storage.
[0170] Example: "A server on AWS receives and stores video data."
[0171] Step 4:
[0172] The server analyzes the video
[0173] The server uses a deep learning model (e.g., TensorFlow model) to analyze the video data and extract key motion data for each video frame. The input is the saved video file, and the output is the extracted keypoint data (e.g., shoulder position, elbow angle, etc.).
[0174] Example: "The server analyzes the motion using a TensorFlow model and extracts the shoulder position and elbow angle."
[0175] Step 5:
[0176] The server compares the behavior data with the best form data
[0177] The server compares the analyzed motion data with the ideal form data in a database. The comparison method is to match each frame of motion data with the ideal data and identify discrepancies. The input is the analyzed motion data and the ideal form data, and the output is specific discrepancy information (e.g., elbow angle is 15 degrees lower).
[0178] Example: "Compare analyzed behavior data with optimal form data to detect discrepancies."
[0179] Step 6:
[0180] Server generates corrective instructions
[0181] The server generates correction instructions based on the specific difference information. The instructions are generated using a natural language generation model, and are generated as specific content that is easy for users to understand. The input is the difference information, and the output is specific correction instructions (e.g., "Lift your elbow 15 degrees").
[0182] Example: Generate specific instructions such as "Keep your elbow 15 degrees higher on your next pitch."
[0183] Step 7:
[0184] The server sends correction instructions to the device
[0185] The server sends the generated correction instructions to the user's device. The transmission is performed using secure HTTPS communication. The input is the generated correction instructions, and the output is a transmission completion notification to the user's device.
[0186] Example: "The server sends correction instructions to the user's device using HTTPS."
[0187] Step 8:
[0188] The device displays correction instructions
[0189] The device displays the received correction instructions to the user. The app uses an intuitive interface to show instructions using text and animation. The input is the received correction instructions, and the output is the displayed instruction information.
[0190] Example: "The user's smartphone displays correction instructions and animations demonstrating the movements."
[0191] Through these concrete steps, the system will efficiently improve athletes' performance.
[0192] (Application example 1)
[0193] Next, a description will be given of Application Example 1. In the following description, the data processing device 12 will be referred to as a "server" and the smart device 14 will be referred to as a "terminal."
[0194] Current sports training systems and industrial robot control systems have difficulty optimizing the movement efficiency of athletes and work machines in real time. Furthermore, there is a lack of means to improve the accuracy of robot movements while integrating them with human motion analysis technology. Therefore, there is a need to simultaneously improve the performance of athletes and the efficiency of robots in industrial settings.
[0195] The specific processing by the specific processing unit 290 of the data processing device 12 in the application example 1 is realized by the following means.
[0196] In this invention, the server includes a device for filming the performance of athletes and industrial machinery, an analysis means for receiving and analyzing the video data and extracting movement data, a comparison means for comparing the movement data with optimal form data or work efficiency data, and an instruction generation means for generating specific correction instructions based on the differences. This allows for real-time analysis of the movements of athletes and industrial machinery, providing efficient feedback, and enabling performance improvement.
[0197] An "athlete" is a person who participates in a sport or particular performance competition with the aim of improving their performance.
[0198] "Industrial machinery" refers to robots and machines in general used in factories and production facilities, and is a device used to efficiently perform specific tasks.
[0199] "Motion data" is data that records in detail the movements of athletes and industrial machinery, and contains information that includes important key points extracted through analysis.
[0200] "Analysis means" refers to an algorithm or system for receiving video data, analyzing it, and extracting motion data.
[0201] The "comparison means" is a system or method that has the function of comparing the motion data extracted by the analysis means with optimal form data or work efficiency data.
[0202] The "instruction generation means" refers to an algorithm or system for generating specific corrective instructions based on the differences obtained by the comparison means.
[0203] The "transmission means" is a communication means having a function for transmitting the generated correction instruction to the terminal.
[0204] The "display means" is an interface or device for visually displaying to the user the corrective instructions received at the terminal.
[0205] The "control means" is a system or algorithm for applying corrective instructions received by the terminal to the robot and correcting its behavior in real time.
[0206] The present invention relates to a system for efficiently optimizing the movements of athletes and industrial machines. DETAILED DESCRIPTION OF THE INVENTION ... Hereinafter, specific embodiments of the present invention will be described.
[0207] definition
[0208] An "athlete" is a person who participates in a sport or particular performance competition with the aim of improving their performance.
[0209] "Industrial machinery" refers to robots and machines in general used in factories and production facilities, and is a device used to efficiently perform specific tasks.
[0210] "Motion data" is data that records in detail the movements of athletes and industrial machinery, and contains information that includes important key points extracted through analysis.
[0211] "Analysis means" refers to an algorithm or system for receiving video data, analyzing it, and extracting motion data.
[0212] The "comparison means" is a system or method that has the function of comparing the motion data extracted by the analysis means with optimal form data or work efficiency data.
[0213] The "instruction generation means" refers to an algorithm or system for generating specific corrective instructions based on the differences obtained by the comparison means.
[0214] The "transmission means" is a communication means having a function for transmitting the generated correction instruction to the terminal.
[0215] The "display means" is an interface or device for visually displaying to the user the corrective instructions received at the terminal.
[0216] The "control means" is a system or algorithm for applying corrective instructions received by the terminal to the robot and correcting its behavior in real time.
[0217] System Overview
[0218] 1. Use of photography devices
[0219] The server records the movements of athletes and industrial machines using a device that records the movements of the athletes and industrial machines, such as a smartphone or a dedicated high-resolution camera.
[0220] 2. Uploading video data
[0221] Users upload the videos they have taken to the server using a dedicated app. This allows for intuitive operation, making it easy to select and upload videos.
[0222] 3. Data Receipt and Analysis
[0223] The server analyzes the received video data and extracts motion data of athletes and industrial machines using deep learning models such as TensorFlow.
[0224] 4. Comparison with optimal form
[0225] The server compares the extracted behavioral data with optimal form or performance data, which identifies specific deviations between the behavior and optimal performance.
[0226] 5. Generate corrective instructions
[0227] The server generates specific corrective instructions based on the differences obtained by the comparison means, which instruct the athlete or industrial machine in detail how and where to correct the athlete or industrial machine.
[0228] 6. Sending Corrective Instructions
[0229] The server sends the generated correction instructions to the terminal, which may include instructions in textual and visual form.
[0230] 7. Display of orthodontic instructions
[0231] The device receives correction instructions from the server and displays them to the user. The app has an easy-to-use interface that allows for intuitive understanding.
[0232] 8. Real-time behavior correction
[0233] Using the control means, the terminal is able to apply the received corrective instructions to the robot to modify its behavior in real time.
[0234] Hardware and software used
[0235] Hardware: Smartphones, dedicated cameras, industrial robots
[0236] Software: TensorFlow (deep learning model), dedicated app, server
[0237] Adding concrete examples and prompt sentence examples
[0238] For example, to optimize the behavior of a robot used in a factory installing parts, the following prompts could be input to a generative AI model:
[0239] plaintext
[0240] Analyze the robot arm's motion using video and compare it with the optimal arm motion pattern to generate corrective instructions. The optimal arm motion pattern includes the arm height, movement method, and speed. For example, if the arm is raised too high, issue the instruction "Lower the arm."
[0241] This will enable the operational efficiency of robots within factories to be optimized in real time, which is expected to improve productivity.
[0242] The flow of the specific processing in the application example 1 will be described with reference to FIG.
[0243] Step 1:
[0244] The user films the operation of industrial machinery in operation in a factory. Using a smartphone or a dedicated camera, the operation is recorded as high-resolution video data. The filming is done so that the movement of each part of the industrial machinery is clearly visible.
[0245] Input: Video taken with a smartphone or dedicated camera
[0246] Output: High-resolution video data
[0247] Step 2:
[0248] Users upload the video data they have taken to the server using a dedicated app. The app's intuitive operation makes it easy to select and upload videos.
[0249] Input: High-resolution video data
[0250] Output: Video data uploaded to the server
[0251] Step 3:
[0252] The server receives the video data and begins analysis. Using deep learning models such as TensorFlow, the server analyzes the movements of the industrial machinery in the video and extracts key points (such as the position of the arm and the angle of movement).
[0253] Input: Video data uploaded to the server
[0254] Output: Extracted keypoint data
[0255] Step 4:
[0256] The server compares the extracted key point data with optimal motion pattern data in a database, and the comparison identifies differences between the motion data and the optimal motion pattern.
[0257] Input: Extracted keypoint data, optimal motion pattern data in the database
[0258] Output: Difference between motion data and optimal motion pattern
[0259] Step 5:
[0260] The server generates specific correction instructions based on the difference between the operational data and the optimal operational pattern. The generated instructions provide detailed instructions on how and where the industrial machinery should be corrected.
[0261] Input: Difference between motion data and optimal motion pattern
[0262] Output: Specific correction instructions
[0263] Step 6:
[0264] The server transmits the generated correction instructions to the terminal, which may include instructions in textual and visual form.
[0265] Input: Specific correction instructions
[0266] Output: Corrective instructions sent to the terminal
[0267] Step 7:
[0268] The device receives correction instructions from the server and displays them to the user. The app has an easy-to-use interface and is designed to allow users to intuitively understand the instructions.
[0269] Input: Corrective instructions sent to the terminal
[0270] Output: Corrective instructions displayed to the user
[0271] Step 8:
[0272] The corrective instructions confirmed by the user are applied to the industrial machine using the control means to correct its operation in real time, thereby improving the operational efficiency of the industrial machine.
[0273] Input: User confirmed corrective instructions
[0274] Output: Real-time corrected industrial machine behavior
[0275] Furthermore, an emotion engine that estimates the user's emotion may be combined. That is, the identification processing unit 290 may estimate the user's emotion using the emotion identification model 59 and perform identification processing using the user's emotion.
[0276] This invention relates to an AI sports trainer system for efficiently improving athletes' performance, and by combining it with an emotion engine, provides instruction according to the user's emotional state. This system includes a device for recording the athletes' performance, a server for receiving and analyzing the video data, a means for generating and transmitting corrective instructions, a terminal for displaying the instructions to the user, and the emotion engine.
[0277] System Overview
[0278] 1. Use of photography devices
[0279] Users use their smartphone or dedicated camera to film their pitching, batting, or other performances. Once they've finished filming, they launch the app and select the video.
[0280] 2. Uploading data
[0281] The user uploads the video they have taken to the server using a dedicated app. The app sends the user ID and video data together and confirms that the server has received the video.
[0282] 3. Data Receipt and Analysis
[0283] The server receives the uploaded video data and stores it in a database. The video file is linked to a user ID and can be accessed later for analysis.
[0284] 4. Movement analysis and form extraction
[0285] The server retrieves the stored video data and uses deep learning to analyze the player's movements in the video. Specifically, an AI algorithm analyzes each frame and extracts the positions of the player's joints and body parts.
[0286] 5. Comparison with optimal form
[0287] The server compares the extracted motion data with the optimal form data, using ideal form data stored in a database and matching the current motion data point by point.
[0288] 6. Generating Corrective Instructions
[0289] The server analyzes the specific differences based on the comparison results and generates specific corrective instructions based on the differences, which may be in text format or, if necessary, in visual format (images or animations).
[0290] 7. Emotion Engine Operation
[0291] The emotion engine extracts emotional data from the user's facial expressions and voice. For example, it uses a camera to recognize the user's face and an expression analysis algorithm to determine their emotion. It also uses voice recognition technology to analyze emotions from the user's tone of voice and speaking style.
[0292] 8. Use of Emotional Data
[0293] The server receives the emotion data and adjusts the generated corrective instructions, providing standard instructions if the emotion is positive, and adding encouraging or supportive messages if the emotion is negative.
[0294] Emotional data is stored historically and used to monitor training progress.
[0295] 9. Sending and Displaying Corrective Instructions
[0296] The server sends tailored corrective instructions to the terminal, which may include motivational or encouraging messages.
[0297] The device receives the correction instructions and motivational messages from the server and displays them to the user. The app provides an interface that allows the user to easily check the instructions and intuitively understand specific ways to improve.
[0298] Specific examples
[0299] 1. Users film and upload pitching videos
[0300] Users record their pitching on their smartphones and upload the video to the server using a dedicated app.
[0301] 2. The server receives and analyzes the video
[0302] The server receives the video and uses a deep learning model to analyze and extract key points in the pitching motion.
[0303] 3. Comparison with optimal form
[0304] The server compares the analyzed movement data with optimal form data and finds specific differences such as "elbow position is too low" or "timing is too early."
[0305] 4. Generate corrective instructions
[0306] Based on the comparison results, the server generates specific corrective instructions, such as "keep your elbow 5 degrees higher on your next pitch."
[0307] 5. Operation of the Emotion Engine and Use of Emotion Data
[0308] The emotion engine recognizes emotions from the user's facial expressions and voice and extracts emotion data. For example, if the user has a dissatisfied expression, the emotion data is determined to be negative.
[0309] The server receives the emotion data and adds a motivational message to the corrective instruction to reduce the negative emotion.
[0310] 6. Sending and Displaying Instructions
[0311] The server sends the generated correction instructions and motivation messages to the terminal, and the terminal displays the instructions to the user, allowing the user to get effective feedback and correct the form immediately.
[0312] In this way, this system, which combines an emotion engine, can efficiently correct athletes' form and provide mental support in real time, which is expected to improve their overall performance.
[0313] The processing flow will be explained below.
[0314] Step 1:
[0315] Users use a smartphone or dedicated camera to record their sports performance (e.g., pitching or batting). After finishing recording, users launch the app and select the video.
[0316] Step 2:
[0317] The user uploads the video they have taken using a dedicated app to the server. The app sends the user ID and video data together and confirms that the server has received the video.
[0318] Step 3:
[0319] The server receives the uploaded video data and stores it in a database. The video file is linked to a user ID and can be accessed later for analysis.
[0320] Step 4:
[0321] The server retrieves the stored video data and uses deep learning to analyze the player's movements in the video. Specifically, an AI algorithm analyzes each frame and extracts the positions of the player's joints and body parts.
[0322] Step 5:
[0323] The server compares the extracted motion data with the optimal form data, using ideal form data stored in a database and matching the current motion data point by point.
[0324] Step 6:
[0325] The server analyzes the comparison results and identifies specific differences, such as shoulders being too high or elbows being at an incorrect angle.
[0326] Step 7:
[0327] The server generates specific corrective instructions based on an analysis of the differences, which are specific and detailed instructions on how to modify the behavior.
[0328] Step 8:
[0329] The emotion engine extracts emotional data from the user's facial expressions and voice. For example, it uses a camera to recognize the user's face and an expression analysis algorithm to determine their emotion. It also uses voice recognition technology to analyze emotions from the user's tone of voice and speaking style.
[0330] Step 9:
[0331] The server receives the emotion data and adjusts the generated corrective instructions. If the emotion is positive, it issues a standard instruction, but if it is negative, it adds a message of encouragement or support. For example, an instruction such as "Keep your elbow 5 degrees higher on your next pitch" could include an encouraging message such as "You can do it!"
[0332] Step 10:
[0333] The server sends the adjusted correction instructions to the device, which may be in text format or, if necessary, in visual format (images or animations).
[0334] Step 11:
[0335] The device receives the correction instructions from the server and displays them to the user. The app has an easy-to-use interface, allowing users to easily check the instructions. Specific improvement methods and encouraging messages are also displayed.
[0336] Step 12:
[0337] Users can correct their form based on the correction instructions and encouraging messages they receive through the app, and can then film and upload the video again to see the results of their corrections.
[0338] In this way, each step works in tandem, allowing players to efficiently correct their form and receive mental support in real time, which is expected to improve their overall performance.
[0339] Example 2
[0340] Next, a description will be given of Example 2. In the following description, the data processing device 12 will be referred to as a "server" and the smart device 14 will be referred to as a "terminal."
[0341] In recent years, there has been a demand for systems that can efficiently improve athletes' performance in sports training. Conventional systems analyze movements and improve form, but rarely provide guidance that takes into account the user's emotional state, resulting in problems such as a decrease in motivation and a lack of psychological support. To solve these issues, a system that can provide real-time guidance that corresponds to the user's emotional state is needed.
[0342] The specific processing by the specific processing unit 290 of the data processing device 12 in the second embodiment is realized by the following means.
[0343] In this invention, the server includes a device for filming the player's performance, a computer unit that receives the video filmed by the device, analysis means that analyzes the video data received by the computer unit and extracts the player's movement data, comparison means that compares the movement data extracted by the analysis means with optimal form data, instruction generation means that generates specific correction instructions based on the differences obtained by the comparison means, transmission means that transmits the correction instructions generated by the instruction generation means to a terminal, display means that displays the correction instructions received by the terminal to the user, an emotion recognition device that extracts emotion data from the user's facial expressions and voice, and adjustment means that adjusts the correction instructions based on the emotion data obtained from the emotion recognition device. This enables appropriate feedback and instruction tailored to the user's emotional state.
[0344] An "athlete" is a person who participates in a sport or competition and seeks to improve their performance.
[0345] "Performance" is a general term for the actions and techniques that athletes perform in sports or competitions.
[0346] "Devices" refers to equipment used to film athletes' performances, including smartphones and dedicated cameras.
[0347] "Video" refers to video data captured by the device.
[0348] The "computer section" is a section that includes a computer system for receiving and processing video data.
[0349] "Analysis means" refers to a function for analyzing the video data received by the computer unit and extracting the movement data of the players.
[0350] "Movement data" refers to information extracted by analytical means, such as the position of a player's joints and body, and the timing of their movements.
[0351] "Form data" is data that serves as a model of ideal behavior or technology and serves as a standard for comparison.
[0352] "Comparison means" refers to functionality for comparing motion data with optimal form data and identifying differences.
[0353] "Difference" refers to the specific differences observed between the behavior data and the form data.
[0354] "Corrective instructions" refer to instructions given to the athlete to improve their movements based on the differences obtained by the comparison means.
[0355] The "instruction generating means" refers to a function for generating specific corrective instructions based on the difference.
[0356] The "transmitting means" refers to a function for transmitting the generated correction instruction to the terminal.
[0357] A "terminal" is a device used by a user to check correction instructions, and includes a smartphone, tablet, etc.
[0358] The "display means" refers to a function for displaying the received correction instructions to the user at the terminal.
[0359] An "emotion recognition device" refers to a device for extracting emotional data from a user's facial expressions and voice.
[0360] "Emotion data" refers to information that indicates the emotional state of a user extracted by an emotion recognition device.
[0361] The "adjustment means" refers to a function for adjusting the correction instruction based on the emotion data to match the user's emotional state.
[0362] The present invention relates to an AI sports trainer system for efficiently improving athletes' performance, and by combining it with an emotion engine, the system provides instruction according to the user's emotional state.
[0363] The system includes the following elements:
[0364] 1. Imaging equipment
[0365] Users use a smartphone or dedicated camera to film performances such as pitching and batting, and this device can capture high-resolution images.
[0366] 2. Upload a video
[0367] The user uploads the video they have taken to the server using a dedicated app. This app sends the user ID and video data together and confirms that the server has received the video.
[0368] 3. Receiving and storing video data
[0369] The server receives the uploaded video data and stores it in a database. The video files are linked to user IDs and can be accessed for analysis.
[0370] 4. Motion analysis
[0371] The server retrieves the stored video data and uses deep learning to analyze the player's movements in the video. Specifically, it analyzes each frame and extracts the positions of the player's joints and body parts.
[0372] 5. Comparison with optimal form
[0373] The server compares the analyzed movement data with ideal form data stored in a database, checking points such as elbow angle and weight transfer.
[0374] 6. Generating Corrective Instructions
[0375] Based on the comparison results, the server analyzes the specific differences and generates corrective instructions, such as "keep your elbow 5 degrees higher on your next pitch."
[0376] 7. Emotion recognition
[0377] The emotion engine uses a camera to recognize the user's facial expressions, and uses a facial expression analysis algorithm to determine the user's emotions. It also uses voice recognition technology to analyze emotions from the user's tone of voice and speaking style.
[0378] 8. Use of Emotional Data
[0379] The server receives the emotion data sent from the emotion engine and adjusts the corrective instructions: if the user is in a positive emotional state, it provides normal instructions, and if the user is in a negative emotional state, it adds encouraging or supportive messages.
[0380] 9. Sending and Displaying Corrective Instructions
[0381] The server sends the generated correction instructions and motivation messages to the terminal.
[0382] The terminal displays the correction instructions and motivational messages received from the server to the user, who then works to correct the form.
[0383] Specific examples
[0384] 1. Users film and upload pitching videos
[0385] Users record their pitching on their smartphones and upload the video to the server using a dedicated app.
[0386] 2. The server receives and analyzes the video
[0387] The server receives the video and uses a deep learning model to analyze each key point of the pitching motion.
[0388] 3. Comparison with optimal form
[0389] The server compares the analyzed movement data with ideal form data and finds specific differences such as "elbow position is too low" or "timing is too early."
[0390] 4. Generate corrective instructions
[0391] The server generates specific corrective instructions such as "keep your elbow 5 degrees higher on your next pitch."
[0392] 5. Operation of the Emotion Engine and Use of Emotion Data
[0393] The emotion engine recognizes emotions from the user's facial expressions and voice, and if negative emotions are detected, the server adds a motivational message.
[0394] 6. Sending and Displaying Instructions
[0395] The server sends the generated correction instructions and motivation messages to the terminal, and the terminal displays the instructions to the user, allowing the user to get effective feedback and correct the form immediately.
[0396] Prompt Sentence Examples
[0397] "Upload a pitching video and analyze elbow position and timing. If emotions are negative, generate improvement instructions with encouraging messages."
[0398] The flow of the identification process in the second embodiment will be described with reference to FIG.
[0399] Step 1:
[0400] Users can film sports performances
[0401] Input: Smartphone or dedicated camera
[0402] How it works: A user records their own sports performance (e.g., pitching or batting) in high resolution, which allows the user to obtain video data.
[0403] Output: Recorded video file
[0404] Step 2:
[0405] User uploads video
[0406] Input: Recorded video file, dedicated app
[0407] How it works: The user launches the app, selects the video they have taken, and uploads it to the server. At this time, the app sends the user ID and the video data together.
[0408] Output: Video data sent to the server
[0409] Step 3:
[0410] The server receives and stores video data
[0411] Input: User ID, video data
[0412] Operation: The server receives the uploaded video data, checks the data integrity, and stores it in the database if there are no problems.
[0413] Output: Video data stored in a database
[0414] Step 4:
[0415] The server retrieves and analyzes the video data
[0416] Input: Video data stored in a database
[0417] Movement: The server retrieves the video data and analyzes each frame using a deep learning algorithm. By extracting the positions of the player's joints and body parts, various movement data is generated.
[0418] Output: Extracted behavioral data
[0419] Step 5:
[0420] Comparison with server-optimized form
[0421] Input: extracted behavioral data, ideal form data
[0422] Movement: The server compares the analyzed movement data with the ideal form data stored in a database, matching each point in detail, such as elbow angle and body timing.
[0423] Output: Difference between behavior data and form data
[0424] Step 6:
[0425] Server generates corrective instructions
[0426] Input: Differences between behavioral data and form data
[0427] Action: The server generates specific corrective instructions based on the discrepancy. A typical example might be, "Keep your elbow 5 degrees higher on your next pitch." Visual feedback (images or animations) is also generated if necessary.
[0428] Output: Generated correction instructions
[0429] Step 7:
[0430] Emotion Engine Operation
[0431] Input: Camera video, audio data
[0432] How it works: The emotion engine uses a camera to recognize the user's facial expressions, and uses a facial expression analysis algorithm to determine the user's emotions. It also uses voice recognition technology to analyze emotions from the user's tone of voice and speaking style.
[0433] Output: Extracted emotion data
[0434] Step 8:
[0435] The server uses emotion data to adjust correction instructions.
[0436] Input: Generated correction instructions, extracted emotion data
[0437] How it works: The server receives the emotion data and adjusts the corrective instructions depending on the user's emotional state: normal instructions for positive emotions, and encouraging or supportive messages for negative emotions.
[0438] Output: Adjusted straightening instructions
[0439] Step 9:
[0440] The server sends correction instructions to the terminal
[0441] Input: Adjusted orthodontic instructions
[0442] Operation: The server sends the generated correction instructions and motivation messages to the device. This transmission process is low-latency.
[0443] Output: Corrective instructions and motivation messages received by the device
[0444] Step 10:
[0445] The device displays correction instructions and messages
[0446] Input: Corrective instructions and motivation messages received by the device
[0447] How it works: The device displays correction instructions and motivational messages to the user. Through the app interface, users can intuitively understand specific ways to improve and immediately start correcting their form.
[0448] Output: Improvement instructions and motivational messages seen by the user
[0449] (Application example 2)
[0450] Next, a description will be given of Application Example 2. In the following description, the data processing device 12 will be referred to as a "server" and the smart device 14 will be referred to as a "terminal."
[0451] In conventional factory robot operation, there was a lack of guidance that took into account the optimization of robot operations and the emotional state of the operator. This made it difficult to provide effective corrections when the robot's performance deteriorated, often leading to stress and a decline in operator motivation. In particular, there was a need to simultaneously improve the robot's operational efficiency and the operator's mental health.
[0452] The specific processing by the specific processing unit 290 of the data processing device 12 in Application Example 2 is realized by the following means. In this invention, the server includes a device for filming the robot's movements, a means for receiving the video, a means for analyzing the video data and extracting movement data, a means for comparing the movement data with optimal form data, and a means for generating specific correction instructions. This enables optimization of the robot's movements and effective instruction that takes into account the operator's emotional state.
[0453] A "device" is an electronic device used to capture the action of interest.
[0454] A "server" is a computer system that receives, stores, and analyzes data over a network.
[0455] The "analysis means" refers to algorithms or software for extracting motion data from received video data.
[0456] The "comparison means" is an algorithm or software that has the functionality to compare extracted motion data with optimal form data.
[0457] The "instruction generating means" is an algorithm or software for generating specific correction instructions based on the results of the comparison means.
[0458] The "transmitting means" is a component having a function for transmitting the generated correction instruction to the terminal.
[0459] The "display means" refers to a device or software for visually displaying to the user the correction instructions received at the terminal.
[0460] "Emotion analysis means" refers to algorithms or software for extracting emotional data from a user's facial expressions and voice.
[0461] The "adjustment means" is software having a function for appropriately adjusting correction instructions based on the emotion data acquired by the emotion analysis means.
[0462] "Robot operator" refers to a worker who operates a robot within a factory.
[0463] System Configuration
[0464] This invention is an AI system for improving the performance of factory robot operations, providing optimal instructions by filming and analyzing the robot's movements and achieving effective feedback by taking into account the emotional state of the operator. The system consists of the following components:
[0465] shooting device
[0466] The user (robot operator) uses a smartphone or dedicated camera to record the factory robot's work movements. The camera has high resolution and can clearly capture even the smallest movements. Once the recording is complete, the user launches the dedicated app, selects the video, and uploads it to the server.
[0467] server
[0468] The server stores the received video data in a database and analyzes it using deep learning. Specifically, it uses the following methods:
[0469] Video analysis method: Using TensorFlow and PyTorch, motion data is extracted from the video data. Each frame is analyzed to extract the robot's joints and motion patterns.
[0470] Comparison method: The analyzed motion data is compared with the optimal form data, which represents the ideal robot motion and is stored in a database.
[0471] Instruction generation means: Based on the comparison results, appropriate corrective instructions are generated. For example, specific instructions such as "Make the arm move 5 degrees faster" are generated.
[0472] Emotion analysis method: Using OpenCV and Google Cloud Vision API, emotional data is extracted from the robot operator's facial expressions and voice. The operator's face is recognized by a camera, and emotions are determined using an expression analysis algorithm. Speech recognition technology is also used to analyze emotions from the operator's tone of voice and speaking style.
[0473] Adjustment means: Receives emotion data and adjusts the generated corrective instructions. In the case of positive emotions, it issues normal instructions, and in the case of negative emotions, it adds encouraging messages.
[0474] Terminal
[0475] The terminal is used to provide feedback to the operator, displaying corrective instructions and motivational messages received from the server to help the operator comply with instructions promptly.
[0476] Specific examples
[0477] Consider a situation in a factory where a robot arm is not moving smoothly, resulting in a drop in efficiency. An operator records the robot's movements on a smartphone and uploads the video to a server using a dedicated app. The server analyzes the video, identifies the problem area, and generates specific instructions such as "Make the arm move five degrees faster." At the same time, if the operator looks dissatisfied, the emotion analysis means determines this to be a negative emotion and provides an encouraging message such as "Try a little harder to improve efficiency!" These instructions and messages are displayed on the terminal, allowing the operator to respond immediately.
[0478] Prompt Sentence Examples
[0479] An example of a prompt would be:
[0480] "Today, the robot's arm movements were slow, causing work efficiency to drop. Please film the robot's arm movements and provide appropriate instructions for improvement based on the analysis results."
[0481] As described above, the system of the present invention can simultaneously optimize the operation of factory robots and provide psychological support to operators, thereby contributing to improving overall work efficiency.
[0482] The flow of the specific processing in the application example 2 will be described with reference to FIG.
[0483] Step 1:
[0484] The user films the robot's movements. The user uses a smartphone or a dedicated camera to record the robot's movements, making sure to clearly capture the operating environment and specific movements. The input data is the filmed video. The output data is a video file.
[0485] Step 2:
[0486] The server receives the video. The user uses a dedicated app to upload the video they have taken to the server. The app sends the user ID and video data together, and the server receives the video. The input data is the uploaded video file and user ID, and the output data is the video file saved on the server.
[0487] Step 3:
[0488] The server analyzes the video and extracts motion data. Using a deep learning model (such as TensorFlow or PyTorch), it analyzes the robot's motion in the video frame by frame. It extracts the robot's joints and motion patterns from each frame. The input data is the saved video file, and the output data is the extracted motion data.
[0489] Step 4:
[0490] The server compares the motion data with the optimal form data. The server then matches the extracted motion data with the ideal robot motion data stored in a database. Specifically, it evaluates the degree of match for each frame and identifies discrepancies. The input data is the extracted motion data and the optimal form data, and the output data is the comparison result.
[0491] Step 5:
[0492] The server generates the corrective instructions. Based on the comparison results, the algorithm generates appropriate corrective instructions. For example, specific instructions such as "Make the arm move 5 degrees faster" are generated. The input data is the comparison result, and the output data is the corrective instructions.
[0493] Step 6:
[0494] The server analyzes the user's emotions. An emotion analysis engine (such as OpenCV or Google Cloud Vision API) is used to extract emotion data from the user's facial expressions and voice. The camera recognizes the user's face, and an emotion analysis algorithm is used to determine their emotion. Speech recognition technology is also used to analyze emotions from the tone of voice and speaking style. The input data is the user's facial expression images and voice data, and the output data is emotion data.
[0495] Step 7:
[0496] The server adjusts the correction instructions based on the emotion data. It receives the emotion data and adjusts the generated correction instructions according to the user's emotional state. For example, in the case of negative emotions, it adds an encouraging message such as "Let's try a little harder to improve efficiency!" The input data are emotion data and correction instructions, and the output data are the adjusted correction instructions.
[0497] Step 8:
[0498] The server sends the adjusted correction instruction to the terminal. The sending means sends the adjusted correction instruction and the motivation message to the terminal. The input data is the adjusted correction instruction, and the output data is the instruction and the message sent to the terminal.
[0499] Step 9:
[0500] The terminal displays the corrective instructions to the user. The terminal displays the received corrective instructions and motivational messages to the user. The user checks the instructions on the terminal and immediately understands the specific improvement methods. The input data are the instructions and messages sent to the terminal, and the output data are the display to the user.
[0501] The specific processing unit 290 transmits the result of the specific processing to the smart device 14. In the smart device 14, the control unit 46A causes the output device 40 to output the result of the specific processing. The microphone 38B acquires audio indicating a user input regarding the result of the specific processing. The control unit 46A transmits audio data indicating the user input acquired by the microphone 38B to the data processing device 12. In the data processing device 12, the specific processing unit 290 acquires the audio data.
[0502] The data generation model 58 is a so-called generative AI (Artificial Intelligence). An example of the data generation model 58 is ChatGPT (Internet Search<URL: https: / / openai.com / blog / chatgpt> ), Gemini (Internet search <url: https: gemini.google.com ?hl="ja">) and other generation AIs. The data generation model 58 is obtained by performing deep learning on a neural network. A prompt including an instruction is input to the data generation model 58, and inference data such as voice data indicating voice, text data indicating text, and image data indicating an image is also input. The data generation model 58 performs inference on the input inference data in accordance with the instruction indicated by the prompt, and outputs the inference result in a data format such as voice data and text data. Here, inference refers to, for example, analysis, classification, prediction, and / or summarization.
[0503] In the above embodiment, an example in which the specific process is performed by the data processing device 12 has been given, but the technology of the present disclosure is not limited to this, and the specific process may be performed by the smart device 14.
[0504] [Second embodiment]
[0505] FIG. 3 shows an example of the configuration of a data processing system 210 according to the second embodiment.
[0506] 3, the data processing system 210 includes the data processing device 12 and smart glasses 214. An example of the data processing device 12 is a server.
[0507] The data processing device 12 includes a computer 22, a database 24, and a communication I / F 26. The computer 22 is an example of a "computer" according to the technology of the present disclosure. The computer 22 includes a processor 28, a RAM 30, and a storage 32. The processor 28, the RAM 30, and the storage 32 are connected to a bus 34. The database 24 and the communication I / F 26 are also connected to the bus 34. The communication I / F 26 is connected to a network 54. Examples of the network 54 include a WAN (Wide Area Network) and / or a LAN (Local Area Network).
[0508] The smart glasses 214 include a computer 36, a microphone 238, a speaker 240, a camera 42, and a communication I / F 44. The computer 36 includes a processor 46, a RAM 48, and a storage 50. The processor 46, the RAM 48, and the storage 50 are connected to a bus 52. The microphone 238, the speaker 240, and the camera 42 are also connected to the bus 52.
[0509] The microphone 238 receives instructions and the like from the user 20 by receiving voice uttered by the user 20. The microphone 238 captures the voice uttered by the user 20, converts the captured voice into audio data, and outputs it to the processor 46. The speaker 240 outputs audio in accordance with instructions from the processor 46.
[0510] Camera 42 is a small digital camera equipped with an optical system including a lens, aperture, and shutter, and an imaging element such as a CMOS (Complementary Metal-Oxide-Semiconductor) image sensor or a CCD (Charge Coupled Device) image sensor, and captures images of the surroundings of user 20 (for example, an imaging range defined by an angle of view equivalent to the field of vision of a typical healthy person).
[0511] The communication I / F 44 is connected to a network 54. The communication I / Fs 44 and 26 are responsible for the exchange of various information between the processor 46 and the processor 28 via the network 54. The exchange of various information between the processor 46 and the processor 28 using the communication I / Fs 44 and 26 is carried out in a secure state.
[0512] Fig. 4 shows an example of the main functions of the data processing device 12 and the smart glasses 214. As shown in Fig. 4, in the data processing device 12, a specific process is performed by the processor 28. A specific process program 56 is stored in the storage 32.
[0513] The specific processing program 56 is an example of a "program" according to the technology of the present disclosure. The processor 28 reads the specific processing program 56 from the storage 32 and executes the read specific processing program 56 on the RAM 30. The specific processing is realized by the processor 28 operating as a specific processing unit 290 in accordance with the specific processing program 56 executed on the RAM 30.
[0514] The storage 32 stores a data generation model 58 and an emotion identification model 59. The data generation model 58 and the emotion identification model 59 are used by the identification processing unit 290.
[0515] In the smart glasses 214, the reception output process is performed by the processor 46. A reception output program 60 is stored in the storage 50. The processor 46 reads the reception output program 60 from the storage 50 and executes the read reception output program 60 on the RAM 48. The reception output process is realized by the processor 46 operating as the control unit 46A in accordance with the reception output program 60 executed on the RAM 48.
[0516] Next, a description will be given of the identification process performed by the identification processing unit 290 of the data processing device 12. In the following description, the data processing device 12 will be referred to as the "server" and the smart glasses 214 will be referred to as the "terminal."
[0517] The present invention relates to an AI sports trainer system for efficiently improving athletes' performance. The system includes a device for recording athletes' performance, a server for receiving and analyzing the video data, a means for generating and transmitting corrective instructions, and a terminal for displaying the instructions to the user.
[0518] System Overview
[0519] 1. Use of photography devices
[0520] Users can use their smartphones or dedicated cameras to film their pitching, batting, and other performances.
[0521] The video data is shot in high resolution, clearly recording the players' detailed movements.
[0522] 2. Uploading data
[0523] Users upload the videos they have taken to the server through a dedicated app. The app's intuitive operation makes it easy to select and upload videos.
[0524] 3. Data Receipt and Analysis
[0525] The server receives the video data and begins analysis. It uses powerful AI algorithms to model the player's movements in the video and extract each key element (e.g., shoulder position, elbow angle, etc.).
[0526] In this case, a deep learning model operates as an analytical tool to automatically generate motion data.
[0527] 4. Comparison with optimal form
[0528] The server compares the extracted motion data with optimal form data stored in a database, which is based on model athletes and their best past performances.
[0529] The comparison means relatively matches the motion data and the optimal form data for each frame and identifies specific differences.
[0530] 5. Generate corrective instructions
[0531] Based on the analysis of the discrepancies, the server generates specific corrective instructions, detailing what the player needs to correct and how.
[0532] For example, specific instructions for correcting the movement such as "pull your shoulders back a little more" are generated.
[0533] 6. Sending Corrective Instructions
[0534] The server generates and transmits corrective instructions to the terminal, which may include instructions in textual and visual formats.
[0535] In addition to text instructions, animations and illustrations may also be sent.
[0536] 7. Display of orthodontic instructions
[0537] The device receives correction instructions from the server and displays them to the user. The app has an easy-to-use interface and is designed to allow users to intuitively understand the instructions.
[0538] This allows users to learn how to correct their form in real time.
[0539] Specific examples
[0540] 1. Users film and upload pitching videos
[0541] Users record their pitching on their smartphones and upload the video to the server using a dedicated app.
[0542] 2. The server receives and analyzes the video
[0543] The server receives the video and uses a deep learning model to analyze and extract key points in the pitching motion.
[0544] 3. Comparison with optimal form
[0545] The server compares the analyzed movement data with optimal form data and detects differences such as "elbow position is too low" or "timing is too early."
[0546] 4. Generate corrective instructions
[0547] Based on the comparison results, the server generates specific corrective instructions, such as "keep your elbow 5 degrees higher on your next pitch."
[0548] 5. Sending and Displaying Instructions
[0549] The server generates and sends the correction instructions to the device, which then displays them to the user, allowing the user to receive effective feedback and instantly correct their form.
[0550] Such a system will enable players to efficiently correct their form in real time, which is expected to improve their overall performance.
[0551] The processing flow will be explained below.
[0552] Step 1:
[0553] Users use a smartphone or dedicated camera to film their own sports performance (pitching, batting, etc.) After finishing filming, they launch the app and select the video.
[0554] Step 2:
[0555] The user uploads the video they took using the app to the server. The app sends the user ID and the video data together, and confirms that the server has received the video.
[0556] Step 3:
[0557] The server receives the uploaded video data and stores it in a database. The video file is linked to a user ID and can be accessed later for analysis.
[0558] Step 4:
[0559] The server retrieves the stored video data and uses deep learning to analyze the player's movements in the video. Specifically, an AI algorithm analyzes each frame and extracts the positions of the player's joints and body parts.
[0560] Step 5:
[0561] The server compares the extracted motion data with the optimal form data, using ideal form data stored in a database and matching the current motion data point by point.
[0562] Step 6:
[0563] The server analyzes the comparison results and identifies specific differences, such as shoulders being too high or elbows being at an incorrect angle.
[0564] Step 7:
[0565] The server generates specific corrective instructions based on an analysis of the differences, which are specific and detailed instructions on how to modify the behavior.
[0566] Step 8:
[0567] The server sends the generated correction instructions to the terminal, which may be in text format or, if necessary, in visual format (images or animations).
[0568] Step 9:
[0569] The device receives the correction instructions from the server and displays them to the user. The app provides an interface that allows the user to easily check the instructions and intuitively understand specific improvement methods.
[0570] Step 10:
[0571] Users can correct their form based on the correction instructions they receive through the app, and then film and upload the video again to check the effectiveness of their corrections.
[0572] In this way, each step works in tandem to create a system that efficiently analyzes athletes' performance and provides specific methods for improvement.
[0573] Example 1
[0574] Next, a description will be given of Example 1. In the following description, the data processing device 12 will be referred to as a "server" and the smart glasses 214 will be referred to as a "terminal."
[0575] With conventional sports training systems, it was difficult to analyze athletes' movements and provide instructions for correcting their form quickly and efficiently. Furthermore, the process of uploading and analyzing video data was complicated, making it difficult for users to operate. Furthermore, conventional systems had a low level of expressiveness in providing correction instructions, making it difficult for users to intuitively understand them.
[0576] The specific processing by the specific processing unit 290 of the data processing device 12 in the first embodiment is realized by the following means.
[0577] In this invention, the server includes a device means for filming a player's performance, a server means for receiving the video filmed by the device, an analysis means for analyzing the video data received by the server and extracting the player's motion data, a comparison means for comparing the extracted motion data with optimal form data, an instruction generation means for generating specific correction instructions based on the differences obtained by the comparison means, a transmission means for transmitting the correction instructions generated by the instruction generation means to a terminal, a display means for displaying the correction instructions received by the terminal to a user, a filming means for filming video data at high resolution and recording the player's detailed movements, a means for easily uploading the video to the server via a dedicated app, a means for analyzing the player's movements in the video using a deep learning model and automatically generating movement data, and a means for sending correction instructions in text or visual format and displaying them in a format that the user can intuitively understand. This makes it possible to quickly and efficiently provide a player's motion analysis and form correction instructions, and to provide a training environment that is easy for users to operate and understand.
[0578] "Athlete" means a person engaged in sports activities and who participates in training and competition.
[0579] "Performance" refers to the overall sporting movements performed by athletes, including all physical movements related to the competition.
[0580] "Device" refers to equipment used to record athletes' performances, including smartphones and dedicated cameras.
[0581] A "server" is a computer system that receives video data sent from devices via a network and analyzes and processes the data.
[0582] "Analysis means" refers to the technology and algorithms used to process video data and extract and analyze player movement data.
[0583] "Movement data" is specific information about the player's movements extracted by the analysis means, and includes the positions and angles of the joints.
[0584] The "comparison means" refers to the technology and algorithm that compares the motion data with the pre-established optimal form data.
[0585] "Optimal form data" refers to ideal movement information set based on model athletes and their best past performances.
[0586] The "instruction generation means" refers to the technology and algorithms that generate corrective instructions for the player based on the analysis results of the movement data.
[0587] The "transmission means" is a communication technology for transmitting the generated correction instructions to the user's terminal.
[0588] A "terminal" is a device that receives correction instructions and displays them to the user, and includes a smartphone, tablet, etc.
[0589] "Display means" refers to the technology and interface for visually displaying to the user the corrective instructions received at the terminal.
[0590] "High resolution" refers to a level of image quality where the video data is highly detailed and the players' movements can be clearly recorded.
[0591] The "dedicated app" is application software that allows users to easily operate and upload video data.
[0592] A "deep learning model" is an artificial intelligence model that uses deep learning technology to analyze the movements of players in video and automatically generate movement data.
[0593] "Text format" refers to a format in which correction instructions are expressed as text.
[0594] "Visual format" refers to a format in which corrective instructions are conveyed through visual representations such as diagrams or animations.
[0595] This invention relates to an AI sports trainer system for efficiently improving athletes' performance. The system includes a device for recording athletes' performance, a server for receiving and analyzing the video data, a means for generating and transmitting corrective instructions, and a terminal for displaying the instructions to the user.
[0596] The system consists of the following main components:
[0597] 1. Device
[0598] The devices, which can include smartphones or dedicated cameras, are used to capture high-resolution footage of athletes' performances. For example, iPhones and dedicated sports cameras have high-resolution capabilities that allow for clear recording of athletes' detailed movements.
[0599] 2. Server
[0600] The server is a computer system that receives and analyzes video data, and often uses powerful cloud infrastructure such as Amazon Web Services (AWS) or Google Cloud Platform (GCP).
[0601] The server also uses deep learning frameworks such as TensorFlow and PyTorch to run deep learning models to analyze the player's movements in the received video.
[0602] As a specific example, a user can record their batting movements using an iPhone and then upload the video to a server on AWS via a dedicated app.
[0603] 3. Analysis method
[0604] The analysis method is integrated into the server and is a technology for extracting movement data from video data. It uses a deep learning model to analyze each key point of a player's movement.
[0605] For example, a TensorFlow model analyzes each frame in a video and extracts key data points such as the position of a player's shoulders and the angle of their elbows.
[0606] 4. Means of comparison
[0607] The server also includes a comparison mechanism that compares the extracted motion data with optimal form data, using the motion data of past best performances and model athletes.
[0608] For example, the server detects specific differences such as "elbows are positioned too low."
[0609] 5. Instruction generation means
[0610] The instruction generation means generates corrective instructions based on the differences obtained from the analysis and comparison means. The generated instructions are specific and show the player what parts to correct and how.
[0611] For example, specific instructions such as "Keep your elbow 5 degrees higher on your next pitch" are generated.
[0612] 6. Means of transmission
[0613] The transmission means is a technique for transmitting the server-generated correction instructions to the user's terminal using the secure HTTPS protocol.
[0614] 7. Terminal
[0615] The terminal is a device such as a smartphone or tablet that displays the orthodontic instructions received from the server to the user. A dedicated app is installed on the terminal, and it is designed to allow the user to intuitively understand the instructions.
[0616] For example, the smartphone app screen displays text instructions such as "Pull your shoulders back a little more," along with animations and illustrations.
[0617] These components enable the system to provide an environment for efficiently and quickly improving athlete performance.
[0618] Specific examples
[0619] 1. Users film and upload their performances
[0620] For example, a user can record their pitching motion using an iPhone and then upload the video to a server on AWS via a dedicated app.
[0621] Example prompt: "Upload a pitching video"
[0622] 2. The server receives and analyzes the video
[0623] The server receives the video data and uses TensorFlow to analyze and extract each key point of the player (e.g., shoulder position, elbow angle, etc.).
[0624] Example prompt: "Start behavior analysis using deep learning models."
[0625] 3. Comparison of motion data and optimal form data
[0626] The server compares the movement data with optimal form data in a database and detects specific differences, such as "elbow position too low."
[0627] Example prompt: "Your elbow position is 15 degrees off from optimal."
[0628] 4. Generate and send correction instructions
[0629] The server generates specific corrective instructions, such as "keep your elbow 5 degrees higher on your next pitch," and sends them to the user's device using the HTTPS protocol.
[0630] Example prompt: "Keep your elbow 5 degrees higher on your next pitch."
[0631] 5. The device will display correction instructions.
[0632] The user's smartphone app receives the correction instructions, displays "Keep your elbow 5 degrees higher on your next pitch," and plays an animation showing the movement.
[0633] Example prompt: "Push your shoulders back a little more."
[0634] In this way, the system provides a comprehensive solution for efficiently improving athletes' performance in sports.
[0635] The flow of the identification process in the first embodiment will be described with reference to FIG.
[0636] Step 1:
[0637] Users film their performances
[0638] Specifically, a user uses a smartphone or a dedicated camera to capture high-resolution footage of a sports performance (e.g., batting or pitching). The input is actual footage of the player's movements, and the output is a high-resolution video file.
[0639] Example: "User records batting action on iPhone"
[0640] Step 2:
[0641] User uploads video to server
[0642] Users upload the recorded video to the server through a dedicated app. They launch the app, select the recorded video file, and click the "Upload" button. The input is the recorded video file, and the output is the video data saved on the server.
[0643] Example: "Use the dedicated app to upload a video of your batting practice."
[0644] Step 3:
[0645] Server receives video
[0646] The server receives videos uploaded by users. The server saves the video files in cloud storage (e.g., AWS S3 bucket). The input is the uploaded video data, and the output is the video file saved in the storage.
[0647] Example: "A server on AWS receives and stores video data."
[0648] Step 4:
[0649] The server analyzes the video
[0650] The server uses a deep learning model (e.g., TensorFlow model) to analyze the video data and extract key motion data for each video frame. The input is the saved video file, and the output is the extracted keypoint data (e.g., shoulder position, elbow angle, etc.).
[0651] Example: "The server analyzes the motion using a TensorFlow model and extracts the shoulder position and elbow angle."
[0652] Step 5:
[0653] The server compares the behavior data with the best form data
[0654] The server compares the analyzed motion data with the ideal form data in a database. The comparison method is to match each frame of motion data with the ideal data and identify discrepancies. The input is the analyzed motion data and the ideal form data, and the output is specific discrepancy information (e.g., elbow angle is 15 degrees lower).
[0655] Example: "Compare analyzed behavior data with optimal form data to detect discrepancies."
[0656] Step 6:
[0657] Server generates corrective instructions
[0658] The server generates correction instructions based on the specific difference information. The instructions are generated using a natural language generation model, and are generated as specific content that is easy for users to understand. The input is the difference information, and the output is specific correction instructions (e.g., "Lift your elbow 15 degrees").
[0659] Example: Generate specific instructions such as "Keep your elbow 15 degrees higher on your next pitch."
[0660] Step 7:
[0661] The server sends correction instructions to the device
[0662] The server sends the generated correction instructions to the user's device. The transmission is performed using secure HTTPS communication. The input is the generated correction instructions, and the output is a transmission completion notification to the user's device.
[0663] Example: "The server sends correction instructions to the user's device using HTTPS."
[0664] Step 8:
[0665] The device displays correction instructions
[0666] The device displays the received correction instructions to the user. The app uses an intuitive interface to show instructions using text and animation. The input is the received correction instructions, and the output is the displayed instruction information.
[0667] Example: "The user's smartphone displays correction instructions and animations demonstrating the movements."
[0668] Through these concrete steps, the system will efficiently improve athletes' performance.
[0669] (Application example 1)
[0670] Next, a description will be given of Application Example 1. In the following description, the data processing device 12 will be referred to as a "server" and the smart glasses 214 will be referred to as a "terminal."
[0671] Current sports training systems and industrial robot control systems have difficulty optimizing the movement efficiency of athletes and work machines in real time. Furthermore, there is a lack of means to improve the accuracy of robot movements while integrating them with human motion analysis technology. Therefore, there is a need to simultaneously improve the performance of athletes and the efficiency of robots in industrial settings.
[0672] The specific processing by the specific processing unit 290 of the data processing device 12 in the application example 1 is realized by the following means.
[0673] In this invention, the server includes a device for filming the performance of athletes and industrial machinery, an analysis means for receiving and analyzing the video data and extracting movement data, a comparison means for comparing the movement data with optimal form data or work efficiency data, and an instruction generation means for generating specific correction instructions based on the differences. This allows for real-time analysis of the movements of athletes and industrial machinery, providing efficient feedback, and enabling performance improvement.
[0674] An "athlete" is a person who participates in a sport or particular performance competition with the aim of improving their performance.
[0675] "Industrial machinery" refers to robots and machines in general used in factories and production facilities, and is a device used to efficiently perform specific tasks.
[0676] "Motion data" is data that records in detail the movements of athletes and industrial machinery, and contains information that includes important key points extracted through analysis.
[0677] "Analysis means" refers to an algorithm or system for receiving video data, analyzing it, and extracting motion data.
[0678] The "comparison means" is a system or method that has the function of comparing the motion data extracted by the analysis means with optimal form data or work efficiency data.
[0679] The "instruction generation means" refers to an algorithm or system for generating specific corrective instructions based on the differences obtained by the comparison means.
[0680] The "transmission means" is a communication means having a function for transmitting the generated correction instruction to the terminal.
[0681] The "display means" is an interface or device for visually displaying to the user the corrective instructions received at the terminal.
[0682] The "control means" is a system or algorithm for applying corrective instructions received by the terminal to the robot and correcting its behavior in real time.
[0683] The present invention relates to a system for efficiently optimizing the movements of athletes and industrial machines. DETAILED DESCRIPTION OF THE INVENTION ... Hereinafter, specific embodiments of the present invention will be described.
[0684] definition
[0685] An "athlete" is a person who participates in a sport or particular performance competition with the aim of improving their performance.
[0686] "Industrial machinery" refers to robots and machines in general used in factories and production facilities, and is a device used to efficiently perform specific tasks.
[0687] "Motion data" is data that records in detail the movements of athletes and industrial machinery, and contains information that includes important key points extracted through analysis.
[0688] "Analysis means" refers to an algorithm or system for receiving video data, analyzing it, and extracting motion data.
[0689] The "comparison means" is a system or method that has the function of comparing the motion data extracted by the analysis means with optimal form data or work efficiency data.
[0690] The "instruction generation means" refers to an algorithm or system for generating specific corrective instructions based on the differences obtained by the comparison means.
[0691] The "transmission means" is a communication means having a function for transmitting the generated correction instruction to the terminal.
[0692] The "display means" is an interface or device for visually displaying to the user the corrective instructions received at the terminal.
[0693] The "control means" is a system or algorithm for applying corrective instructions received by the terminal to the robot and correcting its behavior in real time.
[0694] System Overview
[0695] 1. Use of photography devices
[0696] The server records the movements of athletes and industrial machines using a device that records the movements of the athletes and industrial machines, such as a smartphone or a dedicated high-resolution camera.
[0697] 2. Uploading video data
[0698] Users upload the videos they have taken to the server using a dedicated app. This allows for intuitive operation, making it easy to select and upload videos.
[0699] 3. Data Receipt and Analysis
[0700] The server analyzes the received video data and extracts motion data of athletes and industrial machines using deep learning models such as TensorFlow.
[0701] 4. Comparison with optimal form
[0702] The server compares the extracted behavioral data with optimal form or performance data, which identifies specific deviations between the behavior and optimal performance.
[0703] 5. Generate corrective instructions
[0704] The server generates specific corrective instructions based on the differences obtained by the comparison means, which instruct the athlete or industrial machine in detail how and where to correct the athlete or industrial machine.
[0705] 6. Sending Corrective Instructions
[0706] The server sends the generated correction instructions to the terminal, which may include instructions in textual and visual form.
[0707] 7. Display of orthodontic instructions
[0708] The device receives correction instructions from the server and displays them to the user. The app has an easy-to-use interface that allows for intuitive understanding.
[0709] 8. Real-time behavior correction
[0710] Using the control means, the terminal is able to apply the received corrective instructions to the robot to modify its behavior in real time.
[0711] Hardware and software used
[0712] Hardware: Smartphones, dedicated cameras, industrial robots
[0713] Software: TensorFlow (deep learning model), dedicated app, server
[0714] Adding concrete examples and prompt sentence examples
[0715] For example, to optimize the behavior of a robot used in a factory installing parts, the following prompts could be input to a generative AI model:
[0716] plaintext
[0717] Analyze the robot arm's motion using video and compare it with the optimal arm motion pattern to generate corrective instructions. The optimal arm motion pattern includes the arm height, movement method, and speed. For example, if the arm is raised too high, issue the instruction "Lower the arm."
[0718] This will enable the operational efficiency of robots within factories to be optimized in real time, which is expected to improve productivity.
[0719] The flow of the specific processing in the application example 1 will be described with reference to FIG.
[0720] Step 1:
[0721] The user films the operation of industrial machinery in operation in a factory. Using a smartphone or a dedicated camera, the operation is recorded as high-resolution video data. The filming is done so that the movement of each part of the industrial machinery is clearly visible.
[0722] Input: Video taken with a smartphone or dedicated camera
[0723] Output: High-resolution video data
[0724] Step 2:
[0725] Users upload the video data they have taken to the server using a dedicated app. The app's intuitive operation makes it easy to select and upload videos.
[0726] Input: High-resolution video data
[0727] Output: Video data uploaded to the server
[0728] Step 3:
[0729] The server receives the video data and begins analysis. Using deep learning models such as TensorFlow, the server analyzes the movements of the industrial machinery in the video and extracts key points (such as the position of the arm and the angle of movement).
[0730] Input: Video data uploaded to the server
[0731] Output: Extracted keypoint data
[0732] Step 4:
[0733] The server compares the extracted key point data with optimal motion pattern data in a database, and the comparison identifies differences between the motion data and the optimal motion pattern.
[0734] Input: Extracted keypoint data, optimal motion pattern data in the database
[0735] Output: Difference between motion data and optimal motion pattern
[0736] Step 5:
[0737] The server generates specific correction instructions based on the difference between the operational data and the optimal operational pattern. The generated instructions provide detailed instructions on how and where the industrial machinery should be corrected.
[0738] Input: Difference between motion data and optimal motion pattern
[0739] Output: Specific correction instructions
[0740] Step 6:
[0741] The server transmits the generated correction instructions to the terminal, which may include instructions in textual and visual form.
[0742] Input: Specific correction instructions
[0743] Output: Corrective instructions sent to the terminal
[0744] Step 7:
[0745] The device receives correction instructions from the server and displays them to the user. The app has an easy-to-use interface and is designed to allow users to intuitively understand the instructions.
[0746] Input: Corrective instructions sent to the terminal
[0747] Output: Corrective instructions displayed to the user
[0748] Step 8:
[0749] The corrective instructions confirmed by the user are applied to the industrial machine using the control means to correct its operation in real time, thereby improving the operational efficiency of the industrial machine.
[0750] Input: User confirmed corrective instructions
[0751] Output: Real-time corrected industrial machine behavior
[0752] Furthermore, an emotion engine that estimates the user's emotion may be further combined. That is, the identification processing unit 290 may estimate the user's emotion using the emotion identification model 59, and perform identification processing using the user's emotion.
[0753] This invention relates to an AI sports trainer system for efficiently improving athletes' performance, and by combining it with an emotion engine, provides instruction according to the user's emotional state. This system includes a device for recording the athletes' performance, a server for receiving and analyzing the video data, a means for generating and transmitting corrective instructions, a terminal for displaying the instructions to the user, and the emotion engine.
[0754] System Overview
[0755] 1. Use of photography devices
[0756] Users use their smartphone or dedicated camera to film their pitching, batting, or other performances. Once they've finished filming, they launch the app and select the video.
[0757] 2. Uploading data
[0758] The user uploads the video they have taken to the server using a dedicated app. The app sends the user ID and video data together and confirms that the server has received the video.
[0759] 3. Data Receipt and Analysis
[0760] The server receives the uploaded video data and stores it in a database. The video file is linked to a user ID and can be accessed later for analysis.
[0761] 4. Movement analysis and form extraction
[0762] The server retrieves the stored video data and uses deep learning to analyze the player's movements in the video. Specifically, an AI algorithm analyzes each frame and extracts the positions of the player's joints and body parts.
[0763] 5. Comparison with optimal form
[0764] The server compares the extracted motion data with the optimal form data, using ideal form data stored in a database and matching the current motion data point by point.
[0765] 6. Generating Corrective Instructions
[0766] The server analyzes the specific differences based on the comparison results and generates specific corrective instructions based on the differences, which may be in text format or, if necessary, in visual format (images or animations).
[0767] 7. Emotion Engine Operation
[0768] The emotion engine extracts emotional data from the user's facial expressions and voice. For example, it uses a camera to recognize the user's face and an expression analysis algorithm to determine their emotion. It also uses voice recognition technology to analyze emotions from the user's tone of voice and speaking style.
[0769] 8. Use of Emotional Data
[0770] The server receives the emotion data and adjusts the generated corrective instructions, providing standard instructions if the emotion is positive, and adding encouraging or supportive messages if the emotion is negative.
[0771] Emotional data is stored historically and used to monitor training progress.
[0772] 9. Sending and Displaying Corrective Instructions
[0773] The server sends tailored corrective instructions to the terminal, which may include motivational or encouraging messages.
[0774] The device receives the correction instructions and motivational messages from the server and displays them to the user. The app provides an interface that allows the user to easily check the instructions and intuitively understand specific ways to improve.
[0775] Specific examples
[0776] 1. Users film and upload pitching videos
[0777] Users record their pitching on their smartphones and upload the video to the server using a dedicated app.
[0778] 2. The server receives and analyzes the video
[0779] The server receives the video and uses a deep learning model to analyze and extract key points in the pitching motion.
[0780] 3. Comparison with optimal form
[0781] The server compares the analyzed movement data with optimal form data and finds specific differences such as "elbow position is too low" or "timing is too early."
[0782] 4. Generate corrective instructions
[0783] Based on the comparison results, the server generates specific corrective instructions, such as "keep your elbow 5 degrees higher on your next pitch."
[0784] 5. Operation of the Emotion Engine and Use of Emotion Data
[0785] The emotion engine recognizes emotions from the user's facial expressions and voice and extracts emotion data. For example, if the user has a dissatisfied expression, the emotion data is determined to be negative.
[0786] The server receives the emotion data and adds a motivational message to the corrective instruction to reduce the negative emotion.
[0787] 6. Sending and Displaying Instructions
[0788] The server sends the generated correction instructions and motivation messages to the terminal, and the terminal displays the instructions to the user, allowing the user to get effective feedback and correct the form immediately.
[0789] In this way, this system, which combines an emotion engine, can efficiently correct athletes' form and provide mental support in real time, which is expected to improve their overall performance.
[0790] The processing flow will be explained below.
[0791] Step 1:
[0792] Users use a smartphone or dedicated camera to record their sports performance (e.g., pitching or batting). After finishing recording, users launch the app and select the video.
[0793] Step 2:
[0794] The user uploads the video they have taken using a dedicated app to the server. The app sends the user ID and video data together and confirms that the server has received the video.
[0795] Step 3:
[0796] The server receives the uploaded video data and stores it in a database. The video file is linked to a user ID and can be accessed later for analysis.
[0797] Step 4:
[0798] The server retrieves the stored video data and uses deep learning to analyze the player's movements in the video. Specifically, an AI algorithm analyzes each frame and extracts the positions of the player's joints and body parts.
[0799] Step 5:
[0800] The server compares the extracted motion data with the optimal form data, using ideal form data stored in a database and matching the current motion data point by point.
[0801] Step 6:
[0802] The server analyzes the comparison results and identifies specific differences, such as shoulders being too high or elbows being at an incorrect angle.
[0803] Step 7:
[0804] The server generates specific corrective instructions based on an analysis of the differences, which are specific and detailed instructions on how to modify the behavior.
[0805] Step 8:
[0806] The emotion engine extracts emotional data from the user's facial expressions and voice. For example, it uses a camera to recognize the user's face and an expression analysis algorithm to determine their emotion. It also uses voice recognition technology to analyze emotions from the user's tone of voice and speaking style.
[0807] Step 9:
[0808] The server receives the emotion data and adjusts the generated corrective instructions. If the emotion is positive, it issues a standard instruction, but if it is negative, it adds a message of encouragement or support. For example, an instruction such as "Keep your elbow 5 degrees higher on your next pitch" could include an encouraging message such as "You can do it!"
[0809] Step 10:
[0810] The server sends the adjusted correction instructions to the device, which may be in text format or, if necessary, in visual format (images or animations).
[0811] Step 11:
[0812] The device receives the correction instructions from the server and displays them to the user. The app has an easy-to-use interface, allowing users to easily check the instructions. Specific improvement methods and encouraging messages are also displayed.
[0813] Step 12:
[0814] Users can correct their form based on the correction instructions and encouraging messages they receive through the app, and can then film and upload the video again to see the results of their corrections.
[0815] In this way, each step works in tandem, allowing players to efficiently correct their form and receive mental support in real time, which is expected to improve their overall performance.
[0816] Example 2
[0817] Next, a description will be given of Example 2. In the following description, the data processing device 12 will be referred to as a "server" and the smart glasses 214 will be referred to as a "terminal."
[0818] In recent years, there has been a demand for systems that can efficiently improve athletes' performance in sports training. Conventional systems analyze movements and improve form, but rarely provide guidance that takes into account the user's emotional state, resulting in problems such as a decrease in motivation and a lack of psychological support. To solve these issues, a system that can provide real-time guidance that corresponds to the user's emotional state is needed.
[0819] The specific processing by the specific processing unit 290 of the data processing device 12 in the second embodiment is realized by the following means.
[0820] In this invention, the server includes a device for filming the player's performance, a computer unit that receives the video filmed by the device, analysis means that analyzes the video data received by the computer unit and extracts the player's movement data, comparison means that compares the movement data extracted by the analysis means with optimal form data, instruction generation means that generates specific correction instructions based on the differences obtained by the comparison means, transmission means that transmits the correction instructions generated by the instruction generation means to a terminal, display means that displays the correction instructions received by the terminal to the user, an emotion recognition device that extracts emotion data from the user's facial expressions and voice, and adjustment means that adjusts the correction instructions based on the emotion data obtained from the emotion recognition device. This enables appropriate feedback and instruction tailored to the user's emotional state.
[0821] An "athlete" is a person who participates in a sport or competition and seeks to improve their performance.
[0822] "Performance" is a general term for the actions and techniques that athletes perform in sports or competitions.
[0823] "Devices" refers to equipment used to film athletes' performances, including smartphones and dedicated cameras.
[0824] "Video" refers to video data captured by the device.
[0825] The "computer section" is a section that includes a computer system for receiving and processing video data.
[0826] "Analysis means" refers to a function for analyzing the video data received by the computer unit and extracting the movement data of the players.
[0827] "Movement data" refers to information extracted by analytical means, such as the position of a player's joints and body, and the timing of their movements.
[0828] "Form data" is data that serves as a model of ideal behavior or technology and serves as a standard for comparison.
[0829] "Comparison means" refers to functionality for comparing motion data with optimal form data and identifying differences.
[0830] "Difference" refers to the specific differences observed between the behavior data and the form data.
[0831] "Corrective instructions" refer to instructions given to the athlete to improve their movements based on the differences obtained by the comparison means.
[0832] The "instruction generating means" refers to a function for generating specific corrective instructions based on the difference.
[0833] The "transmitting means" refers to a function for transmitting the generated correction instruction to the terminal.
[0834] A "terminal" is a device used by a user to check correction instructions, and includes a smartphone, tablet, etc.
[0835] The "display means" refers to a function for displaying the received correction instructions to the user at the terminal.
[0836] An "emotion recognition device" refers to a device for extracting emotional data from a user's facial expressions and voice.
[0837] "Emotion data" refers to information that indicates the emotional state of a user extracted by an emotion recognition device.
[0838] The "adjustment means" refers to a function for adjusting the correction instruction based on the emotion data to match the user's emotional state.
[0839] The present invention relates to an AI sports trainer system for efficiently improving athletes' performance, and by combining it with an emotion engine, the system provides instruction according to the user's emotional state.
[0840] The system includes the following elements:
[0841] 1. Imaging equipment
[0842] Users use a smartphone or dedicated camera to film performances such as pitching and batting, and this device can capture high-resolution images.
[0843] 2. Upload a video
[0844] The user uploads the video they have taken to the server using a dedicated app. This app sends the user ID and video data together and confirms that the server has received the video.
[0845] 3. Receiving and storing video data
[0846] The server receives the uploaded video data and stores it in a database. The video files are linked to user IDs and can be accessed for analysis.
[0847] 4. Motion analysis
[0848] The server retrieves the stored video data and uses deep learning to analyze the player's movements in the video. Specifically, it analyzes each frame and extracts the positions of the player's joints and body parts.
[0849] 5. Comparison with optimal form
[0850] The server compares the analyzed movement data with ideal form data stored in a database, checking points such as elbow angle and weight transfer.
[0851] 6. Generating Corrective Instructions
[0852] Based on the comparison results, the server analyzes the specific differences and generates corrective instructions, such as "keep your elbow 5 degrees higher on your next pitch."
[0853] 7. Emotion recognition
[0854] The emotion engine uses a camera to recognize the user's facial expressions, and uses a facial expression analysis algorithm to determine the user's emotions. It also uses voice recognition technology to analyze emotions from the user's tone of voice and speaking style.
[0855] 8. Use of Emotional Data
[0856] The server receives the emotion data sent from the emotion engine and adjusts the corrective instructions: if the user is in a positive emotional state, it provides normal instructions, and if the user is in a negative emotional state, it adds encouraging or supportive messages.
[0857] 9. Sending and Displaying Corrective Instructions
[0858] The server sends the generated correction instructions and motivation messages to the terminal.
[0859] The terminal displays the correction instructions and motivational messages received from the server to the user, who then works to correct the form.
[0860] Specific examples
[0861] 1. Users film and upload pitching videos
[0862] Users record their pitching on their smartphones and upload the video to the server using a dedicated app.
[0863] 2. The server receives and analyzes the video
[0864] The server receives the video and uses a deep learning model to analyze each key point of the pitching motion.
[0865] 3. Comparison with optimal form
[0866] The server compares the analyzed movement data with ideal form data and finds specific differences such as "elbow position is too low" or "timing is too early."
[0867] 4. Generate corrective instructions
[0868] The server generates specific corrective instructions such as "keep your elbow 5 degrees higher on your next pitch."
[0869] 5. Operation of the Emotion Engine and Use of Emotion Data
[0870] The emotion engine recognizes emotions from the user's facial expressions and voice, and if negative emotions are detected, the server adds a motivational message.
[0871] 6. Sending and Displaying Instructions
[0872] The server sends the generated correction instructions and motivation messages to the terminal, and the terminal displays the instructions to the user, allowing the user to get effective feedback and correct the form immediately.
[0873] Prompt Sentence Examples
[0874] "Upload a pitching video and analyze elbow position and timing. If emotions are negative, generate improvement instructions with encouraging messages."
[0875] The flow of the identification process in the second embodiment will be described with reference to FIG.
[0876] Step 1:
[0877] Users can film sports performances
[0878] Input: Smartphone or dedicated camera
[0879] How it works: A user records their own sports performance (e.g., pitching or batting) in high resolution, which allows the user to obtain video data.
[0880] Output: Recorded video file
[0881] Step 2:
[0882] User uploads video
[0883] Input: Recorded video file, dedicated app
[0884] How it works: The user launches the app, selects the video they have taken, and uploads it to the server. At this time, the app sends the user ID and the video data together.
[0885] Output: Video data sent to the server
[0886] Step 3:
[0887] The server receives and stores video data
[0888] Input: User ID, video data
[0889] Operation: The server receives the uploaded video data, checks the data integrity, and stores it in the database if there are no problems.
[0890] Output: Video data stored in a database
[0891] Step 4:
[0892] The server retrieves and analyzes the video data
[0893] Input: Video data stored in a database
[0894] Movement: The server retrieves the video data and analyzes each frame using a deep learning algorithm. By extracting the positions of the player's joints and body parts, various movement data is generated.
[0895] Output: Extracted behavioral data
[0896] Step 5:
[0897] Comparison with server-optimized form
[0898] Input: extracted behavioral data, ideal form data
[0899] Movement: The server compares the analyzed movement data with the ideal form data stored in a database, matching each point in detail, such as elbow angle and body timing.
[0900] Output: Difference between behavior data and form data
[0901] Step 6:
[0902] Server generates corrective instructions
[0903] Input: Differences between behavioral data and form data
[0904] Action: The server generates specific corrective instructions based on the discrepancy. A typical example might be, "Keep your elbow 5 degrees higher on your next pitch." Visual feedback (images or animations) is also generated if necessary.
[0905] Output: Generated correction instructions
[0906] Step 7:
[0907] Emotion Engine Operation
[0908] Input: Camera video, audio data
[0909] How it works: The emotion engine uses a camera to recognize the user's facial expressions, and uses a facial expression analysis algorithm to determine the user's emotions. It also uses voice recognition technology to analyze emotions from the user's tone of voice and speaking style.
[0910] Output: Extracted emotion data
[0911] Step 8:
[0912] The server uses emotion data to adjust correction instructions.
[0913] Input: Generated correction instructions, extracted emotion data
[0914] How it works: The server receives the emotion data and adjusts the corrective instructions depending on the user's emotional state: normal instructions for positive emotions, and encouraging or supportive messages for negative emotions.
[0915] Output: Adjusted straightening instructions
[0916] Step 9:
[0917] The server sends correction instructions to the terminal
[0918] Input: Adjusted orthodontic instructions
[0919] Operation: The server sends the generated correction instructions and motivation messages to the device. This transmission process is low-latency.
[0920] Output: Corrective instructions and motivation messages received by the device
[0921] Step 10:
[0922] The device displays correction instructions and messages
[0923] Input: Corrective instructions and motivation messages received by the device
[0924] How it works: The device displays correction instructions and motivational messages to the user. Through the app interface, users can intuitively understand specific ways to improve and immediately start correcting their form.
[0925] Output: Improvement instructions and motivational messages seen by the user
[0926] (Application example 2)
[0927] Next, a description will be given of Application Example 2. In the following description, the data processing device 12 will be referred to as a "server" and the smart glasses 214 will be referred to as a "terminal."
[0928] In conventional factory robot operation, there was a lack of guidance that took into account the optimization of robot operations and the emotional state of the operator. This made it difficult to provide effective corrections when the robot's performance deteriorated, often leading to stress and a decline in operator motivation. In particular, there was a need to simultaneously improve the robot's operational efficiency and the operator's mental health.
[0929] The specific processing by the specific processing unit 290 of the data processing device 12 in Application Example 2 is realized by the following means. In this invention, the server includes a device for filming the robot's movements, a means for receiving the video, a means for analyzing the video data and extracting movement data, a means for comparing the movement data with optimal form data, and a means for generating specific correction instructions. This enables optimization of the robot's movements and effective instruction that takes into account the operator's emotional state.
[0930] A "device" is an electronic device used to capture the action of interest.
[0931] A "server" is a computer system that receives, stores, and analyzes data over a network.
[0932] The "analysis means" refers to algorithms or software for extracting motion data from received video data.
[0933] The "comparison means" is an algorithm or software that has the functionality to compare extracted motion data with optimal form data.
[0934] The "instruction generating means" is an algorithm or software for generating specific correction instructions based on the results of the comparison means.
[0935] The "transmitting means" is a component having a function for transmitting the generated correction instruction to the terminal.
[0936] The "display means" refers to a device or software for visually displaying to the user the correction instructions received at the terminal.
[0937] "Emotion analysis means" refers to algorithms or software for extracting emotional data from a user's facial expressions and voice.
[0938] The "adjustment means" is software having a function for appropriately adjusting correction instructions based on the emotion data acquired by the emotion analysis means.
[0939] "Robot operator" refers to a worker who operates a robot within a factory.
[0940] System Configuration
[0941] This invention is an AI system for improving the performance of factory robot operations, providing optimal instructions by filming and analyzing the robot's movements and achieving effective feedback by taking into account the emotional state of the operator. The system consists of the following components:
[0942] shooting device
[0943] The user (robot operator) uses a smartphone or dedicated camera to record the factory robot's work movements. The camera has high resolution and can clearly capture even the smallest movements. Once the recording is complete, the user launches the dedicated app, selects the video, and uploads it to the server.
[0944] server
[0945] The server stores the received video data in a database and analyzes it using deep learning. Specifically, it uses the following methods:
[0946] Video analysis method: Using TensorFlow and PyTorch, motion data is extracted from the video data. Each frame is analyzed to extract the robot's joints and motion patterns.
[0947] Comparison method: The analyzed motion data is compared with the optimal form data, which represents the ideal robot motion and is stored in a database.
[0948] Instruction generation means: Based on the comparison results, appropriate corrective instructions are generated. For example, specific instructions such as "Make the arm move 5 degrees faster" are generated.
[0949] Emotion analysis method: Using OpenCV and Google Cloud Vision API, emotional data is extracted from the robot operator's facial expressions and voice. The operator's face is recognized by a camera, and emotions are determined using an expression analysis algorithm. Speech recognition technology is also used to analyze emotions from the operator's tone of voice and speaking style.
[0950] Adjustment means: Receives emotion data and adjusts the generated corrective instructions. In the case of positive emotions, it issues normal instructions, and in the case of negative emotions, it adds encouraging messages.
[0951] Terminal
[0952] The terminal is used to provide feedback to the operator, displaying corrective instructions and motivational messages received from the server to help the operator comply with instructions promptly.
[0953] Specific examples
[0954] Consider a situation in a factory where a robot arm is not moving smoothly, resulting in a drop in efficiency. An operator records the robot's movements on a smartphone and uploads the video to a server using a dedicated app. The server analyzes the video, identifies the problem area, and generates specific instructions such as "Make the arm move five degrees faster." At the same time, if the operator looks dissatisfied, the emotion analysis means determines this to be a negative emotion and provides an encouraging message such as "Try a little harder to improve efficiency!" These instructions and messages are displayed on the terminal, allowing the operator to respond immediately.
[0955] Prompt Sentence Examples
[0956] An example of a prompt would be:
[0957] "Today, the robot's arm movements were slow, causing work efficiency to drop. Please film the robot's arm movements and provide appropriate instructions for improvement based on the analysis results."
[0958] As described above, the system of the present invention can simultaneously optimize the operation of factory robots and provide psychological support to operators, thereby contributing to improving overall work efficiency.
[0959] The flow of the specific processing in the application example 2 will be described with reference to FIG.
[0960] Step 1:
[0961] The user films the robot's movements. The user uses a smartphone or a dedicated camera to record the robot's movements, making sure to clearly capture the operating environment and specific movements. The input data is the filmed video. The output data is a video file.
[0962] Step 2:
[0963] The server receives the video. The user uses a dedicated app to upload the video they have taken to the server. The app sends the user ID and video data together, and the server receives the video. The input data is the uploaded video file and user ID, and the output data is the video file saved on the server.
[0964] Step 3:
[0965] The server analyzes the video and extracts motion data. Using a deep learning model (such as TensorFlow or PyTorch), it analyzes the robot's motion in the video frame by frame. It extracts the robot's joints and motion patterns from each frame. The input data is the saved video file, and the output data is the extracted motion data.
[0966] Step 4:
[0967] The server compares the motion data with the optimal form data. The server then matches the extracted motion data with the ideal robot motion data stored in a database. Specifically, it evaluates the degree of match for each frame and identifies discrepancies. The input data is the extracted motion data and the optimal form data, and the output data is the comparison result.
[0968] Step 5:
[0969] The server generates the corrective instructions. Based on the comparison results, the algorithm generates appropriate corrective instructions. For example, specific instructions such as "Make the arm move 5 degrees faster" are generated. The input data is the comparison result, and the output data is the corrective instructions.
[0970] Step 6:
[0971] The server analyzes the user's emotions. An emotion analysis engine (such as OpenCV or Google Cloud Vision API) is used to extract emotion data from the user's facial expressions and voice. The camera recognizes the user's face, and an emotion analysis algorithm is used to determine their emotion. Speech recognition technology is also used to analyze emotions from the tone of voice and speaking style. The input data is the user's facial expression images and voice data, and the output data is emotion data.
[0972] Step 7:
[0973] The server adjusts the correction instructions based on the emotion data. It receives the emotion data and adjusts the generated correction instructions according to the user's emotional state. For example, in the case of negative emotions, it adds an encouraging message such as "Let's try a little harder to improve efficiency!" The input data are emotion data and correction instructions, and the output data are the adjusted correction instructions.
[0974] Step 8:
[0975] The server sends the adjusted correction instruction to the terminal. The sending means sends the adjusted correction instruction and the motivation message to the terminal. The input data is the adjusted correction instruction, and the output data is the instruction and the message sent to the terminal.
[0976] Step 9:
[0977] The terminal displays the corrective instructions to the user. The terminal displays the received corrective instructions and motivational messages to the user. The user checks the instructions on the terminal and immediately understands the specific improvement methods. The input data are the instructions and messages sent to the terminal, and the output data are the display to the user.
[0978] The specific processing unit 290 transmits the result of the specific processing to the smart glasses 214. In the smart glasses 214, the control unit 46A causes the speaker 240 to output the result of the specific processing. The microphone 238 acquires audio indicating a user input regarding the result of the specific processing. The control unit 46A transmits audio data indicating the user input acquired by the microphone 238 to the data processing device 12. In the data processing device 12, the specific processing unit 290 acquires the audio data.
[0979] The data generation model 58 is a so-called generative AI (Artificial Intelligence). An example of the data generation model 58 is ChatGPT (Internet Search<URL: https: / / openai.com / blog / chatgpt> ), Gemini (Internet search <url: https: gemini.google.com ?hl="ja">) and other generation AIs. The data generation model 58 is obtained by performing deep learning on a neural network. A prompt including an instruction is input to the data generation model 58, and inference data such as voice data indicating voice, text data indicating text, and image data indicating an image is also input. The data generation model 58 performs inference on the input inference data in accordance with the instruction indicated by the prompt, and outputs the inference result in a data format such as voice data and text data. Here, inference refers to, for example, analysis, classification, prediction, and / or summarization.
[0980] In the above embodiment, an example in which the specific processing is performed by the data processing device 12 has been given, but the technology of the present disclosure is not limited to this, and the specific processing may be performed by the smart glasses 214.
[0981] [Third embodiment]
[0982] FIG. 5 shows an example of the configuration of a data processing system 310 according to the third embodiment.
[0983] 5, the data processing system 310 includes the data processing device 12 and a headset type terminal 314. An example of the data processing device 12 is a server.
[0984] The data processing device 12 includes a computer 22, a database 24, and a communication I / F 26. The computer 22 is an example of a "computer" according to the technology of the present disclosure. The computer 22 includes a processor 28, a RAM 30, and a storage 32. The processor 28, the RAM 30, and the storage 32 are connected to a bus 34. The database 24 and the communication I / F 26 are also connected to the bus 34. The communication I / F 26 is connected to a network 54. Examples of the network 54 include a WAN (Wide Area Network) and / or a LAN (Local Area Network).
[0985] The headset type terminal 314 includes a computer 36, a microphone 238, a speaker 240, a camera 42, a communication I / F 44, and a display 343. The computer 36 includes a processor 46, a RAM 48, and a storage 50. The processor 46, the RAM 48, and the storage 50 are connected to a bus 52. The microphone 238, the speaker 240, the camera 42, and the display 343 are also connected to the bus 52.
[0986] The microphone 238 receives instructions and the like from the user 20 by receiving voice uttered by the user 20. The microphone 238 captures the voice uttered by the user 20, converts the captured voice into audio data, and outputs it to the processor 46. The speaker 240 outputs audio in accordance with instructions from the processor 46.
[0987] Camera 42 is a small digital camera equipped with an optical system including a lens, aperture, and shutter, and an imaging element such as a CMOS (Complementary Metal-Oxide-Semiconductor) image sensor or a CCD (Charge Coupled Device) image sensor, and captures images of the surroundings of user 20 (for example, an imaging range defined by an angle of view equivalent to the field of vision of a typical healthy person).
[0988] The communication I / F 44 is connected to a network 54. The communication I / Fs 44 and 26 are responsible for the exchange of various information between the processor 46 and the processor 28 via the network 54. The exchange of various information between the processor 46 and the processor 28 using the communication I / Fs 44 and 26 is carried out in a secure state.
[0989] Fig. 6 shows an example of the main functions of the data processing device 12 and the headset type terminal 314. As shown in Fig. 6, in the data processing device 12, a specific process is performed by the processor 28. A specific process program 56 is stored in the storage 32.
[0990] The specific processing program 56 is an example of a "program" according to the technology of the present disclosure. The processor 28 reads the specific processing program 56 from the storage 32 and executes the read specific processing program 56 on the RAM 30. The specific processing is realized by the processor 28 operating as a specific processing unit 290 in accordance with the specific processing program 56 executed on the RAM 30.
[0991] The storage 32 stores a data generation model 58 and an emotion identification model 59. The data generation model 58 and the emotion identification model 59 are used by the identification processing unit 290.
[0992] In the headset type terminal 314, a reception output process is performed by the processor 46. A reception output program 60 is stored in the storage 50. The processor 46 reads the reception output program 60 from the storage 50 and executes the read reception output program 60 on the RAM 48. The reception output process is realized by the processor 46 operating as the control unit 46A in accordance with the reception output program 60 executed on the RAM 48.
[0993] Next, a description will be given of the identification process performed by the identification processing unit 290 of the data processing device 12. In the following description, the data processing device 12 will be referred to as the "server" and the headset type terminal 314 will be referred to as the "terminal."
[0994] The present invention relates to an AI sports trainer system for efficiently improving athletes' performance. The system includes a device for recording athletes' performance, a server for receiving and analyzing the video data, a means for generating and transmitting corrective instructions, and a terminal for displaying the instructions to the user.
[0995] System Overview
[0996] 1. Use of photography devices
[0997] Users can use their smartphones or dedicated cameras to film their pitching, batting, and other performances.
[0998] The video data is shot in high resolution, clearly recording the players' detailed movements.
[0999] 2. Uploading data
[1000] Users upload the videos they have taken to the server through a dedicated app. The app's intuitive operation makes it easy to select and upload videos.
[1001] 3. Data Receipt and Analysis
[1002] The server receives the video data and begins analysis. It uses powerful AI algorithms to model the player's movements in the video and extract each key element (e.g., shoulder position, elbow angle, etc.).
[1003] In this case, a deep learning model operates as an analytical tool to automatically generate motion data.
[1004] 4. Comparison with optimal form
[1005] The server compares the extracted motion data with optimal form data stored in a database, which is based on model athletes and their best past performances.
[1006] The comparison means relatively matches the motion data and the optimal form data for each frame and identifies specific differences.
[1007] 5. Generate corrective instructions
[1008] Based on the analysis of the discrepancies, the server generates specific corrective instructions, detailing what the player needs to correct and how.
[1009] For example, specific instructions for correcting the movement such as "pull your shoulders back a little more" are generated.
[1010] 6. Sending Corrective Instructions
[1011] The server generates and transmits corrective instructions to the terminal, which may include instructions in textual and visual formats.
[1012] In addition to text instructions, animations and illustrations may also be sent.
[1013] 7. Display of orthodontic instructions
[1014] The device receives correction instructions from the server and displays them to the user. The app has an easy-to-use interface and is designed to allow users to intuitively understand the instructions.
[1015] This allows users to learn how to correct their form in real time.
[1016] Specific examples
[1017] 1. Users film and upload pitching videos
[1018] Users record their pitching on their smartphones and upload the video to the server using a dedicated app.
[1019] 2. The server receives and analyzes the video
[1020] The server receives the video and uses a deep learning model to analyze and extract key points in the pitching motion.
[1021] 3. Comparison with optimal form
[1022] The server compares the analyzed movement data with optimal form data and detects differences such as "elbow position is too low" or "timing is too early."
[1023] 4. Generate corrective instructions
[1024] Based on the comparison results, the server generates specific corrective instructions, such as "keep your elbow 5 degrees higher on your next pitch."
[1025] 5. Sending and Displaying Instructions
[1026] The server generates and sends the correction instructions to the device, which then displays them to the user, allowing the user to receive effective feedback and instantly correct their form.
[1027] Such a system will enable players to efficiently correct their form in real time, which is expected to improve their overall performance.
[1028] The processing flow will be explained below.
[1029] Step 1:
[1030] Users use a smartphone or dedicated camera to film their own sports performance (pitching, batting, etc.) After finishing filming, they launch the app and select the video.
[1031] Step 2:
[1032] The user uploads the video they took using the app to the server. The app sends the user ID and the video data together, and confirms that the server has received the video.
[1033] Step 3:
[1034] The server receives the uploaded video data and stores it in a database. The video file is linked to a user ID and can be accessed later for analysis.
[1035] Step 4:
[1036] The server retrieves the stored video data and uses deep learning to analyze the player's movements in the video. Specifically, an AI algorithm analyzes each frame and extracts the positions of the player's joints and body parts.
[1037] Step 5:
[1038] The server compares the extracted motion data with the optimal form data, using ideal form data stored in a database and matching the current motion data point by point.
[1039] Step 6:
[1040] The server analyzes the comparison results and identifies specific differences, such as shoulders being too high or elbows being at an incorrect angle.
[1041] Step 7:
[1042] The server generates specific corrective instructions based on an analysis of the differences, which are specific and detailed instructions on how to modify the behavior.
[1043] Step 8:
[1044] The server sends the generated correction instructions to the terminal, which may be in text format or, if necessary, in visual format (images or animations).
[1045] Step 9:
[1046] The device receives the correction instructions from the server and displays them to the user. The app provides an interface that allows the user to easily check the instructions and intuitively understand specific improvement methods.
[1047] Step 10:
[1048] Users can correct their form based on the correction instructions they receive through the app, and then film and upload the video again to check the effectiveness of their corrections.
[1049] In this way, each step works in tandem to create a system that efficiently analyzes athletes' performance and provides specific methods for improvement.
[1050] Example 1
[1051] Next, a description will be given of Example 1. In the following description, the data processing device 12 will be referred to as a "server" and the headset type terminal 314 will be referred to as a "terminal."
[1052] With conventional sports training systems, it was difficult to analyze athletes' movements and provide instructions for correcting their form quickly and efficiently. Furthermore, the process of uploading and analyzing video data was complicated, making it difficult for users to operate. Furthermore, conventional systems had a low level of expressiveness in providing correction instructions, making it difficult for users to intuitively understand them.
[1053] The specific processing by the specific processing unit 290 of the data processing device 12 in the first embodiment is realized by the following means.
[1054] In this invention, the server includes a device means for filming a player's performance, a server means for receiving the video filmed by the device, an analysis means for analyzing the video data received by the server and extracting the player's motion data, a comparison means for comparing the extracted motion data with optimal form data, an instruction generation means for generating specific correction instructions based on the differences obtained by the comparison means, a transmission means for transmitting the correction instructions generated by the instruction generation means to a terminal, a display means for displaying the correction instructions received by the terminal to a user, a filming means for filming video data at high resolution and recording the player's detailed movements, a means for easily uploading the video to the server via a dedicated app, a means for analyzing the player's movements in the video using a deep learning model and automatically generating movement data, and a means for sending correction instructions in text or visual format and displaying them in a format that the user can intuitively understand. This makes it possible to quickly and efficiently provide a player's motion analysis and form correction instructions, and to provide a training environment that is easy for users to operate and understand.
[1055] "Athlete" means a person engaged in sports activities and who participates in training and competition.
[1056] "Performance" refers to the overall sporting movements performed by athletes, including all physical movements related to the competition.
[1057] "Device" refers to equipment used to record athletes' performances, including smartphones and dedicated cameras.
[1058] A "server" is a computer system that receives video data sent from devices via a network and analyzes and processes the data.
[1059] "Analysis means" refers to the technology and algorithms used to process video data and extract and analyze player movement data.
[1060] "Movement data" is specific information about the player's movements extracted by the analysis means, and includes the positions and angles of the joints.
[1061] The "comparison means" refers to the technology and algorithm that compares the motion data with the pre-established optimal form data.
[1062] "Optimal form data" refers to ideal movement information set based on model athletes and their best past performances.
[1063] The "instruction generation means" refers to the technology and algorithms that generate corrective instructions for the player based on the analysis results of the movement data.
[1064] The "transmission means" is a communication technology for transmitting the generated correction instructions to the user's terminal.
[1065] A "terminal" is a device that receives correction instructions and displays them to the user, and includes a smartphone, tablet, etc.
[1066] "Display means" refers to the technology and interface for visually displaying to the user the corrective instructions received at the terminal.
[1067] "High resolution" refers to a level of image quality where the video data is highly detailed and the players' movements can be clearly recorded.
[1068] The "dedicated app" is application software that allows users to easily operate and upload video data.
[1069] A "deep learning model" is an artificial intelligence model that uses deep learning technology to analyze the movements of players in video and automatically generate movement data.
[1070] "Text format" refers to a format in which correction instructions are expressed as text.
[1071] "Visual format" refers to a format in which corrective instructions are conveyed through visual representations such as diagrams or animations.
[1072] This invention relates to an AI sports trainer system for efficiently improving athletes' performance. The system includes a device for recording athletes' performance, a server for receiving and analyzing the video data, a means for generating and transmitting corrective instructions, and a terminal for displaying the instructions to the user.
[1073] The system consists of the following main components:
[1074] 1. Device
[1075] The devices, which can include smartphones or dedicated cameras, are used to capture high-resolution footage of athletes' performances. For example, iPhones and dedicated sports cameras have high-resolution capabilities that allow for clear recording of athletes' detailed movements.
[1076] 2. Server
[1077] The server is a computer system that receives and analyzes video data, and often uses powerful cloud infrastructure such as Amazon Web Services (AWS) or Google Cloud Platform (GCP).
[1078] The server also uses deep learning frameworks such as TensorFlow and PyTorch to run deep learning models to analyze the player's movements in the received video.
[1079] As a specific example, a user can record their batting movements using an iPhone and then upload the video to a server on AWS via a dedicated app.
[1080] 3. Analysis method
[1081] The analysis method is integrated into the server and is a technology for extracting movement data from video data. It uses a deep learning model to analyze each key point of a player's movement.
[1082] For example, a TensorFlow model analyzes each frame in a video and extracts key data points such as the position of a player's shoulders and the angle of their elbows.
[1083] 4. Means of comparison
[1084] The server also includes a comparison mechanism that compares the extracted motion data with optimal form data, using the motion data of past best performances and model athletes.
[1085] For example, the server detects specific differences such as "elbows are positioned too low."
[1086] 5. Instruction generation means
[1087] The instruction generation means generates corrective instructions based on the differences obtained from the analysis and comparison means. The generated instructions are specific and show the player what parts to correct and how.
[1088] For example, specific instructions such as "Keep your elbow 5 degrees higher on your next pitch" are generated.
[1089] 6. Means of transmission
[1090] The transmission means is a technique for transmitting the server-generated correction instructions to the user's terminal using the secure HTTPS protocol.
[1091] 7. Terminal
[1092] The terminal is a device such as a smartphone or tablet that displays the orthodontic instructions received from the server to the user. A dedicated app is installed on the terminal, and it is designed to allow the user to intuitively understand the instructions.
[1093] For example, the smartphone app screen displays text instructions such as "Pull your shoulders back a little more," along with animations and illustrations.
[1094] These components enable the system to provide an environment for efficiently and quickly improving athlete performance.
[1095] Specific examples
[1096] 1. Users film and upload their performances
[1097] For example, a user can record their pitching motion using an iPhone and then upload the video to a server on AWS via a dedicated app.
[1098] Example prompt: "Upload a pitching video"
[1099] 2. The server receives and analyzes the video
[1100] The server receives the video data and uses TensorFlow to analyze and extract each key point of the player (e.g., shoulder position, elbow angle, etc.).
[1101] Example prompt: "Start behavior analysis using deep learning models."
[1102] 3. Comparison of motion data and optimal form data
[1103] The server compares the movement data with optimal form data in a database and detects specific differences, such as "elbow position too low."
[1104] Example prompt: "Your elbow position is 15 degrees off from optimal."
[1105] 4. Generate and send correction instructions
[1106] The server generates specific corrective instructions, such as "keep your elbow 5 degrees higher on your next pitch," and sends them to the user's device using the HTTPS protocol.
[1107] Example prompt: "Keep your elbow 5 degrees higher on your next pitch."
[1108] 5. The device will display correction instructions.
[1109] The user's smartphone app receives the correction instructions, displays "Keep your elbow 5 degrees higher on your next pitch," and plays an animation showing the movement.
[1110] Example prompt: "Push your shoulders back a little more."
[1111] In this way, the system provides a comprehensive solution for efficiently improving athletes' performance in sports.
[1112] The flow of the identification process in the first embodiment will be described with reference to FIG.
[1113] Step 1:
[1114] Users film their performances
[1115] Specifically, a user uses a smartphone or a dedicated camera to capture high-resolution footage of a sports performance (e.g., batting or pitching). The input is actual footage of the player's movements, and the output is a high-resolution video file.
[1116] Example: "User records batting action on iPhone"
[1117] Step 2:
[1118] User uploads video to server
[1119] Users upload the recorded video to the server through a dedicated app. They launch the app, select the recorded video file, and click the "Upload" button. The input is the recorded video file, and the output is the video data saved on the server.
[1120] Example: "Use the dedicated app to upload a video of your batting practice."
[1121] Step 3:
[1122] Server receives video
[1123] The server receives videos uploaded by users. The server saves the video files in cloud storage (e.g., AWS S3 bucket). The input is the uploaded video data, and the output is the video file saved in the storage.
[1124] Example: "A server on AWS receives and stores video data."
[1125] Step 4:
[1126] The server analyzes the video
[1127] The server uses a deep learning model (e.g., TensorFlow model) to analyze the video data and extract key motion data for each video frame. The input is the saved video file, and the output is the extracted keypoint data (e.g., shoulder position, elbow angle, etc.).
[1128] Example: "The server analyzes the motion using a TensorFlow model and extracts the shoulder position and elbow angle."
[1129] Step 5:
[1130] The server compares the behavior data with the best form data
[1131] The server compares the analyzed motion data with the ideal form data in a database. The comparison method is to match each frame of motion data with the ideal data and identify discrepancies. The input is the analyzed motion data and the ideal form data, and the output is specific discrepancy information (e.g., elbow angle is 15 degrees lower).
[1132] Example: "Compare analyzed behavior data with optimal form data to detect discrepancies."
[1133] Step 6:
[1134] Server generates corrective instructions
[1135] The server generates correction instructions based on the specific difference information. The instructions are generated using a natural language generation model, and are generated as specific content that is easy for users to understand. The input is the difference information, and the output is specific correction instructions (e.g., "Lift your elbow 15 degrees").
[1136] Example: Generate specific instructions such as "Keep your elbow 15 degrees higher on your next pitch."
[1137] Step 7:
[1138] The server sends correction instructions to the device
[1139] The server sends the generated correction instructions to the user's device. The transmission is performed using secure HTTPS communication. The input is the generated correction instructions, and the output is a transmission completion notification to the user's device.
[1140] Example: "The server sends correction instructions to the user's device using HTTPS."
[1141] Step 8:
[1142] The device displays correction instructions
[1143] The device displays the received correction instructions to the user. The app uses an intuitive interface to show instructions using text and animation. The input is the received correction instructions, and the output is the displayed instruction information.
[1144] Example: "The user's smartphone displays correction instructions and animations demonstrating the movements."
[1145] Through these concrete steps, the system will efficiently improve athletes' performance.
[1146] (Application example 1)
[1147] Next, a description will be given of Application Example 1. In the following description, the data processing device 12 will be referred to as a "server" and the headset type terminal 314 will be referred to as a "terminal."
[1148] Current sports training systems and industrial robot control systems have difficulty optimizing the movement efficiency of athletes and work machines in real time. Furthermore, there is a lack of means to improve the accuracy of robot movements while integrating them with human motion analysis technology. Therefore, there is a need to simultaneously improve the performance of athletes and the efficiency of robots in industrial settings.
[1149] The specific processing by the specific processing unit 290 of the data processing device 12 in the application example 1 is realized by the following means.
[1150] In this invention, the server includes a device for filming the performance of athletes and industrial machinery, an analysis means for receiving and analyzing the video data and extracting movement data, a comparison means for comparing the movement data with optimal form data or work efficiency data, and an instruction generation means for generating specific correction instructions based on the differences. This allows for real-time analysis of the movements of athletes and industrial machinery, providing efficient feedback, and enabling performance improvement.
[1151] An "athlete" is a person who participates in a sport or particular performance competition with the aim of improving their performance.
[1152] "Industrial machinery" refers to robots and machines in general used in factories and production facilities, and is a device used to efficiently perform specific tasks.
[1153] "Motion data" is data that records in detail the movements of athletes and industrial machinery, and contains information that includes important key points extracted through analysis.
[1154] "Analysis means" refers to an algorithm or system for receiving video data, analyzing it, and extracting motion data.
[1155] The "comparison means" is a system or method that has the function of comparing the motion data extracted by the analysis means with optimal form data or work efficiency data.
[1156] The "instruction generation means" refers to an algorithm or system for generating specific corrective instructions based on the differences obtained by the comparison means.
[1157] The "transmission means" is a communication means having a function for transmitting the generated correction instruction to the terminal.
[1158] The "display means" is an interface or device for visually displaying to the user the corrective instructions received at the terminal.
[1159] The "control means" is a system or algorithm for applying corrective instructions received by the terminal to the robot and correcting its behavior in real time.
[1160] The present invention relates to a system for efficiently optimizing the movements of athletes and industrial machines. DETAILED DESCRIPTION OF THE INVENTION ... Hereinafter, specific embodiments of the present invention will be described.
[1161] definition
[1162] An "athlete" is a person who participates in a sport or particular performance competition with the aim of improving their performance.
[1163] "Industrial machinery" refers to robots and machines in general used in factories and production facilities, and is a device used to efficiently perform specific tasks.
[1164] "Motion data" is data that records in detail the movements of athletes and industrial machinery, and contains information that includes important key points extracted through analysis.
[1165] "Analysis means" refers to an algorithm or system for receiving video data, analyzing it, and extracting motion data.
[1166] The "comparison means" is a system or method that has the function of comparing the motion data extracted by the analysis means with optimal form data or work efficiency data.
[1167] The "instruction generation means" refers to an algorithm or system for generating specific corrective instructions based on the differences obtained by the comparison means.
[1168] The "transmission means" is a communication means having a function for transmitting the generated correction instruction to the terminal.
[1169] The "display means" is an interface or device for visually displaying to the user the corrective instructions received at the terminal.
[1170] The "control means" is a system or algorithm for applying corrective instructions received by the terminal to the robot and correcting its behavior in real time.
[1171] System Overview
[1172] 1. Use of photography devices
[1173] The server records the movements of athletes and industrial machines using a device that records the movements of the athletes and industrial machines, such as a smartphone or a dedicated high-resolution camera.
[1174] 2. Uploading video data
[1175] Users upload the videos they have taken to the server using a dedicated app. This allows for intuitive operation, making it easy to select and upload videos.
[1176] 3. Data Receipt and Analysis
[1177] The server analyzes the received video data and extracts motion data of athletes and industrial machines using deep learning models such as TensorFlow.
[1178] 4. Comparison with optimal form
[1179] The server compares the extracted behavioral data with optimal form or performance data, which identifies specific deviations between the behavior and optimal performance.
[1180] 5. Generate corrective instructions
[1181] The server generates specific corrective instructions based on the differences obtained by the comparison means, which instruct the athlete or industrial machine in detail how and where to correct the athlete or industrial machine.
[1182] 6. Sending Corrective Instructions
[1183] The server sends the generated correction instructions to the terminal, which may include instructions in textual and visual form.
[1184] 7. Display of orthodontic instructions
[1185] The device receives correction instructions from the server and displays them to the user. The app has an easy-to-use interface that allows for intuitive understanding.
[1186] 8. Real-time behavior correction
[1187] Using the control means, the terminal is able to apply the received corrective instructions to the robot to modify its behavior in real time.
[1188] Hardware and software used
[1189] Hardware: Smartphones, dedicated cameras, industrial robots
[1190] Software: TensorFlow (deep learning model), dedicated app, server
[1191] Adding concrete examples and prompt sentence examples
[1192] For example, to optimize the behavior of a robot used in a factory installing parts, the following prompts could be input to a generative AI model:
[1193] plaintext
[1194] Analyze the robot arm's motion using video and compare it with the optimal arm motion pattern to generate corrective instructions. The optimal arm motion pattern includes the arm height, movement method, and speed. For example, if the arm is raised too high, issue the instruction "Lower the arm."
[1195] This will enable the operational efficiency of robots within factories to be optimized in real time, which is expected to improve productivity.
[1196] The flow of the specific processing in the application example 1 will be described with reference to FIG.
[1197] Step 1:
[1198] The user films the operation of industrial machinery in operation in a factory. Using a smartphone or a dedicated camera, the operation is recorded as high-resolution video data. The filming is done so that the movement of each part of the industrial machinery is clearly visible.
[1199] Input: Video taken with a smartphone or dedicated camera
[1200] Output: High-resolution video data
[1201] Step 2:
[1202] Users upload the video data they have taken to the server using a dedicated app. The app's intuitive operation makes it easy to select and upload videos.
[1203] Input: High-resolution video data
[1204] Output: Video data uploaded to the server
[1205] Step 3:
[1206] The server receives the video data and begins analysis. Using deep learning models such as TensorFlow, the server analyzes the movements of the industrial machinery in the video and extracts key points (such as the position of the arm and the angle of movement).
[1207] Input: Video data uploaded to the server
[1208] Output: Extracted keypoint data
[1209] Step 4:
[1210] The server compares the extracted key point data with optimal motion pattern data in a database, and the comparison identifies differences between the motion data and the optimal motion pattern.
[1211] Input: Extracted keypoint data, optimal motion pattern data in the database
[1212] Output: Difference between motion data and optimal motion pattern
[1213] Step 5:
[1214] The server generates specific correction instructions based on the difference between the operational data and the optimal operational pattern. The generated instructions provide detailed instructions on how and where the industrial machinery should be corrected.
[1215] Input: Difference between motion data and optimal motion pattern
[1216] Output: Specific correction instructions
[1217] Step 6:
[1218] The server transmits the generated correction instructions to the terminal, which may include instructions in textual and visual form.
[1219] Input: Specific correction instructions
[1220] Output: Corrective instructions sent to the terminal
[1221] Step 7:
[1222] The device receives correction instructions from the server and displays them to the user. The app has an easy-to-use interface and is designed to allow users to intuitively understand the instructions.
[1223] Input: Corrective instructions sent to the terminal
[1224] Output: Corrective instructions displayed to the user
[1225] Step 8:
[1226] The corrective instructions confirmed by the user are applied to the industrial machine using the control means to correct its operation in real time, thereby improving the operational efficiency of the industrial machine.
[1227] Input: User confirmed corrective instructions
[1228] Output: Real-time corrected industrial machine behavior
[1229] Furthermore, an emotion engine that estimates the user's emotion may be further combined. That is, the identification processing unit 290 may estimate the user's emotion using the emotion identification model 59, and perform identification processing using the user's emotion.
[1230] This invention relates to an AI sports trainer system for efficiently improving athletes' performance, and by combining it with an emotion engine, provides instruction according to the user's emotional state. This system includes a device for recording the athletes' performance, a server for receiving and analyzing the video data, a means for generating and transmitting corrective instructions, a terminal for displaying the instructions to the user, and the emotion engine.
[1231] System Overview
[1232] 1. Use of photography devices
[1233] Users use their smartphone or dedicated camera to film their pitching, batting, or other performances. Once they've finished filming, they launch the app and select the video.
[1234] 2. Uploading data
[1235] The user uploads the video they have taken to the server using a dedicated app. The app sends the user ID and video data together and confirms that the server has received the video.
[1236] 3. Data Receipt and Analysis
[1237] The server receives the uploaded video data and stores it in a database. The video file is linked to a user ID and can be accessed later for analysis.
[1238] 4. Movement analysis and form extraction
[1239] The server retrieves the stored video data and uses deep learning to analyze the player's movements in the video. Specifically, an AI algorithm analyzes each frame and extracts the positions of the player's joints and body parts.
[1240] 5. Comparison with optimal form
[1241] The server compares the extracted motion data with the optimal form data, using ideal form data stored in a database and matching the current motion data point by point.
[1242] 6. Generating Corrective Instructions
[1243] The server analyzes the specific differences based on the comparison results and generates specific corrective instructions based on the differences, which may be in text format or, if necessary, in visual format (images or animations).
[1244] 7. Emotion Engine Operation
[1245] The emotion engine extracts emotional data from the user's facial expressions and voice. For example, it uses a camera to recognize the user's face and an expression analysis algorithm to determine their emotion. It also uses voice recognition technology to analyze emotions from the user's tone of voice and speaking style.
[1246] 8. Use of Emotional Data
[1247] The server receives the emotion data and adjusts the generated corrective instructions, providing standard instructions if the emotion is positive, and adding encouraging or supportive messages if the emotion is negative.
[1248] Emotional data is stored historically and used to monitor training progress.
[1249] 9. Sending and Displaying Corrective Instructions
[1250] The server sends tailored corrective instructions to the terminal, which may include motivational or encouraging messages.
[1251] The device receives the correction instructions and motivational messages from the server and displays them to the user. The app provides an interface that allows the user to easily check the instructions and intuitively understand specific ways to improve.
[1252] Specific examples
[1253] 1. Users film and upload pitching videos
[1254] Users record their pitching on their smartphones and upload the video to the server using a dedicated app.
[1255] 2. The server receives and analyzes the video
[1256] The server receives the video and uses a deep learning model to analyze and extract key points in the pitching motion.
[1257] 3. Comparison with optimal form
[1258] The server compares the analyzed movement data with optimal form data and finds specific differences such as "elbow position is too low" or "timing is too early."
[1259] 4. Generate corrective instructions
[1260] Based on the comparison results, the server generates specific corrective instructions, such as "keep your elbow 5 degrees higher on your next pitch."
[1261] 5. Operation of the Emotion Engine and Use of Emotion Data
[1262] The emotion engine recognizes emotions from the user's facial expressions and voice and extracts emotion data. For example, if the user has a dissatisfied expression, the emotion data is determined to be negative.
[1263] The server receives the emotion data and adds a motivational message to the corrective instruction to reduce the negative emotion.
[1264] 6. Sending and Displaying Instructions
[1265] The server sends the generated correction instructions and motivation messages to the terminal, and the terminal displays the instructions to the user, allowing the user to get effective feedback and correct the form immediately.
[1266] In this way, this system, which combines an emotion engine, can efficiently correct athletes' form and provide mental support in real time, which is expected to improve their overall performance.
[1267] The processing flow will be explained below.
[1268] Step 1:
[1269] Users use a smartphone or dedicated camera to record their sports performance (e.g., pitching or batting). After finishing recording, users launch the app and select the video.
[1270] Step 2:
[1271] The user uploads the video they have taken using a dedicated app to the server. The app sends the user ID and video data together and confirms that the server has received the video.
[1272] Step 3:
[1273] The server receives the uploaded video data and stores it in a database. The video file is linked to a user ID and can be accessed later for analysis.
[1274] Step 4:
[1275] The server retrieves the stored video data and uses deep learning to analyze the player's movements in the video. Specifically, an AI algorithm analyzes each frame and extracts the positions of the player's joints and body parts.
[1276] Step 5:
[1277] The server compares the extracted motion data with the optimal form data, using ideal form data stored in a database and matching the current motion data point by point.
[1278] Step 6:
[1279] The server analyzes the comparison results and identifies specific differences, such as shoulders being too high or elbows being at an incorrect angle.
[1280] Step 7:
[1281] The server generates specific corrective instructions based on an analysis of the differences, which are specific and detailed instructions on how to modify the behavior.
[1282] Step 8:
[1283] The emotion engine extracts emotional data from the user's facial expressions and voice. For example, it uses a camera to recognize the user's face and an expression analysis algorithm to determine their emotion. It also uses voice recognition technology to analyze emotions from the user's tone of voice and speaking style.
[1284] Step 9:
[1285] The server receives the emotion data and adjusts the generated corrective instructions. If the emotion is positive, it issues a standard instruction, but if it is negative, it adds a message of encouragement or support. For example, an instruction such as "Keep your elbow 5 degrees higher on your next pitch" could include an encouraging message such as "You can do it!"
[1286] Step 10:
[1287] The server sends the adjusted correction instructions to the device, which may be in text format or, if necessary, in visual format (images or animations).
[1288] Step 11:
[1289] The device receives the correction instructions from the server and displays them to the user. The app has an easy-to-use interface, allowing users to easily check the instructions. Specific improvement methods and encouraging messages are also displayed.
[1290] Step 12:
[1291] Users can correct their form based on the correction instructions and encouraging messages they receive through the app, and can then film and upload the video again to see the results of their corrections.
[1292] In this way, each step works in tandem, allowing players to efficiently correct their form and receive mental support in real time, which is expected to improve their overall performance.
[1293] Example 2
[1294] Next, a description will be given of Example 2. In the following description, the data processing device 12 will be referred to as a "server" and the headset type terminal 314 will be referred to as a "terminal."
[1295] In recent years, there has been a demand for systems that can efficiently improve athletes' performance in sports training. Conventional systems analyze movements and improve form, but rarely provide guidance that takes into account the user's emotional state, resulting in problems such as a decrease in motivation and a lack of psychological support. To solve these issues, a system that can provide real-time guidance that corresponds to the user's emotional state is needed.
[1296] The specific processing by the specific processing unit 290 of the data processing device 12 in the second embodiment is realized by the following means.
[1297] In this invention, the server includes a device for filming the player's performance, a computer unit that receives the video filmed by the device, analysis means that analyzes the video data received by the computer unit and extracts the player's movement data, comparison means that compares the movement data extracted by the analysis means with optimal form data, instruction generation means that generates specific correction instructions based on the differences obtained by the comparison means, transmission means that transmits the correction instructions generated by the instruction generation means to a terminal, display means that displays the correction instructions received by the terminal to the user, an emotion recognition device that extracts emotion data from the user's facial expressions and voice, and adjustment means that adjusts the correction instructions based on the emotion data obtained from the emotion recognition device. This enables appropriate feedback and instruction tailored to the user's emotional state.
[1298] An "athlete" is a person who participates in a sport or competition and seeks to improve their performance.
[1299] "Performance" is a general term for the actions and techniques that athletes perform in sports or competitions.
[1300] "Devices" refers to equipment used to film athletes' performances, including smartphones and dedicated cameras.
[1301] "Video" refers to video data captured by the device.
[1302] The "computer section" is a section that includes a computer system for receiving and processing video data.
[1303] "Analysis means" refers to a function for analyzing the video data received by the computer unit and extracting the movement data of the players.
[1304] "Movement data" refers to information extracted by analytical means, such as the position of a player's joints and body, and the timing of their movements.
[1305] "Form data" is data that serves as a model of ideal behavior or technology and serves as a standard for comparison.
[1306] "Comparison means" refers to functionality for comparing motion data with optimal form data and identifying differences.
[1307] "Difference" refers to the specific differences observed between the behavior data and the form data.
[1308] "Corrective instructions" refer to instructions given to the athlete to improve their movements based on the differences obtained by the comparison means.
[1309] The "instruction generating means" refers to a function for generating specific corrective instructions based on the difference.
[1310] The "transmitting means" refers to a function for transmitting the generated correction instruction to the terminal.
[1311] A "terminal" is a device used by a user to check correction instructions, and includes a smartphone, tablet, etc.
[1312] The "display means" refers to a function for displaying the received correction instructions to the user at the terminal.
[1313] An "emotion recognition device" refers to a device for extracting emotional data from a user's facial expressions and voice.
[1314] "Emotion data" refers to information that indicates the emotional state of a user extracted by an emotion recognition device.
[1315] The "adjustment means" refers to a function for adjusting the correction instruction based on the emotion data to match the user's emotional state.
[1316] The present invention relates to an AI sports trainer system for efficiently improving athletes' performance, and by combining it with an emotion engine, the system provides instruction according to the user's emotional state.
[1317] The system includes the following elements:
[1318] 1. Imaging equipment
[1319] Users use a smartphone or dedicated camera to film performances such as pitching and batting, and this device can capture high-resolution images.
[1320] 2. Upload a video
[1321] The user uploads the video they have taken to the server using a dedicated app. This app sends the user ID and video data together and confirms that the server has received the video.
[1322] 3. Receiving and storing video data
[1323] The server receives the uploaded video data and stores it in a database. The video files are linked to user IDs and can be accessed for analysis.
[1324] 4. Motion analysis
[1325] The server retrieves the stored video data and uses deep learning to analyze the player's movements in the video. Specifically, it analyzes each frame and extracts the positions of the player's joints and body parts.
[1326] 5. Comparison with optimal form
[1327] The server compares the analyzed movement data with ideal form data stored in a database, checking points such as elbow angle and weight transfer.
[1328] 6. Generating Corrective Instructions
[1329] Based on the comparison results, the server analyzes the specific differences and generates corrective instructions, such as "keep your elbow 5 degrees higher on your next pitch."
[1330] 7. Emotion recognition
[1331] The emotion engine uses a camera to recognize the user's facial expressions, and uses a facial expression analysis algorithm to determine the user's emotions. It also uses voice recognition technology to analyze emotions from the user's tone of voice and speaking style.
[1332] 8. Use of Emotional Data
[1333] The server receives the emotion data sent from the emotion engine and adjusts the corrective instructions: if the user is in a positive emotional state, it provides normal instructions, and if the user is in a negative emotional state, it adds encouraging or supportive messages.
[1334] 9. Sending and Displaying Corrective Instructions
[1335] The server sends the generated correction instructions and motivation messages to the terminal.
[1336] The terminal displays the correction instructions and motivational messages received from the server to the user, who then works to correct the form.
[1337] Specific examples
[1338] 1. Users film and upload pitching videos
[1339] Users record their pitching on their smartphones and upload the video to the server using a dedicated app.
[1340] 2. The server receives and analyzes the video
[1341] The server receives the video and uses a deep learning model to analyze each key point of the pitching motion.
[1342] 3. Comparison with optimal form
[1343] The server compares the analyzed movement data with ideal form data and finds specific differences such as "elbow position is too low" or "timing is too early."
[1344] 4. Generate corrective instructions
[1345] The server generates specific corrective instructions such as "keep your elbow 5 degrees higher on your next pitch."
[1346] 5. Operation of the Emotion Engine and Use of Emotion Data
[1347] The emotion engine recognizes emotions from the user's facial expressions and voice, and if negative emotions are detected, the server adds a motivational message.
[1348] 6. Sending and Displaying Instructions
[1349] The server sends the generated correction instructions and motivation messages to the terminal, and the terminal displays the instructions to the user, allowing the user to get effective feedback and correct the form immediately.
[1350] Prompt Sentence Examples
[1351] "Upload a pitching video and analyze elbow position and timing. If emotions are negative, generate improvement instructions with encouraging messages."
[1352] The flow of the identification process in the second embodiment will be described with reference to FIG.
[1353] Step 1:
[1354] Users can film sports performances
[1355] Input: Smartphone or dedicated camera
[1356] How it works: A user records their own sports performance (e.g., pitching or batting) in high resolution, which allows the user to obtain video data.
[1357] Output: Recorded video file
[1358] Step 2:
[1359] User uploads video
[1360] Input: Recorded video file, dedicated app
[1361] How it works: The user launches the app, selects the video they have taken, and uploads it to the server. At this time, the app sends the user ID and the video data together.
[1362] Output: Video data sent to the server
[1363] Step 3:
[1364] The server receives and stores video data
[1365] Input: User ID, video data
[1366] Operation: The server receives the uploaded video data, checks the data integrity, and stores it in the database if there are no problems.
[1367] Output: Video data stored in a database
[1368] Step 4:
[1369] The server retrieves and analyzes the video data
[1370] Input: Video data stored in a database
[1371] Movement: The server retrieves the video data and analyzes each frame using a deep learning algorithm. By extracting the positions of the player's joints and body parts, various movement data is generated.
[1372] Output: Extracted behavioral data
[1373] Step 5:
[1374] Comparison with server-optimized form
[1375] Input: extracted behavioral data, ideal form data
[1376] Movement: The server compares the analyzed movement data with the ideal form data stored in a database, matching each point in detail, such as elbow angle and body timing.
[1377] Output: Difference between behavior data and form data
[1378] Step 6:
[1379] Server generates corrective instructions
[1380] Input: Differences between behavioral data and form data
[1381] Action: The server generates specific corrective instructions based on the discrepancy. A typical example might be, "Keep your elbow 5 degrees higher on your next pitch." Visual feedback (images or animations) is also generated if necessary.
[1382] Output: Generated correction instructions
[1383] Step 7:
[1384] Emotion Engine Operation
[1385] Input: Camera video, audio data
[1386] How it works: The emotion engine uses a camera to recognize the user's facial expressions, and uses a facial expression analysis algorithm to determine the user's emotions. It also uses voice recognition technology to analyze emotions from the user's tone of voice and speaking style.
[1387] Output: Extracted emotion data
[1388] Step 8:
[1389] The server uses emotion data to adjust correction instructions.
[1390] Input: Generated correction instructions, extracted emotion data
[1391] How it works: The server receives the emotion data and adjusts the corrective instructions depending on the user's emotional state: normal instructions for positive emotions, and encouraging or supportive messages for negative emotions.
[1392] Output: Adjusted straightening instructions
[1393] Step 9:
[1394] The server sends correction instructions to the terminal
[1395] Input: Adjusted orthodontic instructions
[1396] Operation: The server sends the generated correction instructions and motivation messages to the device. This transmission process is low-latency.
[1397] Output: Corrective instructions and motivation messages received by the device
[1398] Step 10:
[1399] The device displays correction instructions and messages
[1400] Input: Corrective instructions and motivation messages received by the device
[1401] How it works: The device displays correction instructions and motivational messages to the user. Through the app interface, users can intuitively understand specific ways to improve and immediately start correcting their form.
[1402] Output: Improvement instructions and motivational messages seen by the user
[1403] (Application example 2)
[1404] Next, a description will be given of Application Example 2. In the following description, the data processing device 12 will be referred to as a "server" and the headset type terminal 314 will be referred to as a "terminal."
[1405] In conventional factory robot operation, there was a lack of guidance that took into account the optimization of robot operations and the emotional state of the operator. This made it difficult to provide effective corrections when the robot's performance deteriorated, often leading to stress and a decline in operator motivation. In particular, there was a need to simultaneously improve the robot's operational efficiency and the operator's mental health.
[1406] The specific processing by the specific processing unit 290 of the data processing device 12 in Application Example 2 is realized by the following means. In this invention, the server includes a device for filming the robot's movements, a means for receiving the video, a means for analyzing the video data and extracting movement data, a means for comparing the movement data with optimal form data, and a means for generating specific correction instructions. This enables optimization of the robot's movements and effective instruction that takes into account the operator's emotional state.
[1407] A "device" is an electronic device used to capture the action of interest.
[1408] A "server" is a computer system that receives, stores, and analyzes data over a network.
[1409] The "analysis means" refers to algorithms or software for extracting motion data from received video data.
[1410] The "comparison means" is an algorithm or software that has the functionality to compare extracted motion data with optimal form data.
[1411] The "instruction generating means" is an algorithm or software for generating specific correction instructions based on the results of the comparison means.
[1412] The "transmitting means" is a component having a function for transmitting the generated correction instruction to the terminal.
[1413] The "display means" refers to a device or software for visually displaying to the user the correction instructions received at the terminal.
[1414] "Emotion analysis means" refers to algorithms or software for extracting emotional data from a user's facial expressions and voice.
[1415] The "adjustment means" is software having a function for appropriately adjusting correction instructions based on the emotion data acquired by the emotion analysis means.
[1416] "Robot operator" refers to a worker who operates a robot within a factory.
[1417] System Configuration
[1418] This invention is an AI system for improving the performance of factory robot operations, providing optimal instructions by filming and analyzing the robot's movements and achieving effective feedback by taking into account the emotional state of the operator. The system consists of the following components:
[1419] shooting device
[1420] The user (robot operator) uses a smartphone or dedicated camera to record the factory robot's work movements. The camera has high resolution and can clearly capture even the smallest movements. Once the recording is complete, the user launches the dedicated app, selects the video, and uploads it to the server.
[1421] server
[1422] The server stores the received video data in a database and analyzes it using deep learning. Specifically, it uses the following methods:
[1423] Video analysis method: Using TensorFlow and PyTorch, motion data is extracted from the video data. Each frame is analyzed to extract the robot's joints and motion patterns.
[1424] Comparison method: The analyzed motion data is compared with the optimal form data, which represents the ideal robot motion and is stored in a database.
[1425] Instruction generation means: Based on the comparison results, appropriate corrective instructions are generated. For example, specific instructions such as "Make the arm move 5 degrees faster" are generated.
[1426] Emotion analysis method: Using OpenCV and Google Cloud Vision API, emotional data is extracted from the robot operator's facial expressions and voice. The operator's face is recognized by a camera, and emotions are determined using an expression analysis algorithm. Speech recognition technology is also used to analyze emotions from the operator's tone of voice and speaking style.
[1427] Adjustment means: Receives emotion data and adjusts the generated corrective instructions. In the case of positive emotions, it issues normal instructions, and in the case of negative emotions, it adds encouraging messages.
[1428] Terminal
[1429] The terminal is used to provide feedback to the operator, displaying corrective instructions and motivational messages received from the server to help the operator comply with instructions promptly.
[1430] Specific examples
[1431] Consider a situation in a factory where a robot arm is not moving smoothly, resulting in a drop in efficiency. An operator records the robot's movements on a smartphone and uploads the video to a server using a dedicated app. The server analyzes the video, identifies the problem area, and generates specific instructions such as "Make the arm move five degrees faster." At the same time, if the operator looks dissatisfied, the emotion analysis means determines this to be a negative emotion and provides an encouraging message such as "Try a little harder to improve efficiency!" These instructions and messages are displayed on the terminal, allowing the operator to respond immediately.
[1432] Prompt Sentence Examples
[1433] An example of a prompt would be:
[1434] "Today, the robot's arm movements were slow, causing work efficiency to drop. Please film the robot's arm movements and provide appropriate instructions for improvement based on the analysis results."
[1435] As described above, the system of the present invention can simultaneously optimize the operation of factory robots and provide psychological support to operators, thereby contributing to improving overall work efficiency.
[1436] The flow of the specific processing in the application example 2 will be described with reference to FIG.
[1437] Step 1:
[1438] The user films the robot's movements. The user uses a smartphone or a dedicated camera to record the robot's movements, making sure to clearly capture the operating environment and specific movements. The input data is the filmed video. The output data is a video file.
[1439] Step 2:
[1440] The server receives the video. The user uses a dedicated app to upload the video they have taken to the server. The app sends the user ID and video data together, and the server receives the video. The input data is the uploaded video file and user ID, and the output data is the video file saved on the server.
[1441] Step 3:
[1442] The server analyzes the video and extracts motion data. Using a deep learning model (such as TensorFlow or PyTorch), it analyzes the robot's motion in the video frame by frame. It extracts the robot's joints and motion patterns from each frame. The input data is the saved video file, and the output data is the extracted motion data.
[1443] Step 4:
[1444] The server compares the motion data with the optimal form data. The server then matches the extracted motion data with the ideal robot motion data stored in a database. Specifically, it evaluates the degree of match for each frame and identifies discrepancies. The input data is the extracted motion data and the optimal form data, and the output data is the comparison result.
[1445] Step 5:
[1446] The server generates the corrective instructions. Based on the comparison results, the algorithm generates appropriate corrective instructions. For example, specific instructions such as "Make the arm move 5 degrees faster" are generated. The input data is the comparison result, and the output data is the corrective instructions.
[1447] Step 6:
[1448] The server analyzes the user's emotions. An emotion analysis engine (such as OpenCV or Google Cloud Vision API) is used to extract emotion data from the user's facial expressions and voice. The camera recognizes the user's face, and an emotion analysis algorithm is used to determine their emotion. Speech recognition technology is also used to analyze emotions from the tone of voice and speaking style. The input data is the user's facial expression images and voice data, and the output data is emotion data.
[1449] Step 7:
[1450] The server adjusts the correction instructions based on the emotion data. It receives the emotion data and adjusts the generated correction instructions according to the user's emotional state. For example, in the case of negative emotions, it adds an encouraging message such as "Let's try a little harder to improve efficiency!" The input data are emotion data and correction instructions, and the output data are the adjusted correction instructions.
[1451] Step 8:
[1452] The server sends the adjusted correction instruction to the terminal. The sending means sends the adjusted correction instruction and the motivation message to the terminal. The input data is the adjusted correction instruction, and the output data is the instruction and the message sent to the terminal.
[1453] Step 9:
[1454] The terminal displays the corrective instructions to the user. The terminal displays the received corrective instructions and motivational messages to the user. The user checks the instructions on the terminal and immediately understands the specific improvement methods. The input data are the instructions and messages sent to the terminal, and the output data are the display to the user.
[1455] The specific processing unit 290 transmits the result of the specific processing to the headset type terminal 314. In the headset type terminal 314, the control unit 46A causes the speaker 240 and the display 343 to output the result of the specific processing. The microphone 238 acquires audio indicating a user input regarding the result of the specific processing. The control unit 46A transmits audio data indicating the user input acquired by the microphone 238 to the data processing device 12. In the data processing device 12, the specific processing unit 290 acquires the audio data.
[1456] The data generation model 58 is a so-called generative AI (Artificial Intelligence). An example of the data generation model 58 is ChatGPT (Internet Search<URL: https: / / openai.com / blog / chatgpt> ), Gemini (Internet search <url: https: gemini.google.com ?hl="ja">) and other generation AIs. The data generation model 58 is obtained by performing deep learning on a neural network. A prompt including an instruction is input to the data generation model 58, and inference data such as voice data indicating voice, text data indicating text, and image data indicating an image is also input. The data generation model 58 performs inference on the input inference data in accordance with the instruction indicated by the prompt, and outputs the inference result in a data format such as voice data and text data. Here, inference refers to, for example, analysis, classification, prediction, and / or summarization.
[1457] In the above embodiment, an example was given in which the specific processing is performed by the data processing device 12, but the technology of the present disclosure is not limited to this, and the specific processing may be performed by the headset type terminal 314.
[1458] [Fourth embodiment]
[1459] FIG. 7 shows an example of the configuration of a data processing system 410 according to the fourth embodiment.
[1460] 7, a data processing system 410 includes a data processing device 12 and a robot 414. An example of the data processing device 12 is a server.
[1461] The data processing device 12 includes a computer 22, a database 24, and a communication I / F 26. The computer 22 is an example of a "computer" according to the technology of the present disclosure. The computer 22 includes a processor 28, a RAM 30, and a storage 32. The processor 28, the RAM 30, and the storage 32 are connected to a bus 34. The database 24 and the communication I / F 26 are also connected to the bus 34. The communication I / F 26 is connected to a network 54. Examples of the network 54 include a WAN (Wide Area Network) and / or a LAN (Local Area Network).
[1462] The robot 414 includes a computer 36, a microphone 238, a speaker 240, a camera 42, a communication I / F 44, and a control target 443. The computer 36 includes a processor 46, a RAM 48, and a storage 50. The processor 46, the RAM 48, and the storage 50 are connected to a bus 52. The microphone 238, the speaker 240, the camera 42, and the control target 443 are also connected to the bus 52.
[1463] The microphone 238 receives instructions and the like from the user 20 by receiving voice uttered by the user 20. The microphone 238 captures the voice uttered by the user 20, converts the captured voice into audio data, and outputs it to the processor 46. The speaker 240 outputs audio in accordance with instructions from the processor 46.
[1464] Camera 42 is a small digital camera equipped with an optical system including a lens, aperture, and shutter, and an imaging element such as a CMOS (Complementary Metal-Oxide-Semiconductor) image sensor or a CCD (Charge Coupled Device) image sensor, and captures images of the surroundings of user 20 (for example, an imaging range defined by an angle of view equivalent to the field of vision of a typical healthy person).
[1465] The communication I / F 44 is connected to a network 54. The communication I / Fs 44 and 26 are responsible for the exchange of various information between the processor 46 and the processor 28 via the network 54. The exchange of various information between the processor 46 and the processor 28 using the communication I / Fs 44 and 26 is carried out in a secure state.
[1466] The control object 443 includes a display device, LEDs in the eyes, and motors for driving the arms, hands, and feet. The posture and gestures of the robot 414 are controlled by controlling the motors of the arms, hands, and feet. Some of the emotions of the robot 414 can be expressed by controlling these motors. In addition, the facial expressions of the robot 414 can also be expressed by controlling the light emission state of the LEDs in the eyes of the robot 414.
[1467] Fig. 8 shows an example of the main functions of the data processing device 12 and the robot 414. As shown in Fig. 8, in the data processing device 12, a specific process is performed by the processor 28. A specific process program 56 is stored in the storage 32.
[1468] The specific processing program 56 is an example of a "program" according to the technology of the present disclosure. The processor 28 reads the specific processing program 56 from the storage 32 and executes the read specific processing program 56 on the RAM 30. The specific processing is realized by the processor 28 operating as a specific processing unit 290 in accordance with the specific processing program 56 executed on the RAM 30.
[1469] The storage 32 stores a data generation model 58 and an emotion identification model 59. The data generation model 58 and the emotion identification model 59 are used by the identification processing unit 290.
[1470] In the robot 414, the processor 46 performs the reception output process. A reception output program 60 is stored in the storage 50. The processor 46 reads the reception output program 60 from the storage 50 and executes the read reception output program 60 on the RAM 48. The reception output process is realized by the processor 46 operating as the control unit 46A in accordance with the reception output program 60 executed on the RAM 48.
[1471] Next, a description will be given of the specific processing performed by the specific processing unit 290 of the data processing device 12. In the following description, the data processing device 12 will be referred to as a "server" and the robot 414 will be referred to as a "terminal."
[1472] The present invention relates to an AI sports trainer system for efficiently improving athletes' performance. The system includes a device for recording athletes' performance, a server for receiving and analyzing the video data, a means for generating and transmitting corrective instructions, and a terminal for displaying the instructions to the user.
[1473] System Overview
[1474] 1. Use of photography devices
[1475] Users can use their smartphones or dedicated cameras to film their pitching, batting, and other performances.
[1476] The video data is shot in high resolution, clearly recording the players' detailed movements.
[1477] 2. Uploading data
[1478] Users upload the videos they have taken to the server through a dedicated app. The app's intuitive operation makes it easy to select and upload videos.
[1479] 3. Data Receipt and Analysis
[1480] The server receives the video data and begins analysis. It uses powerful AI algorithms to model the player's movements in the video and extract each key element (e.g., shoulder position, elbow angle, etc.).
[1481] In this case, a deep learning model operates as an analytical tool to automatically generate motion data.
[1482] 4. Comparison with optimal form
[1483] The server compares the extracted motion data with optimal form data stored in a database, which is based on model athletes and their best past performances.
[1484] The comparison means relatively matches the motion data and the optimal form data for each frame and identifies specific differences.
[1485] 5. Generate corrective instructions
[1486] Based on the analysis of the discrepancies, the server generates specific corrective instructions, detailing what the player needs to correct and how.
[1487] For example, specific instructions for correcting the movement such as "pull your shoulders back a little more" are generated.
[1488] 6. Sending Corrective Instructions
[1489] The server generates and transmits corrective instructions to the terminal, which may include instructions in textual and visual formats.
[1490] In addition to text instructions, animations and illustrations may also be sent.
[1491] 7. Display of orthodontic instructions
[1492] The device receives correction instructions from the server and displays them to the user. The app has an easy-to-use interface and is designed to allow users to intuitively understand the instructions.
[1493] This allows users to learn how to correct their form in real time.
[1494] Specific examples
[1495] 1. Users film and upload pitching videos
[1496] Users record their pitching on their smartphones and upload the video to the server using a dedicated app.
[1497] 2. The server receives and analyzes the video
[1498] The server receives the video and uses a deep learning model to analyze and extract key points in the pitching motion.
[1499] 3. Comparison with optimal form
[1500] The server compares the analyzed movement data with optimal form data and detects differences such as "elbow position is too low" or "timing is too early."
[1501] 4. Generate corrective instructions
[1502] Based on the comparison results, the server generates specific corrective instructions, such as "keep your elbow 5 degrees higher on your next pitch."
[1503] 5. Sending and Displaying Instructions
[1504] The server generates and sends the correction instructions to the device, which then displays them to the user, allowing the user to receive effective feedback and instantly correct their form.
[1505] Such a system will enable players to efficiently correct their form in real time, which is expected to improve their overall performance.
[1506] The processing flow will be explained below.
[1507] Step 1:
[1508] Users use a smartphone or dedicated camera to film their own sports performance (pitching, batting, etc.) After finishing filming, they launch the app and select the video.
[1509] Step 2:
[1510] The user uploads the video they took using the app to the server. The app sends the user ID and the video data together, and confirms that the server has received the video.
[1511] Step 3:
[1512] The server receives the uploaded video data and stores it in a database. The video file is linked to a user ID and can be accessed later for analysis.
[1513] Step 4:
[1514] The server retrieves the stored video data and uses deep learning to analyze the player's movements in the video. Specifically, an AI algorithm analyzes each frame and extracts the positions of the player's joints and body parts.
[1515] Step 5:
[1516] The server compares the extracted motion data with the optimal form data, using ideal form data stored in a database and matching the current motion data point by point.
[1517] Step 6:
[1518] The server analyzes the comparison results and identifies specific differences, such as shoulders being too high or elbows being at an incorrect angle.
[1519] Step 7:
[1520] The server generates specific corrective instructions based on an analysis of the differences, which are specific and detailed instructions on how to modify the behavior.
[1521] Step 8:
[1522] The server sends the generated correction instructions to the terminal, which may be in text format or, if necessary, in visual format (images or animations).
[1523] Step 9:
[1524] The device receives the correction instructions from the server and displays them to the user. The app provides an interface that allows the user to easily check the instructions and intuitively understand specific improvement methods.
[1525] Step 10:
[1526] Users can correct their form based on the correction instructions they receive through the app, and then film and upload the video again to check the effectiveness of their corrections.
[1527] In this way, each step works in tandem to create a system that efficiently analyzes athletes' performance and provides specific methods for improvement.
[1528] Example 1
[1529] Next, a description will be given of Example 1. In the following description, the data processing device 12 will be referred to as a "server" and the robot 414 will be referred to as a "terminal."
[1530] With conventional sports training systems, it was difficult to analyze athletes' movements and provide instructions for correcting their form quickly and efficiently. Furthermore, the process of uploading and analyzing video data was complicated, making it difficult for users to operate. Furthermore, conventional systems had a low level of expressiveness in providing correction instructions, making it difficult for users to intuitively understand them.
[1531] The specific processing by the specific processing unit 290 of the data processing device 12 in the first embodiment is realized by the following means.
[1532] In this invention, the server includes a device means for filming a player's performance, a server means for receiving the video filmed by the device, an analysis means for analyzing the video data received by the server and extracting the player's motion data, a comparison means for comparing the extracted motion data with optimal form data, an instruction generation means for generating specific correction instructions based on the differences obtained by the comparison means, a transmission means for transmitting the correction instructions generated by the instruction generation means to a terminal, a display means for displaying the correction instructions received by the terminal to a user, a filming means for filming video data at high resolution and recording the player's detailed movements, a means for easily uploading the video to the server via a dedicated app, a means for analyzing the player's movements in the video using a deep learning model and automatically generating movement data, and a means for sending correction instructions in text or visual format and displaying them in a format that the user can intuitively understand. This makes it possible to quickly and efficiently provide a player's motion analysis and form correction instructions, and to provide a training environment that is easy for users to operate and understand.
[1533] "Athlete" means a person engaged in sports activities and who participates in training and competition.
[1534] "Performance" refers to the overall sporting movements performed by athletes, including all physical movements related to the competition.
[1535] "Device" refers to equipment used to record athletes' performances, including smartphones and dedicated cameras.
[1536] A "server" is a computer system that receives video data sent from devices via a network and analyzes and processes the data.
[1537] "Analysis means" refers to the technology and algorithms used to process video data and extract and analyze player movement data.
[1538] "Movement data" is specific information about the player's movements extracted by the analysis means, and includes the positions and angles of the joints.
[1539] The "comparison means" refers to the technology and algorithm that compares the motion data with the pre-established optimal form data.
[1540] "Optimal form data" refers to ideal movement information set based on model athletes and their best past performances.
[1541] The "instruction generation means" refers to the technology and algorithms that generate corrective instructions for the player based on the analysis results of the movement data.
[1542] The "transmission means" is a communication technology for transmitting the generated correction instructions to the user's terminal.
[1543] A "terminal" is a device that receives correction instructions and displays them to the user, and includes a smartphone, tablet, etc.
[1544] "Display means" refers to the technology and interface for visually displaying to the user the corrective instructions received at the terminal.
[1545] "High resolution" refers to a level of image quality where the video data is highly detailed and the players' movements can be clearly recorded.
[1546] The "dedicated app" is application software that allows users to easily operate and upload video data.
[1547] A "deep learning model" is an artificial intelligence model that uses deep learning technology to analyze the movements of players in video and automatically generate movement data.
[1548] "Text format" refers to a format in which correction instructions are expressed as text.
[1549] "Visual format" refers to a format in which corrective instructions are conveyed through visual representations such as diagrams or animations.
[1550] This invention relates to an AI sports trainer system for efficiently improving athletes' performance. The system includes a device for recording athletes' performance, a server for receiving and analyzing the video data, a means for generating and transmitting corrective instructions, and a terminal for displaying the instructions to the user.
[1551] The system consists of the following main components:
[1552] 1. Device
[1553] The devices, which can include smartphones or dedicated cameras, are used to capture high-resolution footage of athletes' performances. For example, iPhones and dedicated sports cameras have high-resolution capabilities that allow for clear recording of athletes' detailed movements.
[1554] 2. Server
[1555] The server is a computer system that receives and analyzes video data, and often uses powerful cloud infrastructure such as Amazon Web Services (AWS) or Google Cloud Platform (GCP).
[1556] The server also uses deep learning frameworks such as TensorFlow and PyTorch to run deep learning models to analyze the player's movements in the received video.
[1557] As a specific example, a user can record their batting movements using an iPhone and then upload the video to a server on AWS via a dedicated app.
[1558] 3. Analysis method
[1559] The analysis method is integrated into the server and is a technology for extracting movement data from video data. It uses a deep learning model to analyze each key point of a player's movement.
[1560] For example, a TensorFlow model analyzes each frame in a video and extracts key data points such as the position of a player's shoulders and the angle of their elbows.
[1561] 4. Means of comparison
[1562] The server also includes a comparison mechanism that compares the extracted motion data with optimal form data, using the motion data of past best performances and model athletes.
[1563] For example, the server detects specific differences such as "elbows are positioned too low."
[1564] 5. Instruction generation means
[1565] The instruction generation means generates corrective instructions based on the differences obtained from the analysis and comparison means. The generated instructions are specific and show the player what parts to correct and how.
[1566] For example, specific instructions such as "Keep your elbow 5 degrees higher on your next pitch" are generated.
[1567] 6. Means of transmission
[1568] The transmission means is a technique for transmitting the server-generated correction instructions to the user's terminal using the secure HTTPS protocol.
[1569] 7. Terminal
[1570] The terminal is a device such as a smartphone or tablet that displays the orthodontic instructions received from the server to the user. A dedicated app is installed on the terminal, and it is designed to allow the user to intuitively understand the instructions.
[1571] For example, the smartphone app screen displays text instructions such as "Pull your shoulders back a little more," along with animations and illustrations.
[1572] These components enable the system to provide an environment for efficiently and quickly improving athlete performance.
[1573] Specific examples
[1574] 1. Users film and upload their performances
[1575] For example, a user can record their pitching motion using an iPhone and then upload the video to a server on AWS via a dedicated app.
[1576] Example prompt: "Upload a pitching video"
[1577] 2. The server receives and analyzes the video
[1578] The server receives the video data and uses TensorFlow to analyze and extract each key point of the player (e.g., shoulder position, elbow angle, etc.).
[1579] Example prompt: "Start behavior analysis using deep learning models."
[1580] 3. Comparison of motion data and optimal form data
[1581] The server compares the movement data with optimal form data in a database and detects specific differences, such as "elbow position too low."
[1582] Example prompt: "Your elbow position is 15 degrees off from optimal."
[1583] 4. Generate and send correction instructions
[1584] The server generates specific corrective instructions, such as "keep your elbow 5 degrees higher on your next pitch," and sends them to the user's device using the HTTPS protocol.
[1585] Example prompt: "Keep your elbow 5 degrees higher on your next pitch."
[1586] 5. The device will display correction instructions.
[1587] The user's smartphone app receives the correction instructions, displays "Keep your elbow 5 degrees higher on your next pitch," and plays an animation showing the movement.
[1588] Example prompt: "Push your shoulders back a little more."
[1589] In this way, the system provides a comprehensive solution for efficiently improving athletes' performance in sports.
[1590] The flow of the identification process in the first embodiment will be described with reference to FIG.
[1591] Step 1:
[1592] Users film their performances
[1593] Specifically, a user uses a smartphone or a dedicated camera to capture high-resolution footage of a sports performance (e.g., batting or pitching). The input is actual footage of the player's movements, and the output is a high-resolution video file.
[1594] Example: "User records batting action on iPhone"
[1595] Step 2:
[1596] User uploads video to server
[1597] Users upload the recorded video to the server through a dedicated app. They launch the app, select the recorded video file, and click the "Upload" button. The input is the recorded video file, and the output is the video data saved on the server.
[1598] Example: "Use the dedicated app to upload a video of your batting practice."
[1599] Step 3:
[1600] Server receives video
[1601] The server receives videos uploaded by users. The server saves the video files in cloud storage (e.g., AWS S3 bucket). The input is the uploaded video data, and the output is the video file saved in the storage.
[1602] Example: "A server on AWS receives and stores video data."
[1603] Step 4:
[1604] The server analyzes the video
[1605] The server uses a deep learning model (e.g., TensorFlow model) to analyze the video data and extract key motion data for each video frame. The input is the saved video file, and the output is the extracted keypoint data (e.g., shoulder position, elbow angle, etc.).
[1606] Example: "The server analyzes the motion using a TensorFlow model and extracts the shoulder position and elbow angle."
[1607] Step 5:
[1608] The server compares the behavior data with the best form data
[1609] The server compares the analyzed motion data with the ideal form data in a database. The comparison method is to match each frame of motion data with the ideal data and identify discrepancies. The input is the analyzed motion data and the ideal form data, and the output is specific discrepancy information (e.g., elbow angle is 15 degrees lower).
[1610] Example: "Compare analyzed behavior data with optimal form data to detect discrepancies."
[1611] Step 6:
[1612] Server generates corrective instructions
[1613] The server generates correction instructions based on the specific difference information. The instructions are generated using a natural language generation model, and are generated as specific content that is easy for users to understand. The input is the difference information, and the output is specific correction instructions (e.g., "Lift your elbow 15 degrees").
[1614] Example: Generate specific instructions such as "Keep your elbow 15 degrees higher on your next pitch."
[1615] Step 7:
[1616] The server sends correction instructions to the device
[1617] The server sends the generated correction instructions to the user's device. The transmission is performed using secure HTTPS communication. The input is the generated correction instructions, and the output is a transmission completion notification to the user's device.
[1618] Example: "The server sends correction instructions to the user's device using HTTPS."
[1619] Step 8:
[1620] The device displays correction instructions
[1621] The device displays the received correction instructions to the user. The app uses an intuitive interface to show instructions using text and animation. The input is the received correction instructions, and the output is the displayed instruction information.
[1622] Example: "The user's smartphone displays correction instructions and animations demonstrating the movements."
[1623] Through these concrete steps, the system will efficiently improve athletes' performance.
[1624] (Application example 1)
[1625] Next, a description will be given of Application Example 1. In the following description, the data processing device 12 will be referred to as a "server" and the robot 414 will be referred to as a "terminal."
[1626] Current sports training systems and industrial robot control systems have difficulty optimizing the movement efficiency of athletes and work machines in real time. Furthermore, there is a lack of means to improve the accuracy of robot movements while integrating them with human motion analysis technology. Therefore, there is a need to simultaneously improve the performance of athletes and the efficiency of robots in industrial settings.
[1627] The specific processing by the specific processing unit 290 of the data processing device 12 in the application example 1 is realized by the following means.
[1628] In this invention, the server includes a device for filming the performance of athletes and industrial machinery, an analysis means for receiving and analyzing the video data and extracting movement data, a comparison means for comparing the movement data with optimal form data or work efficiency data, and an instruction generation means for generating specific correction instructions based on the differences. This allows for real-time analysis of the movements of athletes and industrial machinery, providing efficient feedback, and enabling performance improvement.
[1629] An "athlete" is a person who participates in a sport or particular performance competition with the aim of improving their performance.
[1630] "Industrial machinery" refers to robots and machines in general used in factories and production facilities, and is a device used to efficiently perform specific tasks.
[1631] "Motion data" is data that records in detail the movements of athletes and industrial machinery, and contains information that includes important key points extracted through analysis.
[1632] "Analysis means" refers to an algorithm or system for receiving video data, analyzing it, and extracting motion data.
[1633] The "comparison means" is a system or method that has the function of comparing the motion data extracted by the analysis means with optimal form data or work efficiency data.
[1634] The "instruction generation means" refers to an algorithm or system for generating specific corrective instructions based on the differences obtained by the comparison means.
[1635] The "transmission means" is a communication means having a function for transmitting the generated correction instruction to the terminal.
[1636] The "display means" is an interface or device for visually displaying to the user the corrective instructions received at the terminal.
[1637] The "control means" is a system or algorithm for applying corrective instructions received by the terminal to the robot and correcting its behavior in real time.
[1638] The present invention relates to a system for efficiently optimizing the movements of athletes and industrial machines. DETAILED DESCRIPTION OF THE INVENTION ... Hereinafter, specific embodiments of the present invention will be described.
[1639] definition
[1640] An "athlete" is a person who participates in a sport or particular performance competition with the aim of improving their performance.
[1641] "Industrial machinery" refers to robots and machines in general used in factories and production facilities, and is a device used to efficiently perform specific tasks.
[1642] "Motion data" is data that records in detail the movements of athletes and industrial machinery, and contains information that includes important key points extracted through analysis.
[1643] "Analysis means" refers to an algorithm or system for receiving video data, analyzing it, and extracting motion data.
[1644] The "comparison means" is a system or method that has the function of comparing the motion data extracted by the analysis means with optimal form data or work efficiency data.
[1645] The "instruction generation means" refers to an algorithm or system for generating specific corrective instructions based on the differences obtained by the comparison means.
[1646] The "transmission means" is a communication means having a function for transmitting the generated correction instruction to the terminal.
[1647] The "display means" is an interface or device for visually displaying to the user the corrective instructions received at the terminal.
[1648] The "control means" is a system or algorithm for applying corrective instructions received by the terminal to the robot and correcting its behavior in real time.
[1649] System Overview
[1650] 1. Use of photography devices
[1651] The server records the movements of athletes and industrial machines using a device that records the movements of the athletes and industrial machines, such as a smartphone or a dedicated high-resolution camera.
[1652] 2. Uploading video data
[1653] Users upload the videos they have taken to the server using a dedicated app. This allows for intuitive operation, making it easy to select and upload videos.
[1654] 3. Data Receipt and Analysis
[1655] The server analyzes the received video data and extracts motion data of athletes and industrial machines using deep learning models such as TensorFlow.
[1656] 4. Comparison with optimal form
[1657] The server compares the extracted behavioral data with optimal form or performance data, which identifies specific deviations between the behavior and optimal performance.
[1658] 5. Generate corrective instructions
[1659] The server generates specific corrective instructions based on the differences obtained by the comparison means, which instruct the athlete or industrial machine in detail how and where to correct the athlete or industrial machine.
[1660] 6. Sending Corrective Instructions
[1661] The server sends the generated correction instructions to the terminal, which may include instructions in textual and visual form.
[1662] 7. Display of orthodontic instructions
[1663] The device receives correction instructions from the server and displays them to the user. The app has an easy-to-use interface that allows for intuitive understanding.
[1664] 8. Real-time behavior correction
[1665] Using the control means, the terminal is able to apply the received corrective instructions to the robot to modify its behavior in real time.
[1666] Hardware and software used
[1667] Hardware: Smartphones, dedicated cameras, industrial robots
[1668] Software: TensorFlow (deep learning model), dedicated app, server
[1669] Adding concrete examples and prompt sentence examples
[1670] For example, to optimize the behavior of a robot used in a factory installing parts, the following prompts could be input to a generative AI model:
[1671] plaintext
[1672] Analyze the robot arm's motion using video and compare it with the optimal arm motion pattern to generate corrective instructions. The optimal arm motion pattern includes the arm height, movement method, and speed. For example, if the arm is raised too high, issue the instruction "Lower the arm."
[1673] This will enable the operational efficiency of robots within factories to be optimized in real time, which is expected to improve productivity.
[1674] The flow of the specific processing in the application example 1 will be described with reference to FIG.
[1675] Step 1:
[1676] The user films the operation of industrial machinery in operation in a factory. Using a smartphone or a dedicated camera, the operation is recorded as high-resolution video data. The filming is done so that the movement of each part of the industrial machinery is clearly visible.
[1677] Input: Video taken with a smartphone or dedicated camera
[1678] Output: High-resolution video data
[1679] Step 2:
[1680] Users upload the video data they have taken to the server using a dedicated app. The app's intuitive operation makes it easy to select and upload videos.
[1681] Input: High-resolution video data
[1682] Output: Video data uploaded to the server
[1683] Step 3:
[1684] The server receives the video data and begins analysis. Using deep learning models such as TensorFlow, the server analyzes the movements of the industrial machinery in the video and extracts key points (such as the position of the arm and the angle of movement).
[1685] Input: Video data uploaded to the server
[1686] Output: Extracted keypoint data
[1687] Step 4:
[1688] The server compares the extracted key point data with optimal motion pattern data in a database, and the comparison identifies differences between the motion data and the optimal motion pattern.
[1689] Input: Extracted keypoint data, optimal motion pattern data in the database
[1690] Output: Difference between motion data and optimal motion pattern
[1691] Step 5:
[1692] The server generates specific correction instructions based on the difference between the operational data and the optimal operational pattern. The generated instructions provide detailed instructions on how and where the industrial machinery should be corrected.
[1693] Input: Difference between motion data and optimal motion pattern
[1694] Output: Specific correction instructions
[1695] Step 6:
[1696] The server transmits the generated correction instructions to the terminal, which may include instructions in textual and visual form.
[1697] Input: Specific correction instructions
[1698] Output: Corrective instructions sent to the terminal
[1699] Step 7:
[1700] The device receives correction instructions from the server and displays them to the user. The app has an easy-to-use interface and is designed to allow users to intuitively understand the instructions.
[1701] Input: Corrective instructions sent to the terminal
[1702] Output: Corrective instructions displayed to the user
[1703] Step 8:
[1704] The corrective instructions confirmed by the user are applied to the industrial machine using the control means to correct its operation in real time, thereby improving the operational efficiency of the industrial machine.
[1705] Input: User confirmed corrective instructions
[1706] Output: Real-time corrected industrial machine behavior
[1707] Furthermore, an emotion engine that estimates the user's emotion may be further combined. That is, the identification processing unit 290 may estimate the user's emotion using the emotion identification model 59, and perform identification processing using the user's emotion.
[1708] This invention relates to an AI sports trainer system for efficiently improving athletes' performance, and by combining it with an emotion engine, provides instruction according to the user's emotional state. This system includes a device for recording the athletes' performance, a server for receiving and analyzing the video data, a means for generating and transmitting corrective instructions, a terminal for displaying the instructions to the user, and the emotion engine.
[1709] System Overview
[1710] 1. Use of photography devices
[1711] Users use their smartphone or dedicated camera to film their pitching, batting, or other performances. Once they've finished filming, they launch the app and select the video.
[1712] 2. Uploading data
[1713] The user uploads the video they have taken to the server using a dedicated app. The app sends the user ID and video data together and confirms that the server has received the video.
[1714] 3. Data Receipt and Analysis
[1715] The server receives the uploaded video data and stores it in a database. The video file is linked to a user ID and can be accessed later for analysis.
[1716] 4. Movement analysis and form extraction
[1717] The server retrieves the stored video data and uses deep learning to analyze the player's movements in the video. Specifically, an AI algorithm analyzes each frame and extracts the positions of the player's joints and body parts.
[1718] 5. Comparison with optimal form
[1719] The server compares the extracted motion data with the optimal form data, using ideal form data stored in a database and matching the current motion data point by point.
[1720] 6. Generating Corrective Instructions
[1721] The server analyzes the specific differences based on the comparison results and generates specific corrective instructions based on the differences, which may be in text format or, if necessary, in visual format (images or animations).
[1722] 7. Emotion Engine Operation
[1723] The emotion engine extracts emotional data from the user's facial expressions and voice. For example, it uses a camera to recognize the user's face and an expression analysis algorithm to determine their emotion. It also uses voice recognition technology to analyze emotions from the user's tone of voice and speaking style.
[1724] 8. Use of Emotional Data
[1725] The server receives the emotion data and adjusts the generated corrective instructions, providing standard instructions if the emotion is positive, and adding encouraging or supportive messages if the emotion is negative.
[1726] Emotional data is stored historically and used to monitor training progress.
[1727] 9. Sending and Displaying Corrective Instructions
[1728] The server sends tailored corrective instructions to the terminal, which may include motivational or encouraging messages.
[1729] The device receives the correction instructions and motivational messages from the server and displays them to the user. The app provides an interface that allows the user to easily check the instructions and intuitively understand specific ways to improve.
[1730] Specific examples
[1731] 1. Users film and upload pitching videos
[1732] Users record their pitching on their smartphones and upload the video to the server using a dedicated app.
[1733] 2. The server receives and analyzes the video
[1734] The server receives the video and uses a deep learning model to analyze and extract key points in the pitching motion.
[1735] 3. Comparison with optimal form
[1736] The server compares the analyzed movement data with optimal form data and finds specific differences such as "elbow position is too low" or "timing is too early."
[1737] 4. Generate corrective instructions
[1738] Based on the comparison results, the server generates specific corrective instructions, such as "keep your elbow 5 degrees higher on your next pitch."
[1739] 5. Operation of the Emotion Engine and Use of Emotion Data
[1740] The emotion engine recognizes emotions from the user's facial expressions and voice and extracts emotion data. For example, if the user has a dissatisfied expression, the emotion data is determined to be negative.
[1741] The server receives the emotion data and adds a motivational message to the corrective instruction to reduce the negative emotion.
[1742] 6. Sending and Displaying Instructions
[1743] The server sends the generated correction instructions and motivation messages to the terminal, and the terminal displays the instructions to the user, allowing the user to get effective feedback and correct the form immediately.
[1744] In this way, this system, which combines an emotion engine, can efficiently correct athletes' form and provide mental support in real time, which is expected to improve their overall performance.
[1745] The processing flow will be explained below.
[1746] Step 1:
[1747] Users use a smartphone or dedicated camera to record their sports performance (e.g., pitching or batting). After finishing recording, users launch the app and select the video.
[1748] Step 2:
[1749] The user uploads the video they have taken using a dedicated app to the server. The app sends the user ID and video data together and confirms that the server has received the video.
[1750] Step 3:
[1751] The server receives the uploaded video data and stores it in a database. The video file is linked to a user ID and can be accessed later for analysis.
[1752] Step 4:
[1753] The server retrieves the stored video data and uses deep learning to analyze the player's movements in the video. Specifically, an AI algorithm analyzes each frame and extracts the positions of the player's joints and body parts.
[1754] Step 5:
[1755] The server compares the extracted motion data with the optimal form data, using ideal form data stored in a database and matching the current motion data point by point.
[1756] Step 6:
[1757] The server analyzes the comparison results and identifies specific differences, such as shoulders being too high or elbows being at an incorrect angle.
[1758] Step 7:
[1759] The server generates specific corrective instructions based on an analysis of the differences, which are specific and detailed instructions on how to modify the behavior.
[1760] Step 8:
[1761] The emotion engine extracts emotional data from the user's facial expressions and voice. For example, it uses a camera to recognize the user's face and an expression analysis algorithm to determine their emotion. It also uses voice recognition technology to analyze emotions from the user's tone of voice and speaking style.
[1762] Step 9:
[1763] The server receives the emotion data and adjusts the generated corrective instructions. If the emotion is positive, it issues a standard instruction, but if it is negative, it adds a message of encouragement or support. For example, an instruction such as "Keep your elbow 5 degrees higher on your next pitch" could include an encouraging message such as "You can do it!"
[1764] Step 10:
[1765] The server sends the adjusted correction instructions to the device, which may be in text format or, if necessary, in visual format (images or animations).
[1766] Step 11:
[1767] The device receives the correction instructions from the server and displays them to the user. The app has an easy-to-use interface, allowing users to easily check the instructions. Specific improvement methods and encouraging messages are also displayed.
[1768] Step 12:
[1769] Users can correct their form based on the correction instructions and encouraging messages they receive through the app, and can then film and upload the video again to see the results of their corrections.
[1770] In this way, each step works in tandem, allowing players to efficiently correct their form and receive mental support in real time, which is expected to improve their overall performance.
[1771] Example 2
[1772] Next, a description will be given of Example 2. In the following description, the data processing device 12 will be referred to as a "server" and the robot 414 will be referred to as a "terminal."
[1773] In recent years, there has been a demand for systems that can efficiently improve athletes' performance in sports training. Conventional systems analyze movements and improve form, but rarely provide guidance that takes into account the user's emotional state, resulting in problems such as a decrease in motivation and a lack of psychological support. To solve these issues, a system that can provide real-time guidance that corresponds to the user's emotional state is needed.
[1774] The specific processing by the specific processing unit 290 of the data processing device 12 in the second embodiment is realized by the following means.
[1775] In this invention, the server includes a device for filming the player's performance, a computer unit that receives the video filmed by the device, analysis means that analyzes the video data received by the computer unit and extracts the player's movement data, comparison means that compares the movement data extracted by the analysis means with optimal form data, instruction generation means that generates specific correction instructions based on the differences obtained by the comparison means, transmission means that transmits the correction instructions generated by the instruction generation means to a terminal, display means that displays the correction instructions received by the terminal to the user, an emotion recognition device that extracts emotion data from the user's facial expressions and voice, and adjustment means that adjusts the correction instructions based on the emotion data obtained from the emotion recognition device. This enables appropriate feedback and instruction tailored to the user's emotional state.
[1776] An "athlete" is a person who participates in a sport or competition and seeks to improve their performance.
[1777] "Performance" is a general term for the actions and techniques that athletes perform in sports or competitions.
[1778] "Devices" refers to equipment used to film athletes' performances, including smartphones and dedicated cameras.
[1779] "Video" refers to video data captured by the device.
[1780] The "computer section" is a section that includes a computer system for receiving and processing video data.
[1781] "Analysis means" refers to a function for analyzing the video data received by the computer unit and extracting the movement data of the players.
[1782] "Movement data" refers to information extracted by analytical means, such as the position of a player's joints and body, and the timing of their movements.
[1783] "Form data" is data that serves as a model of ideal behavior or technology and serves as a standard for comparison.
[1784] "Comparison means" refers to functionality for comparing motion data with optimal form data and identifying differences.
[1785] "Difference" refers to the specific differences observed between the behavior data and the form data.
[1786] "Corrective instructions" refer to instructions given to the athlete to improve their movements based on the differences obtained by the comparison means.
[1787] The "instruction generating means" refers to a function for generating specific corrective instructions based on the difference.
[1788] The "transmitting means" refers to a function for transmitting the generated correction instruction to the terminal.
[1789] A "terminal" is a device used by a user to check correction instructions, and includes a smartphone, tablet, etc.
[1790] The "display means" refers to a function for displaying the received correction instructions to the user at the terminal.
[1791] An "emotion recognition device" refers to a device for extracting emotional data from a user's facial expressions and voice.
[1792] "Emotion data" refers to information that indicates the emotional state of a user extracted by an emotion recognition device.
[1793] The "adjustment means" refers to a function for adjusting the correction instruction based on the emotion data to match the user's emotional state.
[1794] The present invention relates to an AI sports trainer system for efficiently improving athletes' performance, and by combining it with an emotion engine, the system provides instruction according to the user's emotional state.
[1795] The system includes the following elements:
[1796] 1. Imaging equipment
[1797] Users use a smartphone or dedicated camera to film performances such as pitching and batting, and this device can capture high-resolution images.
[1798] 2. Upload a video
[1799] The user uploads the video they have taken to the server using a dedicated app. This app sends the user ID and video data together and confirms that the server has received the video.
[1800] 3. Receiving and storing video data
[1801] The server receives the uploaded video data and stores it in a database. The video files are linked to user IDs and can be accessed for analysis.
[1802] 4. Motion analysis
[1803] The server retrieves the stored video data and uses deep learning to analyze the player's movements in the video. Specifically, it analyzes each frame and extracts the positions of the player's joints and body parts.
[1804] 5. Comparison with optimal form
[1805] The server compares the analyzed movement data with ideal form data stored in a database, checking points such as elbow angle and weight transfer.
[1806] 6. Generating Corrective Instructions
[1807] Based on the comparison results, the server analyzes the specific differences and generates corrective instructions, such as "keep your elbow 5 degrees higher on your next pitch."
[1808] 7. Emotion recognition
[1809] The emotion engine uses a camera to recognize the user's facial expressions, and uses a facial expression analysis algorithm to determine the user's emotions. It also uses voice recognition technology to analyze emotions from the user's tone of voice and speaking style.
[1810] 8. Use of Emotional Data
[1811] The server receives the emotion data sent from the emotion engine and adjusts the corrective instructions: if the user is in a positive emotional state, it provides normal instructions, and if the user is in a negative emotional state, it adds encouraging or supportive messages.
[1812] 9. Sending and Displaying Corrective Instructions
[1813] The server sends the generated correction instructions and motivation messages to the terminal.
[1814] The terminal displays the correction instructions and motivational messages received from the server to the user, who then works to correct the form.
[1815] Specific examples
[1816] 1. Users film and upload pitching videos
[1817] Users record their pitching on their smartphones and upload the video to the server using a dedicated app.
[1818] 2. The server receives and analyzes the video
[1819] The server receives the video and uses a deep learning model to analyze each key point of the pitching motion.
[1820] 3. Comparison with optimal form
[1821] The server compares the analyzed movement data with ideal form data and finds specific differences such as "elbow position is too low" or "timing is too early."
[1822] 4. Generate corrective instructions
[1823] The server generates specific corrective instructions such as "keep your elbow 5 degrees higher on your next pitch."
[1824] 5. Operation of the Emotion Engine and Use of Emotion Data
[1825] The emotion engine recognizes emotions from the user's facial expressions and voice, and if negative emotions are detected, the server adds a motivational message.
[1826] 6. Sending and Displaying Instructions
[1827] The server sends the generated correction instructions and motivation messages to the terminal, and the terminal displays the instructions to the user, allowing the user to get effective feedback and correct the form immediately.
[1828] Prompt Sentence Examples
[1829] "Upload a pitching video and analyze elbow position and timing. If emotions are negative, generate improvement instructions with encouraging messages."
[1830] The flow of the identification process in the second embodiment will be described with reference to FIG.
[1831] Step 1:
[1832] Users can film sports performances
[1833] Input: Smartphone or dedicated camera
[1834] How it works: A user records their own sports performance (e.g., pitching or batting) in high resolution, which allows the user to obtain video data.
[1835] Output: Recorded video file
[1836] Step 2:
[1837] User uploads video
[1838] Input: Recorded video file, dedicated app
[1839] How it works: The user launches the app, selects the video they have taken, and uploads it to the server. At this time, the app sends the user ID and the video data together.
[1840] Output: Video data sent to the server
[1841] Step 3:
[1842] The server receives and stores video data
[1843] Input: User ID, video data
[1844] Operation: The server receives the uploaded video data, checks the data integrity, and stores it in the database if there are no problems.
[1845] Output: Video data stored in a database
[1846] Step 4:
[1847] The server retrieves and analyzes the video data
[1848] Input: Video data stored in a database
[1849] Movement: The server retrieves the video data and analyzes each frame using a deep learning algorithm. By extracting the positions of the player's joints and body parts, various movement data is generated.
[1850] Output: Extracted behavioral data
[1851] Step 5:
[1852] Comparison with server-optimized form
[1853] Input: extracted behavioral data, ideal form data
[1854] Movement: The server compares the analyzed movement data with the ideal form data stored in a database, matching each point in detail, such as elbow angle and body timing.
[1855] Output: Difference between behavior data and form data
[1856] Step 6:
[1857] Server generates corrective instructions
[1858] Input: Differences between behavioral data and form data
[1859] Action: The server generates specific corrective instructions based on the discrepancy. A typical example might be, "Keep your elbow 5 degrees higher on your next pitch." Visual feedback (images or animations) is also generated if necessary.
[1860] Output: Generated correction instructions
[1861] Step 7:
[1862] Emotion Engine Operation
[1863] Input: Camera video, audio data
[1864] How it works: The emotion engine uses a camera to recognize the user's facial expressions, and uses a facial expression analysis algorithm to determine the user's emotions. It also uses voice recognition technology to analyze emotions from the user's tone of voice and speaking style.
[1865] Output: Extracted emotion data
[1866] Step 8:
[1867] The server uses emotion data to adjust correction instructions.
[1868] Input: Generated correction instructions, extracted emotion data
[1869] How it works: The server receives the emotion data and adjusts the corrective instructions depending on the user's emotional state: normal instructions for positive emotions, and encouraging or supportive messages for negative emotions.
[1870] Output: Adjusted straightening instructions
[1871] Step 9:
[1872] The server sends correction instructions to the terminal
[1873] Input: Adjusted orthodontic instructions
[1874] Operation: The server sends the generated correction instructions and motivation messages to the device. This transmission process is low-latency.
[1875] Output: Corrective instructions and motivation messages received by the device
[1876] Step 10:
[1877] The device displays correction instructions and messages
[1878] Input: Corrective instructions and motivation messages received by the device
[1879] How it works: The device displays correction instructions and motivational messages to the user. Through the app interface, users can intuitively understand specific ways to improve and immediately start correcting their form.
[1880] Output: Improvement instructions and motivational messages seen by the user
[1881] (Application example 2)
[1882] Next, a description will be given of Application Example 2. In the following description, the data processing device 12 will be referred to as a "server" and the robot 414 will be referred to as a "terminal."
[1883] In conventional factory robot operation, there was a lack of guidance that took into account the optimization of robot operations and the emotional state of the operator. This made it difficult to provide effective corrections when the robot's performance deteriorated, often leading to stress and a decline in operator motivation. In particular, there was a need to simultaneously improve the robot's operational efficiency and the operator's mental health.
[1884] The specific processing by the specific processing unit 290 of the data processing device 12 in Application Example 2 is realized by the following means. In this invention, the server includes a device for filming the robot's movements, a means for receiving the video, a means for analyzing the video data and extracting movement data, a means for comparing the movement data with optimal form data, and a means for generating specific correction instructions. This enables optimization of the robot's movements and effective instruction that takes into account the operator's emotional state.
[1885] A "device" is an electronic device used to capture the action of interest.
[1886] A "server" is a computer system that receives, stores, and analyzes data over a network.
[1887] The "analysis means" refers to algorithms or software for extracting motion data from received video data.
[1888] The "comparison means" is an algorithm or software that has the functionality to compare extracted motion data with optimal form data.
[1889] The "instruction generating means" is an algorithm or software for generating specific correction instructions based on the results of the comparison means.
[1890] The "transmitting means" is a component having a function for transmitting the generated correction instruction to the terminal.
[1891] The "display means" refers to a device or software for visually displaying to the user the correction instructions received at the terminal.
[1892] "Emotion analysis means" refers to algorithms or software for extracting emotional data from a user's facial expressions and voice.
[1893] The "adjustment means" is software having a function for appropriately adjusting correction instructions based on the emotion data acquired by the emotion analysis means.
[1894] "Robot operator" refers to a worker who operates a robot within a factory.
[1895] System Configuration
[1896] This invention is an AI system for improving the performance of factory robot operations, providing optimal instructions by filming and analyzing the robot's movements and achieving effective feedback by taking into account the emotional state of the operator. The system consists of the following components:
[1897] shooting device
[1898] The user (robot operator) uses a smartphone or dedicated camera to record the factory robot's work movements. The camera has high resolution and can clearly capture even the smallest movements. Once the recording is complete, the user launches the dedicated app, selects the video, and uploads it to the server.
[1899] server
[1900] The server stores the received video data in a database and analyzes it using deep learning. Specifically, it uses the following methods:
[1901] Video analysis method: Using TensorFlow and PyTorch, motion data is extracted from the video data. Each frame is analyzed to extract the robot's joints and motion patterns.
[1902] Comparison method: The analyzed motion data is compared with the optimal form data, which represents the ideal robot motion and is stored in a database.
[1903] Instruction generation means: Based on the comparison results, appropriate corrective instructions are generated. For example, specific instructions such as "Make the arm move 5 degrees faster" are generated.
[1904] Emotion analysis method: Using OpenCV and Google Cloud Vision API, emotional data is extracted from the robot operator's facial expressions and voice. The operator's face is recognized by a camera, and emotions are determined using an expression analysis algorithm. Speech recognition technology is also used to analyze emotions from the operator's tone of voice and speaking style.
[1905] Adjustment means: Receives emotion data and adjusts the generated corrective instructions. In the case of positive emotions, it issues normal instructions, and in the case of negative emotions, it adds encouraging messages.
[1906] Terminal
[1907] The terminal is used to provide feedback to the operator, displaying corrective instructions and motivational messages received from the server to help the operator comply with instructions promptly.
[1908] Specific examples
[1909] Consider a situation in a factory where a robot arm is not moving smoothly, resulting in a drop in efficiency. An operator records the robot's movements on a smartphone and uploads the video to a server using a dedicated app. The server analyzes the video, identifies the problem area, and generates specific instructions such as "Make the arm move five degrees faster." At the same time, if the operator looks dissatisfied, the emotion analysis means determines this to be a negative emotion and provides an encouraging message such as "Try a little harder to improve efficiency!" These instructions and messages are displayed on the terminal, allowing the operator to respond immediately.
[1910] Prompt Sentence Examples
[1911] An example of a prompt would be:
[1912] "Today, the robot's arm movements were slow, causing work efficiency to drop. Please film the robot's arm movements and provide appropriate instructions for improvement based on the analysis results."
[1913] As described above, the system of the present invention can simultaneously optimize the operation of factory robots and provide psychological support to operators, thereby contributing to improving overall work efficiency.
[1914] The flow of the specific processing in the application example 2 will be described with reference to FIG.
[1915] Step 1:
[1916] The user films the robot's movements. The user uses a smartphone or a dedicated camera to record the robot's movements, making sure to clearly capture the operating environment and specific movements. The input data is the filmed video. The output data is a video file.
[1917] Step 2:
[1918] The server receives the video. The user uses a dedicated app to upload the video they have taken to the server. The app sends the user ID and video data together, and the server receives the video. The input data is the uploaded video file and user ID, and the output data is the video file saved on the server.
[1919] Step 3:
[1920] The server analyzes the video and extracts motion data. Using a deep learning model (such as TensorFlow or PyTorch), it analyzes the robot's motion in the video frame by frame. It extracts the robot's joints and motion patterns from each frame. The input data is the saved video file, and the output data is the extracted motion data.
[1921] Step 4:
[1922] The server compares the motion data with the optimal form data. The server then matches the extracted motion data with the ideal robot motion data stored in a database. Specifically, it evaluates the degree of match for each frame and identifies discrepancies. The input data is the extracted motion data and the optimal form data, and the output data is the comparison result.
[1923] Step 5:
[1924] The server generates the corrective instructions. Based on the comparison results, the algorithm generates appropriate corrective instructions. For example, specific instructions such as "Make the arm move 5 degrees faster" are generated. The input data is the comparison result, and the output data is the corrective instructions.
[1925] Step 6:
[1926] The server analyzes the user's emotions. An emotion analysis engine (such as OpenCV or Google Cloud Vision API) is used to extract emotion data from the user's facial expressions and voice. The camera recognizes the user's face, and an emotion analysis algorithm is used to determine their emotion. Speech recognition technology is also used to analyze emotions from the tone of voice and speaking style. The input data is the user's facial expression images and voice data, and the output data is emotion data.
[1927] Step 7:
[1928] The server adjusts the correction instructions based on the emotion data. It receives the emotion data and adjusts the generated correction instructions according to the user's emotional state. For example, in the case of negative emotions, it adds an encouraging message such as "Let's try a little harder to improve efficiency!" The input data are emotion data and correction instructions, and the output data are the adjusted correction instructions.
[1929] Step 8:
[1930] The server sends the adjusted correction instruction to the terminal. The sending means sends the adjusted correction instruction and the motivation message to the terminal. The input data is the adjusted correction instruction, and the output data is the instruction and the message sent to the terminal.
[1931] Step 9:
[1932] The terminal displays the corrective instructions to the user. The terminal displays the received corrective instructions and motivational messages to the user. The user checks the instructions on the terminal and immediately understands the specific improvement methods. The input data are the instructions and messages sent to the terminal, and the output data are the display to the user.
[1933] The specific processing unit 290 transmits the result of the specific processing to the robot 414. In the robot 414, the control unit 46A causes the speaker 240 and the control target 443 to output the result of the specific processing. The microphone 238 acquires voice indicating a user input regarding the result of the specific processing. The control unit 46A transmits voice data indicating the user input acquired by the microphone 238 to the data processing device 12. In the data processing device 12, the specific processing unit 290 acquires the voice data.
[1934] The data generation model 58 is a so-called generative AI (Artificial Intelligence). An example of the data generation model 58 is ChatGPT (Internet Search<URL: https: / / openai.com / blog / chatgpt> ), Gemini (Internet search <url: https: gemini.google.com ?hl="ja">) and other generation AIs. The data generation model 58 is obtained by performing deep learning on a neural network. A prompt including an instruction is input to the data generation model 58, and inference data such as voice data indicating voice, text data indicating text, and image data indicating an image is also input. The data generation model 58 performs inference on the input inference data in accordance with the instruction indicated by the prompt, and outputs the inference result in a data format such as voice data and text data. Here, inference refers to, for example, analysis, classification, prediction, and / or summarization.
[1935] In the above embodiment, an example was given in which the specific processing is performed by the data processing device 12, but the technology of the present disclosure is not limited to this, and the specific processing may be performed by the robot 414.
[1936] The emotion identification model 59 as an emotion engine may determine the user's emotion according to a specific mapping. Specifically, the emotion identification model 59 may determine the user's emotion according to an emotion map (see FIG. 9), which is a specific mapping. Similarly, the emotion identification model 59 may determine the robot's emotion, and the identification processing unit 290 may perform identification processing using the robot's emotion.
[1937] FIG. 9 illustrates an emotion map 400 on which multiple emotions are mapped. In the emotion map 400, emotions are arranged in concentric circles radiating from the center. Emotions closer to the center of the concentric circles are more primitive. Emotions representing states and behaviors arising from a state of mind are arranged on the outer edges of the concentric circles. The concept of emotion includes both affect and mental states. Emotions generally generated from reactions occurring in the brain are arranged on the left side of the concentric circles. Emotions generally induced by situational judgment are arranged on the right side of the concentric circles. Emotions generally generated from reactions occurring in the brain and induced by situational judgment are arranged on the upper and lower sides of the concentric circles. Furthermore, the emotion of "pleasure" is arranged on the upper side of the concentric circles, and the emotion of "discomfort" is arranged on the lower side. In this way, in the emotion map 400, multiple emotions are mapped based on the structure by which emotions are generated, and emotions that tend to occur simultaneously are mapped close to each other.
[1938] These emotions are distributed in the 3 o'clock direction on emotion map 400, and typically fluctuate between relief and anxiety. In the right half of emotion map 400, situational awareness dominates over internal sensations, resulting in a sense of calm.
[1939] The inside of emotion map 400 represents what is going on in the mind, and the outside of emotion map 400 represents behavior, so the further you go outside emotion map 400, the more visible the emotions become (the more they are expressed in behavior).
[1940] Human emotions are based on various balances, such as posture and blood sugar levels. When these balances deviate from the ideal, a state of discomfort is indicated, and when they approach the ideal, a state of pleasure is indicated. Emotions can also be created for robots, automobiles, and motorcycles, based on various balances, such as posture and remaining battery life. When these balances deviate from the ideal, a state of discomfort is indicated, and when they approach the ideal, a state of pleasure is indicated. An emotion map can be generated, for example, based on Dr. Mitsuyoshi's emotion map (Research on Voice Emotion Recognition and Emotional Brain Physiological Signal Analysis Systems, Tokushima University, Doctoral Dissertation: https: / / ci.nii.ac.jp / naid / 500000375379). The left half of the emotion map lists emotions belonging to the "reaction" domain, where sensation is dominant. The right half of the emotion map lists emotions belonging to the "situation" domain, where situational awareness is dominant.
[1941] The emotion map defines two emotions that promote learning. One is a negative emotion on the situation side, around the middle of "repentance" or "reflection." In other words, this occurs when the robot experiences negative emotions such as "I never want to feel this way again" or "I don't want to be scolded again." The other is a positive emotion on the response side, around "desire." In other words, this occurs when the robot experiences positive feelings such as "I want more" or "I want to know more."
[1942] The emotion identification model 59 inputs user input into a pre-trained neural network, obtains emotion values indicating each emotion shown in the emotion map 400, and determines the user's emotion. This neural network is pre-trained based on multiple pieces of training data that are combinations of user input and emotion values indicating each emotion shown in the emotion map 400. Furthermore, this neural network is trained so that emotions that are located close to each other have similar values, as in the emotion map 900 shown in FIG. 10. FIG. 10 shows an example in which multiple emotions, "relieved," "calm," and "reassuring," have similar emotion values.
[1943] The system according to the present disclosure has been described above mainly with respect to the functions of the data processing device 12, but the system according to the present disclosure is not necessarily implemented on a server. The system according to the present disclosure may be implemented as a general information processing system. The present disclosure may be implemented, for example, as a software program running on a personal computer or an application running on a smartphone, etc. The method according to the present disclosure may be provided to users in the form of SaaS (Software as a Service).
[1944] In the above embodiment, an example was given in which the specific processing is performed by one computer 22, but the technology of the present disclosure is not limited to this, and the specific processing may be distributed and performed by a plurality of computers including the computer 22. For example, the data generation model 58 may be provided in an external device of the data processing device 12, and data may be generated in the external device in accordance with input data.
[1945] In the above embodiment, an example in which the specific processing program 56 is stored in the storage 32 has been described, but the technology of the present disclosure is not limited to this. For example, the specific processing program 56 may be stored in a portable, computer-readable, non-transitory storage medium such as a USB (Universal Serial Bus) memory. The specific processing program 56 stored in the non-transitory storage medium is installed in the computer 22 of the data processing device 12. The processor 28 executes the specific processing in accordance with the specific processing program 56.
[1946] Alternatively, the specific processing program 56 may be stored in a storage device such as a server connected to the data processing device 12 via the network 54, and the specific processing program 56 may be downloaded and installed on the computer 22 in response to a request from the data processing device 12.
[1947] It is not necessary to store all of the specific processing program 56 in a storage device such as a server connected to the data processing device 12 via the network 54, or to store all of the specific processing program 56 in the storage 32; only a portion of the specific processing program 56 may be stored.
[1948] The hardware resource for executing a specific process can be any of the following processors: An example of a processor is a CPU, which is a general-purpose processor that functions as a hardware resource for executing a specific process by executing software, i.e., a program. Another example of a processor is a dedicated electrical circuit, such as an FPGA (Field-Programmable Gate Array), a PLD (Programmable Logic Device), or an ASIC (Application Specific Integrated Circuit), which is a processor with a circuit configuration designed specifically for executing a specific process. Each processor has built-in or connected memory, and each processor uses the memory to execute the specific process.
[1949] The hardware resource that executes the specific processing may be configured with one of these various processors, or may be configured with a combination of two or more processors of the same or different types (for example, a combination of multiple FPGAs, or a combination of a CPU and an FPGA). Also, the hardware resource that executes the specific processing may be a single processor.
[1950] As an example of a system configured with a single processor, first, one processor is configured by combining one or more CPUs and software, and this processor functions as a hardware resource that executes a specific process. Second, there is a system that uses a processor that realizes the functions of an entire system including multiple hardware resources that execute a specific process on a single IC chip, as typified by SoC (System-on-a-chip). In this way, a specific process is realized using one or more of the above-mentioned various processors as hardware resources.
[1951] Furthermore, the hardware structure of these various processors can be, more specifically, an electric circuit that combines circuit elements such as semiconductor devices. The specific processing described above is merely an example. Therefore, it goes without saying that unnecessary steps may be deleted, new steps may be added, or the processing order may be rearranged, without departing from the spirit of the invention.
[1952] The above-described description and illustrations are a detailed explanation of the parts related to the technology of the present disclosure and are merely an example of the technology of the present disclosure. For example, the above description of the configuration, functions, actions, and effects is an explanation of an example of the configuration, functions, actions, and effects of the parts related to the technology of the present disclosure. Therefore, it goes without saying that unnecessary parts may be deleted, new elements may be added, or replacements may be made to the above-described description and illustrations within the scope of the gist of the technology of the present disclosure. Furthermore, to avoid confusion and facilitate understanding of the parts related to the technology of the present disclosure, the above-described description and illustrations omit explanations of common technical knowledge that do not require particular explanation to enable the implementation of the technology of the present disclosure.
[1953] All publications, patent applications, and technical standards mentioned in this specification are herein incorporated by reference to the same extent as if each individual publication, patent application, or technical standard was specifically and individually indicated to be incorporated by reference.
[1954] The following is further disclosed regarding the above embodiment.
[1955] (Claim 1)
[1956] A device for filming the athletes' performances,
[1957] a server that receives video captured by the device;
[1958] analysis means for analyzing the video data received by the server and extracting player movement data;
[1959] a comparison means for comparing the motion data extracted by the analysis means with optimal form data;
[1960] an instruction generating means for generating specific corrective instructions based on the difference obtained by the comparing means;
[1961] a transmitting means for transmitting the correction instruction generated by the instruction generating means to a terminal;
[1962] a display means for displaying the received correction instruction to a user at the terminal;
[1963] A system including:
[1964] (Claim 2)
[1965] 2. The system according to claim 1, wherein the analysis means has a function of automatically extracting key points of each movement of the player.
[1966] (Claim 3)
[1967] 2. The system of claim 1, wherein the comparison means is operable to compare the motion data with optimal form data for each frame.
[1968] "Example 1"
[1969] (Claim 1)
[1970] A device for filming the athletes' performances,
[1971] a server that receives video captured by the device;
[1972] analysis means for analyzing the video data received by the server and extracting player movement data;
[1973] a comparison means for comparing the motion data extracted by the analysis means with optimal form data;
[1974] an instruction generating means for generating specific corrective instructions based on the difference obtained by the comparing means;
[1975] a transmitting means for transmitting the correction instruction generated by the instruction generating means to a terminal;
[1976] a display means for displaying the received correction instruction to a user at the terminal;
[1977] a photographing means for photographing video data at high resolution and recording the detailed movements of the players;
[1978] A way to easily upload videos to the server through a dedicated app,
[1979] A method for analyzing the movements of athletes in videos using deep learning models and automatically generating movement data;
[1980] a means for delivering corrective instructions in textual and visual formats and displaying them in a manner that is intuitively understandable to the user;
[1981] A system including:
[1982] (Claim 2)
[1983] 2. The system according to claim 1, wherein the analysis means has a function of automatically extracting key points of each movement of the player.
[1984] (Claim 3)
[1985] 2. The system of claim 1, wherein the comparison means is operable to compare the motion data with optimal form data for each frame.
[1986] "Application Example 1"
[1987] (Claim 1)
[1988] A device for filming the athletes' performances,
[1989] a server that receives video captured by the device;
[1990] analysis means for analyzing the video data received by the server and extracting player movement data;
[1991] a comparison means for comparing the motion data extracted by the analysis means with optimal form data;
[1992] an instruction generating means for generating specific corrective instructions based on the difference obtained by the comparing means;
[1993] a transmitting means for transmitting the correction instruction generated by the instruction generating means to a terminal;
[1994] a display means for displaying the received correction instruction to a user at the terminal;
[1995] a control means for applying the corrective instructions received by the terminal to the robot and correcting its motion in real time;
[1996] A system including:
[1997] (Claim 2)
[1998] 2. The system according to claim 1, wherein the analysis means has a function of automatically extracting key points of each movement of the athlete and the industrial machine.
[1999] (Claim 3)
[2000] 2. The system according to claim 1, wherein the comparison means has a function of comparing the motion data with optimal form data or work efficiency data for each frame.
[2001] "Example 2: Combining Emotion Engines"
[2002] (Claim 1)
[2003] Equipment for filming the athletes' performances;
[2004] a computer unit that receives the video captured by the device;
[2005] analysis means for analyzing the video data received by the computer unit and extracting player movement data;
[2006] a comparison means for comparing the motion data extracted by the analysis means with optimal form data;
[2007] an instruction generating means for generating specific corrective instructions based on the difference obtained by the comparing means;
[2008] a transmitting means for transmitting the correction instruction generated by the instruction generating means to a terminal;
[2009] a display means for displaying the received correction instruction to a user at the terminal;
[2010] an emotion recognition device that extracts emotion data from a user's facial expression and voice;
[2011] an adjustment means for adjusting a correction instruction based on emotion data obtained from the emotion recognition device;
[2012] A system including:
[2013] (Claim 2)
[2014] 2. The system according to claim 1, wherein the analysis means has a function of automatically extracting key points of each movement of the player.
[2015] (Claim 3)
[2016] 2. The system of claim 1, wherein the comparison means is operable to compare the motion data with optimal form data for each frame.
[2017] "Application example 2 when combining emotion engines"
[2018] (Claim 1)
[2019] A device for filming the athletes' performances,
[2020] a server that receives video captured by the device;
[2021] analysis means for analyzing the video data received by the server and extracting player movement data;
[2022] a comparison means for comparing the motion data extracted by the analysis means with optimal form data;
[2023] an instruction generating means for generating specific corrective instructions based on the difference obtained by the comparing means;
[2024] a transmitting means for transmitting the correction instruction generated by the instruction generating means to a terminal;
[2025] a display means for displaying the received correction instruction to a user at the terminal;
[2026] An emotion analysis means that combines an emotion engine to extract emotion data from the user's facial expressions and voice;
[2027] a means for adjusting a correction instruction based on the emotion data extracted by the emotion analysis means;
[2028] A system including:
[2029] (Claim 2)
[2030] 2. The system according to claim 1, wherein the analysis means has a function of automatically extracting key points of each movement of t...
Claims
1. A device for filming the athletes' performances, a server that receives video captured by the device; analysis means for analyzing the video data received by the server and extracting player movement data; a comparison means for comparing the motion data extracted by the analysis means with optimal form data; an instruction generating means for generating specific corrective instructions based on the difference obtained by the comparing means; a transmitting means for transmitting the correction instruction generated by the instruction generating means to a terminal; a display means for displaying the received correction instruction to a user at the terminal; A system including:
2. 2. The system according to claim 1, wherein said analysis means has a function of automatically extracting key points of each movement of a player.
3. 2. The system of claim 1, wherein said comparison means has the function of comparing said motion data with optimal form data for each frame.
Citation Information
Patent Citations
Persona chatbot control method and system
JP2022180282A