System
A system with a camera, server, and display provides real-time, detailed feedback on golf swing form, addressing the lack of efficient professional instruction in golf training.
Patent Information
- Application Number
- JP2024125357
- Authority / Receiving Office
- JP · JP
- Patent Type
- Applications
- Current Assignee / Owner
- Filing Date
- 2024-07-31
- Publication Date
- 2026-02-13
AI Technical Summary
Golf training often lacks efficient systems for providing specific professional instruction, limiting players' opportunities for form improvement, and existing technologies struggle to offer real-time, detailed feedback on swing form.
A system that includes a camera for capturing high-frame-rate video of a player's swing, a server for multimodal AI analysis, and a display for immediate feedback on form improvements and practice methods.
Enables players to receive specific and timely feedback, allowing for efficient and cost-effective improvement of their golf swing form.
Smart Images

Figure 2026023422000001_ABST
Abstract
Description
[Technical Field]
[0001] The technology of the present disclosure relates to a system. [Background technology]
[0002] Patent document 1 discloses a persona chatbot control method performed by at least one processor, the method including the steps of receiving a user utterance, adding the user utterance to a prompt including an instruction sentence related to a description of the chatbot character, encoding the prompt, and inputting the encoded prompt into a language model to generate a chatbot utterance in response to the user utterance. [Prior art documents] [Patent documents]
[0003] [Patent Document 1] Japanese Patent Publication No. 2022-180282 Summary of the Invention [Problem to be solved by the invention]
[0004] In golf training, professional instruction is necessary to accurately improve a player's form, but many players have limited opportunities to receive such instruction. Also, specific instruction methods are required for self-improvement, but the lack of a system to efficiently provide this is a problem. [Means for solving the problem]
[0005] The present invention is a system that includes a camera for capturing a player's movements in real time, a processing unit for receiving and analyzing the captured video data, and a display unit for providing feedback information to the player based on the analysis results. Specifically, a camera is used to capture the player's swing form at a high frame rate, and multimodal AI analyzes the video data to identify areas for form improvement and displays them in real time on the player's device. This system enables many players to improve their form efficiently and at low cost.
[0006] A "player" is a person who uses the system to train in golf.
[0007] An "action" is an act in which a player performs an exercise such as a golf swing.
[0008] "Real-time" refers to processing and feedback that occurs immediately, without delay.
[0009] "Filming means" refers to a device such as a camera for recording the player's actions as video.
[0010] "Video data" refers to video information recorded by a photographing means.
[0011] "Processing means" refers to a computer system or software that analyzes the video data and extracts information about the player's actions.
[0012] "Analysis" is the act of detecting and evaluating movement characteristics and problems based on video data.
[0013] "Feedback information" refers to specific instructional content for improving actions that is provided to the player as a result of the analysis.
[0014] "Display means" refers to a device such as a display or monitor that visually presents feedback information to the player.
[0015] The "system" refers to the entire device and software set that combines the above-mentioned imaging means, processing means, and display means.
[0016] The "camera" is a photographing device for recording the player's swing form at a high frame rate.
[0017] "Multimodal AI" is an artificial intelligence technology that uses video and other data in an integrated manner to perform analysis.
[0018] A "server" is a remote computer system for processing video data and generating analytical results.
[0019] A "terminal" is a personal device that receives feedback information sent from the server and displays it to the player. [Brief explanation of the drawings]
[0020] [Figure 1] 1 is a conceptual diagram showing an example of the configuration of a data processing system according to a first embodiment. [Figure 2] 1 is a conceptual diagram showing an example of main functions of a data processing device and a smart device according to a first embodiment. [Figure 3] FIG. 10 is a conceptual diagram showing an example of the configuration of a data processing system according to a second embodiment. [Figure 4] FIG. 10 is a conceptual diagram showing an example of main functions of a data processing device and smart glasses according to a second embodiment. [Figure 5] FIG. 10 is a conceptual diagram showing an example of the configuration of a data processing system according to a third embodiment. [Figure 6] FIG. 11 is a conceptual diagram showing an example of main functions of a data processing device and a headset-type terminal according to a third embodiment. [Figure 7]FIG. 10 is a conceptual diagram showing an example of the configuration of a data processing system according to a fourth embodiment. [Figure 8] FIG. 10 is a conceptual diagram showing an example of main functions of a data processing device and a robot according to a fourth embodiment. [Figure 9] 1 shows an emotion map onto which multiple emotions are mapped. [Figure 10] 1 shows an emotion map onto which multiple emotions are mapped. [Figure 11] FIG. 3 is a sequence diagram showing a processing flow of the data processing system according to the first embodiment. [Figure 12] FIG. 10 is a sequence diagram showing the flow of processing in the data processing system in Application Example 1. [Figure 13] FIG. 10 is a sequence diagram showing the flow of processing in the data processing system according to the second embodiment when an emotion engine is combined. [Figure 14] FIG. 10 is a sequence diagram showing the flow of processing in the data processing system in Application Example 2 when an emotion engine is combined. DETAILED DESCRIPTION OF THE INVENTION
[0021] An example of an embodiment of a system according to the technology of the present disclosure will be described below with reference to the accompanying drawings.
[0022] First, the terms used in the following description will be explained.
[0023] In the following embodiments, a coded processor (hereinafter simply referred to as a "processor") may be a single arithmetic device or a combination of multiple arithmetic devices. Furthermore, a processor may be a single type of arithmetic device or a combination of multiple types of arithmetic devices. Examples of arithmetic devices include a CPU (Central Processing Unit), a GPU (Graphics Processing Unit), a GPGPU (General-Purpose computing on Graphics Processing Units), and an APU (Accelerated Processing Unit).
[0024] In the following embodiments, a coded RAM (Random Access Memory) is a memory in which information is temporarily stored and is used as a working memory by a processor.
[0025] In the following embodiments, the coded storage is one or more non-volatile storage devices that store various programs, various parameters, etc. Examples of non-volatile storage devices include flash memory (SSD (Solid State Drive)), magnetic disks (e.g., hard disks), and magnetic tapes.
[0026] In the following embodiments, a communication I / F (Interface) with a symbol is an interface including a communication processor, an antenna, etc. The communication I / F controls communication between multiple computers. Examples of communication standards applied to the communication I / F include wireless communication standards including 5G (5th Generation Mobile Communication System), Wi-Fi (registered trademark), Bluetooth (registered trademark), etc.
[0027] In the following embodiments, "A and / or B" is synonymous with "at least one of A and B." In other words, "A and / or B" means that it may be only A, only B, or a combination of A and B. Furthermore, in this specification, the same concept as "A and / or B" is also applied when three or more things are expressed connected by "and / or."
[0028] [First embodiment]
[0029] FIG. 1 shows an example of the configuration of a data processing system 10 according to the first embodiment.
[0030] 1, a data processing system 10 includes a data processing device 12 and a smart device 14. An example of the data processing device 12 is a server.
[0031] The data processing device 12 includes a computer 22, a database 24, and a communication I / F 26. The computer 22 is an example of a "computer" according to the technology of the present disclosure. The computer 22 includes a processor 28, a RAM 30, and a storage 32. The processor 28, the RAM 30, and the storage 32 are connected to a bus 34. The database 24 and the communication I / F 26 are also connected to the bus 34. The communication I / F 26 is connected to a network 54. Examples of the network 54 include a WAN (Wide Area Network) and / or a LAN (Local Area Network).
[0032] The smart device 14 includes a computer 36, a reception device 38, an output device 40, a camera 42, and a communication I / F 44. The computer 36 includes a processor 46, a RAM 48, and a storage 50. The processor 46, the RAM 48, and the storage 50 are connected to a bus 52. The reception device 38, the output device 40, and the camera 42 are also connected to the bus 52.
[0033] The reception device 38 includes a touch panel 38A, a microphone 38B, and the like, and receives user input. The touch panel 38A detects contact with an indicator (for example, a pen or a finger) to receive user input by the touch of the indicator. The microphone 38B detects the user's voice to receive user input by voice. The control unit 46A transmits data indicating the user input received by the touch panel 38A and the microphone 38B to the data processing device 12. In the data processing device 12, the specific processing unit 290 acquires the data indicating the user input.
[0034] The output device 40 includes a display 40A and a speaker 40B, and presents data to the user 20 by outputting the data in a form of expression that the user 20 can perceive (for example, audio and / or text). The display 40A displays visible information such as text and images in accordance with instructions from the processor 46. The speaker 40B outputs audio in accordance with instructions from the processor 46. The camera 42 is a compact digital camera equipped with an optical system including a lens, aperture, and shutter, and an imaging element such as a CMOS (Complementary Metal-Oxide-Semiconductor) image sensor or a CCD (Charge Coupled Device) image sensor.
[0035] The communication I / F 44 is connected to a network 54. The communication I / Fs 44 and 26 control the exchange of various information between the processor 46 and the processor 28 via the network 54.
[0036] FIG. 2 shows an example of the main functions of the data processing device 12 and the smart device 14.
[0037] 2, in the data processing device 12, a specific process is performed by the processor 28. A specific processing program 56 is stored in the storage 32. The specific processing program 56 is an example of a "program" according to the technology of the present disclosure. The processor 28 reads the specific processing program 56 from the storage 32 and executes the read specific processing program 56 on the RAM 30. The specific process is realized by the processor 28 operating as a specific processing unit 290 in accordance with the specific processing program 56 executed on the RAM 30.
[0038] The storage 32 stores a data generation model 58 and an emotion identification model 59. The data generation model 58 and the emotion identification model 59 are used by the identification processing unit 290.
[0039] In the smart device 14, the processor 46 performs the reception output process. The storage 50 stores a reception output program 60. The reception output program 60 is used in conjunction with the specific processing program 56 by the data processing system 10. The processor 46 reads the reception output program 60 from the storage 50 and executes the read reception output program 60 on the RAM 48. The reception output process is realized by the processor 46 operating as the control unit 46A in accordance with the reception output program 60 executed on the RAM 48.
[0040] Next, a description will be given of the specific processing performed by the specific processing unit 290 of the data processing device 12. In the following description, the data processing device 12 will be referred to as a "server" and the smart device 14 will be referred to as a "terminal."
[0041] The present invention is a system that combines a photographing means, a processing means, and a display means to efficiently improve a player's golf form. Below, the processing of the program of this system will be explained in natural language, and an embodiment of the invention will be described in detail with concrete examples.
[0042] System Overview
[0043] This system uses a camera linked to a golf simulator to capture the player's swing form in real time. The captured video data is sent to a server, which then analyzes the data using multimodal AI. The analysis results are fed back to the player in real time via their device, suggesting areas for form improvement and optimal practice methods.
[0044] Program processing
[0045] Step 1: Capture video with a camera
[0046] The user stands at the starting position of the golf simulator and begins swinging.
[0047] The device detects the start of the swing and activates the camera in conjunction with the golf simulator.
[0048] The camera captures the player's swing form in real time and acquires video data at a high frame rate.
[0049] Step 2: Sending video data
[0050] The device temporarily stores the captured video data and performs pre-processing, after which the data is compressed and encrypted for transmission to the server.
[0051] Step 3: AI-powered video analysis
[0052] The server inputs the received video data into the multimodal AI and begins analysis.
[0053] The AI extracts features from the video data, such as swing speed, angle, and wrist movement, to detect problems with the swing form. For example, if the arms are positioned too high, this problem will be identified.
[0054] Step 4: Identify areas for improvement and provide feedback
[0055] The server then generates feedback based on the analysis, identifying specific areas for improvement and the reasons for them. For example, the feedback could be, "Your arms are positioned too high, making it difficult to make proper impact with the ball."
[0056] Specific practice methods (e.g., specific drills to correct arm position) are also generated as feedback information.
[0057] Step 5: View your feedback
[0058] The server transmits the generated feedback information to the terminal.
[0059] The device receives the feedback information and displays it to the user, using graphics and animations to make it visually easy to understand.
[0060] Specific examples
[0061] Consider a case where a player (user) is practicing on a golf simulator.
[0062] 1. The user starts swinging. The device detects the start of the swing and the camera captures the swing form.
[0063] 2. The device sends this video data to the server.
[0064] 3. The server receives the video data and analyzes it using multimodal AI, detecting, for example, problems such as the arm being positioned slightly too high.
[0065] 4. The server generates feedback such as "Your arms are positioned too high, making it difficult to make a proper impact," and sends it to the device along with specific practice instructions.
[0066] 5. The device displays feedback information to the user in an intuitive format, for example, by providing a graphic indication of arm position and visually demonstrating correct form.
[0067] This allows users to receive instant feedback and use it to improve their form. By repeating this process, players can efficiently improve their golfing skills.
[0068] The processing flow will be explained below.
[0069] Step 1:
[0070] The user stands at the starting position of the golf simulator and begins swinging.
[0071] The device detects the start of the swing and activates the camera in conjunction with the golf simulator.
[0072] The camera captures the player's swing form in real time and acquires video data at a high frame rate.
[0073] Step 2:
[0074] The device temporarily stores the captured video data and performs preprocessing, which includes noise reduction and image correction.
[0075] The terminal performs data compression and encryption to send the pre-processed data to the server.
[0076] Step 3:
[0077] The server inputs the received video data into the multimodal AI and begins analysis.
[0078] Multimodal AI extracts features such as swing speed, angle, and wrist movement from video data to detect problems with swing form. For example, if the arms are positioned too high, this will be identified.
[0079] Step 4:
[0080] The server then uses the analysis results to identify specific areas for improvement and the reasons for them, such as determining that "your arms are positioned too high, making it difficult to make proper impact with the ball."
[0081] The server generates specific feedback based on the identified areas for improvement, including specific practice methods (e.g., specific drills to correct arm position).
[0082] Step 5:
[0083] The server transmits the generated feedback information to the terminal.
[0084] The device receives the feedback information and displays it to the user, using graphics and animations to make it visually easy to understand.
[0085] Step 6:
[0086] The user reviews the displayed feedback information and practices to correct their next swing form. They adjust their form based on the specific improvements indicated in the feedback.
[0087] This series of steps allows users to efficiently improve their swing form, and real-time feedback can help them improve their technique quickly.
[0088] Example 1
[0089] Next, a description will be given of Example 1. In the following description, the data processing device 12 will be referred to as a "server" and the smart device 14 will be referred to as a "terminal."
[0090] For players to efficiently improve their form, it is important to accurately understand their current form and immediately identify specific areas for improvement. However, with conventional technology, analyzing and evaluating form takes time, making it difficult to provide real-time feedback. In addition, the analysis results are not specific, making it difficult for players to understand how to improve.
[0091] The specific processing by the specific processing unit 290 of the data processing device 12 in the first embodiment is realized by the following means.
[0092] In this invention, the server includes a camera for capturing the player's movements in real time, a terminal for temporarily storing and preprocessing the captured video data, a transmission means for compressing and encrypting the preprocessed data and transmitting it to the server, a processing means for analyzing the video data received by the server using multimodal AI to detect problems with the player's swing form, a feedback generation means for identifying specific areas for improvement based on the analysis results and generating feedback information, and a means for transmitting the generated feedback information to the terminal and displaying it to the user. This allows the player to receive specific and detailed feedback in real time, enabling them to improve their form efficiently.
[0093] A "player" is someone who uses the system to improve their own movements and form.
[0094] "Movement" refers to the physical movements of the player related to their swing and form.
[0095] "Real-time" refers to immediate processing and feedback without delay.
[0096] "Photography means" refers to a device used to capture the player's actions, such as a camera.
[0097] "Terminal" refers to a device that temporarily stores, pre-processes, and transmits data, typically a computer or smart device.
[0098] "Transmission means" refers to the process or function that transmits the pre-processed data to the server.
[0099] "Server" refers to a device or system that analyzes video data and generates feedback information.
[0100] "Multimodal AI" refers to artificial intelligence technology that comprehensively analyzes multiple input data (video, audio, sensors, etc.).
[0101] "Processing means" refers to a device or process that has the function of analyzing received video data and detecting problems.
[0102] "Feedback generation means" refers to the process or function that generates specific improvements and practice methods based on the analysis results.
[0103] The "display means" refers to a device or system that has the function of visually displaying the generated feedback information to the user.
[0104] "Video data" refers to video data that captures the player's actions.
[0105] The present invention is a system that captures a player's actions in real time and provides feedback based on the analysis results. Specifically, it is realized by linking a camera, a terminal, a server, and a display.
[0106] The system overview includes the following hardware and software:
[0107] Filming method
[0108] The recording method consists of a camera that captures the player's swing form at a high frame rate. This camera operates at a high frame rate such as 120 fps, allowing for detailed capture of the player's movements. For example, it is preferable to use a high-performance sports camera or a dedicated motion capture camera rather than a regular webcam.
[0109] Terminal
[0110] The terminal is a device that temporarily stores and pre-processes the video data captured by the camera. This pre-processing includes noise reduction and color correction. The video data is then compressed into MPEG format and encrypted with AES. This allows the data to be sent securely to the server while maintaining its quality. The terminal requires a high-performance processor and large-capacity storage.
[0111] Transmission method
[0112] The transmission means has the function of sending preprocessed data from the terminal to the server. It is important to use the HTTPS protocol for data transmission, ensuring safety and speed.
[0113] server
[0114] The server includes a processing means for analyzing the received video data using multimodal AI. This AI is built using deep learning frameworks such as TensorFlow and PyTorch, and performs detailed analysis of swing speed, angle, wrist movement, and other parameters. As a result of the analysis, it detects problems with the swing form and generates feedback. The server then generates feedback that includes specific practice methods and corrections.
[0115] Feedback Generation Method
[0116] The feedback generator explains specific areas for improvement and the reasons for them based on the analysis results from the server. For example, it may give specific indications such as, "Your right arm is too high, making your impact unstable." It also suggests practice methods for keeping the right arm low.
[0117] Display means
[0118] The display means is a device or system that visually displays the generated feedback information to the user. The terminal receives the feedback information and uses graphics or animations to show it to the user in an easy-to-understand manner. For example, by visually displaying the difference between incorrect and correct form, the user can understand specifically which parts need to be corrected.
[0119] Specific examples
[0120] For example, when a player practices on a golf simulator, the system operates in the following steps:
[0121] 1. The user starts swinging, the device detects the movement, and the camera captures the swing form at a high frame rate.
[0122] 2. The device temporarily stores the video data, performs preprocessing, compresses it into MPEG format, encrypts it using AES, and sends it to the server.
[0123] 3. The server receives the video data and analyzes it using multimodal AI, detecting problems such as "the right arm is raised too high."
[0124] 4. Based on the analysis results, the server generates feedback recommending "practice swinging with both arms fixed at waist height to practice keeping the right arm low" and sends it to the device.
[0125] 5. The device displays the feedback information in a graphical format, visually showing the user specific areas for improvement and correct form.
[0126] Prompt Sentence Examples
[0127] The user can enter prompts for the generative AI model, such as:
[0128] "Analyze the swing footage of the player and if there are any problems with their swing form, tell them how to improve it."
[0129] Thus, the present invention is a system that can provide detailed feedback in real time to help players improve their form efficiently.
[0130] The flow of the identification process in the first embodiment will be described with reference to FIG.
[0131] The flow of this system's program processing
[0132] Step 1:
[0133] The user stands at the starting position of the golf simulator and begins swinging.
[0134] Input: User swing motion
[0135] Operation: The device uses sensors to recognize when the user starts swinging.
[0136] Output: Swing start detection signal
[0137] Step 2:
[0138] The device activates the camera and captures the player's swing form in real time.
[0139] Input: Swing start detection signal
[0140] How it works: The device controls the camera and captures video data of the swing form at a high frame rate such as 120 fps.
[0141] Output: Captured video data
[0142] Step 3:
[0143] The device temporarily stores the captured video data and performs preprocessing.
[0144] Input: Captured video data
[0145] How it works: The device breaks down the video data into frames and performs noise reduction and color correction.
[0146] Output: Pre-processed video data
[0147] Step 4:
[0148] The terminal compresses and encrypts the pre-processed video data.
[0149] Input: Preprocessed video data
[0150] Operation: Video data is compressed into MPEG format and encrypted with AES.
[0151] Output: Compressed and encrypted video data
[0152] Step 5:
[0153] The terminal transmits the compressed and encrypted video data to the server.
[0154] Input: Compressed and encrypted video data
[0155] How it works: The device uses the HTTPS protocol to securely upload data to the server.
[0156] Output: Video data sent to the server
[0157] Step 6:
[0158] The video data received by the server is analyzed using multimodal AI.
[0159] Input: Received video data
[0160] How it works: The server inputs video data into the AI model and extracts swing form characteristics (e.g., speed, angle, wrist movement).
[0161] Output: Swing form feature data
[0162] Step 7:
[0163] The server detects problems based on characteristic data of the swing form.
[0164] Input: Swing form characteristics data
[0165] How it works: The server compares the extracted features with predefined standards and identifies form flaws (e.g., the right arm is too high).
[0166] Output: Detected issue data
[0167] Step 8:
[0168] The server generates feedback information based on the analysis results.
[0169] Input: Detected issue data
[0170] How it works: The server generates feedback with specific areas for improvement and reasons for doing so, and also suggests practice methods.
[0171] Output: Generated feedback information
[0172] Step 9:
[0173] The server transmits the generated feedback information to the terminal.
[0174] Input: Generated feedback information
[0175] Operation: The server sends feedback information to the device in JSON format or similar.
[0176] Output: Feedback information sent to the terminal
[0177] Step 10:
[0178] The terminal displays the feedback information to the user.
[0179] Input: Feedback information sent to the device
[0180] Action: The device visually displays feedback information to the user using graphics and animations.
[0181] Output: Feedback information displayed to the user
[0182] (Application example 1)
[0183] Next, a description will be given of Application Example 1. In the following description, the data processing device 12 will be referred to as a "server" and the smart device 14 will be referred to as a "terminal."
[0184] Conventional technologies exist that provide real-time feedback on player movements and form improvements. However, there is a need for a system that can optimize the movement accuracy and efficiency of robots and equipment in real time, even in factory operations. The present invention provides technology that improves movement accuracy and efficiency by capturing, analyzing, and providing feedback on the movements of factory robots.
[0185] The specific processing by the specific processing unit 290 of the data processing device 12 in the application example 1 is realized by the following means.
[0186] In this invention, the server includes a camera for capturing the actions of the player or the actions of the equipment in real time, a processing unit for receiving and analyzing the captured video data, and a display unit for providing feedback information to the player or operator based on the analysis results, thereby enabling efficient improvement of the operations of factory robots and equipment.
[0187] "Photographing means" refers to a device for capturing the actions of a player or device in real time.
[0188] A "processing means" is a device or system for receiving and analyzing captured video data.
[0189] "Display means" refers to a device or system for providing feedback information to a player or operator based on the analysis results.
[0190] "Multimodal AI" is an artificial intelligence technology that integrates and analyzes multiple data modalities (e.g., video, audio, text).
[0191] "Feedback" is information based on the analysis results that provides improvements to movements and form, as well as optimal practice methods.
[0192] "Capture" refers to the process of converting a subject's movements and form into data in real time using a photographic device.
[0193] A "server" is a computer system that receives data, analyzes it, and generates feedback information.
[0194] The present invention is a system for optimizing the operation of a factory robot, and includes a photographing means for capturing the operation of a player or equipment in real time, a processing means for receiving and analyzing the photographed video data, and a display means for providing feedback information to a player or operator based on the analysis results.
[0195] System Program
[0196] The main hardware and software components of this system include smartphones, factory robots, smartphone apps (e.g., cross-platform development with Flutter or React Native), servers, multimodal AI (e.g., TensorFlow or PyTorch), and communication protocols (e.g., HTTPS or WebSocket).
[0197] Natural language explanation of program processing
[0198] Step 1: Capture video with a camera
[0199] The smartphone camera captures the robot's movements in the factory in real time, capturing the video data at a high frame rate.
[0200] Step 2: Sending video data
[0201] The captured video data is temporarily stored on the smartphone and then encrypted and sent to a server.
[0202] Step 3: AI-powered video analysis
[0203] The server inputs the received video data into the multimodal AI for analysis. The AI extracts characteristics of the robot's movements from the video data and detects problems with the movements. For example, if the robot's arm is moving too slowly, this problem will be identified.
[0204] Step 4: Identify areas for improvement and provide feedback
[0205] Based on the analysis results, the server identifies specific areas for improvement and the reasons for them, and generates feedback information. For example, the feedback might say, "The arm's movements are too slow, reducing work efficiency." This feedback might also include specific drills to improve the arm's speed by 20%.
[0206] Step 5: View your feedback
[0207] The server sends the generated feedback information to a smartphone and displays it for the user or operator to check. This display uses graphics and animations to make it visually easy to understand.
[0208] Specific examples
[0209] For example, if the robot arm is analyzed as moving too slowly, the app will display the following feedback:
[0210] "The arm is moving too slowly, reducing work efficiency. Please perform drills to increase the arm's movement speed by 20%."
[0211] AI prompt examples
[0212] "Analyze the robot arm's movement pattern in this video, identify the speed, angle, and errors, and suggest improvements based on that."
[0213] The flow of the specific processing in the application example 1 will be described with reference to FIG.
[0214] Step 1:
[0215] A user captures the operation of a factory robot in real time using a smartphone camera. The smartphone camera acquires video data at a high frame rate and temporarily stores the data within the smartphone. The input is the captured video data, and the output is the temporarily stored video data.
[0216] Step 2:
[0217] The device preprocesses the stored video data. Preprocessing includes compressing and encrypting the data. The compressed and encrypted data is then sent to the server. The input is the stored video data, and the output is the compressed and encrypted data.
[0218] Step 3:
[0219] The server decompresses the received video data and prepares it for analysis. The decompressed data is input into the multimodal AI, and data analysis begins. The input is compressed and encrypted data, and the output is analyzable data.
[0220] Step 4:
[0221] The server's multimodal AI extracts features related to the robot's movements from the video data and analyzes any problems with the movements. For example, it analyzes the robot's arm's movement speed, angle, and identifies any errors. The input is analyzable data, and the output is the analysis results, including any problems with the movements.
[0222] Step 5:
[0223] Based on the analysis results, the server identifies specific areas for improvement and the reasons for them. It then generates feedback information to present the areas for improvement to the user. For example, it generates feedback information such as "The arm's movements are too slow, reducing work efficiency." The input is the analysis results, and the output is feedback information.
[0224] Step 6:
[0225] The server generates feedback information and sends it to the terminal. The terminal receives the feedback information and displays it in a visually understandable format for the user, for example, using graphics or animations to demonstrate correct actions or forms. The input is the feedback information, and the output is the feedback content displayed to the user.
[0226] Step 7:
[0227] Based on the displayed feedback information, the user can take specific actions to improve the factory robot's operation. For example, they can adjust the robot's parameters to improve its speed and accuracy. In this step, the feedback information is directly used for on-site improvement actions. The input is the feedback displayed to the user, and the output is the improved robot's operation.
[0228] Furthermore, an emotion engine that estimates the user's emotion may be combined. That is, the identification processing unit 290 may estimate the user's emotion using the emotion identification model 59 and perform identification processing using the user's emotion.
[0229] The present invention relates to a system that combines a photographing means, a processing means, a display means, and an emotion engine to efficiently improve a player's golf form. Below, the processing of the program of this system will be explained in natural language, and an embodiment of the invention will be described in detail with concrete examples.
[0230] System Overview
[0231] This system uses a camera linked to a golf simulator to capture the player's swing form in real time. The captured video data is sent to a server, which then analyzes it using multimodal AI. It also recognizes the user's emotions using an emotion engine and generates feedback information based on the analysis results and emotional state. The feedback information is provided to the user in real time via their device, suggesting areas for improvement in form, optimal practice methods, and even encouraging messages based on their emotions.
[0232] Program processing
[0233] Step 1: Capture video with a camera
[0234] The user stands at the starting position of the golf simulator and begins swinging.
[0235] The device detects the start of the swing and activates the camera in conjunction with the golf simulator.
[0236] The camera captures the player's swing form in real time and acquires video data at a high frame rate.
[0237] Step 2: Sending video data
[0238] The device temporarily stores the captured video data and performs preprocessing, which includes noise reduction and image correction.
[0239] The terminal performs data compression and encryption to send the pre-processed data to the server.
[0240] Step 3: AI-powered video analysis
[0241] The server inputs the received video data into the multimodal AI and begins analysis.
[0242] Multimodal AI extracts features such as swing speed, angle, and wrist movement from video data to detect problems with swing form. For example, if the arms are positioned too high, this will be identified.
[0243] Step 4: Emotion Recognition with the Emotion Engine
[0244] It uses the device's built-in camera and microphone to capture the user's facial expressions and tone of voice.
[0245] An emotion engine analyzes this data to identify the user's emotional state, for example, recognizing whether the user is focused or stressed.
[0246] Step 5: Identify areas for improvement and generate feedback
[0247] The server generates specific feedback information based on the analysis results and the emotional state identified by the emotion engine.
[0248] The feedback information includes not only improvements to form, but also practice methods and encouraging messages based on the user's emotions.
[0249] Step 6: View your feedback
[0250] The server transmits the generated feedback information to the terminal.
[0251] The device receives the feedback information and displays it to the user. The display uses graphics and animations to make it visually easy to understand. It also displays encouraging messages that match the user's emotions.
[0252] Specific examples
[0253] Consider a case where a player (user) is practicing on a golf simulator.
[0254] 1. The user starts swinging. The device detects the start of the swing and the camera captures the swing form.
[0255] 2. The device sends this video data to the server.
[0256] 3. The server receives the video data and analyzes it using multimodal AI, detecting, for example, problems such as the arm being positioned slightly too high.
[0257] 4. The device captures the user's facial expressions and tone of voice, and the emotion engine recognizes that the user is feeling a little stressed.
[0258] 5. The server generates feedback such as "Your arms are too high, making it difficult to make a good impact," and includes an encouraging message such as "Calm down, try lowering your arms a little next time."
[0259] 6. The device receives the feedback information and visually displays it to the user.
[0260] This process allows users to receive immediate feedback and improve their form. Furthermore, encouraging messages that take emotional information into account improve the user's training experience and enable more effective learning.
[0261] The processing flow will be explained below.
[0262] Step 1:
[0263] The user stands at the starting position of the golf simulator and begins swinging.
[0264] The device detects the start of the swing and activates the camera in conjunction with the golf simulator.
[0265] The camera captures the player's swing form in real time and acquires video data at a high frame rate.
[0266] Step 2:
[0267] The device temporarily stores the captured video data and performs preprocessing, which includes noise reduction and image correction.
[0268] The terminal performs data compression and encryption to send the pre-processed data to the server.
[0269] Step 3:
[0270] The server inputs the received video data into the multimodal AI and begins analysis.
[0271] Multimodal AI extracts features such as swing speed, angle, and wrist movement from video data to detect problems with swing form. For example, if the arms are positioned too high, this will be identified.
[0272] Step 4:
[0273] It uses the device's built-in camera and microphone to capture the user's facial expressions and tone of voice.
[0274] The device processes the captured emotion data in real time and sends it to the emotion engine.
[0275] The emotion engine analyzes the user's emotional state from facial expressions and tone of voice, and recognizes, for example, that the user is feeling stressed.
[0276] Step 5:
[0277] The server generates improvement points and feedback information based on the results of video data analysis and the emotional state determined by the emotion engine.
[0278] The server generates feedback such as "Your arm is positioned too high, making it difficult to achieve proper impact," and depending on the user's stress level, may also include encouraging messages such as "Calm down, try lowering your arm a bit next time."
[0279] Step 6:
[0280] The server transmits the generated feedback information to the terminal.
[0281] The device receives the feedback information and displays it visually to the user. Graphics and animations are used to make the display easy to understand. Encouraging messages tailored to the user's emotions are also displayed.
[0282] Step 7:
[0283] The user checks the displayed feedback information and encouraging messages, and practices to correct their swing form next time. They then adjust their form based on the specific areas for improvement indicated in the feedback. Through this series of processes, users can efficiently improve their form and receive support tailored to their emotional state.
[0284] Example 2
[0285] Next, a description will be given of Example 2. In the following description, the data processing device 12 will be referred to as a "server" and the smart device 14 will be referred to as a "terminal."
[0286] To effectively improve golf swing form, it is important to capture the player's movements in real time and provide detailed analysis and immediate feedback. However, conventional systems have had problems with form improvement due to the low accuracy of video data analysis and the low quality of feedback. Furthermore, they lack the ability to provide feedback that takes into account the player's emotional state, which can lead to a loss of motivation.
[0287] The specific processing by the specific processing unit 290 of the data processing device 12 in the second embodiment is realized by the following means.
[0288] In this invention, the server includes a camera for capturing the player's movements in real time, a processor for receiving the captured video data and performing preprocessing such as noise reduction and image correction, a transmitter for compressing and encrypting the preprocessed data and transmitting it to the server, an analyzer for analyzing the received data using multimodal AI and identifying areas for form improvement, an emotion recognition processor for capturing the user's facial expressions and tone of voice based on the analysis results to identify their emotional state, a generator for generating specific feedback information based on the analysis results and their emotional state, and a displayer for providing the generated feedback information to the user. This makes it possible to not only instantly analyze the player's form in detail and provide effective feedback, but also to send appropriate encouraging messages according to the user's emotional state.
[0289] "Photographing means" is a device for capturing the player's movements in real time.
[0290] The "processing means" is a device that receives the captured video data and performs pre-processing such as noise removal and image correction.
[0291] The "transmitting means" is a device that compresses and encrypts the preprocessed data and transmits it to the server.
[0292] The "analysis means" is a device that uses multimodal AI to analyze the received data and extract areas for improvement in form.
[0293] The "emotion recognition means" is a device that captures the user's facial expressions and tone of voice based on the analysis results to identify the user's emotional state.
[0294] The "generating means" is a device that generates specific feedback information based on the analysis results and the emotional state.
[0295] The "display means" is a device for providing the generated feedback information to the user.
[0296] "Multimodal AI" is an artificial intelligence technology that analyzes multiple data modes (such as images and audio) to extract meaning.
[0297] "Feedback information" is information such as improvements to the form and encouraging messages that are provided to the user based on the results obtained from the analysis means and emotion recognition means.
[0298] This invention is a system for efficiently improving a player's golf form, combining a photographing means, a processing means, a transmission means, an analysis means, an emotion recognition means, a generation means, and a display means. This system works in conjunction with a golf simulator to capture the player's swing form in real time and provide immediate feedback based on the analysis results. Furthermore, by taking the user's emotional state into consideration, an effective training environment is realized.
[0299] Hardware and software used
[0300] Camera (photography means)
[0301] This system uses a camera that captures video at a high frame rate, for example, a camera capable of capturing 120 fps (frames per second) is suitable.
[0302] Terminal (processing means and transmission means)
[0303] The device performs pre-processing on the captured video data, including noise reduction and image enhancement. The pre-processed data is then compressed in H.264 format and encrypted with AES, allowing the data to be transmitted efficiently and securely to the server.
[0304] Server (analysis means, emotion recognition means, generation means)
[0305] The server inputs the received data into the multimodal AI for detailed analysis. Specifically, it analyzes swing speed, angle, wrist movement, etc. to identify areas for form improvement. It also has an emotion recognition mechanism that captures the user's facial expressions and tone of voice to identify their emotional state. Specific feedback information is generated based on the analysis results and emotional state.
[0306] Terminal (display means)
[0307] The device visually displays the generated feedback to the user, using graphics and animations to highlight problem areas on the form and providing emotionally tailored encouraging messages.
[0308] Specific examples
[0309] Consider a case where a player (user) is practicing on a golf simulator. When the user starts swinging, the device detects the start of the swing and the camera captures the swing form. The camera records video at 120 fps and temporarily stores the high-quality video in memory.
[0310] The device performs preprocessing such as noise reduction and white balance adjustment, compresses the preprocessed video data in H.264 format, encrypts it using AES, and sends it to the server. The server receives the video data and performs detailed analysis using multimodal AI. For example, the AI can detect problems such as the arm being positioned slightly too high.
[0311] The device then captures the user's facial expressions, and the emotion engine recognizes that the user is feeling a little stressed. The server then points out the problem with the swing form, generating feedback such as, "Your arms are positioned too high, making it difficult to make a proper impact." It also generates encouraging messages such as, "Please stay calm. Next time, try lowering your arms a little."
[0312] The device receives feedback, graphically highlights areas of form that are problematic, and even displays encouraging messages like "Put your arms a little lower."
[0313] Prompt Sentence Examples
[0314] Here are some example prompts to input to a generative AI model:
[0315] "Analyze the golf swing footage of the player and point out any problems with their swing form. Additionally, analyze the player's facial expressions and tone of voice to provide feedback based on the player's emotions."
[0316] These prompts allow you to accurately input the necessary information into the generative AI model and get the desired feedback.
[0317] The flow of the identification process in the second embodiment will be described with reference to FIG.
[0318] Step 1:
[0319] The user stands at the starting position of the golf simulator and begins swinging. The device detects the start of the swing and activates the camera in conjunction with the golf simulator. The camera captures the player's swing form in real time. The input at this time is the user's swing motion, and the output is the captured high-frame-rate video data.
[0320] Step 2:
[0321] The device temporarily stores the captured video data and performs preprocessing such as noise reduction and white balance adjustment. For example, preprocessing removes background noise and adjusts the color tone of the video. The input of this preprocessing is the captured video data, and the output is the preprocessed video data.
[0322] Step 3:
[0323] The terminal compresses the preprocessed video data in H.264 format and encrypts it using the AES method. This reduces the data size and the risk of data leaks. The input is preprocessed video data, and the output is compressed and encrypted video data.
[0324] Step 4:
[0325] The device sends data to the server, using the appropriate protocol to ensure the data travels securely across the network. The input is the compressed and encrypted video data, and the output is the data received by the server.
[0326] Step 5:
[0327] The server inputs the received data into a multimodal AI for detailed analysis. The AI extracts features such as swing speed, angle, and wrist movement to identify problems with the golf form. For example, it detects problems with the arm position being too high. The input for this step is the video data received by the server, and the output is the analyzed form improvements.
[0328] Step 6:
[0329] The device captures the user's facial expressions and tone of voice and inputs this data into an emotion engine, which analyzes whether the user is focused, relaxed, or stressed. The input is the user's facial expressions and tone of voice, and the output is the identified emotional state.
[0330] Step 7:
[0331] The server generates feedback information based on the analysis results and the user's emotional state. It specifically points out areas for improvement in the user's form and creates an encouraging message that corresponds to the user's emotions. For example, it generates a specific message such as "Lower your arm position. Stay calm and try your best next time." The input is the areas for improvement in the user's form and the user's emotional state, and the output is the generated feedback information.
[0332] Step 8:
[0333] The server sends the generated feedback information to the device. The device receives the feedback information and visually displays it to the user. The display uses graphics and animations to highlight problem areas in the swing form and provide encouraging messages. The input is the generated feedback information, and the output is the feedback information displayed to the user.
[0334] (Application example 2)
[0335] Next, a description will be given of Application Example 2. In the following description, the data processing device 12 will be referred to as a "server" and the smart device 14 will be referred to as a "terminal."
[0336] While conventional golf swing analysis systems can pinpoint technical issues in a player's swing form, they are unable to provide feedback that takes into account the player's emotional state. As a result, if a player feels stressed or frustrated, they are unable to receive appropriate advice or encouragement, which can lead to training being ineffective.
[0337] The specific processing by the specific processing unit 290 of the data processing device 12 in Application Example 2 is realized by the following means. In this invention, the server includes a shooting means for capturing the player's actions in real time, a processing means for receiving and analyzing the shot video data, a display means for providing feedback information to the player based on the analysis results, an emotion recognition means for recognizing the player's emotional state in real time, and a feedback generation means for generating feedback information based on the analysis results and the player's emotional state. This makes it possible to provide not only technical problems but also appropriate feedback and encouraging messages according to the player's emotional state.
[0338] "Filming means for capturing player's movements in real time" refers to a device for capturing player's movements, particularly golf swings and other exercises, as video in real time.
[0339] The "processing means for receiving and analyzing captured video data" is a system that receives video data captured by the image capture means, analyzes it, and extracts important features and problems.
[0340] The "display means for providing feedback information to the player based on the analysis results" refers to a device such as a monitor or display for providing visual feedback to the player based on the analysis results.
[0341] The "emotion recognition means for recognizing the player's emotional state in real time" is a system that includes a camera and microphone for recognizing emotions in real time from the player's facial expressions and tone of voice, as well as software for analyzing them.
[0342] "Feedback generation means for generating feedback information based on the analysis results and the emotional state" refers to software or algorithms that generate feedback information to be provided to the player, taking into account both the technical analysis results and the player's emotional state.
[0343] A "camera that captures at a high frame rate" is a camera that can capture a large number of frames per second, thereby capturing the player's movements in detail.
[0344] "Multimodal AI" is an artificial intelligence technology that can simultaneously analyze multiple different types of data (e.g., video data and emotional data) and make integrated judgments.
[0345] "Areas for improvement in form" refers to elements or problems in a player's movements or posture that are hindering efficient or effective performance.
[0346] The present invention relates to a system for efficiently improving a player's golf swing form, which includes a camera for capturing the player's movements in real time, a processing means for receiving and analyzing the captured video data, a display means for providing feedback information to the player based on the analysis results, an emotion recognition means for recognizing the player's emotional state in real time, and a feedback generation means for generating feedback information based on the analysis results and the player's emotional state.
[0347] The server uses a camera to capture the player's swing form in real time. When the player starts swinging, the device detects this and activates the camera. The camera captures the player's swing form at a high frame rate, and the video data is temporarily saved on the device. The saved video data is then compressed and encrypted before being sent to the server. The server then analyzes the received video data using multimodal AI to extract features such as swing speed, angle, and wrist movement.
[0348] The device also uses its built-in camera and microphone to capture the player's facial expressions and tone of voice, and the emotion engine analyzes this data to identify the player's emotional state, for example, whether the player is focused or stressed.
[0349] The server uses a feedback generation means to generate specific feedback information based on the analysis results and the emotional state identified by the emotion engine. The feedback information includes not only improvements to form, but also practice methods and encouraging messages that correspond to the player's emotions. The generated feedback information is sent to the device in real time, and the device visually displays it to the player. The display uses graphics and animations to make it visually easy to understand. Encouraging messages that match the player's emotions are also displayed.
[0350] As a concrete example, consider a player practicing on a golf simulator. When the player starts swinging, the device detects this and the camera captures the swing form. This video data is sent to a server and analyzed by multimodal AI. For example, the analysis may detect a problem: "Your arms are positioned a little too high." The device also captures the player's facial expressions and tone of voice, and the emotion engine recognizes that the player is feeling a little stressed. The server generates feedback such as, "Your arms are positioned too high, making it difficult to make a proper impact," and includes an encouraging message such as, "Please calm down. Next time, try lowering your arms a little." This feedback information is sent to the device and displayed visually to the player.
[0351] Examples of prompts to input to a generative AI model include:
[0352] "Based on the player's swing footage and emotional data, generate feedback suggesting form improvements and appropriate practice methods. For example, "Your arms are positioned too high, making it difficult to make a proper impact. Calm down and try lowering your arms a little next time."
[0353] In this way, this system pinpoints a player's technical problems with high accuracy while providing training support that also takes into account their emotional side.
[0354] The flow of the specific processing in the application example 2 will be described with reference to FIG.
[0355] Step 1:
[0356] The player stands at the starting position of the golf simulator and begins swinging. The device detects the start of the swing and activates the camera in conjunction with the golf simulator. The camera captures the player's swing form at a high frame rate and obtains the video data. The input of this step is the player's swing motion, and the output is the captured video data.
[0357] Step 2:
[0358] The device temporarily stores the captured video data and performs preprocessing such as noise reduction and image correction. Before sending the processed data to the server, the data is compressed and encrypted. The input of this step is the captured video data, and the output is the compressed and encrypted video data.
[0359] Step 3:
[0360] The server inputs the received video data into a multimodal AI to extract features such as swing speed, angle, and wrist movement. Through AI analysis, problems with the swing form are identified. The input for this step is compressed and encrypted video data, and the output is the analyzed swing form feature data and problems.
[0361] Step 4:
[0362] The device's built-in camera and microphone are used to capture the user's facial expressions and tone of voice. The emotion engine analyzes this data to identify the user's emotional state. For example, it recognizes whether the user is focused or stressed. The input for this step is facial expression and tone of voice data, and the output is the recognized emotional state.
[0363] Step 5:
[0364] The server generates specific feedback information based on the analysis results and the emotional state identified by the emotion engine. Using the feedback generation means, feedback is generated that includes form improvements, practice methods based on the user's emotions, and encouraging messages. The input for this step is the analyzed swing form feature data and the recognized emotional state, and the output is the generated feedback information.
[0365] Step 6:
[0366] The server sends the generated feedback information to the terminal. The terminal receives the feedback information and visually displays it to the user. This display uses graphics and animations, and also displays encouraging messages that match the user's emotions. The input of this step is the generated feedback information, and the output is the feedback information displayed to the user.
[0367] This series of processing steps allows users to receive instant feedback and effectively improve their form, while encouraging messages that take emotional information into account significantly improve the training experience.
[0368] The specific processing unit 290 transmits the result of the specific processing to the smart device 14. In the smart device 14, the control unit 46A causes the output device 40 to output the result of the specific processing. The microphone 38B acquires audio indicating a user input regarding the result of the specific processing. The control unit 46A transmits audio data indicating the user input acquired by the microphone 38B to the data processing device 12. In the data processing device 12, the specific processing unit 290 acquires the audio data.
[0369] The data generation model 58 is a so-called generative AI (Artificial Intelligence). An example of the data generation model 58 is ChatGPT (Internet Search<URL: https: / / openai.com / blog / chatgpt> ), Gemini (Internet search <url: https: gemini.google.com ?hl="ja">) and other generation AIs. The data generation model 58 is obtained by performing deep learning on a neural network. A prompt including an instruction is input to the data generation model 58, and inference data such as voice data indicating voice, text data indicating text, and image data indicating an image is also input. The data generation model 58 performs inference on the input inference data in accordance with the instruction indicated by the prompt, and outputs the inference result in a data format such as voice data and text data. Here, inference refers to, for example, analysis, classification, prediction, and / or summarization.
[0370] In the above embodiment, an example in which the specific process is performed by the data processing device 12 has been given, but the technology of the present disclosure is not limited to this, and the specific process may be performed by the smart device 14.
[0371] [Second embodiment]
[0372] FIG. 3 shows an example of the configuration of a data processing system 210 according to the second embodiment.
[0373] 3, the data processing system 210 includes the data processing device 12 and smart glasses 214. An example of the data processing device 12 is a server.
[0374] The data processing device 12 includes a computer 22, a database 24, and a communication I / F 26. The computer 22 is an example of a "computer" according to the technology of the present disclosure. The computer 22 includes a processor 28, a RAM 30, and a storage 32. The processor 28, the RAM 30, and the storage 32 are connected to a bus 34. The database 24 and the communication I / F 26 are also connected to the bus 34. The communication I / F 26 is connected to a network 54. Examples of the network 54 include a WAN (Wide Area Network) and / or a LAN (Local Area Network).
[0375] The smart glasses 214 include a computer 36, a microphone 238, a speaker 240, a camera 42, and a communication I / F 44. The computer 36 includes a processor 46, a RAM 48, and a storage 50. The processor 46, the RAM 48, and the storage 50 are connected to a bus 52. The microphone 238, the speaker 240, and the camera 42 are also connected to the bus 52.
[0376] The microphone 238 receives instructions and the like from the user 20 by receiving voice uttered by the user 20. The microphone 238 captures the voice uttered by the user 20, converts the captured voice into audio data, and outputs it to the processor 46. The speaker 240 outputs audio in accordance with instructions from the processor 46.
[0377] Camera 42 is a small digital camera equipped with an optical system including a lens, aperture, and shutter, and an imaging element such as a CMOS (Complementary Metal-Oxide-Semiconductor) image sensor or a CCD (Charge Coupled Device) image sensor, and captures images of the surroundings of user 20 (for example, an imaging range defined by an angle of view equivalent to the field of vision of a typical healthy person).
[0378] The communication I / F 44 is connected to a network 54. The communication I / Fs 44 and 26 are responsible for the exchange of various information between the processor 46 and the processor 28 via the network 54. The exchange of various information between the processor 46 and the processor 28 using the communication I / Fs 44 and 26 is carried out in a secure state.
[0379] Fig. 4 shows an example of the main functions of the data processing device 12 and the smart glasses 214. As shown in Fig. 4, in the data processing device 12, a specific process is performed by the processor 28. A specific process program 56 is stored in the storage 32.
[0380] The specific processing program 56 is an example of a "program" according to the technology of the present disclosure. The processor 28 reads the specific processing program 56 from the storage 32 and executes the read specific processing program 56 on the RAM 30. The specific processing is realized by the processor 28 operating as a specific processing unit 290 in accordance with the specific processing program 56 executed on the RAM 30.
[0381] The storage 32 stores a data generation model 58 and an emotion identification model 59. The data generation model 58 and the emotion identification model 59 are used by the identification processing unit 290.
[0382] In the smart glasses 214, the reception output process is performed by the processor 46. A reception output program 60 is stored in the storage 50. The processor 46 reads the reception output program 60 from the storage 50 and executes the read reception output program 60 on the RAM 48. The reception output process is realized by the processor 46 operating as the control unit 46A in accordance with the reception output program 60 executed on the RAM 48.
[0383] Next, a description will be given of the identification process performed by the identification processing unit 290 of the data processing device 12. In the following description, the data processing device 12 will be referred to as the "server" and the smart glasses 214 will be referred to as the "terminal."
[0384] The present invention is a system that combines a photographing means, a processing means, and a display means to efficiently improve a player's golf form. Below, the processing of the program of this system will be explained in natural language, and an embodiment of the invention will be described in detail with concrete examples.
[0385] System Overview
[0386] This system uses a camera linked to a golf simulator to capture the player's swing form in real time. The captured video data is sent to a server, which then analyzes the data using multimodal AI. The analysis results are fed back to the player in real time via their device, suggesting areas for form improvement and optimal practice methods.
[0387] Program processing
[0388] Step 1: Capture video with a camera
[0389] The user stands at the starting position of the golf simulator and begins swinging.
[0390] The device detects the start of the swing and activates the camera in conjunction with the golf simulator.
[0391] The camera captures the player's swing form in real time and acquires video data at a high frame rate.
[0392] Step 2: Sending video data
[0393] The device temporarily stores the captured video data and performs pre-processing, after which the data is compressed and encrypted for transmission to the server.
[0394] Step 3: AI-powered video analysis
[0395] The server inputs the received video data into the multimodal AI and begins analysis.
[0396] The AI extracts features from the video data, such as swing speed, angle, and wrist movement, to detect problems with the swing form. For example, if the arms are positioned too high, this problem will be identified.
[0397] Step 4: Identify areas for improvement and provide feedback
[0398] The server then generates feedback based on the analysis, identifying specific areas for improvement and the reasons for them. For example, the feedback could be, "Your arms are positioned too high, making it difficult to make proper impact with the ball."
[0399] Specific practice methods (e.g., specific drills to correct arm position) are also generated as feedback information.
[0400] Step 5: View your feedback
[0401] The server transmits the generated feedback information to the terminal.
[0402] The device receives the feedback information and displays it to the user, using graphics and animations to make it visually easy to understand.
[0403] Specific examples
[0404] Consider a case where a player (user) is practicing on a golf simulator.
[0405] 1. The user starts swinging. The device detects the start of the swing and the camera captures the swing form.
[0406] 2. The device sends this video data to the server.
[0407] 3. The server receives the video data and analyzes it using multimodal AI, detecting, for example, problems such as the arm being positioned slightly too high.
[0408] 4. The server generates feedback such as "Your arms are positioned too high, making it difficult to make a proper impact," and sends it to the device along with specific practice instructions.
[0409] 5. The device displays feedback information to the user in an intuitive format, for example, by providing a graphic indication of arm position and visually demonstrating correct form.
[0410] This allows users to receive instant feedback and use it to improve their form. By repeating this process, players can efficiently improve their golfing skills.
[0411] The processing flow will be explained below.
[0412] Step 1:
[0413] The user stands at the starting position of the golf simulator and begins swinging.
[0414] The device detects the start of the swing and activates the camera in conjunction with the golf simulator.
[0415] The camera captures the player's swing form in real time and acquires video data at a high frame rate.
[0416] Step 2:
[0417] The device temporarily stores the captured video data and performs preprocessing, which includes noise reduction and image correction.
[0418] The terminal performs data compression and encryption to send the pre-processed data to the server.
[0419] Step 3:
[0420] The server inputs the received video data into the multimodal AI and begins analysis.
[0421] Multimodal AI extracts features such as swing speed, angle, and wrist movement from video data to detect problems with swing form. For example, if the arms are positioned too high, this will be identified.
[0422] Step 4:
[0423] The server then uses the analysis results to identify specific areas for improvement and the reasons for them, such as determining that "your arms are positioned too high, making it difficult to make proper impact with the ball."
[0424] The server generates specific feedback based on the identified areas for improvement, including specific practice methods (e.g., specific drills to correct arm position).
[0425] Step 5:
[0426] The server transmits the generated feedback information to the terminal.
[0427] The device receives the feedback information and displays it to the user, using graphics and animations to make it visually easy to understand.
[0428] Step 6:
[0429] The user reviews the displayed feedback information and practices to correct their next swing form. They adjust their form based on the specific improvements indicated in the feedback.
[0430] This series of steps allows users to efficiently improve their swing form, and real-time feedback can help them improve their technique quickly.
[0431] Example 1
[0432] Next, a description will be given of Example 1. In the following description, the data processing device 12 will be referred to as a "server" and the smart glasses 214 will be referred to as a "terminal."
[0433] For players to efficiently improve their form, it is important to accurately understand their current form and immediately identify specific areas for improvement. However, with conventional technology, analyzing and evaluating form takes time, making it difficult to provide real-time feedback. In addition, the analysis results are not specific, making it difficult for players to understand how to improve.
[0434] The specific processing by the specific processing unit 290 of the data processing device 12 in the first embodiment is realized by the following means.
[0435] In this invention, the server includes a camera for capturing the player's movements in real time, a terminal for temporarily storing and preprocessing the captured video data, a transmission means for compressing and encrypting the preprocessed data and transmitting it to the server, a processing means for analyzing the video data received by the server using multimodal AI to detect problems with the player's swing form, a feedback generation means for identifying specific areas for improvement based on the analysis results and generating feedback information, and a means for transmitting the generated feedback information to the terminal and displaying it to the user. This allows the player to receive specific and detailed feedback in real time, enabling them to improve their form efficiently.
[0436] A "player" is someone who uses the system to improve their own movements and form.
[0437] "Movement" refers to the physical movements of the player related to their swing and form.
[0438] "Real-time" refers to immediate processing and feedback without delay.
[0439] "Photography means" refers to a device used to capture the player's actions, such as a camera.
[0440] "Terminal" refers to a device that temporarily stores, pre-processes, and transmits data, typically a computer or smart device.
[0441] "Transmission means" refers to the process or function that transmits the pre-processed data to the server.
[0442] "Server" refers to a device or system that analyzes video data and generates feedback information.
[0443] "Multimodal AI" refers to artificial intelligence technology that comprehensively analyzes multiple input data (video, audio, sensors, etc.).
[0444] "Processing means" refers to a device or process that has the function of analyzing received video data and detecting problems.
[0445] "Feedback generation means" refers to the process or function that generates specific improvements and practice methods based on the analysis results.
[0446] The "display means" refers to a device or system that has the function of visually displaying the generated feedback information to the user.
[0447] "Video data" refers to video data that captures the player's actions.
[0448] The present invention is a system that captures a player's actions in real time and provides feedback based on the analysis results. Specifically, it is realized by linking a camera, a terminal, a server, and a display.
[0449] The system overview includes the following hardware and software:
[0450] Filming method
[0451] The recording method consists of a camera that captures the player's swing form at a high frame rate. This camera operates at a high frame rate such as 120 fps, allowing for detailed capture of the player's movements. For example, it is preferable to use a high-performance sports camera or a dedicated motion capture camera rather than a regular webcam.
[0452] Terminal
[0453] The terminal is a device that temporarily stores and pre-processes the video data captured by the camera. This pre-processing includes noise reduction and color correction. The video data is then compressed into MPEG format and encrypted with AES. This allows the data to be sent securely to the server while maintaining its quality. The terminal requires a high-performance processor and large-capacity storage.
[0454] Transmission method
[0455] The transmission means has the function of sending preprocessed data from the terminal to the server. It is important to use the HTTPS protocol for data transmission, ensuring safety and speed.
[0456] server
[0457] The server includes a processing means for analyzing the received video data using multimodal AI. This AI is built using deep learning frameworks such as TensorFlow and PyTorch, and performs detailed analysis of swing speed, angle, wrist movement, and other parameters. As a result of the analysis, it detects problems with the swing form and generates feedback. The server then generates feedback that includes specific practice methods and corrections.
[0458] Feedback Generation Method
[0459] The feedback generator explains specific areas for improvement and the reasons for them based on the analysis results from the server. For example, it may give specific indications such as, "Your right arm is too high, making your impact unstable." It also suggests practice methods for keeping the right arm low.
[0460] Display means
[0461] The display means is a device or system that visually displays the generated feedback information to the user. The terminal receives the feedback information and uses graphics or animations to show it to the user in an easy-to-understand manner. For example, by visually displaying the difference between incorrect and correct form, the user can understand specifically which parts need to be corrected.
[0462] Specific examples
[0463] For example, when a player practices on a golf simulator, the system operates in the following steps:
[0464] 1. The user starts swinging, the device detects the movement, and the camera captures the swing form at a high frame rate.
[0465] 2. The device temporarily stores the video data, performs preprocessing, compresses it into MPEG format, encrypts it using AES, and sends it to the server.
[0466] 3. The server receives the video data and analyzes it using multimodal AI, detecting problems such as "the right arm is raised too high."
[0467] 4. Based on the analysis results, the server generates feedback recommending "practice swinging with both arms fixed at waist height to practice keeping the right arm low" and sends it to the device.
[0468] 5. The device displays the feedback information in a graphical format, visually showing the user specific areas for improvement and correct form.
[0469] Prompt Sentence Examples
[0470] The user can enter prompts for the generative AI model, such as:
[0471] "Analyze the swing footage of the player and if there are any problems with their swing form, tell them how to improve it."
[0472] Thus, the present invention is a system that can provide detailed feedback in real time to help players improve their form efficiently.
[0473] The flow of the identification process in the first embodiment will be described with reference to FIG.
[0474] The flow of this system's program processing
[0475] Step 1:
[0476] The user stands at the starting position of the golf simulator and begins swinging.
[0477] Input: User swing motion
[0478] Operation: The device uses sensors to recognize when the user starts swinging.
[0479] Output: Swing start detection signal
[0480] Step 2:
[0481] The device activates the camera and captures the player's swing form in real time.
[0482] Input: Swing start detection signal
[0483] How it works: The device controls the camera and captures video data of the swing form at a high frame rate such as 120 fps.
[0484] Output: Captured video data
[0485] Step 3:
[0486] The device temporarily stores the captured video data and performs preprocessing.
[0487] Input: Captured video data
[0488] How it works: The device breaks down the video data into frames and performs noise reduction and color correction.
[0489] Output: Pre-processed video data
[0490] Step 4:
[0491] The terminal compresses and encrypts the pre-processed video data.
[0492] Input: Preprocessed video data
[0493] Operation: Video data is compressed into MPEG format and encrypted with AES.
[0494] Output: Compressed and encrypted video data
[0495] Step 5:
[0496] The terminal transmits the compressed and encrypted video data to the server.
[0497] Input: Compressed and encrypted video data
[0498] How it works: The device uses the HTTPS protocol to securely upload data to the server.
[0499] Output: Video data sent to the server
[0500] Step 6:
[0501] The video data received by the server is analyzed using multimodal AI.
[0502] Input: Received video data
[0503] How it works: The server inputs video data into the AI model and extracts swing form characteristics (e.g., speed, angle, wrist movement).
[0504] Output: Swing form feature data
[0505] Step 7:
[0506] The server detects problems based on characteristic data of the swing form.
[0507] Input: Swing form characteristics data
[0508] How it works: The server compares the extracted features with predefined standards and identifies form flaws (e.g., the right arm is too high).
[0509] Output: Detected issue data
[0510] Step 8:
[0511] The server generates feedback information based on the analysis results.
[0512] Input: Detected issue data
[0513] How it works: The server generates feedback with specific areas for improvement and reasons for doing so, and also suggests practice methods.
[0514] Output: Generated feedback information
[0515] Step 9:
[0516] The server transmits the generated feedback information to the terminal.
[0517] Input: Generated feedback information
[0518] Operation: The server sends feedback information to the device in JSON format or similar.
[0519] Output: Feedback information sent to the terminal
[0520] Step 10:
[0521] The terminal displays the feedback information to the user.
[0522] Input: Feedback information sent to the device
[0523] Action: The device visually displays feedback information to the user using graphics and animations.
[0524] Output: Feedback information displayed to the user
[0525] (Application example 1)
[0526] Next, a description will be given of Application Example 1. In the following description, the data processing device 12 will be referred to as a "server" and the smart glasses 214 will be referred to as a "terminal."
[0527] Conventional technologies exist that provide real-time feedback on player movements and form improvements. However, there is a need for a system that can optimize the movement accuracy and efficiency of robots and equipment in real time, even in factory operations. The present invention provides technology that improves movement accuracy and efficiency by capturing, analyzing, and providing feedback on the movements of factory robots.
[0528] The specific processing by the specific processing unit 290 of the data processing device 12 in the application example 1 is realized by the following means.
[0529] In this invention, the server includes a camera for capturing the actions of the player or the actions of the equipment in real time, a processing unit for receiving and analyzing the captured video data, and a display unit for providing feedback information to the player or operator based on the analysis results, thereby enabling efficient improvement of the operations of factory robots and equipment.
[0530] "Photographing means" refers to a device for capturing the actions of a player or device in real time.
[0531] A "processing means" is a device or system for receiving and analyzing captured video data.
[0532] "Display means" refers to a device or system for providing feedback information to a player or operator based on the analysis results.
[0533] "Multimodal AI" is an artificial intelligence technology that integrates and analyzes multiple data modalities (e.g., video, audio, text).
[0534] "Feedback" is information based on the analysis results that provides improvements to movements and form, as well as optimal practice methods.
[0535] "Capture" refers to the process of converting a subject's movements and form into data in real time using a photographic device.
[0536] A "server" is a computer system that receives data, analyzes it, and generates feedback information.
[0537] The present invention is a system for optimizing the operation of a factory robot, and includes a photographing means for capturing the operation of a player or equipment in real time, a processing means for receiving and analyzing the photographed video data, and a display means for providing feedback information to a player or operator based on the analysis results.
[0538] System Program
[0539] The main hardware and software components of this system include smartphones, factory robots, smartphone apps (e.g., cross-platform development with Flutter or React Native), servers, multimodal AI (e.g., TensorFlow or PyTorch), and communication protocols (e.g., HTTPS or WebSocket).
[0540] Natural language explanation of program processing
[0541] Step 1: Capture video with a camera
[0542] The smartphone camera captures the robot's movements in the factory in real time, capturing the video data at a high frame rate.
[0543] Step 2: Sending video data
[0544] The captured video data is temporarily stored on the smartphone and then encrypted and sent to a server.
[0545] Step 3: AI-powered video analysis
[0546] The server inputs the received video data into the multimodal AI for analysis. The AI extracts characteristics of the robot's movements from the video data and detects problems with the movements. For example, if the robot's arm is moving too slowly, this problem will be identified.
[0547] Step 4: Identify areas for improvement and provide feedback
[0548] Based on the analysis results, the server identifies specific areas for improvement and the reasons for them, and generates feedback information. For example, the feedback might say, "The arm's movements are too slow, reducing work efficiency." This feedback might also include specific drills to improve the arm's speed by 20%.
[0549] Step 5: View your feedback
[0550] The server sends the generated feedback information to a smartphone and displays it for the user or operator to check. This display uses graphics and animations to make it visually easy to understand.
[0551] Specific examples
[0552] For example, if the robot arm is analyzed as moving too slowly, the app will display the following feedback:
[0553] "The arm is moving too slowly, reducing work efficiency. Please perform drills to increase the arm's movement speed by 20%."
[0554] AI prompt examples
[0555] "Analyze the robot arm's movement pattern in this video, identify the speed, angle, and errors, and suggest improvements based on that."
[0556] The flow of the specific processing in the application example 1 will be described with reference to FIG.
[0557] Step 1:
[0558] A user captures the operation of a factory robot in real time using a smartphone camera. The smartphone camera acquires video data at a high frame rate and temporarily stores the data within the smartphone. The input is the captured video data, and the output is the temporarily stored video data.
[0559] Step 2:
[0560] The device preprocesses the stored video data. Preprocessing includes compressing and encrypting the data. The compressed and encrypted data is then sent to the server. The input is the stored video data, and the output is the compressed and encrypted data.
[0561] Step 3:
[0562] The server decompresses the received video data and prepares it for analysis. The decompressed data is input into the multimodal AI, and data analysis begins. The input is compressed and encrypted data, and the output is analyzable data.
[0563] Step 4:
[0564] The server's multimodal AI extracts features related to the robot's movements from the video data and analyzes any problems with the movements. For example, it analyzes the robot's arm's movement speed, angle, and identifies any errors. The input is analyzable data, and the output is the analysis results, including any problems with the movements.
[0565] Step 5:
[0566] Based on the analysis results, the server identifies specific areas for improvement and the reasons for them. It then generates feedback information to present the areas for improvement to the user. For example, it generates feedback information such as "The arm's movements are too slow, reducing work efficiency." The input is the analysis results, and the output is feedback information.
[0567] Step 6:
[0568] The server generates feedback information and sends it to the terminal. The terminal receives the feedback information and displays it in a visually understandable format for the user, for example, using graphics or animations to demonstrate correct actions or forms. The input is the feedback information, and the output is the feedback content displayed to the user.
[0569] Step 7:
[0570] Based on the displayed feedback information, the user can take specific actions to improve the factory robot's operation. For example, they can adjust the robot's parameters to improve its speed and accuracy. In this step, the feedback information is directly used for on-site improvement actions. The input is the feedback displayed to the user, and the output is the improved robot's operation.
[0571] Furthermore, an emotion engine that estimates the user's emotion may be further combined. That is, the identification processing unit 290 may estimate the user's emotion using the emotion identification model 59, and perform identification processing using the user's emotion.
[0572] The present invention relates to a system that combines a photographing means, a processing means, a display means, and an emotion engine to efficiently improve a player's golf form. Below, the processing of the program of this system will be explained in natural language, and an embodiment of the invention will be described in detail with concrete examples.
[0573] System Overview
[0574] This system uses a camera linked to a golf simulator to capture the player's swing form in real time. The captured video data is sent to a server, which then analyzes it using multimodal AI. It also recognizes the user's emotions using an emotion engine and generates feedback information based on the analysis results and emotional state. The feedback information is provided to the user in real time via their device, suggesting areas for improvement in form, optimal practice methods, and even encouraging messages based on their emotions.
[0575] Program processing
[0576] Step 1: Capture video with a camera
[0577] The user stands at the starting position of the golf simulator and begins swinging.
[0578] The device detects the start of the swing and activates the camera in conjunction with the golf simulator.
[0579] The camera captures the player's swing form in real time and acquires video data at a high frame rate.
[0580] Step 2: Sending video data
[0581] The device temporarily stores the captured video data and performs preprocessing, which includes noise reduction and image correction.
[0582] The terminal performs data compression and encryption to send the pre-processed data to the server.
[0583] Step 3: AI-powered video analysis
[0584] The server inputs the received video data into the multimodal AI and begins analysis.
[0585] Multimodal AI extracts features such as swing speed, angle, and wrist movement from video data to detect problems with swing form. For example, if the arms are positioned too high, this will be identified.
[0586] Step 4: Emotion Recognition with the Emotion Engine
[0587] It uses the device's built-in camera and microphone to capture the user's facial expressions and tone of voice.
[0588] An emotion engine analyzes this data to identify the user's emotional state, for example, recognizing whether the user is focused or stressed.
[0589] Step 5: Identify areas for improvement and generate feedback
[0590] The server generates specific feedback information based on the analysis results and the emotional state identified by the emotion engine.
[0591] The feedback information includes not only improvements to form, but also practice methods and encouraging messages based on the user's emotions.
[0592] Step 6: View your feedback
[0593] The server transmits the generated feedback information to the terminal.
[0594] The device receives the feedback information and displays it to the user. The display uses graphics and animations to make it visually easy to understand. It also displays encouraging messages that match the user's emotions.
[0595] Specific examples
[0596] Consider a case where a player (user) is practicing on a golf simulator.
[0597] 1. The user starts swinging. The device detects the start of the swing and the camera captures the swing form.
[0598] 2. The device sends this video data to the server.
[0599] 3. The server receives the video data and analyzes it using multimodal AI, detecting, for example, problems such as the arm being positioned slightly too high.
[0600] 4. The device captures the user's facial expressions and tone of voice, and the emotion engine recognizes that the user is feeling a little stressed.
[0601] 5. The server generates feedback such as "Your arms are too high, making it difficult to make a good impact," and includes an encouraging message such as "Calm down, try lowering your arms a little next time."
[0602] 6. The device receives the feedback information and visually displays it to the user.
[0603] This process allows users to receive immediate feedback and improve their form. Furthermore, encouraging messages that take emotional information into account improve the user's training experience and enable more effective learning.
[0604] The processing flow will be explained below.
[0605] Step 1:
[0606] The user stands at the starting position of the golf simulator and begins swinging.
[0607] The device detects the start of the swing and activates the camera in conjunction with the golf simulator.
[0608] The camera captures the player's swing form in real time and acquires video data at a high frame rate.
[0609] Step 2:
[0610] The device temporarily stores the captured video data and performs preprocessing, which includes noise reduction and image correction.
[0611] The terminal performs data compression and encryption to send the pre-processed data to the server.
[0612] Step 3:
[0613] The server inputs the received video data into the multimodal AI and begins analysis.
[0614] Multimodal AI extracts features such as swing speed, angle, and wrist movement from video data to detect problems with swing form. For example, if the arms are positioned too high, this will be identified.
[0615] Step 4:
[0616] It uses the device's built-in camera and microphone to capture the user's facial expressions and tone of voice.
[0617] The device processes the captured emotion data in real time and sends it to the emotion engine.
[0618] The emotion engine analyzes the user's emotional state from facial expressions and tone of voice, and recognizes, for example, that the user is feeling stressed.
[0619] Step 5:
[0620] The server generates improvement points and feedback information based on the results of video data analysis and the emotional state determined by the emotion engine.
[0621] The server generates feedback such as "Your arm is positioned too high, making it difficult to achieve proper impact," and depending on the user's stress level, may also include encouraging messages such as "Calm down, try lowering your arm a bit next time."
[0622] Step 6:
[0623] The server transmits the generated feedback information to the terminal.
[0624] The device receives the feedback information and displays it visually to the user. Graphics and animations are used to make the display easy to understand. Encouraging messages tailored to the user's emotions are also displayed.
[0625] Step 7:
[0626] The user checks the displayed feedback information and encouraging messages, and practices to correct their swing form next time. They then adjust their form based on the specific areas for improvement indicated in the feedback. Through this series of processes, users can efficiently improve their form and receive support tailored to their emotional state.
[0627] Example 2
[0628] Next, a description will be given of Example 2. In the following description, the data processing device 12 will be referred to as a "server" and the smart glasses 214 will be referred to as a "terminal."
[0629] To effectively improve golf swing form, it is important to capture the player's movements in real time and provide detailed analysis and immediate feedback. However, conventional systems have had problems with form improvement due to the low accuracy of video data analysis and the low quality of feedback. Furthermore, they lack the ability to provide feedback that takes into account the player's emotional state, which can lead to a loss of motivation.
[0630] The specific processing by the specific processing unit 290 of the data processing device 12 in the second embodiment is realized by the following means.
[0631] In this invention, the server includes a camera for capturing the player's movements in real time, a processor for receiving the captured video data and performing preprocessing such as noise reduction and image correction, a transmitter for compressing and encrypting the preprocessed data and transmitting it to the server, an analyzer for analyzing the received data using multimodal AI and identifying areas for form improvement, an emotion recognition processor for capturing the user's facial expressions and tone of voice based on the analysis results to identify their emotional state, a generator for generating specific feedback information based on the analysis results and their emotional state, and a displayer for providing the generated feedback information to the user. This makes it possible to not only instantly analyze the player's form in detail and provide effective feedback, but also to send appropriate encouraging messages according to the user's emotional state.
[0632] "Photographing means" is a device for capturing the player's movements in real time.
[0633] The "processing means" is a device that receives the captured video data and performs pre-processing such as noise removal and image correction.
[0634] The "transmitting means" is a device that compresses and encrypts the preprocessed data and transmits it to the server.
[0635] The "analysis means" is a device that uses multimodal AI to analyze the received data and extract areas for improvement in form.
[0636] The "emotion recognition means" is a device that captures the user's facial expressions and tone of voice based on the analysis results to identify the user's emotional state.
[0637] The "generating means" is a device that generates specific feedback information based on the analysis results and the emotional state.
[0638] The "display means" is a device for providing the generated feedback information to the user.
[0639] "Multimodal AI" is an artificial intelligence technology that analyzes multiple data modes (such as images and audio) to extract meaning.
[0640] "Feedback information" is information such as improvements to the form and encouraging messages that are provided to the user based on the results obtained from the analysis means and emotion recognition means.
[0641] This invention is a system for efficiently improving a player's golf form, combining a photographing means, a processing means, a transmission means, an analysis means, an emotion recognition means, a generation means, and a display means. This system works in conjunction with a golf simulator to capture the player's swing form in real time and provide immediate feedback based on the analysis results. Furthermore, by taking the user's emotional state into consideration, an effective training environment is realized.
[0642] Hardware and software used
[0643] Camera (photography means)
[0644] This system uses a camera that captures video at a high frame rate, for example, a camera capable of capturing 120 fps (frames per second) is suitable.
[0645] Terminal (processing means and transmission means)
[0646] The device performs pre-processing on the captured video data, including noise reduction and image enhancement. The pre-processed data is then compressed in H.264 format and encrypted with AES, allowing the data to be transmitted efficiently and securely to the server.
[0647] Server (analysis means, emotion recognition means, generation means)
[0648] The server inputs the received data into the multimodal AI for detailed analysis. Specifically, it analyzes swing speed, angle, wrist movement, etc. to identify areas for form improvement. It also has an emotion recognition mechanism that captures the user's facial expressions and tone of voice to identify their emotional state. Specific feedback information is generated based on the analysis results and emotional state.
[0649] Terminal (display means)
[0650] The device visually displays the generated feedback to the user, using graphics and animations to highlight problem areas on the form and providing emotionally tailored encouraging messages.
[0651] Specific examples
[0652] Consider a case where a player (user) is practicing on a golf simulator. When the user starts swinging, the device detects the start of the swing and the camera captures the swing form. The camera records video at 120 fps and temporarily stores the high-quality video in memory.
[0653] The device performs preprocessing such as noise reduction and white balance adjustment, compresses the preprocessed video data in H.264 format, encrypts it using AES, and sends it to the server. The server receives the video data and performs detailed analysis using multimodal AI. For example, the AI can detect problems such as the arm being positioned slightly too high.
[0654] The device then captures the user's facial expressions, and the emotion engine recognizes that the user is feeling a little stressed. The server then points out the problem with the swing form, generating feedback such as, "Your arms are positioned too high, making it difficult to make a proper impact." It also generates encouraging messages such as, "Please stay calm. Next time, try lowering your arms a little."
[0655] The device receives feedback, graphically highlights areas of form that are problematic, and even displays encouraging messages like "Put your arms a little lower."
[0656] Prompt Sentence Examples
[0657] Here are some example prompts to input to a generative AI model:
[0658] "Analyze the golf swing footage of the player and point out any problems with their swing form. Additionally, analyze the player's facial expressions and tone of voice to provide feedback based on the player's emotions."
[0659] These prompts allow you to accurately input the necessary information into the generative AI model and get the desired feedback.
[0660] The flow of the identification process in the second embodiment will be described with reference to FIG.
[0661] Step 1:
[0662] The user stands at the starting position of the golf simulator and begins swinging. The device detects the start of the swing and activates the camera in conjunction with the golf simulator. The camera captures the player's swing form in real time. The input at this time is the user's swing motion, and the output is the captured high-frame-rate video data.
[0663] Step 2:
[0664] The device temporarily stores the captured video data and performs preprocessing such as noise reduction and white balance adjustment. For example, preprocessing removes background noise and adjusts the color tone of the video. The input of this preprocessing is the captured video data, and the output is the preprocessed video data.
[0665] Step 3:
[0666] The terminal compresses the preprocessed video data in H.264 format and encrypts it using the AES method. This reduces the data size and the risk of data leaks. The input is preprocessed video data, and the output is compressed and encrypted video data.
[0667] Step 4:
[0668] The device sends data to the server, using the appropriate protocol to ensure the data travels securely across the network. The input is the compressed and encrypted video data, and the output is the data received by the server.
[0669] Step 5:
[0670] The server inputs the received data into a multimodal AI for detailed analysis. The AI extracts features such as swing speed, angle, and wrist movement to identify problems with the golf form. For example, it detects problems with the arm position being too high. The input for this step is the video data received by the server, and the output is the analyzed form improvements.
[0671] Step 6:
[0672] The device captures the user's facial expressions and tone of voice and inputs this data into an emotion engine, which analyzes whether the user is focused, relaxed, or stressed. The input is the user's facial expressions and tone of voice, and the output is the identified emotional state.
[0673] Step 7:
[0674] The server generates feedback information based on the analysis results and the user's emotional state. It specifically points out areas for improvement in the user's form and creates an encouraging message that corresponds to the user's emotions. For example, it generates a specific message such as "Lower your arm position. Stay calm and try your best next time." The input is the areas for improvement in the user's form and the user's emotional state, and the output is the generated feedback information.
[0675] Step 8:
[0676] The server sends the generated feedback information to the device. The device receives the feedback information and visually displays it to the user. The display uses graphics and animations to highlight problem areas in the swing form and provide encouraging messages. The input is the generated feedback information, and the output is the feedback information displayed to the user.
[0677] (Application example 2)
[0678] Next, a description will be given of Application Example 2. In the following description, the data processing device 12 will be referred to as a "server" and the smart glasses 214 will be referred to as a "terminal."
[0679] While conventional golf swing analysis systems can pinpoint technical issues in a player's swing form, they are unable to provide feedback that takes into account the player's emotional state. As a result, if a player feels stressed or frustrated, they are unable to receive appropriate advice or encouragement, which can lead to training being ineffective.
[0680] The specific processing by the specific processing unit 290 of the data processing device 12 in Application Example 2 is realized by the following means. In this invention, the server includes a shooting means for capturing the player's actions in real time, a processing means for receiving and analyzing the shot video data, a display means for providing feedback information to the player based on the analysis results, an emotion recognition means for recognizing the player's emotional state in real time, and a feedback generation means for generating feedback information based on the analysis results and the player's emotional state. This makes it possible to provide not only technical problems but also appropriate feedback and encouraging messages according to the player's emotional state.
[0681] "Filming means for capturing player's movements in real time" refers to a device for capturing player's movements, particularly golf swings and other exercises, as video in real time.
[0682] The "processing means for receiving and analyzing captured video data" is a system that receives video data captured by the image capture means, analyzes it, and extracts important features and problems.
[0683] The "display means for providing feedback information to the player based on the analysis results" refers to a device such as a monitor or display for providing visual feedback to the player based on the analysis results.
[0684] The "emotion recognition means for recognizing the player's emotional state in real time" is a system that includes a camera and microphone for recognizing emotions in real time from the player's facial expressions and tone of voice, as well as software for analyzing them.
[0685] "Feedback generation means for generating feedback information based on the analysis results and the emotional state" refers to software or algorithms that generate feedback information to be provided to the player, taking into account both the technical analysis results and the player's emotional state.
[0686] A "camera that captures at a high frame rate" is a camera that can capture a large number of frames per second, thereby capturing the player's movements in detail.
[0687] "Multimodal AI" is an artificial intelligence technology that can simultaneously analyze multiple different types of data (e.g., video data and emotional data) and make integrated judgments.
[0688] "Areas for improvement in form" refers to elements or problems in a player's movements or posture that are hindering efficient or effective performance.
[0689] The present invention relates to a system for efficiently improving a player's golf swing form, which includes a camera for capturing the player's movements in real time, a processing means for receiving and analyzing the captured video data, a display means for providing feedback information to the player based on the analysis results, an emotion recognition means for recognizing the player's emotional state in real time, and a feedback generation means for generating feedback information based on the analysis results and the player's emotional state.
[0690] The server uses a camera to capture the player's swing form in real time. When the player starts swinging, the device detects this and activates the camera. The camera captures the player's swing form at a high frame rate, and the video data is temporarily saved on the device. The saved video data is then compressed and encrypted before being sent to the server. The server then analyzes the received video data using multimodal AI to extract features such as swing speed, angle, and wrist movement.
[0691] The device also uses its built-in camera and microphone to capture the player's facial expressions and tone of voice, and the emotion engine analyzes this data to identify the player's emotional state, for example, whether the player is focused or stressed.
[0692] The server uses a feedback generation means to generate specific feedback information based on the analysis results and the emotional state identified by the emotion engine. The feedback information includes not only improvements to form, but also practice methods and encouraging messages that correspond to the player's emotions. The generated feedback information is sent to the device in real time, and the device visually displays it to the player. The display uses graphics and animations to make it visually easy to understand. Encouraging messages that match the player's emotions are also displayed.
[0693] As a concrete example, consider a player practicing on a golf simulator. When the player starts swinging, the device detects this and the camera captures the swing form. This video data is sent to a server and analyzed by multimodal AI. For example, the analysis may detect a problem: "Your arms are positioned a little too high." The device also captures the player's facial expressions and tone of voice, and the emotion engine recognizes that the player is feeling a little stressed. The server generates feedback such as, "Your arms are positioned too high, making it difficult to make a proper impact," and includes an encouraging message such as, "Please calm down. Next time, try lowering your arms a little." This feedback information is sent to the device and displayed visually to the player.
[0694] Examples of prompts to input to a generative AI model include:
[0695] "Based on the player's swing footage and emotional data, generate feedback suggesting form improvements and appropriate practice methods. For example, "Your arms are positioned too high, making it difficult to make a proper impact. Calm down and try lowering your arms a little next time."
[0696] In this way, this system pinpoints a player's technical problems with high accuracy while providing training support that also takes into account their emotional side.
[0697] The flow of the specific processing in the application example 2 will be described with reference to FIG.
[0698] Step 1:
[0699] The player stands at the starting position of the golf simulator and begins swinging. The device detects the start of the swing and activates the camera in conjunction with the golf simulator. The camera captures the player's swing form at a high frame rate and obtains the video data. The input of this step is the player's swing motion, and the output is the captured video data.
[0700] Step 2:
[0701] The device temporarily stores the captured video data and performs preprocessing such as noise reduction and image correction. Before sending the processed data to the server, the data is compressed and encrypted. The input of this step is the captured video data, and the output is the compressed and encrypted video data.
[0702] Step 3:
[0703] The server inputs the received video data into a multimodal AI to extract features such as swing speed, angle, and wrist movement. Through AI analysis, problems with the swing form are identified. The input for this step is compressed and encrypted video data, and the output is the analyzed swing form feature data and problems.
[0704] Step 4:
[0705] The device's built-in camera and microphone are used to capture the user's facial expressions and tone of voice. The emotion engine analyzes this data to identify the user's emotional state. For example, it recognizes whether the user is focused or stressed. The input for this step is facial expression and tone of voice data, and the output is the recognized emotional state.
[0706] Step 5:
[0707] The server generates specific feedback information based on the analysis results and the emotional state identified by the emotion engine. Using the feedback generation means, feedback is generated that includes form improvements, practice methods based on the user's emotions, and encouraging messages. The input for this step is the analyzed swing form feature data and the recognized emotional state, and the output is the generated feedback information.
[0708] Step 6:
[0709] The server sends the generated feedback information to the terminal. The terminal receives the feedback information and visually displays it to the user. This display uses graphics and animations, and also displays encouraging messages that match the user's emotions. The input of this step is the generated feedback information, and the output is the feedback information displayed to the user.
[0710] This series of processing steps allows users to receive instant feedback and effectively improve their form, while encouraging messages that take emotional information into account significantly improve the training experience.
[0711] The specific processing unit 290 transmits the result of the specific processing to the smart glasses 214. In the smart glasses 214, the control unit 46A causes the speaker 240 to output the result of the specific processing. The microphone 238 acquires audio indicating a user input regarding the result of the specific processing. The control unit 46A transmits audio data indicating the user input acquired by the microphone 238 to the data processing device 12. In the data processing device 12, the specific processing unit 290 acquires the audio data.
[0712] The data generation model 58 is a so-called generative AI (Artificial Intelligence). An example of the data generation model 58 is ChatGPT (Internet Search<URL: https: / / openai.com / blog / chatgpt> ), Gemini (Internet search <url: https: gemini.google.com ?hl="ja">) and other generation AIs. The data generation model 58 is obtained by performing deep learning on a neural network. A prompt including an instruction is input to the data generation model 58, and inference data such as voice data indicating voice, text data indicating text, and image data indicating an image is also input. The data generation model 58 performs inference on the input inference data in accordance with the instruction indicated by the prompt, and outputs the inference result in a data format such as voice data and text data. Here, inference refers to, for example, analysis, classification, prediction, and / or summarization.
[0713] In the above embodiment, an example in which the specific processing is performed by the data processing device 12 has been given, but the technology of the present disclosure is not limited to this, and the specific processing may be performed by the smart glasses 214.
[0714] [Third embodiment]
[0715] FIG. 5 shows an example of the configuration of a data processing system 310 according to the third embodiment.
[0716] 5, the data processing system 310 includes the data processing device 12 and a headset type terminal 314. An example of the data processing device 12 is a server.
[0717] The data processing device 12 includes a computer 22, a database 24, and a communication I / F 26. The computer 22 is an example of a "computer" according to the technology of the present disclosure. The computer 22 includes a processor 28, a RAM 30, and a storage 32. The processor 28, the RAM 30, and the storage 32 are connected to a bus 34. The database 24 and the communication I / F 26 are also connected to the bus 34. The communication I / F 26 is connected to a network 54. Examples of the network 54 include a WAN (Wide Area Network) and / or a LAN (Local Area Network).
[0718] The headset type terminal 314 includes a computer 36, a microphone 238, a speaker 240, a camera 42, a communication I / F 44, and a display 343. The computer 36 includes a processor 46, a RAM 48, and a storage 50. The processor 46, the RAM 48, and the storage 50 are connected to a bus 52. The microphone 238, the speaker 240, the camera 42, and the display 343 are also connected to the bus 52.
[0719] The microphone 238 receives instructions and the like from the user 20 by receiving voice uttered by the user 20. The microphone 238 captures the voice uttered by the user 20, converts the captured voice into audio data, and outputs it to the processor 46. The speaker 240 outputs audio in accordance with instructions from the processor 46.
[0720] Camera 42 is a small digital camera equipped with an optical system including a lens, aperture, and shutter, and an imaging element such as a CMOS (Complementary Metal-Oxide-Semiconductor) image sensor or a CCD (Charge Coupled Device) image sensor, and captures images of the surroundings of user 20 (for example, an imaging range defined by an angle of view equivalent to the field of vision of a typical healthy person).
[0721] The communication I / F 44 is connected to a network 54. The communication I / Fs 44 and 26 are responsible for the exchange of various information between the processor 46 and the processor 28 via the network 54. The exchange of various information between the processor 46 and the processor 28 using the communication I / Fs 44 and 26 is carried out in a secure state.
[0722] Fig. 6 shows an example of the main functions of the data processing device 12 and the headset type terminal 314. As shown in Fig. 6, in the data processing device 12, a specific process is performed by the processor 28. A specific process program 56 is stored in the storage 32.
[0723] The specific processing program 56 is an example of a "program" according to the technology of the present disclosure. The processor 28 reads the specific processing program 56 from the storage 32 and executes the read specific processing program 56 on the RAM 30. The specific processing is realized by the processor 28 operating as a specific processing unit 290 in accordance with the specific processing program 56 executed on the RAM 30.
[0724] The storage 32 stores a data generation model 58 and an emotion identification model 59. The data generation model 58 and the emotion identification model 59 are used by the identification processing unit 290.
[0725] In the headset type terminal 314, a reception output process is performed by the processor 46. A reception output program 60 is stored in the storage 50. The processor 46 reads the reception output program 60 from the storage 50 and executes the read reception output program 60 on the RAM 48. The reception output process is realized by the processor 46 operating as the control unit 46A in accordance with the reception output program 60 executed on the RAM 48.
[0726] Next, a description will be given of the identification process performed by the identification processing unit 290 of the data processing device 12. In the following description, the data processing device 12 will be referred to as the "server" and the headset type terminal 314 will be referred to as the "terminal."
[0727] The present invention is a system that combines a photographing means, a processing means, and a display means to efficiently improve a player's golf form. Below, the processing of the program of this system will be explained in natural language, and an embodiment of the invention will be described in detail with concrete examples.
[0728] System Overview
[0729] This system uses a camera linked to a golf simulator to capture the player's swing form in real time. The captured video data is sent to a server, which then analyzes the data using multimodal AI. The analysis results are fed back to the player in real time via their device, suggesting areas for form improvement and optimal practice methods.
[0730] Program processing
[0731] Step 1: Capture video with a camera
[0732] The user stands at the starting position of the golf simulator and begins swinging.
[0733] The device detects the start of the swing and activates the camera in conjunction with the golf simulator.
[0734] The camera captures the player's swing form in real time and acquires video data at a high frame rate.
[0735] Step 2: Sending video data
[0736] The device temporarily stores the captured video data and performs pre-processing, after which the data is compressed and encrypted for transmission to the server.
[0737] Step 3: AI-powered video analysis
[0738] The server inputs the received video data into the multimodal AI and begins analysis.
[0739] The AI extracts features from the video data, such as swing speed, angle, and wrist movement, to detect problems with the swing form. For example, if the arms are positioned too high, this problem will be identified.
[0740] Step 4: Identify areas for improvement and provide feedback
[0741] The server then generates feedback based on the analysis, identifying specific areas for improvement and the reasons for them. For example, the feedback could be, "Your arms are positioned too high, making it difficult to make proper impact with the ball."
[0742] Specific practice methods (e.g., specific drills to correct arm position) are also generated as feedback information.
[0743] Step 5: View your feedback
[0744] The server transmits the generated feedback information to the terminal.
[0745] The device receives the feedback information and displays it to the user, using graphics and animations to make it visually easy to understand.
[0746] Specific examples
[0747] Consider a case where a player (user) is practicing on a golf simulator.
[0748] 1. The user starts swinging. The device detects the start of the swing and the camera captures the swing form.
[0749] 2. The device sends this video data to the server.
[0750] 3. The server receives the video data and analyzes it using multimodal AI, detecting, for example, problems such as the arm being positioned slightly too high.
[0751] 4. The server generates feedback such as "Your arms are positioned too high, making it difficult to make a proper impact," and sends it to the device along with specific practice instructions.
[0752] 5. The device displays feedback information to the user in an intuitive format, for example, by providing a graphic indication of arm position and visually demonstrating correct form.
[0753] This allows users to receive instant feedback and use it to improve their form. By repeating this process, players can efficiently improve their golfing skills.
[0754] The processing flow will be explained below.
[0755] Step 1:
[0756] The user stands at the starting position of the golf simulator and begins swinging.
[0757] The device detects the start of the swing and activates the camera in conjunction with the golf simulator.
[0758] The camera captures the player's swing form in real time and acquires video data at a high frame rate.
[0759] Step 2:
[0760] The device temporarily stores the captured video data and performs preprocessing, which includes noise reduction and image correction.
[0761] The terminal performs data compression and encryption to send the pre-processed data to the server.
[0762] Step 3:
[0763] The server inputs the received video data into the multimodal AI and begins analysis.
[0764] Multimodal AI extracts features such as swing speed, angle, and wrist movement from video data to detect problems with swing form. For example, if the arms are positioned too high, this will be identified.
[0765] Step 4:
[0766] The server then uses the analysis results to identify specific areas for improvement and the reasons for them, such as determining that "your arms are positioned too high, making it difficult to make proper impact with the ball."
[0767] The server generates specific feedback based on the identified areas for improvement, including specific practice methods (e.g., specific drills to correct arm position).
[0768] Step 5:
[0769] The server transmits the generated feedback information to the terminal.
[0770] The device receives the feedback information and displays it to the user, using graphics and animations to make it visually easy to understand.
[0771] Step 6:
[0772] The user reviews the displayed feedback information and practices to correct their next swing form. They adjust their form based on the specific improvements indicated in the feedback.
[0773] This series of steps allows users to efficiently improve their swing form, and real-time feedback can help them improve their technique quickly.
[0774] Example 1
[0775] Next, a description will be given of Example 1. In the following description, the data processing device 12 will be referred to as a "server" and the headset type terminal 314 will be referred to as a "terminal."
[0776] For players to efficiently improve their form, it is important to accurately understand their current form and immediately identify specific areas for improvement. However, with conventional technology, analyzing and evaluating form takes time, making it difficult to provide real-time feedback. In addition, the analysis results are not specific, making it difficult for players to understand how to improve.
[0777] The specific processing by the specific processing unit 290 of the data processing device 12 in the first embodiment is realized by the following means.
[0778] In this invention, the server includes a camera for capturing the player's movements in real time, a terminal for temporarily storing and preprocessing the captured video data, a transmission means for compressing and encrypting the preprocessed data and transmitting it to the server, a processing means for analyzing the video data received by the server using multimodal AI to detect problems with the player's swing form, a feedback generation means for identifying specific areas for improvement based on the analysis results and generating feedback information, and a means for transmitting the generated feedback information to the terminal and displaying it to the user. This allows the player to receive specific and detailed feedback in real time, enabling them to improve their form efficiently.
[0779] A "player" is someone who uses the system to improve their own movements and form.
[0780] "Movement" refers to the physical movements of the player related to their swing and form.
[0781] "Real-time" refers to immediate processing and feedback without delay.
[0782] "Photography means" refers to a device used to capture the player's actions, such as a camera.
[0783] "Terminal" refers to a device that temporarily stores, pre-processes, and transmits data, typically a computer or smart device.
[0784] "Transmission means" refers to the process or function that transmits the pre-processed data to the server.
[0785] "Server" refers to a device or system that analyzes video data and generates feedback information.
[0786] "Multimodal AI" refers to artificial intelligence technology that comprehensively analyzes multiple input data (video, audio, sensors, etc.).
[0787] "Processing means" refers to a device or process that has the function of analyzing received video data and detecting problems.
[0788] "Feedback generation means" refers to the process or function that generates specific improvements and practice methods based on the analysis results.
[0789] The "display means" refers to a device or system that has the function of visually displaying the generated feedback information to the user.
[0790] "Video data" refers to video data that captures the player's actions.
[0791] The present invention is a system that captures a player's actions in real time and provides feedback based on the analysis results. Specifically, it is realized by linking a camera, a terminal, a server, and a display.
[0792] The system overview includes the following hardware and software:
[0793] Filming method
[0794] The recording method consists of a camera that captures the player's swing form at a high frame rate. This camera operates at a high frame rate such as 120 fps, allowing for detailed capture of the player's movements. For example, it is preferable to use a high-performance sports camera or a dedicated motion capture camera rather than a regular webcam.
[0795] Terminal
[0796] The terminal is a device that temporarily stores and pre-processes the video data captured by the camera. This pre-processing includes noise reduction and color correction. The video data is then compressed into MPEG format and encrypted with AES. This allows the data to be sent securely to the server while maintaining its quality. The terminal requires a high-performance processor and large-capacity storage.
[0797] Transmission method
[0798] The transmission means has the function of sending preprocessed data from the terminal to the server. It is important to use the HTTPS protocol for data transmission, ensuring safety and speed.
[0799] server
[0800] The server includes a processing means for analyzing the received video data using multimodal AI. This AI is built using deep learning frameworks such as TensorFlow and PyTorch, and performs detailed analysis of swing speed, angle, wrist movement, and other parameters. As a result of the analysis, it detects problems with the swing form and generates feedback. The server then generates feedback that includes specific practice methods and corrections.
[0801] Feedback Generation Method
[0802] The feedback generator explains specific areas for improvement and the reasons for them based on the analysis results from the server. For example, it may give specific indications such as, "Your right arm is too high, making your impact unstable." It also suggests practice methods for keeping the right arm low.
[0803] Display means
[0804] The display means is a device or system that visually displays the generated feedback information to the user. The terminal receives the feedback information and uses graphics or animations to show it to the user in an easy-to-understand manner. For example, by visually displaying the difference between incorrect and correct form, the user can understand specifically which parts need to be corrected.
[0805] Specific examples
[0806] For example, when a player practices on a golf simulator, the system operates in the following steps:
[0807] 1. The user starts swinging, the device detects the movement, and the camera captures the swing form at a high frame rate.
[0808] 2. The device temporarily stores the video data, performs preprocessing, compresses it into MPEG format, encrypts it using AES, and sends it to the server.
[0809] 3. The server receives the video data and analyzes it using multimodal AI, detecting problems such as "the right arm is raised too high."
[0810] 4. Based on the analysis results, the server generates feedback recommending "practice swinging with both arms fixed at waist height to practice keeping the right arm low" and sends it to the device.
[0811] 5. The device displays the feedback information in a graphical format, visually showing the user specific areas for improvement and correct form.
[0812] Prompt Sentence Examples
[0813] The user can enter prompts for the generative AI model, such as:
[0814] "Analyze the swing footage of the player and if there are any problems with their swing form, tell them how to improve it."
[0815] Thus, the present invention is a system that can provide detailed feedback in real time to help players improve their form efficiently.
[0816] The flow of the identification process in the first embodiment will be described with reference to FIG.
[0817] The flow of this system's program processing
[0818] Step 1:
[0819] The user stands at the starting position of the golf simulator and begins swinging.
[0820] Input: User swing motion
[0821] Operation: The device uses sensors to recognize when the user starts swinging.
[0822] Output: Swing start detection signal
[0823] Step 2:
[0824] The device activates the camera and captures the player's swing form in real time.
[0825] Input: Swing start detection signal
[0826] How it works: The device controls the camera and captures video data of the swing form at a high frame rate such as 120 fps.
[0827] Output: Captured video data
[0828] Step 3:
[0829] The device temporarily stores the captured video data and performs preprocessing.
[0830] Input: Captured video data
[0831] How it works: The device breaks down the video data into frames and performs noise reduction and color correction.
[0832] Output: Pre-processed video data
[0833] Step 4:
[0834] The terminal compresses and encrypts the pre-processed video data.
[0835] Input: Preprocessed video data
[0836] Operation: Video data is compressed into MPEG format and encrypted with AES.
[0837] Output: Compressed and encrypted video data
[0838] Step 5:
[0839] The terminal transmits the compressed and encrypted video data to the server.
[0840] Input: Compressed and encrypted video data
[0841] How it works: The device uses the HTTPS protocol to securely upload data to the server.
[0842] Output: Video data sent to the server
[0843] Step 6:
[0844] The video data received by the server is analyzed using multimodal AI.
[0845] Input: Received video data
[0846] How it works: The server inputs video data into the AI model and extracts swing form characteristics (e.g., speed, angle, wrist movement).
[0847] Output: Swing form feature data
[0848] Step 7:
[0849] The server detects problems based on characteristic data of the swing form.
[0850] Input: Swing form characteristics data
[0851] How it works: The server compares the extracted features with predefined standards and identifies form flaws (e.g., the right arm is too high).
[0852] Output: Detected issue data
[0853] Step 8:
[0854] The server generates feedback information based on the analysis results.
[0855] Input: Detected issue data
[0856] How it works: The server generates feedback with specific areas for improvement and reasons for doing so, and also suggests practice methods.
[0857] Output: Generated feedback information
[0858] Step 9:
[0859] The server transmits the generated feedback information to the terminal.
[0860] Input: Generated feedback information
[0861] Operation: The server sends feedback information to the device in JSON format or similar.
[0862] Output: Feedback information sent to the terminal
[0863] Step 10:
[0864] The terminal displays the feedback information to the user.
[0865] Input: Feedback information sent to the device
[0866] Action: The device visually displays feedback information to the user using graphics and animations.
[0867] Output: Feedback information displayed to the user
[0868] (Application example 1)
[0869] Next, a description will be given of Application Example 1. In the following description, the data processing device 12 will be referred to as a "server" and the headset type terminal 314 will be referred to as a "terminal."
[0870] Conventional technologies exist that provide real-time feedback on player movements and form improvements. However, there is a need for a system that can optimize the movement accuracy and efficiency of robots and equipment in real time, even in factory operations. The present invention provides technology that improves movement accuracy and efficiency by capturing, analyzing, and providing feedback on the movements of factory robots.
[0871] The specific processing by the specific processing unit 290 of the data processing device 12 in the application example 1 is realized by the following means.
[0872] In this invention, the server includes a camera for capturing the actions of the player or the actions of the equipment in real time, a processing unit for receiving and analyzing the captured video data, and a display unit for providing feedback information to the player or operator based on the analysis results, thereby enabling efficient improvement of the operations of factory robots and equipment.
[0873] "Photographing means" refers to a device for capturing the actions of a player or device in real time.
[0874] A "processing means" is a device or system for receiving and analyzing captured video data.
[0875] "Display means" refers to a device or system for providing feedback information to a player or operator based on the analysis results.
[0876] "Multimodal AI" is an artificial intelligence technology that integrates and analyzes multiple data modalities (e.g., video, audio, text).
[0877] "Feedback" is information based on the analysis results that provides improvements to movements and form, as well as optimal practice methods.
[0878] "Capture" refers to the process of converting a subject's movements and form into data in real time using a photographic device.
[0879] A "server" is a computer system that receives data, analyzes it, and generates feedback information.
[0880] The present invention is a system for optimizing the operation of a factory robot, and includes a photographing means for capturing the operation of a player or equipment in real time, a processing means for receiving and analyzing the photographed video data, and a display means for providing feedback information to a player or operator based on the analysis results.
[0881] System Program
[0882] The main hardware and software components of this system include smartphones, factory robots, smartphone apps (e.g., cross-platform development with Flutter or React Native), servers, multimodal AI (e.g., TensorFlow or PyTorch), and communication protocols (e.g., HTTPS or WebSocket).
[0883] Natural language explanation of program processing
[0884] Step 1: Capture video with a camera
[0885] The smartphone camera captures the robot's movements in the factory in real time, capturing the video data at a high frame rate.
[0886] Step 2: Sending video data
[0887] The captured video data is temporarily stored on the smartphone and then encrypted and sent to a server.
[0888] Step 3: AI-powered video analysis
[0889] The server inputs the received video data into the multimodal AI for analysis. The AI extracts characteristics of the robot's movements from the video data and detects problems with the movements. For example, if the robot's arm is moving too slowly, this problem will be identified.
[0890] Step 4: Identify areas for improvement and provide feedback
[0891] Based on the analysis results, the server identifies specific areas for improvement and the reasons for them, and generates feedback information. For example, the feedback might say, "The arm's movements are too slow, reducing work efficiency." This feedback might also include specific drills to improve the arm's speed by 20%.
[0892] Step 5: View your feedback
[0893] The server sends the generated feedback information to a smartphone and displays it for the user or operator to check. This display uses graphics and animations to make it visually easy to understand.
[0894] Specific examples
[0895] For example, if the robot arm is analyzed as moving too slowly, the app will display the following feedback:
[0896] "The arm is moving too slowly, reducing work efficiency. Please perform drills to increase the arm's movement speed by 20%."
[0897] AI prompt examples
[0898] "Analyze the robot arm's movement pattern in this video, identify the speed, angle, and errors, and suggest improvements based on that."
[0899] The flow of the specific processing in the application example 1 will be described with reference to FIG.
[0900] Step 1:
[0901] A user captures the operation of a factory robot in real time using a smartphone camera. The smartphone camera acquires video data at a high frame rate and temporarily stores the data within the smartphone. The input is the captured video data, and the output is the temporarily stored video data.
[0902] Step 2:
[0903] The device preprocesses the stored video data. Preprocessing includes compressing and encrypting the data. The compressed and encrypted data is then sent to the server. The input is the stored video data, and the output is the compressed and encrypted data.
[0904] Step 3:
[0905] The server decompresses the received video data and prepares it for analysis. The decompressed data is input into the multimodal AI, and data analysis begins. The input is compressed and encrypted data, and the output is analyzable data.
[0906] Step 4:
[0907] The server's multimodal AI extracts features related to the robot's movements from the video data and analyzes any problems with the movements. For example, it analyzes the robot's arm's movement speed, angle, and identifies any errors. The input is analyzable data, and the output is the analysis results, including any problems with the movements.
[0908] Step 5:
[0909] Based on the analysis results, the server identifies specific areas for improvement and the reasons for them. It then generates feedback information to present the areas for improvement to the user. For example, it generates feedback information such as "The arm's movements are too slow, reducing work efficiency." The input is the analysis results, and the output is feedback information.
[0910] Step 6:
[0911] The server generates feedback information and sends it to the terminal. The terminal receives the feedback information and displays it in a visually understandable format for the user, for example, using graphics or animations to demonstrate correct actions or forms. The input is the feedback information, and the output is the feedback content displayed to the user.
[0912] Step 7:
[0913] Based on the displayed feedback information, the user can take specific actions to improve the factory robot's operation. For example, they can adjust the robot's parameters to improve its speed and accuracy. In this step, the feedback information is directly used for on-site improvement actions. The input is the feedback displayed to the user, and the output is the improved robot's operation.
[0914] Furthermore, an emotion engine that estimates the user's emotion may be further combined. That is, the identification processing unit 290 may estimate the user's emotion using the emotion identification model 59, and perform identification processing using the user's emotion.
[0915] The present invention relates to a system that combines a photographing means, a processing means, a display means, and an emotion engine to efficiently improve a player's golf form. Below, the processing of the program of this system will be explained in natural language, and an embodiment of the invention will be described in detail with concrete examples.
[0916] System Overview
[0917] This system uses a camera linked to a golf simulator to capture the player's swing form in real time. The captured video data is sent to a server, which then analyzes it using multimodal AI. It also recognizes the user's emotions using an emotion engine and generates feedback information based on the analysis results and emotional state. The feedback information is provided to the user in real time via their device, suggesting areas for improvement in form, optimal practice methods, and even encouraging messages based on their emotions.
[0918] Program processing
[0919] Step 1: Capture video with a camera
[0920] The user stands at the starting position of the golf simulator and begins swinging.
[0921] The device detects the start of the swing and activates the camera in conjunction with the golf simulator.
[0922] The camera captures the player's swing form in real time and acquires video data at a high frame rate.
[0923] Step 2: Sending video data
[0924] The device temporarily stores the captured video data and performs preprocessing, which includes noise reduction and image correction.
[0925] The terminal performs data compression and encryption to send the pre-processed data to the server.
[0926] Step 3: AI-powered video analysis
[0927] The server inputs the received video data into the multimodal AI and begins analysis.
[0928] Multimodal AI extracts features such as swing speed, angle, and wrist movement from video data to detect problems with swing form. For example, if the arms are positioned too high, this will be identified.
[0929] Step 4: Emotion Recognition with the Emotion Engine
[0930] It uses the device's built-in camera and microphone to capture the user's facial expressions and tone of voice.
[0931] An emotion engine analyzes this data to identify the user's emotional state, for example, recognizing whether the user is focused or stressed.
[0932] Step 5: Identify areas for improvement and generate feedback
[0933] The server generates specific feedback information based on the analysis results and the emotional state identified by the emotion engine.
[0934] The feedback information includes not only improvements to form, but also practice methods and encouraging messages based on the user's emotions.
[0935] Step 6: View your feedback
[0936] The server transmits the generated feedback information to the terminal.
[0937] The device receives the feedback information and displays it to the user. The display uses graphics and animations to make it visually easy to understand. It also displays encouraging messages that match the user's emotions.
[0938] Specific examples
[0939] Consider a case where a player (user) is practicing on a golf simulator.
[0940] 1. The user starts swinging. The device detects the start of the swing and the camera captures the swing form.
[0941] 2. The device sends this video data to the server.
[0942] 3. The server receives the video data and analyzes it using multimodal AI, detecting, for example, problems such as the arm being positioned slightly too high.
[0943] 4. The device captures the user's facial expressions and tone of voice, and the emotion engine recognizes that the user is feeling a little stressed.
[0944] 5. The server generates feedback such as "Your arms are too high, making it difficult to make a good impact," and includes an encouraging message such as "Calm down, try lowering your arms a little next time."
[0945] 6. The device receives the feedback information and visually displays it to the user.
[0946] This process allows users to receive immediate feedback and improve their form. Furthermore, encouraging messages that take emotional information into account improve the user's training experience and enable more effective learning.
[0947] The processing flow will be explained below.
[0948] Step 1:
[0949] The user stands at the starting position of the golf simulator and begins swinging.
[0950] The device detects the start of the swing and activates the camera in conjunction with the golf simulator.
[0951] The camera captures the player's swing form in real time and acquires video data at a high frame rate.
[0952] Step 2:
[0953] The device temporarily stores the captured video data and performs preprocessing, which includes noise reduction and image correction.
[0954] The terminal performs data compression and encryption to send the pre-processed data to the server.
[0955] Step 3:
[0956] The server inputs the received video data into the multimodal AI and begins analysis.
[0957] Multimodal AI extracts features such as swing speed, angle, and wrist movement from video data to detect problems with swing form. For example, if the arms are positioned too high, this will be identified.
[0958] Step 4:
[0959] It uses the device's built-in camera and microphone to capture the user's facial expressions and tone of voice.
[0960] The device processes the captured emotion data in real time and sends it to the emotion engine.
[0961] The emotion engine analyzes the user's emotional state from facial expressions and tone of voice, and recognizes, for example, that the user is feeling stressed.
[0962] Step 5:
[0963] The server generates improvement points and feedback information based on the results of video data analysis and the emotional state determined by the emotion engine.
[0964] The server generates feedback such as "Your arm is positioned too high, making it difficult to achieve proper impact," and depending on the user's stress level, may also include encouraging messages such as "Calm down, try lowering your arm a bit next time."
[0965] Step 6:
[0966] The server transmits the generated feedback information to the terminal.
[0967] The device receives the feedback information and displays it visually to the user. Graphics and animations are used to make the display easy to understand. Encouraging messages tailored to the user's emotions are also displayed.
[0968] Step 7:
[0969] The user checks the displayed feedback information and encouraging messages, and practices to correct their swing form next time. They then adjust their form based on the specific areas for improvement indicated in the feedback. Through this series of processes, users can efficiently improve their form and receive support tailored to their emotional state.
[0970] Example 2
[0971] Next, a description will be given of Example 2. In the following description, the data processing device 12 will be referred to as a "server" and the headset type terminal 314 will be referred to as a "terminal."
[0972] To effectively improve golf swing form, it is important to capture the player's movements in real time and provide detailed analysis and immediate feedback. However, conventional systems have had problems with form improvement due to the low accuracy of video data analysis and the low quality of feedback. Furthermore, they lack the ability to provide feedback that takes into account the player's emotional state, which can lead to a loss of motivation.
[0973] The specific processing by the specific processing unit 290 of the data processing device 12 in the second embodiment is realized by the following means.
[0974] In this invention, the server includes a camera for capturing the player's movements in real time, a processor for receiving the captured video data and performing preprocessing such as noise reduction and image correction, a transmitter for compressing and encrypting the preprocessed data and transmitting it to the server, an analyzer for analyzing the received data using multimodal AI and identifying areas for form improvement, an emotion recognition processor for capturing the user's facial expressions and tone of voice based on the analysis results to identify their emotional state, a generator for generating specific feedback information based on the analysis results and their emotional state, and a displayer for providing the generated feedback information to the user. This makes it possible to not only instantly analyze the player's form in detail and provide effective feedback, but also to send appropriate encouraging messages according to the user's emotional state.
[0975] "Photographing means" is a device for capturing the player's movements in real time.
[0976] The "processing means" is a device that receives the captured video data and performs pre-processing such as noise removal and image correction.
[0977] The "transmitting means" is a device that compresses and encrypts the preprocessed data and transmits it to the server.
[0978] The "analysis means" is a device that uses multimodal AI to analyze the received data and extract areas for improvement in form.
[0979] The "emotion recognition means" is a device that captures the user's facial expressions and tone of voice based on the analysis results to identify the user's emotional state.
[0980] The "generating means" is a device that generates specific feedback information based on the analysis results and the emotional state.
[0981] The "display means" is a device for providing the generated feedback information to the user.
[0982] "Multimodal AI" is an artificial intelligence technology that analyzes multiple data modes (such as images and audio) to extract meaning.
[0983] "Feedback information" is information such as improvements to the form and encouraging messages that are provided to the user based on the results obtained from the analysis means and emotion recognition means.
[0984] This invention is a system for efficiently improving a player's golf form, combining a photographing means, a processing means, a transmission means, an analysis means, an emotion recognition means, a generation means, and a display means. This system works in conjunction with a golf simulator to capture the player's swing form in real time and provide immediate feedback based on the analysis results. Furthermore, by taking the user's emotional state into consideration, an effective training environment is realized.
[0985] Hardware and software used
[0986] Camera (photography means)
[0987] This system uses a camera that captures video at a high frame rate, for example, a camera capable of capturing 120 fps (frames per second) is suitable.
[0988] Terminal (processing means and transmission means)
[0989] The device performs pre-processing on the captured video data, including noise reduction and image enhancement. The pre-processed data is then compressed in H.264 format and encrypted with AES, allowing the data to be transmitted efficiently and securely to the server.
[0990] Server (analysis means, emotion recognition means, generation means)
[0991] The server inputs the received data into the multimodal AI for detailed analysis. Specifically, it analyzes swing speed, angle, wrist movement, etc. to identify areas for form improvement. It also has an emotion recognition mechanism that captures the user's facial expressions and tone of voice to identify their emotional state. Specific feedback information is generated based on the analysis results and emotional state.
[0992] Terminal (display means)
[0993] The device visually displays the generated feedback to the user, using graphics and animations to highlight problem areas on the form and providing emotionally tailored encouraging messages.
[0994] Specific examples
[0995] Consider a case where a player (user) is practicing on a golf simulator. When the user starts swinging, the device detects the start of the swing and the camera captures the swing form. The camera records video at 120 fps and temporarily stores the high-quality video in memory.
[0996] The device performs preprocessing such as noise reduction and white balance adjustment, compresses the preprocessed video data in H.264 format, encrypts it using AES, and sends it to the server. The server receives the video data and performs detailed analysis using multimodal AI. For example, the AI can detect problems such as the arm being positioned slightly too high.
[0997] The device then captures the user's facial expressions, and the emotion engine recognizes that the user is feeling a little stressed. The server then points out the problem with the swing form, generating feedback such as, "Your arms are positioned too high, making it difficult to make a proper impact." It also generates encouraging messages such as, "Please stay calm. Next time, try lowering your arms a little."
[0998] The device receives feedback, graphically highlights areas of form that are problematic, and even displays encouraging messages like "Put your arms a little lower."
[0999] Prompt Sentence Examples
[1000] Here are some example prompts to input to a generative AI model:
[1001] "Analyze the golf swing footage of the player and point out any problems with their swing form. Additionally, analyze the player's facial expressions and tone of voice to provide feedback based on the player's emotions."
[1002] These prompts allow you to accurately input the necessary information into the generative AI model and get the desired feedback.
[1003] The flow of the identification process in the second embodiment will be described with reference to FIG.
[1004] Step 1:
[1005] The user stands at the starting position of the golf simulator and begins swinging. The device detects the start of the swing and activates the camera in conjunction with the golf simulator. The camera captures the player's swing form in real time. The input at this time is the user's swing motion, and the output is the captured high-frame-rate video data.
[1006] Step 2:
[1007] The device temporarily stores the captured video data and performs preprocessing such as noise reduction and white balance adjustment. For example, preprocessing removes background noise and adjusts the color tone of the video. The input of this preprocessing is the captured video data, and the output is the preprocessed video data.
[1008] Step 3:
[1009] The terminal compresses the preprocessed video data in H.264 format and encrypts it using the AES method. This reduces the data size and the risk of data leaks. The input is preprocessed video data, and the output is compressed and encrypted video data.
[1010] Step 4:
[1011] The device sends data to the server, using the appropriate protocol to ensure the data travels securely across the network. The input is the compressed and encrypted video data, and the output is the data received by the server.
[1012] Step 5:
[1013] The server inputs the received data into a multimodal AI for detailed analysis. The AI extracts features such as swing speed, angle, and wrist movement to identify problems with the golf form. For example, it detects problems with the arm position being too high. The input for this step is the video data received by the server, and the output is the analyzed form improvements.
[1014] Step 6:
[1015] The device captures the user's facial expressions and tone of voice and inputs this data into an emotion engine, which analyzes whether the user is focused, relaxed, or stressed. The input is the user's facial expressions and tone of voice, and the output is the identified emotional state.
[1016] Step 7:
[1017] The server generates feedback information based on the analysis results and the user's emotional state. It specifically points out areas for improvement in the user's form and creates an encouraging message that corresponds to the user's emotions. For example, it generates a specific message such as "Lower your arm position. Stay calm and try your best next time." The input is the areas for improvement in the user's form and the user's emotional state, and the output is the generated feedback information.
[1018] Step 8:
[1019] The server sends the generated feedback information to the device. The device receives the feedback information and visually displays it to the user. The display uses graphics and animations to highlight problem areas in the swing form and provide encouraging messages. The input is the generated feedback information, and the output is the feedback information displayed to the user.
[1020] (Application example 2)
[1021] Next, a description will be given of Application Example 2. In the following description, the data processing device 12 will be referred to as a "server" and the headset type terminal 314 will be referred to as a "terminal."
[1022] While conventional golf swing analysis systems can pinpoint technical issues in a player's swing form, they are unable to provide feedback that takes into account the player's emotional state. As a result, if a player feels stressed or frustrated, they are unable to receive appropriate advice or encouragement, which can lead to training being ineffective.
[1023] The specific processing by the specific processing unit 290 of the data processing device 12 in Application Example 2 is realized by the following means. In this invention, the server includes a shooting means for capturing the player's actions in real time, a processing means for receiving and analyzing the shot video data, a display means for providing feedback information to the player based on the analysis results, an emotion recognition means for recognizing the player's emotional state in real time, and a feedback generation means for generating feedback information based on the analysis results and the player's emotional state. This makes it possible to provide not only technical problems but also appropriate feedback and encouraging messages according to the player's emotional state.
[1024] "Filming means for capturing player's movements in real time" refers to a device for capturing player's movements, particularly golf swings and other exercises, as video in real time.
[1025] The "processing means for receiving and analyzing captured video data" is a system that receives video data captured by the image capture means, analyzes it, and extracts important features and problems.
[1026] The "display means for providing feedback information to the player based on the analysis results" refers to a device such as a monitor or display for providing visual feedback to the player based on the analysis results.
[1027] The "emotion recognition means for recognizing the player's emotional state in real time" is a system that includes a camera and microphone for recognizing emotions in real time from the player's facial expressions and tone of voice, as well as software for analyzing them.
[1028] "Feedback generation means for generating feedback information based on the analysis results and the emotional state" refers to software or algorithms that generate feedback information to be provided to the player, taking into account both the technical analysis results and the player's emotional state.
[1029] A "camera that captures at a high frame rate" is a camera that can capture a large number of frames per second, thereby capturing the player's movements in detail.
[1030] "Multimodal AI" is an artificial intelligence technology that can simultaneously analyze multiple different types of data (e.g., video data and emotional data) and make integrated judgments.
[1031] "Areas for improvement in form" refers to elements or problems in a player's movements or posture that are hindering efficient or effective performance.
[1032] The present invention relates to a system for efficiently improving a player's golf swing form, which includes a camera for capturing the player's movements in real time, a processing means for receiving and analyzing the captured video data, a display means for providing feedback information to the player based on the analysis results, an emotion recognition means for recognizing the player's emotional state in real time, and a feedback generation means for generating feedback information based on the analysis results and the player's emotional state.
[1033] The server uses a camera to capture the player's swing form in real time. When the player starts swinging, the device detects this and activates the camera. The camera captures the player's swing form at a high frame rate, and the video data is temporarily saved on the device. The saved video data is then compressed and encrypted before being sent to the server. The server then analyzes the received video data using multimodal AI to extract features such as swing speed, angle, and wrist movement.
[1034] The device also uses its built-in camera and microphone to capture the player's facial expressions and tone of voice, and the emotion engine analyzes this data to identify the player's emotional state, for example, whether the player is focused or stressed.
[1035] The server uses a feedback generation means to generate specific feedback information based on the analysis results and the emotional state identified by the emotion engine. The feedback information includes not only improvements to form, but also practice methods and encouraging messages that correspond to the player's emotions. The generated feedback information is sent to the device in real time, and the device visually displays it to the player. The display uses graphics and animations to make it visually easy to understand. Encouraging messages that match the player's emotions are also displayed.
[1036] As a concrete example, consider a player practicing on a golf simulator. When the player starts swinging, the device detects this and the camera captures the swing form. This video data is sent to a server and analyzed by multimodal AI. For example, the analysis may detect a problem: "Your arms are positioned a little too high." The device also captures the player's facial expressions and tone of voice, and the emotion engine recognizes that the player is feeling a little stressed. The server generates feedback such as, "Your arms are positioned too high, making it difficult to make a proper impact," and includes an encouraging message such as, "Please calm down. Next time, try lowering your arms a little." This feedback information is sent to the device and displayed visually to the player.
[1037] Examples of prompts to input to a generative AI model include:
[1038] "Based on the player's swing footage and emotional data, generate feedback suggesting form improvements and appropriate practice methods. For example, "Your arms are positioned too high, making it difficult to make a proper impact. Calm down and try lowering your arms a little next time."
[1039] In this way, this system pinpoints a player's technical problems with high accuracy while providing training support that also takes into account their emotional side.
[1040] The flow of the specific processing in the application example 2 will be described with reference to FIG.
[1041] Step 1:
[1042] The player stands at the starting position of the golf simulator and begins swinging. The device detects the start of the swing and activates the camera in conjunction with the golf simulator. The camera captures the player's swing form at a high frame rate and obtains the video data. The input of this step is the player's swing motion, and the output is the captured video data.
[1043] Step 2:
[1044] The device temporarily stores the captured video data and performs preprocessing such as noise reduction and image correction. Before sending the processed data to the server, the data is compressed and encrypted. The input of this step is the captured video data, and the output is the compressed and encrypted video data.
[1045] Step 3:
[1046] The server inputs the received video data into a multimodal AI to extract features such as swing speed, angle, and wrist movement. Through AI analysis, problems with the swing form are identified. The input for this step is compressed and encrypted video data, and the output is the analyzed swing form feature data and problems.
[1047] Step 4:
[1048] The device's built-in camera and microphone are used to capture the user's facial expressions and tone of voice. The emotion engine analyzes this data to identify the user's emotional state. For example, it recognizes whether the user is focused or stressed. The input for this step is facial expression and tone of voice data, and the output is the recognized emotional state.
[1049] Step 5:
[1050] The server generates specific feedback information based on the analysis results and the emotional state identified by the emotion engine. Using the feedback generation means, feedback is generated that includes form improvements, practice methods based on the user's emotions, and encouraging messages. The input for this step is the analyzed swing form feature data and the recognized emotional state, and the output is the generated feedback information.
[1051] Step 6:
[1052] The server sends the generated feedback information to the terminal. The terminal receives the feedback information and visually displays it to the user. This display uses graphics and animations, and also displays encouraging messages that match the user's emotions. The input of this step is the generated feedback information, and the output is the feedback information displayed to the user.
[1053] This series of processing steps allows users to receive instant feedback and effectively improve their form, while encouraging messages that take emotional information into account significantly improve the training experience.
[1054] The specific processing unit 290 transmits the result of the specific processing to the headset type terminal 314. In the headset type terminal 314, the control unit 46A causes the speaker 240 and the display 343 to output the result of the specific processing. The microphone 238 acquires audio indicating a user input regarding the result of the specific processing. The control unit 46A transmits audio data indicating the user input acquired by the microphone 238 to the data processing device 12. In the data processing device 12, the specific processing unit 290 acquires the audio data.
[1055] The data generation model 58 is a so-called generative AI (Artificial Intelligence). An example of the data generation model 58 is ChatGPT (Internet Search<URL: https: / / openai.com / blog / chatgpt> ), Gemini (Internet search <url: https: gemini.google.com ?hl="ja">) and other generation AIs. The data generation model 58 is obtained by performing deep learning on a neural network. A prompt including an instruction is input to the data generation model 58, and inference data such as voice data indicating voice, text data indicating text, and image data indicating an image is also input. The data generation model 58 performs inference on the input inference data in accordance with the instruction indicated by the prompt, and outputs the inference result in a data format such as voice data and text data. Here, inference refers to, for example, analysis, classification, prediction, and / or summarization.
[1056] In the above embodiment, an example was given in which the specific processing is performed by the data processing device 12, but the technology of the present disclosure is not limited to this, and the specific processing may be performed by the headset type terminal 314.
[1057] [Fourth embodiment]
[1058] FIG. 7 shows an example of the configuration of a data processing system 410 according to the fourth embodiment.
[1059] 7, a data processing system 410 includes a data processing device 12 and a robot 414. An example of the data processing device 12 is a server.
[1060] The data processing device 12 includes a computer 22, a database 24, and a communication I / F 26. The computer 22 is an example of a "computer" according to the technology of the present disclosure. The computer 22 includes a processor 28, a RAM 30, and a storage 32. The processor 28, the RAM 30, and the storage 32 are connected to a bus 34. The database 24 and the communication I / F 26 are also connected to the bus 34. The communication I / F 26 is connected to a network 54. Examples of the network 54 include a WAN (Wide Area Network) and / or a LAN (Local Area Network).
[1061] The robot 414 includes a computer 36, a microphone 238, a speaker 240, a camera 42, a communication I / F 44, and a control target 443. The computer 36 includes a processor 46, a RAM 48, and a storage 50. The processor 46, the RAM 48, and the storage 50 are connected to a bus 52. The microphone 238, the speaker 240, the camera 42, and the control target 443 are also connected to the bus 52.
[1062] The microphone 238 receives instructions and the like from the user 20 by receiving voice uttered by the user 20. The microphone 238 captures the voice uttered by the user 20, converts the captured voice into audio data, and outputs it to the processor 46. The speaker 240 outputs audio in accordance with instructions from the processor 46.
[1063] Camera 42 is a small digital camera equipped with an optical system including a lens, aperture, and shutter, and an imaging element such as a CMOS (Complementary Metal-Oxide-Semiconductor) image sensor or a CCD (Charge Coupled Device) image sensor, and captures images of the surroundings of user 20 (for example, an imaging range defined by an angle of view equivalent to the field of vision of a typical healthy person).
[1064] The communication I / F 44 is connected to a network 54. The communication I / Fs 44 and 26 are responsible for the exchange of various information between the processor 46 and the processor 28 via the network 54. The exchange of various information between the processor 46 and the processor 28 using the communication I / Fs 44 and 26 is carried out in a secure state.
[1065] The control object 443 includes a display device, LEDs in the eyes, and motors for driving the arms, hands, and feet. The posture and gestures of the robot 414 are controlled by controlling the motors of the arms, hands, and feet. Some of the emotions of the robot 414 can be expressed by controlling these motors. In addition, the facial expressions of the robot 414 can also be expressed by controlling the light emission state of the LEDs in the eyes of the robot 414.
[1066] Fig. 8 shows an example of the main functions of the data processing device 12 and the robot 414. As shown in Fig. 8, in the data processing device 12, a specific process is performed by the processor 28. A specific process program 56 is stored in the storage 32.
[1067] The specific processing program 56 is an example of a "program" according to the technology of the present disclosure. The processor 28 reads the specific processing program 56 from the storage 32 and executes the read specific processing program 56 on the RAM 30. The specific processing is realized by the processor 28 operating as a specific processing unit 290 in accordance with the specific processing program 56 executed on the RAM 30.
[1068] The storage 32 stores a data generation model 58 and an emotion identification model 59. The data generation model 58 and the emotion identification model 59 are used by the identification processing unit 290.
[1069] In the robot 414, the processor 46 performs the reception output process. A reception output program 60 is stored in the storage 50. The processor 46 reads the reception output program 60 from the storage 50 and executes the read reception output program 60 on the RAM 48. The reception output process is realized by the processor 46 operating as the control unit 46A in accordance with the reception output program 60 executed on the RAM 48.
[1070] Next, a description will be given of the specific processing performed by the specific processing unit 290 of the data processing device 12. In the following description, the data processing device 12 will be referred to as a "server" and the robot 414 will be referred to as a "terminal."
[1071] The present invention is a system that combines a photographing means, a processing means, and a display means to efficiently improve a player's golf form. Below, the processing of the program of this system will be explained in natural language, and an embodiment of the invention will be described in detail with concrete examples.
[1072] System Overview
[1073] This system uses a camera linked to a golf simulator to capture the player's swing form in real time. The captured video data is sent to a server, which then analyzes the data using multimodal AI. The analysis results are fed back to the player in real time via their device, suggesting areas for form improvement and optimal practice methods.
[1074] Program processing
[1075] Step 1: Capture video with a camera
[1076] The user stands at the starting position of the golf simulator and begins swinging.
[1077] The device detects the start of the swing and activates the camera in conjunction with the golf simulator.
[1078] The camera captures the player's swing form in real time and acquires video data at a high frame rate.
[1079] Step 2: Sending video data
[1080] The device temporarily stores the captured video data and performs pre-processing, after which the data is compressed and encrypted for transmission to the server.
[1081] Step 3: AI-powered video analysis
[1082] The server inputs the received video data into the multimodal AI and begins analysis.
[1083] The AI extracts features from the video data, such as swing speed, angle, and wrist movement, to detect problems with the swing form. For example, if the arms are positioned too high, this problem will be identified.
[1084] Step 4: Identify areas for improvement and provide feedback
[1085] The server then generates feedback based on the analysis, identifying specific areas for improvement and the reasons for them. For example, the feedback could be, "Your arms are positioned too high, making it difficult to make proper impact with the ball."
[1086] Specific practice methods (e.g., specific drills to correct arm position) are also generated as feedback information.
[1087] Step 5: View your feedback
[1088] The server transmits the generated feedback information to the terminal.
[1089] The device receives the feedback information and displays it to the user, using graphics and animations to make it visually easy to understand.
[1090] Specific examples
[1091] Consider a case where a player (user) is practicing on a golf simulator.
[1092] 1. The user starts swinging. The device detects the start of the swing and the camera captures the swing form.
[1093] 2. The device sends this video data to the server.
[1094] 3. The server receives the video data and analyzes it using multimodal AI, detecting, for example, problems such as the arm being positioned slightly too high.
[1095] 4. The server generates feedback such as "Your arms are positioned too high, making it difficult to make a proper impact," and sends it to the device along with specific practice instructions.
[1096] 5. The device displays feedback information to the user in an intuitive format, for example, by providing a graphic indication of arm position and visually demonstrating correct form.
[1097] This allows users to receive instant feedback and use it to improve their form. By repeating this process, players can efficiently improve their golfing skills.
[1098] The processing flow will be explained below.
[1099] Step 1:
[1100] The user stands at the starting position of the golf simulator and begins swinging.
[1101] The device detects the start of the swing and activates the camera in conjunction with the golf simulator.
[1102] The camera captures the player's swing form in real time and acquires video data at a high frame rate.
[1103] Step 2:
[1104] The device temporarily stores the captured video data and performs preprocessing, which includes noise reduction and image correction.
[1105] The terminal performs data compression and encryption to send the pre-processed data to the server.
[1106] Step 3:
[1107] The server inputs the received video data into the multimodal AI and begins analysis.
[1108] Multimodal AI extracts features such as swing speed, angle, and wrist movement from video data to detect problems with swing form. For example, if the arms are positioned too high, this will be identified.
[1109] Step 4:
[1110] The server then uses the analysis results to identify specific areas for improvement and the reasons for them, such as determining that "your arms are positioned too high, making it difficult to make proper impact with the ball."
[1111] The server generates specific feedback based on the identified areas for improvement, including specific practice methods (e.g., specific drills to correct arm position).
[1112] Step 5:
[1113] The server transmits the generated feedback information to the terminal.
[1114] The device receives the feedback information and displays it to the user, using graphics and animations to make it visually easy to understand.
[1115] Step 6:
[1116] The user reviews the displayed feedback information and practices to correct their next swing form. They adjust their form based on the specific improvements indicated in the feedback.
[1117] This series of steps allows users to efficiently improve their swing form, and real-time feedback can help them improve their technique quickly.
[1118] Example 1
[1119] Next, a description will be given of Example 1. In the following description, the data processing device 12 will be referred to as a "server" and the robot 414 will be referred to as a "terminal."
[1120] For players to efficiently improve their form, it is important to accurately understand their current form and immediately identify specific areas for improvement. However, with conventional technology, analyzing and evaluating form takes time, making it difficult to provide real-time feedback. In addition, the analysis results are not specific, making it difficult for players to understand how to improve.
[1121] The specific processing by the specific processing unit 290 of the data processing device 12 in the first embodiment is realized by the following means.
[1122] In this invention, the server includes a camera for capturing the player's movements in real time, a terminal for temporarily storing and preprocessing the captured video data, a transmission means for compressing and encrypting the preprocessed data and transmitting it to the server, a processing means for analyzing the video data received by the server using multimodal AI to detect problems with the player's swing form, a feedback generation means for identifying specific areas for improvement based on the analysis results and generating feedback information, and a means for transmitting the generated feedback information to the terminal and displaying it to the user. This allows the player to receive specific and detailed feedback in real time, enabling them to improve their form efficiently.
[1123] A "player" is someone who uses the system to improve their own movements and form.
[1124] "Movement" refers to the physical movements of the player related to their swing and form.
[1125] "Real-time" refers to immediate processing and feedback without delay.
[1126] "Photography means" refers to a device used to capture the player's actions, such as a camera.
[1127] "Terminal" refers to a device that temporarily stores, pre-processes, and transmits data, typically a computer or smart device.
[1128] "Transmission means" refers to the process or function that transmits the pre-processed data to the server.
[1129] "Server" refers to a device or system that analyzes video data and generates feedback information.
[1130] "Multimodal AI" refers to artificial intelligence technology that comprehensively analyzes multiple input data (video, audio, sensors, etc.).
[1131] "Processing means" refers to a device or process that has the function of analyzing received video data and detecting problems.
[1132] "Feedback generation means" refers to the process or function that generates specific improvements and practice methods based on the analysis results.
[1133] The "display means" refers to a device or system that has the function of visually displaying the generated feedback information to the user.
[1134] "Video data" refers to video data that captures the player's actions.
[1135] The present invention is a system that captures a player's actions in real time and provides feedback based on the analysis results. Specifically, it is realized by linking a camera, a terminal, a server, and a display.
[1136] The system overview includes the following hardware and software:
[1137] Filming method
[1138] The recording method consists of a camera that captures the player's swing form at a high frame rate. This camera operates at a high frame rate such as 120 fps, allowing for detailed capture of the player's movements. For example, it is preferable to use a high-performance sports camera or a dedicated motion capture camera rather than a regular webcam.
[1139] Terminal
[1140] The terminal is a device that temporarily stores and pre-processes the video data captured by the camera. This pre-processing includes noise reduction and color correction. The video data is then compressed into MPEG format and encrypted with AES. This allows the data to be sent securely to the server while maintaining its quality. The terminal requires a high-performance processor and large-capacity storage.
[1141] Transmission method
[1142] The transmission means has the function of sending preprocessed data from the terminal to the server. It is important to use the HTTPS protocol for data transmission, ensuring safety and speed.
[1143] server
[1144] The server includes a processing means for analyzing the received video data using multimodal AI. This AI is built using deep learning frameworks such as TensorFlow and PyTorch, and performs detailed analysis of swing speed, angle, wrist movement, and other parameters. As a result of the analysis, it detects problems with the swing form and generates feedback. The server then generates feedback that includes specific practice methods and corrections.
[1145] Feedback Generation Method
[1146] The feedback generator explains specific areas for improvement and the reasons for them based on the analysis results from the server. For example, it may give specific indications such as, "Your right arm is too high, making your impact unstable." It also suggests practice methods for keeping the right arm low.
[1147] Display means
[1148] The display means is a device or system that visually displays the generated feedback information to the user. The terminal receives the feedback information and uses graphics or animations to show it to the user in an easy-to-understand manner. For example, by visually displaying the difference between incorrect and correct form, the user can understand specifically which parts need to be corrected.
[1149] Specific examples
[1150] For example, when a player practices on a golf simulator, the system operates in the following steps:
[1151] 1. The user starts swinging, the device detects the movement, and the camera captures the swing form at a high frame rate.
[1152] 2. The device temporarily stores the video data, performs preprocessing, compresses it into MPEG format, encrypts it using AES, and sends it to the server.
[1153] 3. The server receives the video data and analyzes it using multimodal AI, detecting problems such as "the right arm is raised too high."
[1154] 4. Based on the analysis results, the server generates feedback recommending "practice swinging with both arms fixed at waist height to practice keeping the right arm low" and sends it to the device.
[1155] 5. The device displays the feedback information in a graphical format, visually showing the user specific areas for improvement and correct form.
[1156] Prompt Sentence Examples
[1157] The user can enter prompts for the generative AI model, such as:
[1158] "Analyze the swing footage of the player and if there are any problems with their swing form, tell them how to improve it."
[1159] Thus, the present invention is a system that can provide detailed feedback in real time to help players improve their form efficiently.
[1160] The flow of the identification process in the first embodiment will be described with reference to FIG.
[1161] The flow of this system's program processing
[1162] Step 1:
[1163] The user stands at the starting position of the golf simulator and begins swinging.
[1164] Input: User swing motion
[1165] Operation: The device uses sensors to recognize when the user starts swinging.
[1166] Output: Swing start detection signal
[1167] Step 2:
[1168] The device activates the camera and captures the player's swing form in real time.
[1169] Input: Swing start detection signal
[1170] How it works: The device controls the camera and captures video data of the swing form at a high frame rate such as 120 fps.
[1171] Output: Captured video data
[1172] Step 3:
[1173] The device temporarily stores the captured video data and performs preprocessing.
[1174] Input: Captured video data
[1175] How it works: The device breaks down the video data into frames and performs noise reduction and color correction.
[1176] Output: Pre-processed video data
[1177] Step 4:
[1178] The terminal compresses and encrypts the pre-processed video data.
[1179] Input: Preprocessed video data
[1180] Operation: Video data is compressed into MPEG format and encrypted with AES.
[1181] Output: Compressed and encrypted video data
[1182] Step 5:
[1183] The terminal transmits the compressed and encrypted video data to the server.
[1184] Input: Compressed and encrypted video data
[1185] How it works: The device uses the HTTPS protocol to securely upload data to the server.
[1186] Output: Video data sent to the server
[1187] Step 6:
[1188] The video data received by the server is analyzed using multimodal AI.
[1189] Input: Received video data
[1190] How it works: The server inputs video data into the AI model and extracts swing form characteristics (e.g., speed, angle, wrist movement).
[1191] Output: Swing form feature data
[1192] Step 7:
[1193] The server detects problems based on characteristic data of the swing form.
[1194] Input: Swing form characteristics data
[1195] How it works: The server compares the extracted features with predefined standards and identifies form flaws (e.g., the right arm is too high).
[1196] Output: Detected issue data
[1197] Step 8:
[1198] The server generates feedback information based on the analysis results.
[1199] Input: Detected issue data
[1200] How it works: The server generates feedback with specific areas for improvement and reasons for doing so, and also suggests practice methods.
[1201] Output: Generated feedback information
[1202] Step 9:
[1203] The server transmits the generated feedback information to the terminal.
[1204] Input: Generated feedback information
[1205] Operation: The server sends feedback information to the device in JSON format or similar.
[1206] Output: Feedback information sent to the terminal
[1207] Step 10:
[1208] The terminal displays the feedback information to the user.
[1209] Input: Feedback information sent to the device
[1210] Action: The device visually displays feedback information to the user using graphics and animations.
[1211] Output: Feedback information displayed to the user
[1212] (Application example 1)
[1213] Next, a description will be given of Application Example 1. In the following description, the data processing device 12 will be referred to as a "server" and the robot 414 will be referred to as a "terminal."
[1214] Conventional technologies exist that provide real-time feedback on player movements and form improvements. However, there is a need for a system that can optimize the movement accuracy and efficiency of robots and equipment in real time, even in factory operations. The present invention provides technology that improves movement accuracy and efficiency by capturing, analyzing, and providing feedback on the movements of factory robots.
[1215] The specific processing by the specific processing unit 290 of the data processing device 12 in the application example 1 is realized by the following means.
[1216] In this invention, the server includes a camera for capturing the actions of the player or the actions of the equipment in real time, a processing unit for receiving and analyzing the captured video data, and a display unit for providing feedback information to the player or operator based on the analysis results, thereby enabling efficient improvement of the operations of factory robots and equipment.
[1217] "Photographing means" refers to a device for capturing the actions of a player or device in real time.
[1218] A "processing means" is a device or system for receiving and analyzing captured video data.
[1219] "Display means" refers to a device or system for providing feedback information to a player or operator based on the analysis results.
[1220] "Multimodal AI" is an artificial intelligence technology that integrates and analyzes multiple data modalities (e.g., video, audio, text).
[1221] "Feedback" is information based on the analysis results that provides improvements to movements and form, as well as optimal practice methods.
[1222] "Capture" refers to the process of converting a subject's movements and form into data in real time using a photographic device.
[1223] A "server" is a computer system that receives data, analyzes it, and generates feedback information.
[1224] The present invention is a system for optimizing the operation of a factory robot, and includes a photographing means for capturing the operation of a player or equipment in real time, a processing means for receiving and analyzing the photographed video data, and a display means for providing feedback information to a player or operator based on the analysis results.
[1225] System Program
[1226] The main hardware and software components of this system include smartphones, factory robots, smartphone apps (e.g., cross-platform development with Flutter or React Native), servers, multimodal AI (e.g., TensorFlow or PyTorch), and communication protocols (e.g., HTTPS or WebSocket).
[1227] Natural language explanation of program processing
[1228] Step 1: Capture video with a camera
[1229] The smartphone camera captures the robot's movements in the factory in real time, capturing the video data at a high frame rate.
[1230] Step 2: Sending video data
[1231] The captured video data is temporarily stored on the smartphone and then encrypted and sent to a server.
[1232] Step 3: AI-powered video analysis
[1233] The server inputs the received video data into the multimodal AI for analysis. The AI extracts characteristics of the robot's movements from the video data and detects problems with the movements. For example, if the robot's arm is moving too slowly, this problem will be identified.
[1234] Step 4: Identify areas for improvement and provide feedback
[1235] Based on the analysis results, the server identifies specific areas for improvement and the reasons for them, and generates feedback information. For example, the feedback might say, "The arm's movements are too slow, reducing work efficiency." This feedback might also include specific drills to improve the arm's speed by 20%.
[1236] Step 5: View your feedback
[1237] The server sends the generated feedback information to a smartphone and displays it for the user or operator to check. This display uses graphics and animations to make it visually easy to understand.
[1238] Specific examples
[1239] For example, if the robot arm is analyzed as moving too slowly, the app will display the following feedback:
[1240] "The arm is moving too slowly, reducing work efficiency. Please perform drills to increase the arm's movement speed by 20%."
[1241] AI prompt examples
[1242] "Analyze the robot arm's movement pattern in this video, identify the speed, angle, and errors, and suggest improvements based on that."
[1243] The flow of the specific processing in the application example 1 will be described with reference to FIG.
[1244] Step 1:
[1245] A user captures the operation of a factory robot in real time using a smartphone camera. The smartphone camera acquires video data at a high frame rate and temporarily stores the data within the smartphone. The input is the captured video data, and the output is the temporarily stored video data.
[1246] Step 2:
[1247] The device preprocesses the stored video data. Preprocessing includes compressing and encrypting the data. The compressed and encrypted data is then sent to the server. The input is the stored video data, and the output is the compressed and encrypted data.
[1248] Step 3:
[1249] The server decompresses the received video data and prepares it for analysis. The decompressed data is input into the multimodal AI, and data analysis begins. The input is compressed and encrypted data, and the output is analyzable data.
[1250] Step 4:
[1251] The server's multimodal AI extracts features related to the robot's movements from the video data and analyzes any problems with the movements. For example, it analyzes the robot's arm's movement speed, angle, and identifies any errors. The input is analyzable data, and the output is the analysis results, including any problems with the movements.
[1252] Step 5:
[1253] Based on the analysis results, the server identifies specific areas for improvement and the reasons for them. It then generates feedback information to present the areas for improvement to the user. For example, it generates feedback information such as "The arm's movements are too slow, reducing work efficiency." The input is the analysis results, and the output is feedback information.
[1254] Step 6:
[1255] The server generates feedback information and sends it to the terminal. The terminal receives the feedback information and displays it in a visually understandable format for the user, for example, using graphics or animations to demonstrate correct actions or forms. The input is the feedback information, and the output is the feedback content displayed to the user.
[1256] Step 7:
[1257] Based on the displayed feedback information, the user can take specific actions to improve the factory robot's operation. For example, they can adjust the robot's parameters to improve its speed and accuracy. In this step, the feedback information is directly used for on-site improvement actions. The input is the feedback displayed to the user, and the output is the improved robot's operation.
[1258] Furthermore, an emotion engine that estimates the user's emotion may be further combined. That is, the identification processing unit 290 may estimate the user's emotion using the emotion identification model 59, and perform identification processing using the user's emotion.
[1259] The present invention relates to a system that combines a photographing means, a processing means, a display means, and an emotion engine to efficiently improve a player's golf form. Below, the processing of the program of this system will be explained in natural language, and an embodiment of the invention will be described in detail with concrete examples.
[1260] System Overview
[1261] This system uses a camera linked to a golf simulator to capture the player's swing form in real time. The captured video data is sent to a server, which then analyzes it using multimodal AI. It also recognizes the user's emotions using an emotion engine and generates feedback information based on the analysis results and emotional state. The feedback information is provided to the user in real time via their device, suggesting areas for improvement in form, optimal practice methods, and even encouraging messages based on their emotions.
[1262] Program processing
[1263] Step 1: Capture video with a camera
[1264] The user stands at the starting position of the golf simulator and begins swinging.
[1265] The device detects the start of the swing and activates the camera in conjunction with the golf simulator.
[1266] The camera captures the player's swing form in real time and acquires video data at a high frame rate.
[1267] Step 2: Sending video data
[1268] The device temporarily stores the captured video data and performs preprocessing, which includes noise reduction and image correction.
[1269] The terminal performs data compression and encryption to send the pre-processed data to the server.
[1270] Step 3: AI-powered video analysis
[1271] The server inputs the received video data into the multimodal AI and begins analysis.
[1272] Multimodal AI extracts features such as swing speed, angle, and wrist movement from video data to detect problems with swing form. For example, if the arms are positioned too high, this will be identified.
[1273] Step 4: Emotion Recognition with the Emotion Engine
[1274] It uses the device's built-in camera and microphone to capture the user's facial expressions and tone of voice.
[1275] An emotion engine analyzes this data to identify the user's emotional state, for example, recognizing whether the user is focused or stressed.
[1276] Step 5: Identify areas for improvement and generate feedback
[1277] The server generates specific feedback information based on the analysis results and the emotional state identified by the emotion engine.
[1278] The feedback information includes not only improvements to form, but also practice methods and encouraging messages based on the user's emotions.
[1279] Step 6: View your feedback
[1280] The server transmits the generated feedback information to the terminal.
[1281] The device receives the feedback information and displays it to the user. The display uses graphics and animations to make it visually easy to understand. It also displays encouraging messages that match the user's emotions.
[1282] Specific examples
[1283] Consider a case where a player (user) is practicing on a golf simulator.
[1284] 1. The user starts swinging. The device detects the start of the swing and the camera captures the swing form.
[1285] 2. The device sends this video data to the server.
[1286] 3. The server receives the video data and analyzes it using multimodal AI, detecting, for example, problems such as the arm being positioned slightly too high.
[1287] 4. The device captures the user's facial expressions and tone of voice, and the emotion engine recognizes that the user is feeling a little stressed.
[1288] 5. The server generates feedback such as "Your arms are too high, making it difficult to make a good impact," and includes an encouraging message such as "Calm down, try lowering your arms a little next time."
[1289] 6. The device receives the feedback information and visually displays it to the user.
[1290] This process allows users to receive immediate feedback and improve their form. Furthermore, encouraging messages that take emotional information into account improve the user's training experience and enable more effective learning.
[1291] The processing flow will be explained below.
[1292] Step 1:
[1293] The user stands at the starting position of the golf simulator and begins swinging.
[1294] The device detects the start of the swing and activates the camera in conjunction with the golf simulator.
[1295] The camera captures the player's swing form in real time and acquires video data at a high frame rate.
[1296] Step 2:
[1297] The device temporarily stores the captured video data and performs preprocessing, which includes noise reduction and image correction.
[1298] The terminal performs data compression and encryption to send the pre-processed data to the server.
[1299] Step 3:
[1300] The server inputs the received video data into the multimodal AI and begins analysis.
[1301] Multimodal AI extracts features such as swing speed, angle, and wrist movement from video data to detect problems with swing form. For example, if the arms are positioned too high, this will be identified.
[1302] Step 4:
[1303] It uses the device's built-in camera and microphone to capture the user's facial expressions and tone of voice.
[1304] The device processes the captured emotion data in real time and sends it to the emotion engine.
[1305] The emotion engine analyzes the user's emotional state from facial expressions and tone of voice, and recognizes, for example, that the user is feeling stressed.
[1306] Step 5:
[1307] The server generates improvement points and feedback information based on the results of video data analysis and the emotional state determined by the emotion engine.
[1308] The server generates feedback such as "Your arm is positioned too high, making it difficult to achieve proper impact," and depending on the user's stress level, may also include encouraging messages such as "Calm down, try lowering your arm a bit next time."
[1309] Step 6:
[1310] The server transmits the generated feedback information to the terminal.
[1311] The device receives the feedback information and displays it visually to the user. Graphics and animations are used to make the display easy to understand. Encouraging messages tailored to the user's emotions are also displayed.
[1312] Step 7:
[1313] The user checks the displayed feedback information and encouraging messages, and practices to correct their swing form next time. They then adjust their form based on the specific areas for improvement indicated in the feedback. Through this series of processes, users can efficiently improve their form and receive support tailored to their emotional state.
[1314] Example 2
[1315] Next, a description will be given of Example 2. In the following description, the data processing device 12 will be referred to as a "server" and the robot 414 will be referred to as a "terminal."
[1316] To effectively improve golf swing form, it is important to capture the player's movements in real time and provide detailed analysis and immediate feedback. However, conventional systems have had problems with form improvement due to the low accuracy of video data analysis and the low quality of feedback. Furthermore, they lack the ability to provide feedback that takes into account the player's emotional state, which can lead to a loss of motivation.
[1317] The specific processing by the specific processing unit 290 of the data processing device 12 in the second embodiment is realized by the following means.
[1318] In this invention, the server includes a camera for capturing the player's movements in real time, a processor for receiving the captured video data and performing preprocessing such as noise reduction and image correction, a transmitter for compressing and encrypting the preprocessed data and transmitting it to the server, an analyzer for analyzing the received data using multimodal AI and identifying areas for form improvement, an emotion recognition processor for capturing the user's facial expressions and tone of voice based on the analysis results to identify their emotional state, a generator for generating specific feedback information based on the analysis results and their emotional state, and a displayer for providing the generated feedback information to the user. This makes it possible to not only instantly analyze the player's form in detail and provide effective feedback, but also to send appropriate encouraging messages according to the user's emotional state.
[1319] "Photographing means" is a device for capturing the player's movements in real time.
[1320] The "processing means" is a device that receives the captured video data and performs pre-processing such as noise removal and image correction.
[1321] The "transmitting means" is a device that compresses and encrypts the preprocessed data and transmits it to the server.
[1322] The "analysis means" is a device that uses multimodal AI to analyze the received data and extract areas for improvement in form.
[1323] The "emotion recognition means" is a device that captures the user's facial expressions and tone of voice based on the analysis results to identify the user's emotional state.
[1324] The "generating means" is a device that generates specific feedback information based on the analysis results and the emotional state.
[1325] The "display means" is a device for providing the generated feedback information to the user.
[1326] "Multimodal AI" is an artificial intelligence technology that analyzes multiple data modes (such as images and audio) to extract meaning.
[1327] "Feedback information" is information such as improvements to the form and encouraging messages that are provided to the user based on the results obtained from the analysis means and emotion recognition means.
[1328] This invention is a system for efficiently improving a player's golf form, combining a photographing means, a processing means, a transmission means, an analysis means, an emotion recognition means, a generation means, and a display means. This system works in conjunction with a golf simulator to capture the player's swing form in real time and provide immediate feedback based on the analysis results. Furthermore, by taking the user's emotional state into consideration, an effective training environment is realized.
[1329] Hardware and software used
[1330] Camera (photography means)
[1331] This system uses a camera that captures video at a high frame rate, for example, a camera capable of capturing 120 fps (frames per second) is suitable.
[1332] Terminal (processing means and transmission means)
[1333] The device performs pre-processing on the captured video data, including noise reduction and image enhancement. The pre-processed data is then compressed in H.264 format and encrypted with AES, allowing the data to be transmitted efficiently and securely to the server.
[1334] Server (analysis means, emotion recognition means, generation means)
[1335] The server inputs the received data into the multimodal AI for detailed analysis. Specifically, it analyzes swing speed, angle, wrist movement, etc. to identify areas for form improvement. It also has an emotion recognition mechanism that captures the user's facial expressions and tone of voice to identify their emotional state. Specific feedback information is generated based on the analysis results and emotional state.
[1336] Terminal (display means)
[1337] The device visually displays the generated feedback to the user, using graphics and animations to highlight problem areas on the form and providing emotionally tailored encouraging messages.
[1338] Specific examples
[1339] Consider a case where a player (user) is practicing on a golf simulator. When the user starts swinging, the device detects the start of the swing and the camera captures the swing form. The camera records video at 120 fps and temporarily stores the high-quality video in memory.
[1340] The device performs preprocessing such as noise reduction and white balance adjustment, compresses the preprocessed video data in H.264 format, encrypts it using AES, and sends it to the server. The server receives the video data and performs detailed analysis using multimodal AI. For example, the AI can detect problems such as the arm being positioned slightly too high.
[1341] The device then captures the user's facial expressions, and the emotion engine recognizes that the user is feeling a little stressed. The server then points out the problem with the swing form, generating feedback such as, "Your arms are positioned too high, making it difficult to make a proper impact." It also generates encouraging messages such as, "Please stay calm. Next time, try lowering your arms a little."
[1342] The device receives feedback, graphically highlights areas of form that are problematic, and even displays encouraging messages like "Put your arms a little lower."
[1343] Prompt Sentence Examples
[1344] Here are some example prompts to input to a generative AI model:
[1345] "Analyze the golf swing footage of the player and point out any problems with their swing form. Additionally, analyze the player's facial expressions and tone of voice to provide feedback based on the player's emotions."
[1346] These prompts allow you to accurately input the necessary information into the generative AI model and get the desired feedback.
[1347] The flow of the identification process in the second embodiment will be described with reference to FIG.
[1348] Step 1:
[1349] The user stands at the starting position of the golf simulator and begins swinging. The device detects the start of the swing and activates the camera in conjunction with the golf simulator. The camera captures the player's swing form in real time. The input at this time is the user's swing motion, and the output is the captured high-frame-rate video data.
[1350] Step 2:
[1351] The device temporarily stores the captured video data and performs preprocessing such as noise reduction and white balance adjustment. For example, preprocessing removes background noise and adjusts the color tone of the video. The input of this preprocessing is the captured video data, and the output is the preprocessed video data.
[1352] Step 3:
[1353] The terminal compresses the preprocessed video data in H.264 format and encrypts it using the AES method. This reduces the data size and the risk of data leaks. The input is preprocessed video data, and the output is compressed and encrypted video data.
[1354] Step 4:
[1355] The device sends data to the server, using the appropriate protocol to ensure the data travels securely across the network. The input is the compressed and encrypted video data, and the output is the data received by the server.
[1356] Step 5:
[1357] The server inputs the received data into a multimodal AI for detailed analysis. The AI extracts features such as swing speed, angle, and wrist movement to identify problems with the golf form. For example, it detects problems with the arm position being too high. The input for this step is the video data received by the server, and the output is the analyzed form improvements.
[1358] Step 6:
[1359] The device captures the user's facial expressions and tone of voice and inputs this data into an emotion engine, which analyzes whether the user is focused, relaxed, or stressed. The input is the user's facial expressions and tone of voice, and the output is the identified emotional state.
[1360] Step 7:
[1361] The server generates feedback information based on the analysis results and the user's emotional state. It specifically points out areas for improvement in the user's form and creates an encouraging message that corresponds to the user's emotions. For example, it generates a specific message such as "Lower your arm position. Stay calm and try your best next time." The input is the areas for improvement in the user's form and the user's emotional state, and the output is the generated feedback information.
[1362] Step 8:
[1363] The server sends the generated feedback information to the device. The device receives the feedback information and visually displays it to the user. The display uses graphics and animations to highlight problem areas in the swing form and provide encouraging messages. The input is the generated feedback information, and the output is the feedback information displayed to the user.
[1364] (Application example 2)
[1365] Next, a description will be given of Application Example 2. In the following description, the data processing device 12 will be referred to as a "server" and the robot 414 will be referred to as a "terminal."
[1366] While conventional golf swing analysis systems can pinpoint technical issues in a player's swing form, they are unable to provide feedback that takes into account the player's emotional state. As a result, if a player feels stressed or frustrated, they are unable to receive appropriate advice or encouragement, which can lead to training being ineffective.
[1367] The specific processing by the specific processing unit 290 of the data processing device 12 in Application Example 2 is realized by the following means. In this invention, the server includes a shooting means for capturing the player's actions in real time, a processing means for receiving and analyzing the shot video data, a display means for providing feedback information to the player based on the analysis results, an emotion recognition means for recognizing the player's emotional state in real time, and a feedback generation means for generating feedback information based on the analysis results and the player's emotional state. This makes it possible to provide not only technical problems but also appropriate feedback and encouraging messages according to the player's emotional state.
[1368] "Filming means for capturing player's movements in real time" refers to a device for capturing player's movements, particularly golf swings and other exercises, as video in real time.
[1369] The "processing means for receiving and analyzing captured video data" is a system that receives video data captured by the image capture means, analyzes it, and extracts important features and problems.
[1370] The "display means for providing feedback information to the player based on the analysis results" refers to a device such as a monitor or display for providing visual feedback to the player based on the analysis results.
[1371] The "emotion recognition means for recognizing the player's emotional state in real time" is a system that includes a camera and microphone for recognizing emotions in real time from the player's facial expressions and tone of voice, as well as software for analyzing them.
[1372] "Feedback generation means for generating feedback information based on the analysis results and the emotional state" refers to software or algorithms that generate feedback information to be provided to the player, taking into account both the technical analysis results and the player's emotional state.
[1373] A "camera that captures at a high frame rate" is a camera that can capture a large number of frames per second, thereby capturing the player's movements in detail.
[1374] "Multimodal AI" is an artificial intelligence technology that can simultaneously analyze multiple different types of data (e.g., video data and emotional data) and make integrated judgments.
[1375] "Areas for improvement in form" refers to elements or problems in a player's movements or posture that are hindering efficient or effective performance.
[1376] The present invention relates to a system for efficiently improving a player's golf swing form, which includes a camera for capturing the player's movements in real time, a processing means for receiving and analyzing the captured video data, a display means for providing feedback information to the player based on the analysis results, an emotion recognition means for recognizing the player's emotional state in real time, and a feedback generation means for generating feedback information based on the analysis results and the player's emotional state.
[1377] The server uses a camera to capture the player's swing form in real time. When the player starts swinging, the device detects this and activates the camera. The camera captures the player's swing form at a high frame rate, and the video data is temporarily saved on the device. The saved video data is then compressed and encrypted before being sent to the server. The server then analyzes the received video data using multimodal AI to extract features such as swing speed, angle, and wrist movement.
[1378] The device also uses its built-in camera and microphone to capture the player's facial expressions and tone of voice, and the emotion engine analyzes this data to identify the player's emotional state, for example, whether the player is focused or stressed.
[1379] The server uses a feedback generation means to generate specific feedback information based on the analysis results and the emotional state identified by the emotion engine. The feedback information includes not only improvements to form, but also practice methods and encouraging messages that correspond to the player's emotions. The generated feedback information is sent to the device in real time, and the device visually displays it to the player. The display uses graphics and animations to make it visually easy to understand. Encouraging messages that match the player's emotions are also displayed.
[1380] As a concrete example, consider a player practicing on a golf simulator. When the player starts swinging, the device detects this and the camera captures the swing form. This video data is sent to a server and analyzed by multimodal AI. For example, the analysis may detect a problem: "Your arms are positioned a little too high." The device also captures the player's facial expressions and tone of voice, and the emotion engine recognizes that the player is feeling a little stressed. The server generates feedback such as, "Your arms are positioned too high, making it difficult to make a proper impact," and includes an encouraging message such as, "Please calm down. Next time, try lowering your arms a little." This feedback information is sent to the device and displayed visually to the player.
[1381] Examples of prompts to input to a generative AI model include:
[1382] "Based on the player's swing footage and emotional data, generate feedback suggesting form improvements and appropriate practice methods. For example, "Your arms are positioned too high, making it difficult to make a proper impact. Calm down and try lowering your arms a little next time."
[1383] In this way, this system pinpoints a player's technical problems with high accuracy while providing training support that also takes into account their emotional side.
[1384] The flow of the specific processing in the application example 2 will be described with reference to FIG.
[1385] Step 1:
[1386] The player stands at the starting position of the golf simulator and begins swinging. The device detects the start of the swing and activates the camera in conjunction with the golf simulator. The camera captures the player's swing form at a high frame rate and obtains the video data. The input of this step is the player's swing motion, and the output is the captured video data.
[1387] Step 2:
[1388] The device temporarily stores the captured video data and performs preprocessing such as noise reduction and image correction. Before sending the processed data to the server, the data is compressed and encrypted. The input of this step is the captured video data, and the output is the compressed and encrypted video data.
[1389] Step 3:
[1390] The server inputs the received video data into a multimodal AI to extract features such as swing speed, angle, and wrist movement. Through AI analysis, problems with the swing form are identified. The input for this step is compressed and encrypted video data, and the output is the analyzed swing form feature data and problems.
[1391] Step 4:
[1392] The device's built-in camera and microphone are used to capture the user's facial expressions and tone of voice. The emotion engine analyzes this data to identify the user's emotional state. For example, it recognizes whether the user is focused or stressed. The input for this step is facial expression and tone of voice data, and the output is the recognized emotional state.
[1393] Step 5:
[1394] The server generates specific feedback information based on the analysis results and the emotional state identified by the emotion engine. Using the feedback generation means, feedback is generated that includes form improvements, practice methods based on the user's emotions, and encouraging messages. The input for this step is the analyzed swing form feature data and the recognized emotional state, and the output is the generated feedback information.
[1395] Step 6:
[1396] The server sends the generated feedback information to the terminal. The terminal receives the feedback information and visually displays it to the user. This display uses graphics and animations, and also displays encouraging messages that match the user's emotions. The input of this step is the generated feedback information, and the output is the feedback information displayed to the user.
[1397] This series of processing steps allows users to receive instant feedback and effectively improve their form, while encouraging messages that take emotional information into account significantly improve the training experience.
[1398] The specific processing unit 290 transmits the result of the specific processing to the robot 414. In the robot 414, the control unit 46A causes the speaker 240 and the control target 443 to output the result of the specific processing. The microphone 238 acquires voice indicating a user input regarding the result of the specific processing. The control unit 46A transmits voice data indicating the user input acquired by the microphone 238 to the data processing device 12. In the data processing device 12, the specific processing unit 290 acquires the voice data.
[1399] The data generation model 58 is a so-called generative AI (Artificial Intelligence). An example of the data generation model 58 is ChatGPT (Internet Search<URL: https: / / openai.com / blog / chatgpt> ), Gemini (Internet search <url: https: gemini.google.com ?hl="ja">) and other generation AIs. The data generation model 58 is obtained by performing deep learning on a neural network. A prompt including an instruction is input to the data generation model 58, and inference data such as voice data indicating voice, text data indicating text, and image data indicating an image is also input. The data generation model 58 performs inference on the input inference data in accordance with the instruction indicated by the prompt, and outputs the inference result in a data format such as voice data and text data. Here, inference refers to, for example, analysis, classification, prediction, and / or summarization.
[1400] In the above embodiment, an example was given in which the specific processing is performed by the data processing device 12, but the technology of the present disclosure is not limited to this, and the specific processing may be performed by the robot 414.
[1401] The emotion identification model 59 as an emotion engine may determine the user's emotion according to a specific mapping. Specifically, the emotion identification model 59 may determine the user's emotion according to an emotion map (see FIG. 9), which is a specific mapping. Similarly, the emotion identification model 59 may determine the robot's emotion, and the identification processing unit 290 may perform identification processing using the robot's emotion.
[1402] FIG. 9 is a diagram illustrating an emotion map 400 on which multiple emotions are mapped. In the emotion map 400, emotions are arranged in concentric circles radiating from the center. Emotions closer to the center of the concentric circles are more primitive. Emotions representing states and actions arising from a state of mind are arranged on the outer edges of the concentric circles. The concept of emotion includes both affect and mental states. Emotions generally generated from reactions occurring in the brain are arranged on the left side of the concentric circles. Emotions generally induced by situational judgment are arranged on the right side of the concentric circles. Emotions generally generated from reactions occurring in the brain and induced by situational judgment are arranged on the upper and lower sides of the concentric circles. Furthermore, the emotion of "pleasure" is arranged on the upper side of the concentric circles, and the emotion of "discomfort" is arranged on the lower side. In this way, in the emotion map 400, multiple emotions are mapped based on the structure by which emotions are generated, and emotions that tend to occur simultaneously are mapped close to each other.
[1403] These emotions are distributed in the 3 o'clock direction on emotion map 400, and typically fluctuate between relief and anxiety. In the right half of emotion map 400, situational awareness dominates over internal sensations, resulting in a sense of calm.
[1404] The inside of emotion map 400 represents what is going on in the mind, and the outside of emotion map 400 represents behavior, so the further you go outside emotion map 400, the more visible the emotions become (the more they are expressed in behavior).
[1405] Human emotions are based on various balances, such as posture and blood sugar levels. When these balances deviate from the ideal, a state of discomfort is indicated, and when they approach the ideal, a state of pleasure is indicated. Emotions can also be created for robots, automobiles, and motorcycles, based on various balances, such as posture and remaining battery life. When these balances deviate from the ideal, a state of discomfort is indicated, and when they approach the ideal, a state of pleasure is indicated. An emotion map can be generated, for example, based on Dr. Mitsuyoshi's emotion map (Research on Voice Emotion Recognition and Emotional Brain Physiological Signal Analysis Systems, Tokushima University, Doctoral Dissertation: https: / / ci.nii.ac.jp / naid / 500000375379). The left half of the emotion map lists emotions belonging to the "reaction" domain, where sensation is dominant. The right half of the emotion map lists emotions belonging to the "situation" domain, where situational awareness is dominant.
[1406] The emotion map defines two emotions that promote learning. One is a negative emotion on the situation side, around the middle of "repentance" or "reflection." In other words, this occurs when the robot experiences negative emotions such as "I never want to feel this way again" or "I don't want to be scolded again." The other is a positive emotion on the response side, around "desire." In other words, this occurs when the robot experiences positive feelings such as "I want more" or "I want to know more."
[1407] The emotion identification model 59 inputs user input into a pre-trained neural network, obtains emotion values indicating each emotion shown in the emotion map 400, and determines the user's emotion. This neural network is pre-trained based on multiple pieces of training data that are combinations of user input and emotion values indicating each emotion shown in the emotion map 400. Furthermore, this neural network is trained so that emotions that are located close to each other have similar values, as in the emotion map 900 shown in FIG. 10. FIG. 10 shows an example in which multiple emotions, "relieved," "calm," and "reassuring," have similar emotion values.
[1408] The system according to the present disclosure has been described above mainly with respect to the functions of the data processing device 12, but the system according to the present disclosure is not necessarily implemented on a server. The system according to the present disclosure may be implemented as a general information processing system. The present disclosure may be implemented, for example, as a software program running on a personal computer or an application running on a smartphone, etc. The method according to the present disclosure may be provided to users in the form of SaaS (Software as a Service).
[1409] In the above embodiment, an example was given in which the specific processing is performed by one computer 22, but the technology of the present disclosure is not limited to this, and the specific processing may be distributed and performed by a plurality of computers including the computer 22. For example, the data generation model 58 may be provided in an external device of the data processing device 12, and data may be generated in the external device in accordance with input data.
[1410] In the above embodiment, an example in which the specific processing program 56 is stored in the storage 32 has been described, but the technology of the present disclosure is not limited to this. For example, the specific processing program 56 may be stored in a portable, computer-readable, non-transitory storage medium such as a USB (Universal Serial Bus) memory. The specific processing program 56 stored in the non-transitory storage medium is installed in the computer 22 of the data processing device 12. The processor 28 executes the specific processing in accordance with the specific processing program 56.
[1411] Alternatively, the specific processing program 56 may be stored in a storage device such as a server connected to the data processing device 12 via the network 54, and the specific processing program 56 may be downloaded and installed on the computer 22 in response to a request from the data processing device 12.
[1412] It is not necessary to store all of the specific processing program 56 in a storage device such as a server connected to the data processing device 12 via the network 54, or to store all of the specific processing program 56 in the storage 32; only a portion of the specific processing program 56 may be stored.
[1413] The hardware resource for executing a specific process can be any of the following processors: An example of a processor is a CPU, which is a general-purpose processor that functions as a hardware resource for executing a specific process by executing software, i.e., a program. Another example of a processor is a dedicated electrical circuit, such as an FPGA (Field-Programmable Gate Array), a PLD (Programmable Logic Device), or an ASIC (Application Specific Integrated Circuit), which is a processor with a circuit configuration designed specifically for executing a specific process. Each processor has built-in or connected memory, and each processor uses the memory to execute the specific process.
[1414] The hardware resource that executes the specific processing may be configured with one of these various processors, or may be configured with a combination of two or more processors of the same or different types (for example, a combination of multiple FPGAs, or a combination of a CPU and an FPGA). Also, the hardware resource that executes the specific processing may be a single processor.
[1415] As an example of a system configured with a single processor, first, one processor is configured by combining one or more CPUs and software, and this processor functions as a hardware resource that executes a specific process. Second, there is a system that uses a processor that realizes the functions of an entire system including multiple hardware resources that execute a specific process on a single IC chip, as typified by SoC (System-on-a-chip). In this way, a specific process is realized using one or more of the above-mentioned various processors as hardware resources.
[1416] Furthermore, the hardware structure of these various processors can be, more specifically, an electric circuit that combines circuit elements such as semiconductor devices. The specific processing described above is merely an example. Therefore, it goes without saying that unnecessary steps may be deleted, new steps may be added, or the processing order may be rearranged, without departing from the spirit of the invention.
[1417] The above-described description and illustrations are a detailed explanation of the parts related to the technology of the present disclosure and are merely an example of the technology of the present disclosure. For example, the above description of the configuration, functions, actions, and effects is an explanation of an example of the configuration, functions, actions, and effects of the parts related to the technology of the present disclosure. Therefore, it goes without saying that unnecessary parts may be deleted, new elements may be added, or replacements may be made to the above-described description and illustrations within the scope of the gist of the technology of the present disclosure. Furthermore, to avoid confusion and facilitate understanding of the parts related to the technology of the present disclosure, the above-described description and illustrations omit explanations of common technical knowledge that do not require particular explanation to enable the implementation of the technology of the present disclosure.
[1418] All publications, patent applications, and technical standards mentioned in this specification are herein incorporated by reference to the same extent as if each individual publication, patent application, or technical standard was specifically and individually indicated to be incorporated by reference.
[1419] The following is further disclosed regarding the above embodiment.
[1420] (Claim 1)
[1421] A means of capturing the player's movements in real time;
[1422] processing means for receiving and analyzing the captured video data;
[1423] display means for providing feedback information to the player based on the analysis results;
[1424] A system including:
[1425] (Claim 2)
[1426] The camera is used to capture the player's swing form at a high frame rate.
[1427] 10. The system of claim 1.
[1428] (Claim 3)
[1429] The processing means is a server that uses multimodal AI to analyze video data and extract areas for form improvement.
[1430] 10. The system of claim 1.
[1431] (Claim 4)
[1432] The display means is a display that displays the analysis results in real time on the player's device.
[1433] 10. The system of claim 1.
[1434] "Example 1"
[1435] (Claim 1)
[1436] A means of capturing the player's movements in real time;
[1437] A terminal that temporarily stores and pre-processes the captured video data;
[1438] a transmitting means for compressing and encrypting the preprocessed data and transmitting the compressed data to a server;
[1439] A processing means for analyzing the received video data in a server using multimodal AI to detect problems with the swing form;
[1440] a feedback generating means for identifying specific improvements based on the analysis results and generating feedback information;
[1441] means for transmitting the generated feedback information to a terminal and displaying it to a user;
[1442] A system including:
[1443] (Claim 2)
[1444] 2. The system of claim 1, wherein the imaging means is a camera that captures the player's swing form at a high frame rate.
[1445] (Claim 3)
[1446] 2. The system of claim 1, wherein the processing means is a server that uses multimodal AI to analyze the video data and extract areas for improvement in the form.
[1447] "Application Example 1"
[1448] (Claim 1)
[1449] a means for capturing a player's actions or a device's actions in real time;
[1450] processing means for receiving and analyzing the captured video data;
[1451] display means for providing feedback information to a player or operator based on the analysis results;
[1452] A system including:
[1453] (Claim 2)
[1454] The shooting means is a camera that captures the swing form and movements of the player or equipment at a high frame rate.
[1455] 10. The system of claim 1.
[1456] (Claim 3)
[1457] The processing means is a server that uses multimodal AI to analyze video data and extract areas for improvement in form and movement.
[1458] 10. The system of claim 1.
[1459] "Example 2: Combining Emotion Engines"
[1460] (Claim 1)
[1461] A means of capturing the player's movements in real time;
[1462] a processing means for receiving the captured video data and performing pre-processing such as noise removal and image correction;
[1463] a transmitting means for compressing and encrypting the preprocessed data and transmitting the compressed data to a server;
[1464] An analytical method that uses multimodal AI to analyze the received data and extract areas for form improvement.
[1465] an emotion recognition means for capturing a user's facial expression and tone of voice based on the analysis results to identify the user's emotional state;
[1466] generating means for generating specific feedback information based on the analysis result and the emotional state;
[1467] display means for providing the generated feedback information to a user;
[1468] A system including:
[1469] (Claim 2)
[1470] 2. The system of claim 1, wherein the imaging means is a camera that captures the player's swing form at a high frame rate.
[1471] (Claim 3)
[1472] The system of claim 1, wherein the analysis means extracts features such as swing speed, angle, and wrist movement and uses multimodal AI to specifically identify detected form problems.
[1473] "Application example 2 when combining emotion engines"
[1474] (Claim 1)
[1475] A means of capturing the player's movements in real time;
[1476] processing means for receiving and analyzing the captured video data;
[1477] display means for providing feedback information to the player based on the analysis results;
[1478] an emotion recognition means for recognizing the player's emotional state in real time;
[1479] a feedback generating means for generating feedback information based on the analysis result and the emotional state;
[1480] A system including:
[1481] (Claim 2)
[1482] The camera is used to capture the player's swing form at a high frame rate.
[1483] 10. The system of claim 1.
[1484] (Claim 3)
[1485] The processing means is a server that uses multimodal AI to analyze video data and extract areas for form improvement.
[1486] 10. The system of claim 1. [Explanation of symbols]
[1487] 10, 210, 310, 410 Data Processing Systems 12 Data Processing Device 14 Smart Devices 214 Smart Glasses 314 Headset-type terminal 414 Robot< / url:> < / url:> < / url:> < / url:>
Claims
1. A means of capturing the player's movements in real time; processing means for receiving and analyzing the captured video data; display means for providing feedback information to the player based on the analysis results; A system including:
2. The camera is used to capture the player's swing form at a high frame rate. The system of claim 1 .
3. The processing means is a server that uses multimodal AI to analyze video data and extract areas for form improvement. The system of claim 1 .
4. The display means is a display that displays the analysis results in real time on the player's device. The system of claim 1 .
Citation Information
Patent Citations
Persona chatbot control method and system
JP2022180282A