system

A system using a camera, server, and AR devices provides real-time, customized feedback to athletes, addressing the lack of effective skill improvement in existing systems by offering immediate and personalized training assistance.

JP2026103463APending Publication Date: 2026-06-24SOFTBANK GROUP CORP
View PDF 1 Cites 0 Cited by

Patent Information

Authority / Receiving Office
JP · JP
Patent Type
Applications
Current Assignee / Owner
SOFTBANK GROUP CORP
Filing Date
2024-12-12
Publication Date
2026-06-24

Smart Images

  • Figure 2026103463000001_ABST
    Figure 2026103463000001_ABST
Patent Text Reader

Abstract

We provide the system. [Solution] A device for acquiring video information, A device that receives the aforementioned video information, analyzes the information, and evaluates the actions of the subject, Based on the aforementioned evaluation, a device for generating information to provide feedback to the subject, A device that presents the generated information to the target person, A system that includes measures to support the improvement of the subject's movements.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The technology of the present disclosure relates to a system.

Background Art

[0002] Patent Document 1 discloses a method for controlling a persona chatbot, which is performed by at least one processor, including the steps of receiving a user utterance, adding the user utterance to a prompt including an instruction sentence related to the description of the chatbot's character, encoding the prompt, and inputting the encoded prompt into a language model to generate a chatbot utterance in response to the user utterance.

Prior Art Documents

Patent Documents

[0003]

Patent Document 1

Summary of the Invention

Problems to be Solved by the Invention

[0004] In training for sports players to improve their skills, there is a problem that it is difficult to obtain real-time feedback. In particular, since the opportunity to receive world-class coaching is limited, players often practice on their own, and there is a problem that efficient skill improvement is hindered.

Means for Solving the Problems

[0005] This invention provides a system comprising terminal means for acquiring video data, server means for analyzing the video data and evaluating the actions, and server means for generating feedback based on the evaluation. This system makes it possible to efficiently support skill improvement by analyzing the form and actions of athletes in real time and providing customized feedback as visual and audio data.

[0006] "Video data" refers to digital information, including the movement and still images of an object, acquired by a camera or other imaging device.

[0007] "Terminal means" refers to devices used for acquiring, displaying, or outputting video data, and includes devices such as cameras, displays, AR glasses, and earphones.

[0008] A "server system" is a computing device for receiving, analyzing, and processing data, and is a system that includes a processor, memory, storage devices, and the like.

[0009] "Feedback" refers to providing information based on an evaluation of the subject's actions, including areas for improvement and instructions for precise technique.

[0010] "Analysis" is the process of extracting key points of movement based on acquired video data and evaluating the subject's form and technique.

[0011] "Profile information" refers to individual information about the subject, including data such as age, gender, exercise history, and goals. [Brief explanation of the drawing]

[0012] [Figure 1] This is a conceptual diagram showing an example of the configuration of a data processing system according to the first embodiment. [Figure 2] This is a conceptual diagram showing an example of the essential functions of a data processing device and a smart device according to the first embodiment. [Figure 3] It is a conceptual diagram showing an example of the configuration of a data processing system according to the second embodiment. [Figure 4] It is a conceptual diagram showing an example of the main functions of a data processing device and smart glasses according to the second embodiment. [Figure 5] It is a conceptual diagram showing an example of the configuration of a data processing system according to the third embodiment. [Figure 6] It is a conceptual diagram showing an example of the main functions of a data processing device and a headset-type terminal according to the third embodiment. [Figure 7] It is a conceptual diagram showing an example of the configuration of a data processing system according to the fourth embodiment. [Figure 8] It is a conceptual diagram showing an example of the main functions of a data processing device and a robot according to the fourth embodiment. [Figure 9] It shows an emotion map to which a plurality of emotions are mapped. [Figure 10] It shows an emotion map to which a plurality of emotions are mapped. [Figure 11] It is a sequence diagram showing the processing flow of the data processing system in Example 1. [Figure 12] It is a sequence diagram showing the processing flow of the data processing system in Application Example 1. [Figure 13] It is a sequence diagram showing the processing flow of the data processing system in Example 2 when an emotion engine is combined. [Figure 14] It is a sequence diagram showing the processing flow of the data processing system in Application Example 2 when an emotion engine is combined.

Embodiments for Carrying Out the Invention

[0013] Hereinafter, an example of an embodiment of a system according to the technology of the present disclosure will be described with reference to the accompanying drawings.

[0014] First, the language used in the following description will be explained.

[0015] In the following embodiments, the numbered processor (hereinafter simply referred to as "processor") may be a single arithmetic unit or a combination of multiple arithmetic units. Also, the processor may be a single type of arithmetic unit or a combination of multiple types of arithmetic units. Examples of arithmetic units include a CPU (Central Processing Unit), a GPU (Graphics Processing Unit), a GPGPU (General-Purpose computing on Graphics Processing Units), an APU (Accelerated Processing Unit), and the like.

[0016] In the following embodiments, the numbered RAM (Random Access Memory) is a memory in which information is temporarily stored and is used as a work memory by the processor.

[0017] In the following embodiments, the numbered storage is one or more non-volatile storage devices that store various programs, various parameters, and the like. Examples of non-volatile storage devices include flash memory (SSD (Solid State Drive)), magnetic disks (e.g., hard disks), or magnetic tapes, and the like.

[0018] In the following embodiments, the numbered communication I / F (Interface) is an interface including a communication processor, an antenna, and the like. The communication I / F controls communication between multiple computers. Examples of communication standards applied to the communication I / F include wireless communication standards including 5G (5th Generation Mobile Communication System), Wi-Fi (registered trademark), or Bluetooth (registered trademark).

[0019] In the following embodiments, "A and / or B" is synonymous with "at least one of A and B." That is, "A and / or B" means that it may be A alone, or B alone, or a combination of A and B. Furthermore, in this specification, the same concept as "A and / or B" applies when expressing three or more things linked by "and / or."

[0020] [First Embodiment]

[0021] Figure 1 shows an example of the configuration of the data processing system 10 according to the first embodiment.

[0022] As shown in Figure 1, the data processing system 10 includes a data processing device 12 and a smart device 14. An example of the data processing device 12 is a server.

[0023] The data processing device 12 comprises a computer 22, a database 24, and a communication interface 26. The computer 22 is an example of a "computer" related to the technology of this disclosure. The computer 22 comprises a processor 28, RAM 30, and storage 32. The processor 28, RAM 30, and storage 32 are connected to a bus 34. The database 24 and the communication interface 26 are also connected to the bus 34. The communication interface 26 is connected to a network 54. An example of the network 54 is a WAN (Wide Area Network) and / or a LAN (Local Area Network).

[0024] The smart device 14 comprises a computer 36, a reception device 38, an output device 40, a camera 42, and a communication interface 44. The computer 36 comprises a processor 46, RAM 48, and storage 50. The processor 46, RAM 48, and storage 50 are connected to a bus 52. The reception device 38, output device 40, and camera 42 are also connected to the bus 52.

[0025] The reception device 38 is equipped with a touch panel 38A and a microphone 38B, etc., and receives user input. The touch panel 38A receives user input by detecting contact with an object (e.g., a pen or finger). The microphone 38B receives user input by detecting the user's voice. The control unit 46A transmits data indicating the user input received by the touch panel 38A and microphone 38B to the data processing device 12. In the data processing device 12, the specific processing unit 290 acquires the data indicating the user input.

[0026] The output device 40 includes a display 40A and a speaker 40B, and presents data to the user 20 by outputting the data in a form perceptible to the user 20 (e.g., audio and / or text). The display 40A displays visible information such as text and images according to instructions from the processor 46. The speaker 40B outputs audio according to instructions from the processor 46. The camera 42 is a small digital camera equipped with an optical system such as a lens, aperture, and shutter, and an image sensor such as a CMOS (Complementary Metal-Oxide-Semiconductor) image sensor or a CCD (Charge Coupled Device) image sensor.

[0027] Communication interface 44 is connected to network 54. Communication interfaces 44 and 26 are responsible for the exchange of various types of information between processor 46 and processor 28 via network 54.

[0028] Figure 2 shows an example of the main functions of the data processing device 12 and the smart device 14.

[0029] As shown in Figure 2, in the data processing device 12, a specific processing is performed by the processor 28. A specific processing program 56 is stored in the storage 32. The specific processing program 56 is an example of a "program" related to the technology of this disclosure. The processor 28 reads the specific processing program 56 from the storage 32 and executes the read specific processing program 56 on the RAM 30. The specific processing is realized by the processor 28 operating as a specific processing unit 290 according to the specific processing program 56 executed on the RAM 30.

[0030] The storage 32 stores the data generation model 58 and the emotion identification model 59. The data generation model 58 and the emotion identification model 59 are used by the identification processing unit 290.

[0031] In the smart device 14, the processor 46 performs the reception output processing. The storage 50 stores the reception output program 60. The reception output program 60 is used in conjunction with a specific processing program 56 by the data processing system 10. The processor 46 reads the reception output program 60 from the storage 50 and executes the read reception output program 60 on the RAM 48. The reception output processing is realized by the processor 46 operating as a control unit 46A according to the reception output program 60 executed on the RAM 48.

[0032] Next, the specific processing performed by the specific processing unit 290 of the data processing device 12 will be described. In the following description, the data processing device 12 will be referred to as the "server" and the smart device 14 as the "terminal".

[0033] This invention relates to a system for analyzing the movements of athletes in real time and providing feedback. This system consists of a combination of a camera, a server, and devices such as AR glasses or earphones.

[0034] First, the camera, which acts as the terminal, captures the athlete's movements and acquires video data. The acquired video data is transmitted to a server via the network. On the server, the received video data is analyzed using a pre-trained deep learning model. The deep learning model identifies the position of joints and key points of movement, and evaluates the athlete's form based on this.

[0035] Furthermore, the server retrieves the athlete's profile information (age, gender, past athletic history, etc.) from a database and generates customized feedback based on the analysis results. This feedback is formatted as AR data including visual instructions and audio data including voice instructions.

[0036] Next, the server sends the generated feedback data to the terminal device. The AR glasses, acting as the terminal, overlay the feedback content within the player's field of view, and the earphones play audio feedback in real time.

[0037] For example, if a user is practicing tennis, and the server detects that their swing follow-through is insufficient, it will display instructions through the AR glasses such as, "Raise your arms higher to improve your swing follow-through," and also communicate the same information audibly through earphones. In this way, the player can immediately correct their movements based on real-time feedback.

[0038] As described above, the present invention can provide a practical means for athletes to efficiently improve their skills through real-time motion analysis and feedback provision.

[0039] The following describes the processing flow.

[0040] Step 1:

[0041] The device activates a camera to capture the player's movements and acquires high-resolution video data. The camera's position and angle are appropriately adjusted to capture the player's overall form and subtle movements.

[0042] Step 2:

[0043] The terminal transmits the acquired video data to the server in real time. A high-speed network is used for transmission to minimize latency.

[0044] Step 3:

[0045] The server stores the video data received from the camera in a buffer and prepares it for analysis using a deep learning model.

[0046] Step 4:

[0047] The server inputs video data into a deep learning model to identify joint positions and movement trajectories. This allows it to extract characteristic points related to the athlete's movements.

[0048] Step 5:

[0049] The server evaluates the player's form based on the extracted feature points. This evaluation is performed using an algorithm that includes comparison with the ideal form.

[0050] Step 6:

[0051] The server retrieves player profile information from the database and determines the content of the feedback based on the analysis results. The feedback will be tailored to the player's characteristics and goals.

[0052] Step 7:

[0053] The server converts the generated feedback into AR data as visual instructions and into audio data as audio instructions.

[0054] Step 8:

[0055] The server sends this data to the device. Visual data is sent to the AR glasses, and audio data is sent to the earphones.

[0056] Step 9:

[0057] The AR glasses on the device overlay the received visual data onto the player's field of view. This allows the player to see in real time which parts need to be corrected.

[0058] Step 10:

[0059] The device's earphones play back the received audio data, providing players with specific instructions for improvement via voice.

[0060] Step 11:

[0061] Users (players) modify their actions based on real-time visual and audio feedback. This process enables immediate correction and skill improvement.

[0062] (Example 1)

[0063] Next, we will describe Example 1. In the following description, the data processing device 12 will be referred to as the "server," and the smart device 14 will be referred to as the "terminal."

[0064] For athletes to efficiently improve their skills, real-time and individually customized feedback is necessary. However, conventional analysis systems have struggled to provide sufficient real-time and individualized support, making it difficult to immediately reflect improvement results.

[0065] The identification process performed by the identification processing unit 290 of the data processing device 12 in Example 1 is realized by the following means.

[0066] In this invention, the server includes means for receiving video, analyzing it using a generating AI model, and evaluating a person's movements; means for generating customized feedback based on the evaluation and the person's personal information; and means for presenting the generated feedback as visual and auditory information. This allows athletes to receive personalized feedback in real time through sight and hearing, enabling them to immediately correct their movements to improve their skills.

[0067] A "device for acquiring video footage" refers to a device that has the function of filming the movements of athletes in real time and collecting them as video data.

[0068] A "generative AI model" refers to artificial intelligence technology used to analyze human movement by identifying the position of joints and key points of movement from input video data.

[0069] A "processing device" refers to a computer system that performs a series of data processing operations, including receiving video data, analyzing it, and generating feedback.

[0070] "Personal information of individuals" refers to individual data such as the age, gender, and athletic history of athletes, and is information used for analysis and feedback generation.

[0071] "Customized feedback" refers to information that provides visual and auditory instructions optimized for an individual, based on analysis results and personal information.

[0072] "A device that presents information as visual and auditory information" refers to a device that can visually display and audibly reproduce the feedback provided.

[0073] This invention relates to a system that analyzes the movements of athletes in real time and provides feedback. The system is broadly composed of a device for acquiring video, a processing device, and a device for presenting feedback.

[0074] First, the video acquisition device, acting as a terminal, uses a high-resolution camera to capture the players' movements in real time and collects video data frame by frame. This acquired video data is then transmitted to a server using wireless communication technology.

[0075] Next, the server uses a generated AI model based on the received video data to analyze the player's movements in detail. This analysis employs deep learning techniques through multiple neural network layers to identify joint positions and key movement points. Subsequently, the server retrieves the player's personal information from a database and combines it with the analysis results to generate customized feedback. The feedback is converted into AR data as visual instructions and formatted as audio data as voice instructions.

[0076] Finally, the generated feedback data is sent to the AR glasses and earphones, which are the devices used. The AR glasses overlay real-time visual feedback within the player's field of view, and the earphones play audio feedback. This allows the player to receive immediate feedback and correct their actions.

[0077] For example, when a user is practicing tennis, the server analyzes and detects that their swing follow-through is insufficient. In this case, the AR glasses display a visual instruction such as, "Raise your arm higher to improve your swing follow-through," and the same information is conveyed as audio through the earphones. In this way, players can improve their technique based on real-time feedback.

[0078] An example of a prompt message is, "Analyze the tennis swing motion and generate real-time feedback for form improvement." Following this example, the system assists in improving individual movements in sports.

[0079] The flow of the specific processing in Example 1 will be explained using Figure 11.

[0080] Step 1:

[0081] The device (camera) captures video data by filming the movements of athletes in real time at high resolution. The input is a scene of the athlete's movements, and the output is a continuous generation of video frames. In this process, the camera meticulously records the athlete's movements and captures the video at a high frame rate to capture the details of the movements.

[0082] Step 2:

[0083] The terminal reduces the data size of the acquired video data using compression technology and transmits it to the server via a wireless network. Raw video data is the input, and compressed video data is transmitted as the output. Specifically, compression algorithms such as H.264 are used to enable high-speed and efficient data transfer.

[0084] Step 3:

[0085] The server analyzes the received video data using a generative AI model based on deep learning to identify key points in the athlete's movements. Compressed video data is the input, and the output is the analyzed joint positions and movement characteristics of the body. The server extracts important body metrics for each frame through a multi-layered neural network.

[0086] Step 4:

[0087] The server combines analysis results with the athlete's personal profile information (age, gender, exercise history, etc.) to generate customized feedback. Inputs include analysis data and profile information, while outputs include AR data for visual feedback and audio feedback data. Based on the analysis, the server automatically generates individually optimized advice.

[0088] Step 5:

[0089] The server transmits the generated feedback data to terminal devices (AR glasses, earphones) and presents it to the players in real time. The input is feedback data, and the output is information that directly engages the players' sight and hearing. The AR glasses provide visual instructions, and the earphones provide audio guidance, enabling players to immediately work on improving their skills.

[0090] Step 6:

[0091] The user (athlete) modifies their actions based on the feedback provided, improving their practice and performance. Input comes from AR glasses and earphones, and output is optimized sports performance. Based on the feedback received in real time, athletes instantly adjust their skills and progress in their learning.

[0092] (Application Example 1)

[0093] Next, we will explain Application Example 1. In the following explanation, the data processing device 12 will be referred to as the "server," and the smart device 14 will be referred to as the "terminal."

[0094] There is a lack of methods to provide appropriate exercise and feedback tailored to the individual condition and abilities of elderly people and patients undergoing rehabilitation. Traditional rehabilitation methods often lack individualized support, and patients may not receive appropriate guidance. Therefore, a system capable of real-time movement analysis and guidance is needed to achieve effective and rapid rehabilitation.

[0095] The specific processing performed by the specific processing unit 290 of the data processing device 12 in Application Example 1 is realized by the following means.

[0096] In this invention, the server includes a device for acquiring video information, a computing device for evaluating the subject's movements, and a computing device for generating feedback information. This enables individual motion analysis and real-time feedback.

[0097] "Video information" refers to visual data acquired to record the actions of a subject.

[0098] A "device" is hardware used to capture and display video information.

[0099] A "processing unit" is a computing device that analyzes received data and generates evaluations and feedback.

[0100] "Information" refers to data used for performance evaluation and feedback, and is represented visually and audibly.

[0101] "Subjects" are individuals who undergo performance evaluation using the system.

[0102] "Feedback" refers to evaluation results that include guidance for improving the subject's actions.

[0103] "Attribute information" refers to the individual information of the target person and is data used for personalized feedback.

[0104] This system promotes efficient exercise by analyzing the movements of elderly or rehabilitation patients in real time and providing immediate feedback on areas for improvement. The device consists of smart glasses with a built-in camera and earphones. These work together to capture movements and provide feedback through both visual and auditory means.

[0105] The server receives video information via the network. A deep learning model is used to evaluate the subject's movements based on the video information. Python and TENSORFLOW® are used to train the model, identifying joint positions and key points of movement. The evaluation results are used to retrieve the subject's attribute information from a database and generate customized feedback based on this information. The generated information is overlaid as visual data on smart glasses and played back as audio data through earphones.

[0106] For example, when performing arm-raising exercises as part of rehabilitation, if there is a problem with the patient's movement, the smart glasses display will say, "It would be better to raise your arm a little higher," and the same instruction will be conveyed via voice through the earphones. This allows the user to immediately correct their movement and improve the effectiveness of the rehabilitation.

[0107] The following example prompt can be used in the generated AI model: "Analyze the rehabilitation movements of a young male and generate visual and auditory feedback if the arm lift is inappropriate." Using this prompt, it is possible to create training data for specific movement analysis and improve the accuracy of the model.

[0108] The flow of a specific process in Application Example 1 will be explained using Figure 12.

[0109] Step 1:

[0110] The device (smart glasses) uses a camera to capture the subject's movements. In this step, the input is the subject's real-time movements, and the output is raw video footage of those movements. The device transmits this raw video to a server via the network.

[0111] Step 2:

[0112] The server processes video information received via the network. The input is the raw video obtained in step 1, and the output is a dataset in a format suitable for input to a deep learning model. The server preprocesses this data to extract joint positions and key points of movement.

[0113] Step 3:

[0114] The server uses a deep learning model to perform motion analysis. The input is a pre-processed dataset, and the output is an evaluation of the subject's movements. Through this evaluation, the server identifies the appropriateness of the movements and areas that need correction.

[0115] Step 4:

[0116] The server retrieves the subject's attribute information (age, weight, past rehabilitation records, etc.) from the database. Based on this input information, it generates customized feedback. The output is feedback information optimized for the subject.

[0117] Step 5:

[0118] The server formats the generated feedback information into audio and visual data and sends it to the terminal. The input feedback information is converted into AR display data and speech synthesis data, and then prepared into a data format that can be presented to the user as output.

[0119] Step 6:

[0120] The device (smart glasses and earphones) provides feedback to the user using data received from the server. Visual data is overlaid on the smart glasses, and audio data is played through the earphones. The input in this step is formatted feedback data, and the output is real-time guidance for the individual. This allows the user to immediately correct their actions and improve the effectiveness of their rehabilitation.

[0121] Furthermore, an emotion engine that estimates the user's emotions may be incorporated. That is, the identification processing unit 290 may use the emotion identification model 59 to estimate the user's emotions and perform identification processing using the user's emotions.

[0122] This invention relates to a system for analyzing the movements of athletes in real time and providing feedback that takes emotional information into account. This system combines a camera for acquiring video data, a server for analyzing the video and audio, AR glasses or earphones worn by the athlete, and an emotion engine for recognizing emotions.

[0123] First, the device's camera captures the player's movements and sends the video data to a server. Simultaneously, the player's voice is captured via a microphone and input into an emotion engine. This emotion engine has the ability to analyze the player's emotional state from their facial expressions and voice tone.

[0124] The server uses deep learning to analyze the received video data and identify key points and forms of the athlete's movements. The analysis results are compared to ideal forms, and areas for improvement in the movements are evaluated.

[0125] Meanwhile, the emotion engine analyzes the player's emotional data in real time to determine, for example, whether the player is feeling fatigued or frustrated. The server adjusts the feedback based on this emotional data and generates advice tailored to the player's mental state. For example, if a player shows signs of fatigue, the server avoids giving instructions for strenuous movements and provides feedback encouraging rest.

[0126] The generated feedback is formatted as AR data for visual instructions and as audio data for audio instructions, and then sent to the device. The device's AR glasses overlay the feedback onto the athlete's field of view, and the earphones play the audio feedback in real time. This allows athletes to not only instantly correct their form but also receive advice that takes their physical and mental state into consideration, enabling more effective training.

[0127] For example, consider a situation where a user (athlete) is practicing track and field. The server analyzes the athlete's running form and evaluates knee angle and stride length. Furthermore, if the emotion engine determines from the athlete's facial expression that their concentration is declining, it will provide feedback such as "Take a short break and drink some water," and then send specific advice such as "Try running with an awareness of widening your stride." In this way, feedback is provided while taking the athlete's emotional state into consideration, significantly improving the quality of training.

[0128] The following describes the processing flow.

[0129] Step 1:

[0130] The device's camera captures the player's movements and acquires high-resolution video data. This data is set to be captured at the optimal angle to clearly capture the player's posture and movements.

[0131] Step 2:

[0132] The terminal sends the acquired video data to the server. This transmission is performed using a high-speed communication protocol to minimize latency.

[0133] Step 3:

[0134] The server stores the received video data in a buffer and prepares to perform motion analysis.

[0135] Step 4:

[0136] The server inputs video data into a deep learning model to analyze key movement points such as joint positions. This quantifies the athlete's form, making it available for evaluation.

[0137] Step 5:

[0138] Based on the analyzed data, the server compares the player's form to an ideal form and identifies areas for improvement.

[0139] Step 6:

[0140] The device's microphone captures the player's voice and sends it to the emotion engine in real time. This audio data is used to evaluate the player's emotional state.

[0141] Step 7:

[0142] The emotion engine analyzes facial expression information obtained from audio and video data to determine the emotional state of the player. This determination serves as an indicator for evaluating the player's concentration, fatigue, frustration, and other factors.

[0143] Step 8:

[0144] The server generates feedback tailored to the player's emotional state based on data from the emotion engine. This feedback takes into account the player's mental state and includes content designed to maintain motivation.

[0145] Step 9:

[0146] The server formats the generated feedback into AR display data and audio data and sends it to the device.

[0147] Step 10:

[0148] The AR glasses on the device overlay the received visual data onto the player's field of view, visually indicating specific areas for improvement.

[0149] Step 11:

[0150] The earphones in the device transmit received audio data to the players in real time, giving them voice instructions for the necessary actions.

[0151] Step 12:

[0152] Based on the feedback provided, users (athletes) can immediately correct their form and receive emotionally responsive advice, enabling efficient and mentally conscious training.

[0153] (Example 2)

[0154] Next, we will describe Example 2. In the following description, the data processing device 12 will be referred to as the "server" and the smart device 14 as the "terminal".

[0155] Traditional sports training systems separate the analysis of athletes' movements from the assessment of their emotional state, making it difficult to provide integrated feedback. Furthermore, the lack of means to provide real-time feedback tailored to the athlete's real-time mental state made it challenging to maximize performance improvement.

[0156] The identification process performed by the identification processing unit 290 of the data processing device 12 in Example 2 is realized by the following means.

[0157] In this invention, the server includes means for analyzing video data to identify key points and forms of movement, means for analyzing emotional states based on audio data, and means for generating feedback based on movement and emotional states. This makes it possible to provide feedback that takes into account both the player's movement and mental state.

[0158] "Video data" refers to visual information acquired using cameras or other recording devices, and this information is used for motion analysis.

[0159] "Audio data" refers to auditory information acquired using sound-collecting devices such as microphones, and this information is used for emotion analysis.

[0160] "Terminal means" refers to an electronic device used to acquire video and audio data and transmit it to a server.

[0161] "Server equipment" refers to computer equipment used to analyze received data and generate feedback based on actions and emotional states.

[0162] "Deep learning" refers to a technology in which artificial intelligence uses large amounts of data to learn patterns and perform predictions and analyses.

[0163] An "emotion analysis engine" refers to a collection of algorithms and software used to estimate a person's emotional state based on voice data and other input information.

[0164] "Feedback" refers to suggestions for improvement and guidance that are generated by a server and presented to the recipient visually or audibly.

[0165] This invention is a system for analyzing the movements and emotions of athletes in real time and providing effective feedback. The system consists of a terminal equipped with input devices such as a camera for acquiring video data and a microphone for acquiring audio data. The terminal temporarily stores this data and then transmits it to a server.

[0166] The server runs a deep learning model to analyze the video data. This model is built using a common deep learning framework and extracts key points from the athlete's body movements, identifying their form. The analyzed motion data is then compared to the ideal form to identify specific areas for improvement.

[0167] Furthermore, the server is equipped with an emotion analysis engine for performing sentiment analysis. This engine determines the emotional state of the players based on the audio data. The algorithms used here include natural language processing and speech emotion analysis algorithms.

[0168] The generated feedback will take into account both the athlete's actions and emotions. This feedback is then transmitted to the device in real time and communicated to the athlete. Specifically, it is displayed as a visual overlay via the device's AR glasses or provided as audio instruction through earphones.

[0169] For example, let's assume a user is practicing track and field. The server analyzes the video footage taken during the run and determines that the athlete should improve their knee angle. Simultaneously, emotional analysis is performed, and if the emotional data detects that the athlete is feeling fatigued, it recommends "taking a short rest and drinking some water," and then provides specific advice such as "focus on your form and lift your knees higher." In this way, the feedback is designed to maximize its effectiveness in both the athlete's athletic performance and mental state.

[0170] An example of a prompt to input into a generative AI model is: "Please describe a system that analyzes the movements and emotions of athletes and generates feedback. Explain how it analyzes movements in real time and provides advice based on the athlete's emotions." This prompt serves as a design guideline for improving the quality of the generated feedback.

[0171] The flow of the specific processing in Example 2 will be explained using Figure 13.

[0172] Step 1:

[0173] The device uses a camera to acquire video data of athletes and a microphone to capture audio data. This data is used to record both the athletes' movements and voice in real time. The input is the athletes' video and audio, and the output is sent to the server as data packets. Specifically, the device records video at 60 frames per second and converts the continuous audio into a digital format.

[0174] Step 2:

[0175] The server receives data packets sent from the terminal. Using a deep learning model, it identifies key points and form in the athlete's movements based on the received video data. The input is video data, and the output is the motion analysis result. Specifically, calculations are performed to quantify the position and angle of the athlete's joints and compare them to the ideal form.

[0176] Step 3:

[0177] The server simultaneously feeds the received audio data into an emotion analysis engine to estimate the player's emotional state. The input is audio data, and the output is the result of the emotional state analysis. This analysis takes into account factors such as the tone and tempo of the voice and the frequency of interruptions. Specifically, the system performs speech analysis using natural language processing technology.

[0178] Step 4:

[0179] The server integrates both the motion analysis results and the emotion analysis results to generate feedback for the player. The input is the analysis results from Step 2 and Step 3, and the output is feedback including areas for improvement and advice. Here, a generative AI model is used to automatically generate the optimal feedback content.

[0180] Step 5:

[0181] The server sends the generated feedback to the device. The device provides feedback to the user visually through the AR glasses' display and audibly through the earphones. The input is the feedback data sent from the server, and the output is the visual and audible feedback to the user. Specifically, the system overlays areas for form improvement on the AR glasses and plays detailed instructions audibly through the earphones.

[0182] (Application Example 2)

[0183] Next, we will explain application example 2. In the following explanation, the data processing device 12 will be referred to as a "server" and the smart device 14 as a "terminal".

[0184] In modern manufacturing, managing the work efficiency and health of factory workers is a critical challenge. In particular, there is a need to improve workers' operational efficiency while simultaneously evaluating emotional factors such as stress and fatigue in real time and providing accurate feedback based on these assessments. Conventional technologies struggle to integrate operational improvement and emotional management, posing a risk of decreased work efficiency.

[0185] The specific processing performed by the specific processing unit 290 of the data processing device 12 in Application Example 2 is realized by the following means.

[0186] In this invention, the server includes means for receiving and analyzing video information, means for acquiring audio information and analyzing emotional states, and means for generating feedback information based on performance evaluation and emotional states. This makes it possible to provide factory workers with real-time instructions for correcting their movements and suggestions for appropriate rest and work paces.

[0187] "Visual information" refers to visual data acquired from devices such as cameras, and is used to analyze the movements and postures of subjects.

[0188] "Device means" refers to hardware or software components used to acquire, process, or present data for a specific function or purpose.

[0189] "Processing device" refers to a computer or server used to analyze acquired data and perform actions or recognize emotions.

[0190] "Audio information" refers to audio data collected through devices such as microphones, and is used to analyze the emotional state of a subject based on their speech and tone of voice.

[0191] "Emotional analysis device means" refers to a software or hardware system for estimating and analyzing a subject's emotional state from audio information or visual data.

[0192] "Feedback information" refers to instructions and advice generated based on performance evaluations and emotional states, with the aim of improving the subject's performance and maintaining a comfortable work environment.

[0193] This system is used in factories to manage and improve the work efficiency and health of workers. The system acquires data through hardware devices such as smart glasses, Bluetooth earphones, cameras, and microphones. A server processes this data on a cloud platform. Specifically, the server is built on cloud infrastructure such as AWS® or Google® Cloud.

[0194] First, the smart glasses worn by the user capture the worker's movements in real time using their camera. The video information is sent to a server, where motion analysis is performed using software such as TensorFlow and PiTouch. The server uses deep learning technology to extract key points of the movements and compares them to ideal movement forms.

[0195] Meanwhile, voice information is also acquired through the microphone and sent to the server. The server uses an emotion analysis engine, such as Azure® Emotion API, to evaluate emotional states such as stress and fatigue from the voice information. This evaluation result influences feedback in real time.

[0196] The server integrates motion analysis and emotion assessment to generate optimal feedback information for factory workers. This feedback information is visually overlaid on smart glasses, and voice guidance is provided through earphones.

[0197] For example, if an incorrect posture is detected during assembly work, instructions such as "Raise your elbows slightly and apply force safely" will be displayed in real time on the smart glasses. Also, if the emotion analysis engine detects high stress levels, it will suggest "Take a 5-minute break" through the earphones.

[0198] This system optimizes worker productivity and resource utilization, protecting them from excessive stress and fatigue. The overall effect of this implementation makes the factory environment a safer and more efficient place.

[0199] An example of a prompt message would be, "Please tell us about any moments during your work today when you found the feedback provided by your smart glasses particularly helpful."

[0200] The flow of a specific process in Application Example 2 will be explained using Figure 14.

[0201] Step 1:

[0202] The device acquires video information. The camera on the smart glasses worn by the user captures the movements of factory workers in real time and generates video information based on that. The input is video frames captured by the camera, and the output is a data stream of video information.

[0203] Step 2:

[0204] The device sends the acquired video information to the server. The device uploads the captured video information to a cloud-based server in real time. The input is the video information generated in step 1, and the output is the data stream sent to the server.

[0205] Step 3:

[0206] The server analyzes the video information. The server uses TensorFlow and PITach to analyze the video information and identify key points of the worker's movements. The input is the video information sent to the server, and the output is key point data of the movements. Specific movements evaluated include posture, joint angles, and stride length.

[0207] Step 4:

[0208] The terminal acquires audio information. The microphone built into the earphone worn by the user captures the worker's voice. The input is the audio signal picked up by the microphone, and the output is a data stream of the audio information.

[0209] Step 5:

[0210] The device sends the acquired audio information to the server. The device sends the captured audio information to a cloud-based server. The input is the audio information generated in step 4, and the output is the data stream sent to the server.

[0211] Step 6:

[0212] The server analyzes the audio information. The server uses the Azure Emotion API and other tools to analyze the audio information and identify the worker's emotional state. The input is the audio information sent to the server, and the output is emotional state data. Specific values ​​evaluated include stress levels and fatigue levels.

[0213] Step 7:

[0214] The server generates feedback information. The server integrates action keypoint data and emotional state data to generate optimal feedback information for factory workers. The input is the output data from steps 3 and 6, and the output is the feedback information. Specific actions include action modification instructions and break suggestions that take safety and efficiency into consideration.

[0215] Step 8:

[0216] The device presents feedback information to the user. The device visually overlays the generated feedback information on the smart glasses and communicates it audibly through the earphones. The input is the feedback information generated in step 7, and the output is the visual and audible feedback provided to the user.

[0217] The specific processing unit 290 transmits the result of the specific processing to the smart device 14. In the smart device 14, the control unit 46A causes the output device 40 to output the result of the specific processing. The microphone 38B acquires audio indicating user input for the result of the specific processing. The control unit 46A transmits the audio data indicating user input acquired by the microphone 38B to the data processing device 12. In the data processing device 12, the specific processing unit 290 acquires the audio data.

[0218] Data generation model 58 is a so-called generative AI (Artificial Intelligence). An example of data generation model 58 is ChatGPT (registered trademark) (Internet search).<URL: https: / / openai.com / blog / chatgpt> ), Gemini (registered trademark) (Internet search) <url: https: gemini.google.com ?hl="ja">Examples of generative AI include the following. The data generation model 58 is obtained by performing deep learning on a neural network. The data generation model 58 is input with prompts containing instructions, and with inference data such as audio data representing speech, text data representing text, and image data representing images. The data generation model 58 infers from the input inference data according to the instructions indicated by the prompts, and outputs the inference results in data formats such as audio data and text data. Here, inference refers to, for example, analysis, classification, prediction, and / or summarization.

[0219] In the above embodiment, an example was given in which specific processing is performed by the data processing device 12, but the technology of this disclosure is not limited thereto, and the specific processing may also be performed by the smart device 14.

[0220] [Second Embodiment]

[0221] Figure 3 shows an example of the configuration of the data processing system 210 according to the second embodiment.

[0222] As shown in Figure 3, the data processing system 210 includes a data processing device 12 and smart glasses 214. An example of the data processing device 12 is a server.

[0223] The data processing device 12 comprises a computer 22, a database 24, and a communication interface 26. The computer 22 is an example of a "computer" related to the technology of this disclosure. The computer 22 comprises a processor 28, RAM 30, and storage 32. The processor 28, RAM 30, and storage 32 are connected to a bus 34. The database 24 and the communication interface 26 are also connected to the bus 34. The communication interface 26 is connected to a network 54. An example of the network 54 is a WAN (Wide Area Network) and / or a LAN (Local Area Network).

[0224] The smart glasses 214 include a computer 36, a microphone 238, a speaker 240, a camera 42, and a communication interface 44. The computer 36 includes a processor 46, RAM 48, and storage 50. The processor 46, RAM 48, and storage 50 are connected to a bus 52. The microphone 238, speaker 240, and camera 42 are also connected to the bus 52.

[0225] The microphone 238 receives voice signals from the user 20 and receives instructions from the user 20. The microphone 238 captures the voice signals from the user 20, converts the captured voice into audio data, and outputs it to the processor 46. The speaker 240 outputs audio according to the instructions from the processor 46.

[0226] Camera 42 is a small digital camera equipped with an optical system including a lens, aperture, and shutter, and an image sensor such as a CMOS (Complementary Metal-Oxide-Semiconductor) image sensor or a CCD (Charge Coupled Device) image sensor, and captures images of the area around the user 20 (for example, an imaging range defined by a field of view equivalent to the width of a typical healthy person's field of vision).

[0227] Communication interface 44 is connected to network 54. Communication interfaces 44 and 26 are responsible for the exchange of various information between processor 46 and processor 28 via network 54. The exchange of various information between processor 46 and processor 28 using communication interfaces 44 and 26 is performed in a secure manner.

[0228] Figure 4 shows an example of the main functions of the data processing device 12 and the smart glasses 214. As shown in Figure 4, the data processing device 12 performs specific processing using the processor 28. The storage 32 stores the specific processing program 56.

[0229] The specific processing program 56 is an example of a "program" relating to the technology of this disclosure. The processor 28 reads the specific processing program 56 from the storage 32 and executes the read specific processing program 56 on the RAM 30. The specific processing is realized by the processor 28 operating as a specific processing unit 290 in accordance with the specific processing program 56 executed on the RAM 30.

[0230] The storage 32 stores the data generation model 58 and the emotion identification model 59. The data generation model 58 and the emotion identification model 59 are used by the identification processing unit 290.

[0231] In the smart glasses 214, the processor 46 performs the reception output processing. The storage 50 stores the reception output program 60. The processor 46 reads the reception output program 60 from the storage 50 and executes the read reception output program 60 on the RAM 48. The reception output processing is realized by the processor 46 operating as a control unit 46A according to the reception output program 60 executed on the RAM 48.

[0232] Next, the identification processing performed by the identification processing unit 290 of the data processing device 12 will be described. In the following description, the data processing device 12 will be referred to as the "server" and the smart glasses 214 will be referred to as the "terminal".

[0233] This invention relates to a system for analyzing the movements of athletes in real time and providing feedback. This system consists of a combination of a camera, a server, and devices such as AR glasses or earphones.

[0234] First, the camera, which acts as the terminal, captures the athlete's movements and acquires video data. The acquired video data is transmitted to a server via the network. On the server, the received video data is analyzed using a pre-trained deep learning model. The deep learning model identifies the position of joints and key points of movement, and evaluates the athlete's form based on this.

[0235] Furthermore, the server retrieves the athlete's profile information (age, gender, past athletic history, etc.) from a database and generates customized feedback based on the analysis results. This feedback is formatted as AR data including visual instructions and audio data including voice instructions.

[0236] Next, the server sends the generated feedback data to the terminal device. The AR glasses, acting as the terminal, overlay the feedback content within the player's field of view, and the earphones play audio feedback in real time.

[0237] For example, if a user is practicing tennis, and the server detects that their swing follow-through is insufficient, it will display instructions through the AR glasses such as, "Raise your arms higher to improve your swing follow-through," and also communicate the same information audibly through earphones. In this way, the player can immediately correct their movements based on real-time feedback.

[0238] As described above, the present invention can provide a practical means for athletes to efficiently improve their skills through real-time motion analysis and feedback provision.

[0239] The following describes the processing flow.

[0240] Step 1:

[0241] The device activates a camera to capture the player's movements and acquires high-resolution video data. The camera's position and angle are appropriately adjusted to capture the player's overall form and subtle movements.

[0242] Step 2:

[0243] The terminal transmits the acquired video data to the server in real time. A high-speed network is used for transmission to minimize latency.

[0244] Step 3:

[0245] The server stores the video data received from the camera in a buffer and prepares it for analysis using a deep learning model.

[0246] Step 4:

[0247] The server inputs video data into a deep learning model to identify joint positions and movement trajectories. This allows it to extract characteristic points related to the athlete's movements.

[0248] Step 5:

[0249] The server evaluates the player's form based on the extracted feature points. This evaluation is performed using an algorithm that includes comparison with the ideal form.

[0250] Step 6:

[0251] The server retrieves player profile information from the database and determines the content of the feedback based on the analysis results. The feedback will be tailored to the player's characteristics and goals.

[0252] Step 7:

[0253] The server converts the generated feedback into AR data as visual instructions and into audio data as audio instructions.

[0254] Step 8:

[0255] The server sends this data to the device. Visual data is sent to the AR glasses, and audio data is sent to the earphones.

[0256] Step 9:

[0257] The AR glasses on the device overlay the received visual data onto the player's field of view. This allows the player to see in real time which parts need to be corrected.

[0258] Step 10:

[0259] The device's earphones play back the received audio data, providing players with specific instructions for improvement via voice.

[0260] Step 11:

[0261] Users (players) modify their actions based on real-time visual and audio feedback. This process enables immediate correction and skill improvement.

[0262] (Example 1)

[0263] Next, we will describe Example 1. In the following description, the data processing device 12 will be referred to as the "server," and the smart glasses 214 will be referred to as the "terminal."

[0264] For athletes to efficiently improve their skills, real-time and individually customized feedback is necessary. However, conventional analysis systems have struggled to provide sufficient real-time and individualized support, making it difficult to immediately reflect improvement results.

[0265] The identification process performed by the identification processing unit 290 of the data processing device 12 in Example 1 is realized by the following means.

[0266] In this invention, the server includes means for receiving video, analyzing it using a generating AI model, and evaluating a person's movements; means for generating customized feedback based on the evaluation and the person's personal information; and means for presenting the generated feedback as visual and auditory information. This allows athletes to receive personalized feedback in real time through sight and hearing, enabling them to immediately correct their movements to improve their skills.

[0267] A "device for acquiring video footage" refers to a device that has the function of filming the movements of athletes in real time and collecting them as video data.

[0268] A "generative AI model" refers to artificial intelligence technology used to analyze human movement by identifying the position of joints and key points of movement from input video data.

[0269] A "processing device" refers to a computer system that performs a series of data processing operations, including receiving video data, analyzing it, and generating feedback.

[0270] "Personal information of individuals" refers to individual data such as the age, gender, and athletic history of athletes, and is information used for analysis and feedback generation.

[0271] "Customized feedback" refers to information that provides visual and auditory instructions optimized for an individual, based on analysis results and personal information.

[0272] "A device that presents information as visual and auditory information" refers to a device that can visually display and audibly reproduce the feedback provided.

[0273] This invention relates to a system that analyzes the movements of athletes in real time and provides feedback. The system is broadly composed of a device for acquiring video, a processing device, and a device for presenting feedback.

[0274] First, the video acquisition device, acting as a terminal, uses a high-resolution camera to capture the players' movements in real time and collects video data frame by frame. This acquired video data is then transmitted to a server using wireless communication technology.

[0275] Next, the server uses a generated AI model based on the received video data to analyze the player's movements in detail. This analysis employs deep learning techniques through multiple neural network layers to identify joint positions and key movement points. Subsequently, the server retrieves the player's personal information from a database and combines it with the analysis results to generate customized feedback. The feedback is converted into AR data as visual instructions and formatted as audio data as voice instructions.

[0276] Finally, the generated feedback data is sent to the AR glasses and earphones, which are the devices used. The AR glasses overlay real-time visual feedback within the player's field of view, and the earphones play audio feedback. This allows the player to receive immediate feedback and correct their actions.

[0277] For example, when a user is practicing tennis, the server analyzes and detects that their swing follow-through is insufficient. In this case, the AR glasses display a visual instruction such as, "Raise your arm higher to improve your swing follow-through," and the same information is conveyed as audio through the earphones. In this way, players can improve their technique based on real-time feedback.

[0278] An example of a prompt message is, "Analyze the tennis swing motion and generate real-time feedback for form improvement." Following this example, the system assists in improving individual movements in sports.

[0279] The flow of the specific processing in Example 1 will be explained using Figure 11.

[0280] Step 1:

[0281] The terminal (camera) captures the actions of the sports player in real time with high resolution to obtain video data. As input, the action scene of the player is given, and as output, video frames are continuously generated. At this time, the camera records the player's movements in detail and captures the video at a high frame rate to capture the details of the actions.

[0282] Step 2:

[0283] The terminal reduces the data size of the acquired video data using compression technology and transmits it to the server through a wireless network. As input, there is raw video data, and as output, compressed video data is transmitted. As a specific example, a compression algorithm such as H.264 is used, and data transfer is performed quickly and efficiently.

[0284] Step 3:

[0285] The server analyzes the received video data using a generative AI model based on deep learning to identify the key points of the player's actions. As input, there is compressed video data, and as output, the analyzed joint positions and movement characteristics of the body are obtained. The server extracts important body metrics for each frame through a multi-layer neural network.

[0286] Step 4:

[0287] The server combines the analysis results with the player's personal profile information (age, gender, sports history, etc.) to generate customized feedback. As input, there is analysis data and profile information, and as output, AR data for visual feedback and audio feedback data are generated. The server automatically generates individually optimized advice based on the analysis.

[0288] Step 5:

[0289] The server transmits the generated feedback data to terminal devices (AR glasses, earphones) and presents it to the players in real time. The input is feedback data, and the output is information that directly engages the players' sight and hearing. The AR glasses provide visual instructions, and the earphones provide audio guidance, enabling players to immediately work on improving their skills.

[0290] Step 6:

[0291] The user (athlete) modifies their actions based on the feedback provided, improving their practice and performance. Input comes from AR glasses and earphones, and output is optimized sports performance. Based on the feedback received in real time, athletes instantly adjust their skills and progress in their learning.

[0292] (Application Example 1)

[0293] Next, we will explain Application Example 1. In the following explanation, the data processing device 12 will be referred to as the "server," and the smart glasses 214 will be referred to as the "terminal."

[0294] There is a lack of methods to provide appropriate exercise and feedback tailored to the individual condition and abilities of elderly people and patients undergoing rehabilitation. Traditional rehabilitation methods often lack individualized support, and patients may not receive appropriate guidance. Therefore, a system capable of real-time movement analysis and guidance is needed to achieve effective and rapid rehabilitation.

[0295] The specific processing performed by the specific processing unit 290 of the data processing device 12 in Application Example 1 is realized by the following means.

[0296] In this invention, the server includes a device for acquiring video information, a computing device for evaluating the subject's movements, and a computing device for generating feedback information. This enables individual motion analysis and real-time feedback.

[0297] "Video information" refers to visual data acquired to record the actions of a subject.

[0298] A "device" is hardware used to capture and display video information.

[0299] A "processing unit" is a computing device that analyzes received data and generates evaluations and feedback.

[0300] "Information" refers to data used for performance evaluation and feedback, and is represented visually and audibly.

[0301] "Subjects" are individuals who undergo performance evaluation using the system.

[0302] "Feedback" refers to evaluation results that include guidance for improving the subject's actions.

[0303] "Attribute information" refers to the individual information of the target person and is data used for personalized feedback.

[0304] This system promotes efficient exercise by analyzing the movements of elderly or rehabilitation patients in real time and providing immediate feedback on areas for improvement. The device consists of smart glasses with a built-in camera and earphones. These work together to capture movements and provide feedback through both visual and auditory means.

[0305] The server receives video information via a network. Using a deep learning model, it evaluates the actions of the target person based on the video information. Python and TensorFlow are used for model training, and the joint positions and key points of the actions are identified. The evaluation results are obtained by retrieving the target person's attribute information from the database, and customized feedback is generated based on this. The generated information is overlaid and displayed on smart glasses as visual data and played back from earphones as audio data.

[0306] As a specific example, when performing the exercise of raising the arm for rehabilitation, if there is a problem with the patient's movement, the display of the smart glasses shows "It would be better to raise the arm a little more", and a similar instruction is conveyed audibly through the earphones. This enables the user to immediately correct their movement and improve the rehabilitation effect.

[0307] The following example of a prompt sentence can be used for the generated AI model. "Analyze the rehabilitation movements of a young man and generate feedback visually and audibly when the way of raising the arm is inappropriate." By using this prompt, it is possible to create training data for specific movement analysis and improve the accuracy of the model.

[0308] The flow of the specific process in Application Example 1 will be described using FIG. 12.

[0309] Step 1:

[0310] The terminal (smart glasses) uses a camera to capture the actions of the target person. The input in this step is the real-time movement performed by the target person, and the output obtained is the raw video recording that movement. The terminal transmits this raw video to the server via the network.

[0311] Step 2:

[0312] The server processes video information received via the network. The input is the raw video obtained in step 1, and the output is a dataset in a format suitable for input to a deep learning model. The server preprocesses this data to extract joint positions and key points of movement.

[0313] Step 3:

[0314] The server uses a deep learning model to perform motion analysis. The input is a pre-processed dataset, and the output is an evaluation of the subject's movements. Through this evaluation, the server identifies the appropriateness of the movements and areas that need correction.

[0315] Step 4:

[0316] The server retrieves the subject's attribute information (age, weight, past rehabilitation records, etc.) from the database. Based on this input information, it generates customized feedback. The output is feedback information optimized for the subject.

[0317] Step 5:

[0318] The server formats the generated feedback information into audio and visual data and sends it to the terminal. The input feedback information is converted into AR display data and speech synthesis data, and then prepared into a data format that can be presented to the user as output.

[0319] Step 6:

[0320] The device (smart glasses and earphones) provides feedback to the user using data received from the server. Visual data is overlaid on the smart glasses, and audio data is played through the earphones. The input in this step is formatted feedback data, and the output is real-time guidance for the individual. This allows the user to immediately correct their actions and improve the effectiveness of their rehabilitation.

[0321] Furthermore, an emotion engine that estimates the user's emotions may be incorporated. That is, the identification processing unit 290 may use the emotion identification model 59 to estimate the user's emotions and perform identification processing using the user's emotions.

[0322] This invention relates to a system for analyzing the movements of athletes in real time and providing feedback that takes emotional information into account. This system combines a camera for acquiring video data, a server for analyzing the video and audio, AR glasses or earphones worn by the athlete, and an emotion engine for recognizing emotions.

[0323] First, the device's camera captures the player's movements and sends the video data to a server. Simultaneously, the player's voice is captured via a microphone and input into an emotion engine. This emotion engine has the ability to analyze the player's emotional state from their facial expressions and voice tone.

[0324] The server uses deep learning to analyze the received video data and identify key points and forms of the athlete's movements. The analysis results are compared to ideal forms, and areas for improvement in the movements are evaluated.

[0325] Meanwhile, the emotion engine analyzes the player's emotional data in real time to determine, for example, whether the player is feeling fatigued or frustrated. The server adjusts the feedback based on this emotional data and generates advice tailored to the player's mental state. For example, if a player shows signs of fatigue, the server avoids giving instructions for strenuous movements and provides feedback encouraging rest.

[0326] The generated feedback is formatted as AR data for visual instructions and as audio data for audio instructions, and then sent to the device. The device's AR glasses overlay the feedback onto the athlete's field of view, and the earphones play the audio feedback in real time. This allows athletes to not only instantly correct their form but also receive advice that takes their physical and mental state into consideration, enabling more effective training.

[0327] For example, consider a situation where a user (athlete) is practicing track and field. The server analyzes the athlete's running form and evaluates knee angle and stride length. Furthermore, if the emotion engine determines from the athlete's facial expression that their concentration is declining, it will provide feedback such as "Take a short break and drink some water," and then send specific advice such as "Try running with an awareness of widening your stride." In this way, feedback is provided while taking the athlete's emotional state into consideration, significantly improving the quality of training.

[0328] The following describes the processing flow.

[0329] Step 1:

[0330] The device's camera captures the player's movements and acquires high-resolution video data. This data is set to be captured at the optimal angle to clearly capture the player's posture and movements.

[0331] Step 2:

[0332] The terminal sends the acquired video data to the server. This transmission is performed using a high-speed communication protocol to minimize latency.

[0333] Step 3:

[0334] The server stores the received video data in a buffer and prepares to perform motion analysis.

[0335] Step 4:

[0336] The server inputs video data into a deep learning model to analyze key movement points such as joint positions. This quantifies the athlete's form, making it available for evaluation.

[0337] Step 5:

[0338] Based on the analyzed data, the server compares the player's form to an ideal form and identifies areas for improvement.

[0339] Step 6:

[0340] The device's microphone captures the player's voice and sends it to the emotion engine in real time. This audio data is used to evaluate the player's emotional state.

[0341] Step 7:

[0342] The emotion engine analyzes facial expression information obtained from audio and video data to determine the emotional state of the player. This determination serves as an indicator for evaluating the player's concentration, fatigue, frustration, and other factors.

[0343] Step 8:

[0344] The server generates feedback tailored to the player's emotional state based on data from the emotion engine. This feedback takes into account the player's mental state and includes content designed to maintain motivation.

[0345] Step 9:

[0346] The server formats the generated feedback into AR display data and audio data and sends it to the device.

[0347] Step 10:

[0348] The AR glasses on the device overlay the received visual data onto the player's field of view, visually indicating specific areas for improvement.

[0349] Step 11:

[0350] The earphones in the device transmit received audio data to the players in real time, giving them voice instructions for the necessary actions.

[0351] Step 12:

[0352] Based on the feedback provided, users (athletes) can immediately correct their form and receive emotionally responsive advice, enabling efficient and mentally conscious training.

[0353] (Example 2)

[0354] Next, we will describe Example 2. In the following description, the data processing device 12 will be referred to as the "server" and the smart glasses 214 will be referred to as the "terminal".

[0355] Traditional sports training systems separate the analysis of athletes' movements from the assessment of their emotional state, making it difficult to provide integrated feedback. Furthermore, the lack of means to provide real-time feedback tailored to the athlete's real-time mental state made it challenging to maximize performance improvement.

[0356] The identification process performed by the identification processing unit 290 of the data processing device 12 in Example 2 is realized by the following means.

[0357] In this invention, the server includes means for analyzing video data to identify key points and forms of movement, means for analyzing emotional states based on audio data, and means for generating feedback based on movement and emotional states. This makes it possible to provide feedback that takes into account both the player's movement and mental state.

[0358] "Video data" refers to visual information acquired using cameras or other recording devices, and this information is used for motion analysis.

[0359] "Audio data" refers to auditory information acquired using sound-collecting devices such as microphones, and this information is used for emotion analysis.

[0360] "Terminal means" refers to an electronic device used to acquire video and audio data and transmit it to a server.

[0361] "Server equipment" refers to computer equipment used to analyze received data and generate feedback based on actions and emotional states.

[0362] "Deep learning" refers to a technology in which artificial intelligence uses large amounts of data to learn patterns and perform predictions and analyses.

[0363] An "emotion analysis engine" refers to a collection of algorithms and software used to estimate a person's emotional state based on voice data and other input information.

[0364] "Feedback" refers to suggestions for improvement and guidance that are generated by a server and presented to the recipient visually or audibly.

[0365] This invention is a system for analyzing the movements and emotions of athletes in real time and providing effective feedback. The system consists of a terminal equipped with input devices such as a camera for acquiring video data and a microphone for acquiring audio data. The terminal temporarily stores this data and then transmits it to a server.

[0366] The server runs a deep learning model to analyze the video data. This model is built using a common deep learning framework and extracts key points from the athlete's body movements, identifying their form. The analyzed motion data is then compared to the ideal form to identify specific areas for improvement.

[0367] Furthermore, the server is equipped with an emotion analysis engine for performing sentiment analysis. This engine determines the emotional state of the players based on the audio data. The algorithms used here include natural language processing and speech emotion analysis algorithms.

[0368] The generated feedback will take into account both the athlete's actions and emotions. This feedback is then transmitted to the device in real time and communicated to the athlete. Specifically, it is displayed as a visual overlay via the device's AR glasses or provided as audio instruction through earphones.

[0369] For example, let's assume a user is practicing track and field. The server analyzes the video footage taken during the run and determines that the athlete should improve their knee angle. Simultaneously, emotional analysis is performed, and if the emotional data detects that the athlete is feeling fatigued, it recommends "taking a short rest and drinking some water," and then provides specific advice such as "focus on your form and lift your knees higher." In this way, the feedback is designed to maximize its effectiveness in both the athlete's athletic performance and mental state.

[0370] An example of a prompt to input into a generative AI model is: "Please describe a system that analyzes the movements and emotions of athletes and generates feedback. Explain how it analyzes movements in real time and provides advice based on the athlete's emotions." This prompt serves as a design guideline for improving the quality of the generated feedback.

[0371] The flow of the specific processing in Example 2 will be explained using Figure 13.

[0372] Step 1:

[0373] The device uses a camera to acquire video data of athletes and a microphone to capture audio data. This data is used to record both the athletes' movements and voice in real time. The input is the athletes' video and audio, and the output is sent to the server as data packets. Specifically, the device records video at 60 frames per second and converts the continuous audio into a digital format.

[0374] Step 2:

[0375] The server receives data packets sent from the terminal. Using a deep learning model, it identifies key points and form in the athlete's movements based on the received video data. The input is video data, and the output is the motion analysis result. Specifically, calculations are performed to quantify the position and angle of the athlete's joints and compare them to the ideal form.

[0376] Step 3:

[0377] The server simultaneously feeds the received audio data into an emotion analysis engine to estimate the player's emotional state. The input is audio data, and the output is the result of the emotional state analysis. This analysis takes into account factors such as the tone and tempo of the voice and the frequency of interruptions. Specifically, the system performs speech analysis using natural language processing technology.

[0378] Step 4:

[0379] The server integrates both the motion analysis results and the emotion analysis results to generate feedback for the player. The input is the analysis results from Step 2 and Step 3, and the output is feedback including areas for improvement and advice. Here, a generative AI model is used to automatically generate the optimal feedback content.

[0380] Step 5:

[0381] The server sends the generated feedback to the device. The device provides feedback to the user visually through the AR glasses' display and audibly through the earphones. The input is the feedback data sent from the server, and the output is the visual and audible feedback to the user. Specifically, the system overlays areas for form improvement on the AR glasses and plays detailed instructions audibly through the earphones.

[0382] (Application Example 2)

[0383] Next, we will explain application example 2. In the following explanation, the data processing device 12 will be referred to as the "server," and the smart glasses 214 will be referred to as the "terminal."

[0384] In modern manufacturing, managing the work efficiency and health of factory workers is a critical challenge. In particular, there is a need to improve workers' operational efficiency while simultaneously evaluating emotional factors such as stress and fatigue in real time and providing accurate feedback based on these assessments. Conventional technologies struggle to integrate operational improvement and emotional management, posing a risk of decreased work efficiency.

[0385] The specific processing performed by the specific processing unit 290 of the data processing device 12 in Application Example 2 is realized by the following means.

[0386] In this invention, the server includes means for receiving and analyzing video information, means for acquiring audio information and analyzing emotional states, and means for generating feedback information based on performance evaluation and emotional states. This makes it possible to provide factory workers with real-time instructions for correcting their movements and suggestions for appropriate rest and work paces.

[0387] "Visual information" refers to visual data acquired from devices such as cameras, and is used to analyze the movements and postures of subjects.

[0388] "Device means" refers to hardware or software components used to acquire, process, or present data for a specific function or purpose.

[0389] "Processing device" refers to a computer or server used to analyze acquired data and perform actions or recognize emotions.

[0390] "Audio information" refers to audio data collected through devices such as microphones, and is used to analyze the emotional state of a subject based on their speech and tone of voice.

[0391] "Emotion analysis device means" refers to a software or hardware system for estimating and analyzing a subject's emotional state from audio information or visual data.

[0392] "Feedback information" refers to instructions and advice generated based on performance evaluations and emotional states, with the aim of improving the subject's performance and maintaining a comfortable work environment.

[0393] This system is used in factories to manage and improve the work efficiency and health of workers. The system acquires data through hardware devices such as smart glasses, Bluetooth earphones, cameras, and microphones. A server processes this data on a cloud platform. Specifically, the server is built on cloud infrastructure such as AWS or Google Cloud.

[0394] First, the smart glasses worn by the user capture the worker's movements in real time using their camera. The video information is sent to a server, where motion analysis is performed using software such as TensorFlow and PiTouch. The server uses deep learning technology to extract key points of the movements and compares them to ideal movement forms.

[0395] Meanwhile, voice information is also acquired through the microphone and sent to the server. The server uses an emotion analysis engine, such as the Azure Emotion API, to evaluate emotional states such as stress and fatigue from the voice information. This evaluation result influences feedback in real time.

[0396] The server integrates motion analysis and emotion assessment to generate optimal feedback information for factory workers. This feedback information is visually overlaid on smart glasses, and voice guidance is provided through earphones.

[0397] For example, if an incorrect posture is detected during assembly work, instructions such as "Raise your elbows slightly and apply force safely" will be displayed in real time on the smart glasses. Also, if the emotion analysis engine detects high stress levels, it will suggest "Take a 5-minute break" through the earphones.

[0398] This system optimizes worker productivity and resource utilization, protecting them from excessive stress and fatigue. The overall effect of this implementation makes the factory environment a safer and more efficient place.

[0399] An example of a prompt message would be, "Please tell us about any moments during your work today when you found the feedback provided by your smart glasses particularly helpful."

[0400] The flow of a specific process in Application Example 2 will be explained using Figure 14.

[0401] Step 1:

[0402] The device acquires video information. The camera on the smart glasses worn by the user captures the movements of factory workers in real time and generates video information based on that. The input is video frames captured by the camera, and the output is a data stream of video information.

[0403] Step 2:

[0404] The device sends the acquired video information to the server. The device uploads the captured video information to a cloud-based server in real time. The input is the video information generated in step 1, and the output is the data stream sent to the server.

[0405] Step 3:

[0406] The server analyzes the video information. The server uses TensorFlow and PITach to analyze the video information and identify key points of the worker's movements. The input is the video information sent to the server, and the output is key point data of the movements. Specific movements evaluated include posture, joint angles, and stride length.

[0407] Step 4:

[0408] The terminal acquires audio information. The microphone built into the earphone worn by the user captures the worker's voice. The input is the audio signal picked up by the microphone, and the output is a data stream of the audio information.

[0409] Step 5:

[0410] The device sends the acquired audio information to the server. The device sends the captured audio information to a cloud-based server. The input is the audio information generated in step 4, and the output is the data stream sent to the server.

[0411] Step 6:

[0412] The server analyzes the audio information. The server uses the Azure Emotion API and other tools to analyze the audio information and identify the worker's emotional state. The input is the audio information sent to the server, and the output is emotional state data. Specific values ​​evaluated include stress levels and fatigue levels.

[0413] Step 7:

[0414] The server generates feedback information. The server integrates action keypoint data and emotional state data to generate optimal feedback information for factory workers. The input is the output data from steps 3 and 6, and the output is the feedback information. Specific actions include action modification instructions and break suggestions that take safety and efficiency into consideration.

[0415] Step 8:

[0416] The device presents feedback information to the user. The device visually overlays the generated feedback information on the smart glasses and communicates it audibly through the earphones. The input is the feedback information generated in step 7, and the output is the visual and audible feedback provided to the user.

[0417] The specific processing unit 290 transmits the result of the specific processing to the smart glasses 214. In the smart glasses 214, the control unit 46A causes the speaker 240 to output the result of the specific processing. The microphone 238 acquires audio indicating user input for the result of the specific processing. The control unit 46A transmits the audio data indicating user input acquired by the microphone 238 to the data processing unit 12. In the data processing unit 12, the specific processing unit 290 acquires the audio data.

[0418] Data generation model 58 is a type of so-called generative AI (Artificial Intelligence). One example of data generation model 58 is ChatGPT (Internet search<URL: https: / / openai.com / blog / chatgpt> ), Gemini (Internet search) <url: https: gemini.google.com ?hl="ja">Examples of generative AI include the following. The data generation model 58 is obtained by performing deep learning on a neural network. The data generation model 58 is input with prompts containing instructions, and with inference data such as audio data representing speech, text data representing text, and image data representing images. The data generation model 58 infers from the input inference data according to the instructions indicated by the prompts, and outputs the inference results in data formats such as audio data and text data. Here, inference refers to, for example, analysis, classification, prediction, and / or summarization.

[0419] In the above embodiment, an example was given in which specific processing is performed by the data processing device 12, but the technology of this disclosure is not limited thereto, and the specific processing may also be performed by the smart glasses 214.

[0420] [Third Embodiment]

[0421] Figure 5 shows an example of the configuration of the data processing system 310 according to the third embodiment.

[0422] As shown in Figure 5, the data processing system 310 includes a data processing device 12 and a headset terminal 314. An example of the data processing device 12 is a server.

[0423] The data processing device 12 comprises a computer 22, a database 24, and a communication interface 26. The computer 22 is an example of a "computer" related to the technology of this disclosure. The computer 22 comprises a processor 28, RAM 30, and storage 32. The processor 28, RAM 30, and storage 32 are connected to a bus 34. The database 24 and the communication interface 26 are also connected to the bus 34. The communication interface 26 is connected to a network 54. An example of the network 54 is a WAN (Wide Area Network) and / or a LAN (Local Area Network).

[0424] The headset terminal 314 includes a computer 36, a microphone 238, a speaker 240, a camera 42, a communication interface 44, and a display 343. The computer 36 includes a processor 46, RAM 48, and storage 50. The processor 46, RAM 48, and storage 50 are connected to a bus 52. The microphone 238, speaker 240, camera 42, and display 343 are also connected to the bus 52.

[0425] The microphone 238 receives voice signals from the user 20 and receives instructions from the user 20. The microphone 238 captures the voice signals from the user 20, converts the captured voice into audio data, and outputs it to the processor 46. The speaker 240 outputs audio according to the instructions from the processor 46.

[0426] Camera 42 is a small digital camera equipped with an optical system including a lens, aperture, and shutter, and an image sensor such as a CMOS (Complementary Metal-Oxide-Semiconductor) image sensor or a CCD (Charge Coupled Device) image sensor, and captures images of the area around the user 20 (for example, an imaging range defined by a field of view equivalent to the width of a typical healthy person's field of vision).

[0427] Communication interface 44 is connected to network 54. Communication interfaces 44 and 26 are responsible for the exchange of various information between processor 46 and processor 28 via network 54. The exchange of various information between processor 46 and processor 28 using communication interfaces 44 and 26 is performed in a secure manner.

[0428] Figure 6 shows an example of the main functions of the data processing device 12 and the headset terminal 314. As shown in Figure 6, the data processing device 12 performs specific processing using the processor 28. The storage 32 stores the specific processing program 56.

[0429] The specific processing program 56 is an example of a "program" relating to the technology of this disclosure. The processor 28 reads the specific processing program 56 from the storage 32 and executes the read specific processing program 56 on the RAM 30. The specific processing is realized by the processor 28 operating as a specific processing unit 290 in accordance with the specific processing program 56 executed on the RAM 30.

[0430] The storage 32 stores the data generation model 58 and the emotion identification model 59. The data generation model 58 and the emotion identification model 59 are used by the identification processing unit 290.

[0431] In the headset terminal 314, the processor 46 performs the reception output processing. The storage 50 stores the reception output program 60. The processor 46 reads the reception output program 60 from the storage 50 and executes the read reception output program 60 on the RAM 48. The reception output processing is realized by the processor 46 operating as a control unit 46A according to the reception output program 60 executed on the RAM 48.

[0432] Next, the specific processing performed by the specific processing unit 290 of the data processing device 12 will be described. In the following description, the data processing device 12 will be referred to as the "server" and the headset terminal 314 will be referred to as the "terminal".

[0433] This invention relates to a system for analyzing the movements of athletes in real time and providing feedback. This system consists of a combination of a camera, a server, and devices such as AR glasses or earphones.

[0434] First, the camera, which acts as the terminal, captures the athlete's movements and acquires video data. The acquired video data is transmitted to a server via the network. On the server, the received video data is analyzed using a pre-trained deep learning model. The deep learning model identifies the position of joints and key points of movement, and evaluates the athlete's form based on this.

[0435] Furthermore, the server retrieves the athlete's profile information (age, gender, past athletic history, etc.) from a database and generates customized feedback based on the analysis results. This feedback is formatted as AR data including visual instructions and audio data including voice instructions.

[0436] Next, the server sends the generated feedback data to the terminal device. The AR glasses, acting as the terminal, overlay the feedback content within the player's field of view, and the earphones play audio feedback in real time.

[0437] For example, if a user is practicing tennis, and the server detects that their swing follow-through is insufficient, it will display instructions through the AR glasses such as, "Raise your arms higher to improve your swing follow-through," and also communicate the same information audibly through earphones. In this way, the player can immediately correct their movements based on real-time feedback.

[0438] As described above, the present invention can provide a practical means for athletes to efficiently improve their skills through real-time motion analysis and feedback provision.

[0439] The following describes the processing flow.

[0440] Step 1:

[0441] The device activates a camera to capture the player's movements and acquires high-resolution video data. The camera's position and angle are appropriately adjusted to capture the player's overall form and subtle movements.

[0442] Step 2:

[0443] The terminal transmits the acquired video data to the server in real time. A high-speed network is used for transmission to minimize latency.

[0444] Step 3:

[0445] The server stores the video data received from the camera in a buffer and prepares it for analysis using a deep learning model.

[0446] Step 4:

[0447] The server inputs video data into a deep learning model to identify joint positions and movement trajectories. This allows it to extract characteristic points related to the athlete's movements.

[0448] Step 5:

[0449] The server evaluates the player's form based on the extracted feature points. This evaluation is performed using an algorithm that includes comparison with the ideal form.

[0450] Step 6:

[0451] The server retrieves player profile information from the database and determines the content of the feedback based on the analysis results. The feedback will be tailored to the player's characteristics and goals.

[0452] Step 7:

[0453] The server converts the generated feedback into AR data as visual instructions and into audio data as audio instructions.

[0454] Step 8:

[0455] The server sends this data to the device. Visual data is sent to the AR glasses, and audio data is sent to the earphones.

[0456] Step 9:

[0457] The AR glasses on the device overlay the received visual data onto the player's field of view. This allows the player to see in real time which parts need to be corrected.

[0458] Step 10:

[0459] The device's earphones play back the received audio data, providing players with specific instructions for improvement via voice.

[0460] Step 11:

[0461] Users (players) modify their actions based on real-time visual and audio feedback. This process enables immediate correction and skill improvement.

[0462] (Example 1)

[0463] Next, we will describe Example 1. In the following description, the data processing device 12 will be referred to as the "server," and the headset-type terminal 314 will be referred to as the "terminal."

[0464] For athletes to efficiently improve their skills, real-time and individually customized feedback is necessary. However, conventional analysis systems have struggled to provide sufficient real-time and individualized support, making it difficult to immediately reflect improvement results.

[0465] The identification process performed by the identification processing unit 290 of the data processing device 12 in Example 1 is realized by the following means.

[0466] In this invention, the server includes means for receiving video, analyzing it using a generating AI model, and evaluating a person's movements; means for generating customized feedback based on the evaluation and the person's personal information; and means for presenting the generated feedback as visual and auditory information. This allows athletes to receive personalized feedback in real time through sight and hearing, enabling them to immediately correct their movements to improve their skills.

[0467] A "device for acquiring video footage" refers to a device that has the function of filming the movements of athletes in real time and collecting them as video data.

[0468] A "generative AI model" refers to artificial intelligence technology used to analyze human movement by identifying the position of joints and key points of movement from input video data.

[0469] A "processing device" refers to a computer system that performs a series of data processing operations, including receiving video data, analyzing it, and generating feedback.

[0470] "Personal information of individuals" refers to individual data such as the age, gender, and athletic history of athletes, and is information used for analysis and feedback generation.

[0471] "Customized feedback" refers to information that provides visual and auditory instructions optimized for an individual, based on analysis results and personal information.

[0472] "A device that presents information as visual and auditory information" refers to a device that can visually display and audibly reproduce the feedback provided.

[0473] This invention relates to a system that analyzes the movements of athletes in real time and provides feedback. The system is broadly composed of a device for acquiring video, a processing device, and a device for presenting feedback.

[0474] First, the video acquisition device, acting as a terminal, uses a high-resolution camera to capture the players' movements in real time and collects video data frame by frame. This acquired video data is then transmitted to a server using wireless communication technology.

[0475] Next, the server uses a generated AI model based on the received video data to analyze the player's movements in detail. This analysis employs deep learning techniques through multiple neural network layers to identify joint positions and key movement points. Subsequently, the server retrieves the player's personal information from a database and combines it with the analysis results to generate customized feedback. The feedback is converted into AR data as visual instructions and formatted as audio data as voice instructions.

[0476] Finally, the generated feedback data is sent to the AR glasses and earphones, which are the devices used. The AR glasses overlay real-time visual feedback within the player's field of view, and the earphones play audio feedback. This allows the player to receive immediate feedback and correct their actions.

[0477] For example, when a user is practicing tennis, the server analyzes and detects that their swing follow-through is insufficient. In this case, the AR glasses display a visual instruction such as, "Raise your arm higher to improve your swing follow-through," and the same information is conveyed as audio through the earphones. In this way, players can improve their technique based on real-time feedback.

[0478] An example of a prompt message is, "Analyze the tennis swing motion and generate real-time feedback for form improvement." Following this example, the system assists in improving individual movements in sports.

[0479] The flow of the specific processing in Example 1 will be explained using Figure 11.

[0480] Step 1:

[0481] The device (camera) captures video data by filming the movements of athletes in real time at high resolution. The input is a scene of the athlete's movements, and the output is a continuous generation of video frames. In this process, the camera meticulously records the athlete's movements and captures the video at a high frame rate to capture the details of the movements.

[0482] Step 2:

[0483] The terminal reduces the data size of the acquired video data using compression technology and transmits it to the server via a wireless network. Raw video data is the input, and compressed video data is transmitted as the output. Specifically, compression algorithms such as H.264 are used to enable high-speed and efficient data transfer.

[0484] Step 3:

[0485] The server analyzes the received video data using a generative AI model based on deep learning to identify key points in the athlete's movements. Compressed video data is the input, and the output is the analyzed joint positions and movement characteristics of the body. The server extracts important body metrics for each frame through a multi-layered neural network.

[0486] Step 4:

[0487] The server combines analysis results with the athlete's personal profile information (age, gender, exercise history, etc.) to generate customized feedback. Inputs include analysis data and profile information, while outputs include AR data for visual feedback and audio feedback data. Based on the analysis, the server automatically generates individually optimized advice.

[0488] Step 5:

[0489] The server transmits the generated feedback data to terminal devices (AR glasses, earphones) and presents it to the players in real time. The input is feedback data, and the output is information that directly engages the players' sight and hearing. The AR glasses provide visual instructions, and the earphones provide audio guidance, enabling players to immediately work on improving their skills.

[0490] Step 6:

[0491] The user (athlete) modifies their actions based on the feedback provided, improving their practice and performance. Input comes from AR glasses and earphones, and output is optimized sports performance. Based on the feedback received in real time, athletes instantly adjust their skills and progress in their learning.

[0492] (Application Example 1)

[0493] Next, we will explain Application Example 1. In the following explanation, the data processing device 12 will be referred to as the "server," and the headset-type terminal 314 will be referred to as the "terminal."

[0494] There is a lack of methods to provide appropriate exercise and feedback tailored to the individual condition and abilities of elderly people and patients undergoing rehabilitation. Traditional rehabilitation methods often lack individualized support, and patients may not receive appropriate guidance. Therefore, a system capable of real-time movement analysis and guidance is needed to achieve effective and rapid rehabilitation.

[0495] The specific processing performed by the specific processing unit 290 of the data processing device 12 in Application Example 1 is realized by the following means.

[0496] In this invention, the server includes a device for acquiring video information, a computing device for evaluating the subject's movements, and a computing device for generating feedback information. This enables individual motion analysis and real-time feedback.

[0497] "Video information" refers to visual data acquired to record the actions of a subject.

[0498] A "device" is hardware used to capture and display video information.

[0499] A "processing unit" is a computing device that analyzes received data and generates evaluations and feedback.

[0500] "Information" refers to data used for performance evaluation and feedback, and is represented visually and audibly.

[0501] "Subjects" are individuals who undergo performance evaluation using the system.

[0502] "Feedback" refers to evaluation results that include guidance for improving the subject's actions.

[0503] "Attribute information" refers to the individual information of the target person and is data used for personalized feedback.

[0504] This system promotes efficient exercise by analyzing the movements of elderly or rehabilitation patients in real time and providing immediate feedback on areas for improvement. The device consists of smart glasses with a built-in camera and earphones. These work together to capture movements and provide feedback through both visual and auditory means.

[0505] The server receives video information via the network. A deep learning model is used to evaluate the subject's movements based on the video information. Python and TensorFlow are used to train the model, identifying joint positions and key points of movement. The evaluation results are used to retrieve the subject's attribute information from a database, and customized feedback is generated based on this information. The generated information is overlaid as visual data on smart glasses and played back as audio data through earphones.

[0506] For example, when performing arm-raising exercises as part of rehabilitation, if there is a problem with the patient's movement, the smart glasses display will say, "It would be better to raise your arm a little higher," and the same instruction will be conveyed via voice through the earphones. This allows the user to immediately correct their movement and improve the effectiveness of the rehabilitation.

[0507] The following example prompt can be used in the generated AI model: "Analyze the rehabilitation movements of a young male and generate visual and auditory feedback if the arm lift is inappropriate." Using this prompt, it is possible to create training data for specific movement analysis and improve the accuracy of the model.

[0508] The flow of a specific process in Application Example 1 will be explained using Figure 12.

[0509] Step 1:

[0510] The device (smart glasses) uses a camera to capture the subject's movements. In this step, the input is the subject's real-time movements, and the output is raw video footage of those movements. The device transmits this raw video to a server via the network.

[0511] Step 2:

[0512] The server processes video information received via the network. The input is the raw video obtained in step 1, and the output is a dataset in a format suitable for input to a deep learning model. The server preprocesses this data to extract joint positions and key points of movement.

[0513] Step 3:

[0514] The server uses a deep learning model to perform motion analysis. The input is a pre-processed dataset, and the output is an evaluation of the subject's movements. Through this evaluation, the server identifies the appropriateness of the movements and areas that need correction.

[0515] Step 4:

[0516] The server retrieves the subject's attribute information (age, weight, past rehabilitation records, etc.) from the database. Based on this input information, it generates customized feedback. The output is feedback information optimized for the subject.

[0517] Step 5:

[0518] The server formats the generated feedback information into audio and visual data and sends it to the terminal. The input feedback information is converted into AR display data and speech synthesis data, and then prepared into a data format that can be presented to the user as output.

[0519] Step 6:

[0520] The device (smart glasses and earphones) provides feedback to the user using data received from the server. Visual data is overlaid on the smart glasses, and audio data is played through the earphones. The input in this step is formatted feedback data, and the output is real-time guidance for the individual. This allows the user to immediately correct their actions and improve the effectiveness of their rehabilitation.

[0521] Furthermore, an emotion engine that estimates the user's emotions may be incorporated. That is, the identification processing unit 290 may use the emotion identification model 59 to estimate the user's emotions and perform identification processing using the user's emotions.

[0522] This invention relates to a system for analyzing the movements of athletes in real time and providing feedback that takes emotional information into account. This system combines a camera for acquiring video data, a server for analyzing the video and audio, AR glasses or earphones worn by the athlete, and an emotion engine for recognizing emotions.

[0523] First, the device's camera captures the player's movements and sends the video data to a server. Simultaneously, the player's voice is captured via a microphone and input into an emotion engine. This emotion engine has the ability to analyze the player's emotional state from their facial expressions and voice tone.

[0524] The server uses deep learning to analyze the received video data and identify key points and forms of the athlete's movements. The analysis results are compared to ideal forms, and areas for improvement in the movements are evaluated.

[0525] Meanwhile, the emotion engine analyzes the player's emotional data in real time to determine, for example, whether the player is feeling fatigued or frustrated. The server adjusts the feedback based on this emotional data and generates advice tailored to the player's mental state. For example, if a player shows signs of fatigue, the server avoids giving instructions for strenuous movements and provides feedback encouraging rest.

[0526] The generated feedback is formatted as AR data for visual instructions and as audio data for audio instructions, and then sent to the device. The device's AR glasses overlay the feedback onto the athlete's field of view, and the earphones play the audio feedback in real time. This allows athletes to not only instantly correct their form but also receive advice that takes their physical and mental state into consideration, enabling more effective training.

[0527] For example, consider a situation where a user (athlete) is practicing track and field. The server analyzes the athlete's running form and evaluates knee angle and stride length. Furthermore, if the emotion engine determines from the athlete's facial expression that their concentration is declining, it will provide feedback such as "Take a short break and drink some water," and then send specific advice such as "Try running with an awareness of widening your stride." In this way, feedback is provided while taking the athlete's emotional state into consideration, significantly improving the quality of training.

[0528] The following describes the processing flow.

[0529] Step 1:

[0530] The device's camera captures the player's movements and acquires high-resolution video data. This data is set to be captured at the optimal angle to clearly capture the player's posture and movements.

[0531] Step 2:

[0532] The terminal sends the acquired video data to the server. This transmission is performed using a high-speed communication protocol to minimize latency.

[0533] Step 3:

[0534] The server stores the received video data in a buffer and prepares to perform motion analysis.

[0535] Step 4:

[0536] The server inputs video data into a deep learning model to analyze key movement points such as joint positions. This quantifies the athlete's form, making it available for evaluation.

[0537] Step 5:

[0538] Based on the analyzed data, the server compares the player's form to an ideal form and identifies areas for improvement.

[0539] Step 6:

[0540] The device's microphone captures the player's voice and sends it to the emotion engine in real time. This audio data is used to evaluate the player's emotional state.

[0541] Step 7:

[0542] The emotion engine analyzes facial expression information obtained from audio and video data to determine the emotional state of the player. This determination serves as an indicator for evaluating the player's concentration, fatigue, frustration, and other factors.

[0543] Step 8:

[0544] The server generates feedback tailored to the player's emotional state based on data from the emotion engine. This feedback takes into account the player's mental state and includes content designed to maintain motivation.

[0545] Step 9:

[0546] The server formats the generated feedback into AR display data and audio data and sends it to the device.

[0547] Step 10:

[0548] The AR glasses on the device overlay the received visual data onto the player's field of view, visually indicating specific areas for improvement.

[0549] Step 11:

[0550] The earphones in the device transmit received audio data to the players in real time, giving them voice instructions for the necessary actions.

[0551] Step 12:

[0552] Based on the feedback provided, users (athletes) can immediately correct their form and receive emotionally responsive advice, enabling efficient and mentally conscious training.

[0553] (Example 2)

[0554] Next, we will describe Example 2. In the following description, the data processing device 12 will be referred to as the "server," and the headset-type terminal 314 will be referred to as the "terminal."

[0555] Traditional sports training systems separate the analysis of athletes' movements from the assessment of their emotional state, making it difficult to provide integrated feedback. Furthermore, the lack of means to provide real-time feedback tailored to the athlete's real-time mental state made it challenging to maximize performance improvement.

[0556] The identification process performed by the identification processing unit 290 of the data processing device 12 in Example 2 is realized by the following means.

[0557] In this invention, the server includes means for analyzing video data to identify key points and forms of movement, means for analyzing emotional states based on audio data, and means for generating feedback based on movement and emotional states. This makes it possible to provide feedback that takes into account both the player's movement and mental state.

[0558] "Video data" refers to visual information acquired using cameras and other recording devices, and this information is used for motion analysis.

[0559] "Audio data" refers to auditory information acquired using sound-collecting devices such as microphones, and this information is used for emotion analysis.

[0560] "Terminal means" refers to an electronic device used to acquire video and audio data and transmit it to a server.

[0561] "Server equipment" refers to computer equipment used to analyze received data and generate feedback based on actions and emotional states.

[0562] "Deep learning" refers to a technology in which artificial intelligence uses large amounts of data to learn patterns and perform predictions and analyses.

[0563] An "emotion analysis engine" refers to a collection of algorithms and software used to estimate a person's emotional state based on voice data and other input information.

[0564] "Feedback" refers to suggestions for improvement and guidance that are generated by a server and presented to the recipient visually or audibly.

[0565] This invention is a system for analyzing the movements and emotions of athletes in real time and providing effective feedback. The system consists of a terminal equipped with input devices such as a camera for acquiring video data and a microphone for acquiring audio data. The terminal temporarily stores this data and then transmits it to a server.

[0566] The server runs a deep learning model to analyze the video data. This model is built using a common deep learning framework and extracts key points from the athlete's body movements, identifying their form. The analyzed motion data is then compared to the ideal form to identify specific areas for improvement.

[0567] Furthermore, the server is equipped with an emotion analysis engine for performing sentiment analysis. This engine determines the emotional state of the players based on the audio data. The algorithms used here include natural language processing and speech emotion analysis algorithms.

[0568] The generated feedback will take into account both the athlete's actions and emotions. This feedback is then transmitted to the device in real time and communicated to the athlete. Specifically, it is displayed as a visual overlay via the device's AR glasses or provided as audio instruction through earphones.

[0569] For example, let's assume a user is practicing track and field. The server analyzes the video footage taken during the run and determines that the athlete should improve their knee angle. Simultaneously, emotional analysis is performed, and if the emotional data detects that the athlete is feeling fatigued, it recommends "taking a short rest and drinking some water," and then provides specific advice such as "focus on your form and lift your knees higher." In this way, the feedback is designed to maximize its effectiveness in both the athlete's athletic performance and mental state.

[0570] An example of a prompt to input into a generative AI model is: "Please describe a system that analyzes the movements and emotions of athletes and generates feedback. Explain how it analyzes movements in real time and provides advice based on the athlete's emotions." This prompt serves as a design guideline for improving the quality of the generated feedback.

[0571] The flow of the specific processing in Example 2 will be explained using Figure 13.

[0572] Step 1:

[0573] The device uses a camera to acquire video data of athletes and a microphone to capture audio data. This data is used to record both the athletes' movements and voice in real time. The input is the athletes' video and audio, and the output is sent to the server as data packets. Specifically, the device records video at 60 frames per second and converts the continuous audio into a digital format.

[0574] Step 2:

[0575] The server receives data packets sent from the terminal. Using a deep learning model, it identifies key points and form in the athlete's movements based on the received video data. The input is video data, and the output is the motion analysis result. Specifically, calculations are performed to quantify the position and angle of the athlete's joints and compare them to the ideal form.

[0576] Step 3:

[0577] The server simultaneously feeds the received audio data into an emotion analysis engine to estimate the player's emotional state. The input is audio data, and the output is the result of the emotional state analysis. This analysis takes into account factors such as the tone and tempo of the voice and the frequency of interruptions. Specifically, the system performs speech analysis using natural language processing technology.

[0578] Step 4:

[0579] The server integrates both the motion analysis results and the emotion analysis results to generate feedback for the player. The input is the analysis results from Step 2 and Step 3, and the output is feedback including areas for improvement and advice. Here, a generative AI model is used to automatically generate the optimal feedback content.

[0580] Step 5:

[0581] The server sends the generated feedback to the device. The device provides feedback to the user visually through the AR glasses' display and audibly through the earphones. The input is the feedback data sent from the server, and the output is the visual and audible feedback to the user. Specifically, the system overlays areas for form improvement on the AR glasses and plays detailed instructions audibly through the earphones.

[0582] (Application Example 2)

[0583] Next, we will explain Application Example 2. In the following explanation, the data processing device 12 will be referred to as the "server," and the headset-type terminal 314 will be referred to as the "terminal."

[0584] In modern manufacturing, managing the work efficiency and health of factory workers is a critical challenge. In particular, there is a need to improve workers' operational efficiency while simultaneously evaluating emotional factors such as stress and fatigue in real time and providing accurate feedback based on these assessments. Conventional technologies struggle to integrate operational improvement and emotional management, posing a risk of decreased work efficiency.

[0585] The specific processing performed by the specific processing unit 290 of the data processing device 12 in Application Example 2 is realized by the following means.

[0586] In this invention, the server includes means for receiving and analyzing video information, means for acquiring audio information and analyzing emotional states, and means for generating feedback information based on performance evaluation and emotional states. This makes it possible to provide factory workers with real-time instructions for correcting their movements and suggestions for appropriate rest and work paces.

[0587] "Visual information" refers to visual data acquired from devices such as cameras, and is used to analyze the movements and postures of subjects.

[0588] "Device means" refers to hardware or software components used to acquire, process, or present data for a specific function or purpose.

[0589] "Processing device" refers to a computer or server used to analyze acquired data and perform actions or recognize emotions.

[0590] "Audio information" refers to audio data collected through devices such as microphones, and is used to analyze the emotional state of a subject based on their speech and tone of voice.

[0591] "Emotional analysis device means" refers to a software or hardware system for estimating and analyzing a subject's emotional state from audio information or visual data.

[0592] "Feedback information" refers to instructions and advice generated based on performance evaluations and emotional states, with the aim of improving the subject's performance and maintaining a comfortable work environment.

[0593] This system is used in factories to manage and improve the work efficiency and health of workers. The system acquires data through hardware devices such as smart glasses, Bluetooth earphones, cameras, and microphones. A server processes this data on a cloud platform. Specifically, the server is built on cloud infrastructure such as AWS or Google Cloud.

[0594] First, the smart glasses worn by the user capture the worker's movements in real time using their camera. The video information is sent to a server, where motion analysis is performed using software such as TensorFlow and PiTouch. The server uses deep learning technology to extract key points of the movements and compares them to ideal movement forms.

[0595] Meanwhile, voice information is also acquired through the microphone and sent to the server. The server uses an emotion analysis engine, such as the Azure Emotion API, to evaluate emotional states such as stress and fatigue from the voice information. This evaluation result influences feedback in real time.

[0596] The server integrates motion analysis and emotion assessment to generate optimal feedback information for factory workers. This feedback information is visually overlaid on smart glasses, and voice guidance is provided through earphones.

[0597] For example, if an incorrect posture is detected during assembly work, instructions such as "Raise your elbows slightly and apply force safely" will be displayed in real time on the smart glasses. Also, if the emotion analysis engine detects high stress levels, it will suggest "Take a 5-minute break" through the earphones.

[0598] This system optimizes worker productivity and resource utilization, protecting them from excessive stress and fatigue. The overall effect of this implementation makes the factory environment a safer and more efficient place.

[0599] An example of a prompt message would be, "Please tell us about any moments during your work today when you found the feedback provided by your smart glasses particularly helpful."

[0600] The flow of a specific process in Application Example 2 will be explained using Figure 14.

[0601] Step 1:

[0602] The device acquires video information. The camera on the smart glasses worn by the user captures the movements of factory workers in real time and generates video information based on that. The input is video frames captured by the camera, and the output is a data stream of video information.

[0603] Step 2:

[0604] The device sends the acquired video information to the server. The device uploads the captured video information to a cloud-based server in real time. The input is the video information generated in step 1, and the output is the data stream sent to the server.

[0605] Step 3:

[0606] The server analyzes the video information. The server uses TensorFlow and PITach to analyze the video information and identify key points of the worker's movements. The input is the video information sent to the server, and the output is key point data of the movements. Specific movements evaluated include posture, joint angles, and stride length.

[0607] Step 4:

[0608] The terminal acquires audio information. The microphone built into the earphone worn by the user captures the worker's voice. The input is the audio signal picked up by the microphone, and the output is a data stream of the audio information.

[0609] Step 5:

[0610] The device sends the acquired audio information to the server. The device sends the captured audio information to a cloud-based server. The input is the audio information generated in step 4, and the output is the data stream sent to the server.

[0611] Step 6:

[0612] The server analyzes the audio information. The server uses the Azure Emotion API and other tools to analyze the audio information and identify the worker's emotional state. The input is the audio information sent to the server, and the output is emotional state data. Specific values ​​evaluated include stress levels and fatigue levels.

[0613] Step 7:

[0614] The server generates feedback information. The server integrates action keypoint data and emotional state data to generate optimal feedback information for factory workers. The input is the output data from steps 3 and 6, and the output is the feedback information. Specific actions include action modification instructions and break suggestions that take safety and efficiency into consideration.

[0615] Step 8:

[0616] The device presents feedback information to the user. The device visually overlays the generated feedback information on the smart glasses and communicates it audibly through the earphones. The input is the feedback information generated in step 7, and the output is the visual and audible feedback provided to the user.

[0617] The specific processing unit 290 transmits the result of the specific processing to the headset terminal 314. In the headset terminal 314, the control unit 46A causes the speaker 240 and display 343 to output the result of the specific processing. The microphone 238 acquires audio indicating user input for the result of the specific processing. The control unit 46A transmits the audio data indicating user input acquired by the microphone 238 to the data processing unit 12. In the data processing unit 12, the specific processing unit 290 acquires the audio data.

[0618] Data generation model 58 is a type of so-called generative AI (Artificial Intelligence). An example of data generation model 58 is ChatGPT (Internet Search<URL: https: / / openai.com / blog / chatgpt> ), Gemini (Internet search) <url: https: gemini.google.com ?hl="ja">Examples of generative AI include the following. The data generation model 58 is obtained by performing deep learning on a neural network. The data generation model 58 is input with prompts containing instructions, and with inference data such as audio data representing speech, text data representing text, and image data representing images. The data generation model 58 infers from the input inference data according to the instructions indicated by the prompts, and outputs the inference results in data formats such as audio data and text data. Here, inference refers to, for example, analysis, classification, prediction, and / or summarization.

[0619] In the above embodiment, an example was given in which specific processing is performed by the data processing device 12, but the technology of this disclosure is not limited thereto, and specific processing may also be performed by the headset terminal 314.

[0620] [Fourth Embodiment]

[0621] Figure 7 shows an example of the configuration of the data processing system 410 according to the fourth embodiment.

[0622] As shown in Figure 7, the data processing system 410 includes a data processing device 12 and a robot 414. An example of the data processing device 12 is a server.

[0623] The data processing device 12 comprises a computer 22, a database 24, and a communication interface 26. The computer 22 is an example of a "computer" related to the technology of this disclosure. The computer 22 comprises a processor 28, RAM 30, and storage 32. The processor 28, RAM 30, and storage 32 are connected to a bus 34. The database 24 and the communication interface 26 are also connected to the bus 34. The communication interface 26 is connected to a network 54. An example of the network 54 is a WAN (Wide Area Network) and / or a LAN (Local Area Network).

[0624] The robot 414 includes a computer 36, a microphone 238, a speaker 240, a camera 42, a communication interface 44, and a controlled object 443. The computer 36 includes a processor 46, RAM 48, and storage 50. The processor 46, RAM 48, and storage 50 are connected to a bus 52. The microphone 238, speaker 240, camera 42, and controlled object 443 are also connected to the bus 52.

[0625] The microphone 238 receives voice signals from the user 20 and receives instructions from the user 20. The microphone 238 captures the voice signals from the user 20, converts the captured voice into audio data, and outputs it to the processor 46. The speaker 240 outputs audio according to the instructions from the processor 46.

[0626] Camera 42 is a small digital camera equipped with an optical system including a lens, aperture, and shutter, and an image sensor such as a CMOS (Complementary Metal-Oxide-Semiconductor) image sensor or a CCD (Charge Coupled Device) image sensor, and captures images of the area around the user 20 (for example, an imaging range defined by a field of view equivalent to the width of a typical healthy person's field of vision).

[0627] Communication interface 44 is connected to network 54. Communication interfaces 44 and 26 are responsible for the exchange of various information between processor 46 and processor 28 via network 54. The exchange of various information between processor 46 and processor 28 using communication interfaces 44 and 26 is performed in a secure manner.

[0628] The controlled object 443 includes a display device, LEDs in the eyes, and motors that drive the arms, hands, and feet. The posture and gestures of the robot 414 are controlled by controlling the motors of the arms, hands, and feet. Some of the robot 414's emotions can be expressed by controlling these motors. Furthermore, the robot 414's facial expressions can also be expressed by controlling the illumination state of the LEDs in its eyes.

[0629] Figure 8 shows an example of the main functions of the data processing device 12 and the robot 414. As shown in Figure 8, the data processing device 12 performs specific processing using the processor 28. The storage 32 stores the specific processing program 56.

[0630] The specific processing program 56 is an example of a "program" relating to the technology of this disclosure. The processor 28 reads the specific processing program 56 from the storage 32 and executes the read specific processing program 56 on the RAM 30. The specific processing is realized by the processor 28 operating as a specific processing unit 290 in accordance with the specific processing program 56 executed on the RAM 30.

[0631] The storage 32 stores the data generation model 58 and the emotion identification model 59. The data generation model 58 and the emotion identification model 59 are used by the identification processing unit 290.

[0632] In robot 414, the processor 46 performs the reception output processing. The storage 50 stores the reception output program 60. The processor 46 reads the reception output program 60 from the storage 50 and executes the read reception output program 60 on the RAM 48. The reception output processing is realized by the processor 46 operating as a control unit 46A according to the reception output program 60 executed on the RAM 48.

[0633] Next, the specific processing performed by the specific processing unit 290 of the data processing device 12 will be described. In the following description, the data processing device 12 will be referred to as the "server" and the robot 414 as the "terminal".

[0634] This invention relates to a system for analyzing the movements of athletes in real time and providing feedback. This system consists of a combination of a camera, a server, and devices such as AR glasses or earphones.

[0635] First, the camera, which acts as the terminal, captures the athlete's movements and acquires video data. The acquired video data is transmitted to a server via the network. On the server, the received video data is analyzed using a pre-trained deep learning model. The deep learning model identifies the position of joints and key points of movement, and evaluates the athlete's form based on this.

[0636] Furthermore, the server retrieves the athlete's profile information (age, gender, past athletic history, etc.) from a database and generates customized feedback based on the analysis results. This feedback is formatted as AR data including visual instructions and audio data including voice instructions.

[0637] Next, the server sends the generated feedback data to the terminal device. The AR glasses, acting as the terminal, overlay the feedback content within the player's field of view, and the earphones play audio feedback in real time.

[0638] For example, if a user is practicing tennis, and the server detects that their swing follow-through is insufficient, it will display instructions through the AR glasses such as, "Raise your arms higher to improve your swing follow-through," and also communicate the same information audibly through earphones. In this way, the player can immediately correct their movements based on real-time feedback.

[0639] As described above, the present invention can provide a practical means for athletes to efficiently improve their skills through real-time motion analysis and feedback provision.

[0640] The following describes the processing flow.

[0641] Step 1:

[0642] The device activates a camera to capture the player's movements and acquires high-resolution video data. The camera's position and angle are appropriately adjusted to capture the player's overall form and subtle movements.

[0643] Step 2:

[0644] The terminal transmits the acquired video data to the server in real time. A high-speed network is used for transmission to minimize latency.

[0645] Step 3:

[0646] The server stores the video data received from the camera in a buffer and prepares it for analysis using a deep learning model.

[0647] Step 4:

[0648] The server inputs video data into a deep learning model to identify joint positions and movement trajectories. This allows it to extract characteristic points related to the athlete's movements.

[0649] Step 5:

[0650] The server evaluates the player's form based on the extracted feature points. This evaluation is performed using an algorithm that includes comparison with the ideal form.

[0651] Step 6:

[0652] The server retrieves player profile information from the database and determines the content of the feedback based on the analysis results. The feedback will be tailored to the player's characteristics and goals.

[0653] Step 7:

[0654] The server converts the generated feedback into AR data as visual instructions and into audio data as audio instructions.

[0655] Step 8:

[0656] The server sends this data to the device. Visual data is sent to the AR glasses, and audio data is sent to the earphones.

[0657] Step 9:

[0658] The AR glasses on the device overlay the received visual data onto the player's field of view. This allows the player to see in real time which parts need to be corrected.

[0659] Step 10:

[0660] The device's earphones play back the received audio data, providing players with specific instructions for improvement via voice.

[0661] Step 11:

[0662] Users (players) modify their actions based on real-time visual and audio feedback. This process enables immediate correction and skill improvement.

[0663] (Example 1)

[0664] Next, we will describe Example 1. In the following description, the data processing device 12 will be referred to as the "server" and the robot 414 as the "terminal".

[0665] For athletes to efficiently improve their skills, real-time and individually customized feedback is necessary. However, conventional analysis systems have struggled to provide sufficient real-time and individualized support, making it difficult to immediately reflect improvement results.

[0666] The identification process performed by the identification processing unit 290 of the data processing device 12 in Example 1 is realized by the following means.

[0667] In this invention, the server includes means for receiving video, analyzing it using a generating AI model, and evaluating a person's movements; means for generating customized feedback based on the evaluation and the person's personal information; and means for presenting the generated feedback as visual and auditory information. This allows athletes to receive personalized feedback in real time through sight and hearing, enabling them to immediately correct their movements to improve their skills.

[0668] A "device for acquiring video footage" refers to a device that has the function of filming the movements of athletes in real time and collecting them as video data.

[0669] A "generative AI model" refers to artificial intelligence technology used to analyze human movement by identifying the position of joints and key points of movement from input video data.

[0670] A "processing device" refers to a computer system that performs a series of data processing operations, including receiving video data, analyzing it, and generating feedback.

[0671] "Personal information of individuals" refers to individual data such as the age, gender, and athletic history of athletes, and is information used for analysis and feedback generation.

[0672] "Customized feedback" refers to information that provides visual and auditory instructions optimized for an individual, based on analysis results and personal information.

[0673] "A device that presents information as visual and auditory information" refers to a device that can visually display and audibly reproduce the feedback provided.

[0674] This invention relates to a system that analyzes the movements of athletes in real time and provides feedback. The system is broadly composed of a device for acquiring video, a processing device, and a device for presenting feedback.

[0675] First, the video acquisition device, acting as a terminal, uses a high-resolution camera to capture the players' movements in real time and collects video data frame by frame. This acquired video data is then transmitted to a server using wireless communication technology.

[0676] Next, the server uses a generated AI model based on the received video data to analyze the player's movements in detail. This analysis employs deep learning techniques through multiple neural network layers to identify joint positions and key movement points. Subsequently, the server retrieves the player's personal information from a database and combines it with the analysis results to generate customized feedback. The feedback is converted into AR data as visual instructions and formatted as audio data as voice instructions.

[0677] Finally, the generated feedback data is sent to the AR glasses and earphones, which are the devices used. The AR glasses overlay real-time visual feedback within the player's field of view, and the earphones play audio feedback. This allows the player to receive immediate feedback and correct their actions.

[0678] For example, when a user is practicing tennis, the server analyzes and detects that their swing follow-through is insufficient. In this case, the AR glasses display a visual instruction such as, "Raise your arm higher to improve your swing follow-through," and the same information is conveyed as audio through the earphones. In this way, players can improve their technique based on real-time feedback.

[0679] An example of a prompt message is, "Analyze the tennis swing motion and generate real-time feedback for form improvement." Following this example, the system assists in improving individual movements in sports.

[0680] The flow of the specific processing in Example 1 will be explained using Figure 11.

[0681] Step 1:

[0682] The device (camera) captures video data by filming the movements of athletes in real time at high resolution. The input is a scene of the athlete's movements, and the output is a continuous generation of video frames. In this process, the camera meticulously records the athlete's movements and captures the video at a high frame rate to capture the details of the movements.

[0683] Step 2:

[0684] The terminal reduces the data size of the acquired video data using compression technology and transmits it to the server via a wireless network. Raw video data is the input, and compressed video data is transmitted as the output. Specifically, compression algorithms such as H.264 are used to enable high-speed and efficient data transfer.

[0685] Step 3:

[0686] The server analyzes the received video data using a generative AI model based on deep learning to identify key points in the athlete's movements. Compressed video data is the input, and the output is the analyzed joint positions and movement characteristics of the body. The server extracts important body metrics for each frame through a multi-layered neural network.

[0687] Step 4:

[0688] The server combines analysis results with the athlete's personal profile information (age, gender, exercise history, etc.) to generate customized feedback. Inputs include analysis data and profile information, while outputs include AR data for visual feedback and audio feedback data. Based on the analysis, the server automatically generates individually optimized advice.

[0689] Step 5:

[0690] The server transmits the generated feedback data to terminal devices (AR glasses, earphones) and presents it to the players in real time. The input is feedback data, and the output is information that directly engages the players' sight and hearing. The AR glasses provide visual instructions, and the earphones provide audio guidance, enabling players to immediately work on improving their skills.

[0691] Step 6:

[0692] The user (athlete) modifies their actions based on the feedback provided, improving their practice and performance. Input comes from AR glasses and earphones, and output is optimized sports performance. Based on the feedback received in real time, athletes instantly adjust their skills and progress in their learning.

[0693] (Application Example 1)

[0694] Next, we will explain Application Example 1. In the following explanation, the data processing device 12 will be referred to as the "server" and the robot 414 as the "terminal".

[0695] There is a lack of methods to provide appropriate exercise and feedback tailored to the individual condition and abilities of elderly people and patients undergoing rehabilitation. Traditional rehabilitation methods often lack individualized support, and patients may not receive appropriate guidance. Therefore, a system capable of real-time movement analysis and guidance is needed to achieve effective and rapid rehabilitation.

[0696] The specific processing performed by the specific processing unit 290 of the data processing device 12 in Application Example 1 is realized by the following means.

[0697] In this invention, the server includes a device for acquiring video information, a computing device for evaluating the subject's movements, and a computing device for generating feedback information. This enables individual motion analysis and real-time feedback.

[0698] "Video information" refers to visual data acquired to record the actions of a subject.

[0699] A "device" is hardware used to capture and display video information.

[0700] A "processing unit" is a computing device that analyzes received data and generates evaluations and feedback.

[0701] "Information" refers to data used for performance evaluation and feedback, and is represented visually and audibly.

[0702] "Subjects" are individuals who undergo performance evaluation using the system.

[0703] "Feedback" refers to evaluation results that include guidance for improving the subject's actions.

[0704] "Attribute information" refers to the individual information of the target person and is data used for personalized feedback.

[0705] This system promotes efficient exercise by analyzing the movements of elderly or rehabilitation patients in real time and providing immediate feedback on areas for improvement. The device consists of smart glasses with a built-in camera and earphones. These work together to capture movements and provide feedback through both visual and auditory means.

[0706] The server receives video information via the network. A deep learning model is used to evaluate the subject's movements based on the video information. Python and TensorFlow are used to train the model, identifying joint positions and key points of movement. The evaluation results are used to retrieve the subject's attribute information from a database, and customized feedback is generated based on this information. The generated information is overlaid as visual data on smart glasses and played back as audio data through earphones.

[0707] For example, when performing arm-raising exercises as part of rehabilitation, if there is a problem with the patient's movement, the smart glasses display will say, "It would be better to raise your arm a little higher," and the same instruction will be conveyed via voice through the earphones. This allows the user to immediately correct their movement and improve the effectiveness of the rehabilitation.

[0708] The following example prompt can be used in the generated AI model: "Analyze the rehabilitation movements of a young male and generate visual and auditory feedback if the arm lift is inappropriate." Using this prompt, it is possible to create training data for specific movement analysis and improve the accuracy of the model.

[0709] The flow of a specific process in Application Example 1 will be explained using Figure 12.

[0710] Step 1:

[0711] The device (smart glasses) uses a camera to capture the subject's movements. In this step, the input is the subject's real-time movements, and the output is raw video footage of those movements. The device transmits this raw video to a server via the network.

[0712] Step 2:

[0713] The server processes video information received via the network. The input is the raw video obtained in step 1, and the output is a dataset in a format suitable for input to a deep learning model. The server preprocesses this data to extract joint positions and key points of movement.

[0714] Step 3:

[0715] The server uses a deep learning model to perform motion analysis. The input is a pre-processed dataset, and the output is an evaluation of the subject's movements. Through this evaluation, the server identifies the appropriateness of the movements and areas that need correction.

[0716] Step 4:

[0717] The server retrieves the subject's attribute information (age, weight, past rehabilitation records, etc.) from the database. Based on this input information, it generates customized feedback. The output is feedback information optimized for the subject.

[0718] Step 5:

[0719] The server formats the generated feedback information into audio and visual data and sends it to the terminal. The input feedback information is converted into AR display data and speech synthesis data, and then prepared into a data format that can be presented to the user as output.

[0720] Step 6:

[0721] The device (smart glasses and earphones) provides feedback to the user using data received from the server. Visual data is overlaid on the smart glasses, and audio data is played through the earphones. The input in this step is formatted feedback data, and the output is real-time guidance for the individual. This allows the user to immediately correct their actions and improve the effectiveness of their rehabilitation.

[0722] Furthermore, an emotion engine that estimates the user's emotions may be incorporated. That is, the identification processing unit 290 may use the emotion identification model 59 to estimate the user's emotions and perform identification processing using the user's emotions.

[0723] This invention relates to a system for analyzing the movements of athletes in real time and providing feedback that takes emotional information into account. This system combines a camera for acquiring video data, a server for analyzing the video and audio, AR glasses or earphones worn by the athlete, and an emotion engine for recognizing emotions.

[0724] First, the device's camera captures the player's movements and sends the video data to a server. Simultaneously, the player's voice is captured via a microphone and input into an emotion engine. This emotion engine has the ability to analyze the player's emotional state from their facial expressions and voice tone.

[0725] The server uses deep learning to analyze the received video data and identify key points and forms of the athlete's movements. The analysis results are compared to ideal forms, and areas for improvement in the movements are evaluated.

[0726] Meanwhile, the emotion engine analyzes the player's emotional data in real time to determine, for example, whether the player is feeling fatigued or frustrated. The server adjusts the feedback based on this emotional data and generates advice tailored to the player's mental state. For example, if a player shows signs of fatigue, the server avoids giving instructions for strenuous movements and provides feedback encouraging rest.

[0727] The generated feedback is formatted as AR data for visual instructions and as audio data for audio instructions, and then sent to the device. The device's AR glasses overlay the feedback onto the athlete's field of view, and the earphones play the audio feedback in real time. This allows athletes to not only instantly correct their form but also receive advice that takes their physical and mental state into consideration, enabling more effective training.

[0728] For example, consider a situation where a user (athlete) is practicing track and field. The server analyzes the athlete's running form and evaluates knee angle and stride length. Furthermore, if the emotion engine determines from the athlete's facial expression that their concentration is declining, it will provide feedback such as "Take a short break and drink some water," and then send specific advice such as "Try running with an awareness of widening your stride." In this way, feedback is provided while taking the athlete's emotional state into consideration, significantly improving the quality of training.

[0729] The following describes the processing flow.

[0730] Step 1:

[0731] The device's camera captures the player's movements and acquires high-resolution video data. This data is set to be captured at the optimal angle to clearly capture the player's posture and movements.

[0732] Step 2:

[0733] The terminal sends the acquired video data to the server. This transmission is performed using a high-speed communication protocol to minimize latency.

[0734] Step 3:

[0735] The server stores the received video data in a buffer and prepares to perform motion analysis.

[0736] Step 4:

[0737] The server inputs video data into a deep learning model to analyze key movement points such as joint positions. This quantifies the athlete's form, making it available for evaluation.

[0738] Step 5:

[0739] Based on the analyzed data, the server compares the player's form to an ideal form and identifies areas for improvement.

[0740] Step 6:

[0741] The device's microphone captures the player's voice and sends it to the emotion engine in real time. This audio data is used to evaluate the player's emotional state.

[0742] Step 7:

[0743] The emotion engine analyzes facial expression information obtained from audio and video data to determine the emotional state of the player. This determination serves as an indicator for evaluating the player's concentration, fatigue, frustration, and other factors.

[0744] Step 8:

[0745] The server generates feedback tailored to the player's emotional state based on data from the emotion engine. This feedback takes into account the player's mental state and includes content designed to maintain motivation.

[0746] Step 9:

[0747] The server formats the generated feedback into AR display data and audio data and sends it to the device.

[0748] Step 10:

[0749] The AR glasses on the device overlay the received visual data onto the player's field of view, visually indicating specific areas for improvement.

[0750] Step 11:

[0751] The earphones in the device transmit received audio data to the players in real time, giving them voice instructions for the necessary actions.

[0752] Step 12:

[0753] Based on the feedback provided, users (athletes) can immediately correct their form and receive emotionally responsive advice, enabling efficient and mentally conscious training.

[0754] (Example 2)

[0755] Next, we will describe Example 2. In the following description, the data processing device 12 will be referred to as the "server" and the robot 414 as the "terminal".

[0756] Traditional sports training systems separate the analysis of athletes' movements from the assessment of their emotional state, making it difficult to provide integrated feedback. Furthermore, the lack of means to provide real-time feedback tailored to the athlete's real-time mental state made it challenging to maximize performance improvement.

[0757] The identification process performed by the identification processing unit 290 of the data processing device 12 in Example 2 is realized by the following means.

[0758] In this invention, the server includes means for analyzing video data to identify key points and forms of movement, means for analyzing emotional states based on audio data, and means for generating feedback based on movement and emotional states. This makes it possible to provide feedback that takes into account both the player's movement and mental state.

[0759] "Video data" refers to visual information acquired using cameras and other recording devices, and this information is used for motion analysis.

[0760] "Audio data" refers to auditory information acquired using sound-collecting devices such as microphones, and this information is used for emotion analysis.

[0761] "Terminal means" refers to an electronic device used to acquire video and audio data and transmit it to a server.

[0762] "Server equipment" refers to computer equipment used to analyze received data and generate feedback based on actions and emotional states.

[0763] "Deep learning" refers to a technology in which artificial intelligence uses large amounts of data to learn patterns and perform predictions and analyses.

[0764] An "emotion analysis engine" refers to a collection of algorithms and software used to estimate a person's emotional state based on voice data and other input information.

[0765] "Feedback" refers to suggestions for improvement and guidance that are generated by a server and presented to the recipient visually or audibly.

[0766] This invention is a system for analyzing the movements and emotions of athletes in real time and providing effective feedback. The system consists of a terminal equipped with input devices such as a camera for acquiring video data and a microphone for acquiring audio data. The terminal temporarily stores this data and then transmits it to a server.

[0767] The server runs a deep learning model to analyze the video data. This model is built using a common deep learning framework and extracts key points from the athlete's body movements, identifying their form. The analyzed motion data is then compared to the ideal form to identify specific areas for improvement.

[0768] Furthermore, the server is equipped with an emotion analysis engine for performing sentiment analysis. This engine determines the emotional state of the players based on the audio data. The algorithms used here include natural language processing and speech emotion analysis algorithms.

[0769] The generated feedback will take into account both the athlete's actions and emotions. This feedback is then transmitted to the device in real time and communicated to the athlete. Specifically, it is displayed as a visual overlay via the device's AR glasses or provided as audio instruction through earphones.

[0770] For example, let's assume a user is practicing track and field. The server analyzes the video footage taken during the run and determines that the athlete should improve their knee angle. Simultaneously, emotional analysis is performed, and if the emotional data detects that the athlete is feeling fatigued, it recommends "taking a short rest and drinking some water," and then provides specific advice such as "focus on your form and lift your knees higher." In this way, the feedback is designed to maximize its effectiveness in both the athlete's athletic performance and mental state.

[0771] An example of a prompt to input into a generative AI model is: "Please describe a system that analyzes the movements and emotions of athletes and generates feedback. Explain how it analyzes movements in real time and provides advice based on the athlete's emotions." This prompt serves as a design guideline for improving the quality of the generated feedback.

[0772] The flow of the specific processing in Example 2 will be explained using Figure 13.

[0773] Step 1:

[0774] The device uses a camera to acquire video data of athletes and a microphone to capture audio data. This data is used to record both the athletes' movements and voice in real time. The input is the athletes' video and audio, and the output is sent to the server as data packets. Specifically, the device records video at 60 frames per second and converts the continuous audio into a digital format.

[0775] Step 2:

[0776] The server receives data packets sent from the terminal. Using a deep learning model, it identifies key points and form in the athlete's movements based on the received video data. The input is video data, and the output is the motion analysis result. Specifically, calculations are performed to quantify the position and angle of the athlete's joints and compare them to the ideal form.

[0777] Step 3:

[0778] The server simultaneously feeds the received audio data into an emotion analysis engine to estimate the player's emotional state. The input is audio data, and the output is the result of the emotional state analysis. This analysis takes into account factors such as the tone and tempo of the voice and the frequency of interruptions. Specifically, the system performs speech analysis using natural language processing technology.

[0779] Step 4:

[0780] The server integrates both the motion analysis results and the emotion analysis results to generate feedback for the player. The input is the analysis results from Step 2 and Step 3, and the output is feedback including areas for improvement and advice. Here, a generative AI model is used to automatically generate the optimal feedback content.

[0781] Step 5:

[0782] The server sends the generated feedback to the device. The device provides feedback to the user visually through the AR glasses' display and audibly through the earphones. The input is the feedback data sent from the server, and the output is the visual and audible feedback to the user. Specifically, the system overlays areas for form improvement on the AR glasses and plays detailed instructions audibly through the earphones.

[0783] (Application Example 2)

[0784] Next, we will explain application example 2. In the following explanation, the data processing device 12 will be referred to as the "server" and the robot 414 as the "terminal".

[0785] In modern manufacturing, managing the work efficiency and health of factory workers is a critical challenge. In particular, there is a need to improve workers' operational efficiency while simultaneously evaluating emotional factors such as stress and fatigue in real time and providing accurate feedback based on these assessments. Conventional technologies struggle to integrate operational improvement and emotional management, posing a risk of decreased work efficiency.

[0786] The specific processing performed by the specific processing unit 290 of the data processing device 12 in Application Example 2 is realized by the following means.

[0787] In this invention, the server includes means for receiving and analyzing video information, means for acquiring audio information and analyzing emotional states, and means for generating feedback information based on performance evaluation and emotional states. This makes it possible to provide factory workers with real-time instructions for correcting their movements and suggestions for appropriate rest and work paces.

[0788] "Visual information" refers to visual data acquired from devices such as cameras, and is used to analyze the movements and postures of subjects.

[0789] "Device means" refers to hardware or software components used to acquire, process, or present data for a specific function or purpose.

[0790] "Processing device" refers to a computer or server used to analyze acquired data and perform actions or recognize emotions.

[0791] "Audio information" refers to audio data collected through devices such as microphones, and is used to analyze the emotional state of a subject based on their speech and tone of voice.

[0792] "Emotional analysis device means" refers to a software or hardware system for estimating and analyzing a subject's emotional state from audio information or visual data.

[0793] "Feedback information" refers to instructions and advice generated based on performance evaluations and emotional states, with the aim of improving the subject's performance and maintaining a comfortable work environment.

[0794] This system is used in factories to manage and improve the work efficiency and health of workers. The system acquires data through hardware devices such as smart glasses, Bluetooth earphones, cameras, and microphones. A server processes this data on a cloud platform. Specifically, the server is built on cloud infrastructure such as AWS or Google Cloud.

[0795] First, the smart glasses worn by the user capture the worker's movements in real time using their camera. The video information is sent to a server, where motion analysis is performed using software such as TensorFlow and PiTouch. The server uses deep learning technology to extract key points of the movements and compares them to ideal movement forms.

[0796] Meanwhile, voice information is also acquired through the microphone and sent to the server. The server uses an emotion analysis engine, such as the Azure Emotion API, to evaluate emotional states such as stress and fatigue from the voice information. This evaluation result influences feedback in real time.

[0797] The server integrates motion analysis and emotion assessment to generate optimal feedback information for factory workers. This feedback information is visually overlaid on smart glasses, and voice guidance is provided through earphones.

[0798] For example, if an incorrect posture is detected during assembly work, instructions such as "Raise your elbows slightly and apply force safely" will be displayed in real time on the smart glasses. Also, if the emotion analysis engine detects high stress levels, it will suggest "Take a 5-minute break" through the earphones.

[0799] This system optimizes worker productivity and resource utilization, protecting them from excessive stress and fatigue. The overall effect of this implementation makes the factory environment a safer and more efficient place.

[0800] An example of a prompt message would be, "Please tell us about any moments during your work today when you found the feedback provided by your smart glasses particularly helpful."

[0801] The flow of a specific process in Application Example 2 will be explained using Figure 14.

[0802] Step 1:

[0803] The device acquires video information. The camera on the smart glasses worn by the user captures the movements of factory workers in real time and generates video information based on that. The input is video frames captured by the camera, and the output is a data stream of video information.

[0804] Step 2:

[0805] The device sends the acquired video information to the server. The device uploads the captured video information to a cloud-based server in real time. The input is the video information generated in step 1, and the output is the data stream sent to the server.

[0806] Step 3:

[0807] The server analyzes the video information. The server uses TensorFlow and PITach to analyze the video information and identify key points of the worker's movements. The input is the video information sent to the server, and the output is key point data of the movements. Specific movements evaluated include posture, joint angles, and stride length.

[0808] Step 4:

[0809] The terminal acquires audio information. The microphone built into the earphone worn by the user captures the worker's voice. The input is the audio signal picked up by the microphone, and the output is a data stream of the audio information.

[0810] Step 5:

[0811] The device sends the acquired audio information to the server. The device sends the captured audio information to a cloud-based server. The input is the audio information generated in step 4, and the output is the data stream sent to the server.

[0812] Step 6:

[0813] The server analyzes the audio information. The server uses the Azure Emotion API and other tools to analyze the audio information and identify the worker's emotional state. The input is the audio information sent to the server, and the output is emotional state data. Specific values ​​evaluated include stress levels and fatigue levels.

[0814] Step 7:

[0815] The server generates feedback information. The server integrates action keypoint data and emotional state data to generate optimal feedback information for factory workers. The input is the output data from steps 3 and 6, and the output is the feedback information. Specific actions include action modification instructions and break suggestions that take safety and efficiency into consideration.

[0816] Step 8:

[0817] The device presents feedback information to the user. The device visually overlays the generated feedback information on the smart glasses and communicates it audibly through the earphones. The input is the feedback information generated in step 7, and the output is the visual and audible feedback provided to the user.

[0818] The specific processing unit 290 transmits the result of the specific processing to the robot 414. In the robot 414, the control unit 46A causes the speaker 240 and the controlled object 443 to output the result of the specific processing. The microphone 238 acquires audio indicating user input for the result of the specific processing. The control unit 46A transmits the audio data indicating user input acquired by the microphone 238 to the data processing unit 12. In the data processing unit 12, the specific processing unit 290 acquires the audio data.

[0819] Data generation model 58 is a type of so-called generative AI (Artificial Intelligence). An example of data generation model 58 is ChatGPT (Internet Search<URL: https: / / openai.com / blog / chatgpt> ), Gemini (Internet search) <url: https: gemini.google.com ?hl="ja">Examples of generative AI include the following. The data generation model 58 is obtained by performing deep learning on a neural network. The data generation model 58 is input with prompts containing instructions, and with inference data such as audio data representing speech, text data representing text, and image data representing images. The data generation model 58 infers from the input inference data according to the instructions indicated by the prompts, and outputs the inference results in data formats such as audio data and text data. Here, inference refers to, for example, analysis, classification, prediction, and / or summarization.

[0820] In the above embodiment, an example was given in which the specific processing is performed by the data processing device 12, but the technology of this disclosure is not limited thereto, and the specific processing may also be performed by the robot 414.

[0821] Furthermore, the emotion identification model 59, acting as an emotion engine, may determine the user's emotion according to a specific mapping. Specifically, the emotion identification model 59 may determine the user's emotion according to a specific mapping, which is an emotion map (see Figure 9). Similarly, the emotion identification model 59 may also determine the robot's emotion, and the identification processing unit 290 may perform identification processing using the robot's emotion.

[0822] Figure 9 shows an emotion map 400 in which multiple emotions are mapped. In the emotion map 400, emotions are arranged in concentric circles radiating from the center. The closer to the center of the concentric circles, the more primitive the emotions are located. Further out of the concentric circles, emotions representing states and actions arising from mental states are located. Emotion is a concept that includes feelings and mental states. On the left side of the concentric circles, emotions that are generally generated from reactions occurring in the brain are located. On the right side of the concentric circles, emotions that are generally induced by situational judgment are located. Above and below the concentric circles, emotions that are generally generated from reactions occurring in the brain and induced by situational judgment are located. In addition, the emotion of "pleasure" is located on the upper side of the concentric circles, and the emotion of "displeasure" is located on the lower side. Thus, in the emotion map 400, multiple emotions are mapped based on the structure in which emotions arise, and emotions that are likely to occur simultaneously are mapped close together.

[0823] These emotions are distributed at the 3 o'clock position on the Emotion Map 400, and usually fluctuate between feelings of security and anxiety. In the right half of the Emotion Map 400, situational awareness takes precedence over internal feelings, resulting in a calm impression.

[0824] The inside of the Emotion Map 400 represents inner thoughts, while the outside represents actions. Therefore, the further you go from the outside of the Emotion Map 400, the more visible (expressed in actions) your emotions become.

[0825] Here, human emotions are based on various balances, such as posture and blood sugar levels. When these balances deviate from the ideal, it results in discomfort, and when they approach the ideal, it results in pleasure. Similarly, in robots, cars, motorcycles, etc., emotions can be created based on various balances, such as posture and battery level. When these balances deviate from the ideal, it results in discomfort, and when they approach the ideal, it results in pleasure. The emotion map can be generated, for example, based on Dr. Mitsuyoshi's emotion map (Research on a system for analyzing brain physiological signals of speech emotion recognition and emotion, Tokushima University, doctoral dissertation: https: / / ci.nii.ac.jp / naid / 500000375379). The left half of the emotion map contains emotions belonging to a region called "response," where sensation is dominant. The right half of the emotion map contains emotions belonging to a region called "situation," where situational awareness is dominant.

[0826] The emotion map defines two emotions that promote learning. One is the emotion around the middle of the negative "repentance" and "reflection" on the situation side. In other words, it is when the robot experiences negative emotions such as "I never want to feel this way again" or "I don't want to be scolded again." The other is the emotion around the positive "desire" on the reaction side. In other words, it is when the robot has positive feelings such as "I want more" or "I want to know more."

[0827] The emotion identification model 59 inputs user input into a pre-trained neural network, obtains emotion values ​​representing each emotion shown in the emotion map 400, and determines the user's emotion. This neural network is pre-trained based on multiple training data sets, which are combinations of user input and emotion values ​​representing each emotion shown in the emotion map 400. Furthermore, this neural network is trained so that emotions located close together have similar values, as shown in the emotion map 900 in Figure 10. Figure 10 shows an example where multiple emotions such as "reassured," "calm," and "confident" have similar emotion values.

[0828] The above description primarily focuses on the functions of the data processing device 12 in relation to this disclosure. However, the system related to this disclosure is not necessarily implemented on a server. The system related to this disclosure may be implemented as a general information processing system. This disclosure may be implemented, for example, as a software program that runs on a personal computer or as an application that runs on a smartphone. The method related to this disclosure may be provided to users in SaaS (Software as a Service) format.

[0829] In the above embodiment, an example was given in which a specific process is performed by a single computer 22. However, the technology of this disclosure is not limited thereto, and a distributed processing of the specific process may be performed by multiple computers, including computer 22. For example, a data generation model 58 may be provided in an external device of the data processing device 12, and the external device may generate data according to the input data.

[0830] In the above embodiment, an example was given in which the specific processing program 56 is stored in the storage 32, but the technology of this disclosure is not limited thereto. For example, the specific processing program 56 may be stored in a portable, computer-readable, non-temporary storage medium such as a USB (Universal Serial Bus) memory. The specific processing program 56 stored in the non-temporary storage medium is installed in the computer 22 of the data processing device 12. The processor 28 executes specific processing according to the specific processing program 56.

[0831] Alternatively, the specific processing program 56 may be stored in a storage device such as a server connected to the data processing device 12 via the network 54, and the specific processing program 56 may be downloaded and installed on the computer 22 in response to a request from the data processing device 12.

[0832] Furthermore, it is not necessary to store the entirety of the specific processing program 56 in a storage device such as a server connected to the data processing device 12 via the network 54, or to store the entirety of the specific processing program 56 in the storage 32; it is acceptable to store only a portion of the specific processing program 56.

[0833] The following types of processors can be used as hardware resources to perform specific processing. Examples of processors include a CPU, a general-purpose processor that functions as a hardware resource to perform specific processing by executing software, i.e., a program. Other examples of processors include dedicated electrical circuits, such as FPGAs (Field-Programmable Gate Arrays), PLDs (Programmable Logic Devices), or ASICs (Application Specific Integrated Circuits), which have circuit configurations specifically designed to perform specific processing. All of these processors have built-in or connected memory, and all of them perform specific processing by using memory.

[0834] The hardware resource that performs a specific process may consist of one of these various processors, or it may consist of a combination of two or more processors of the same or different types (for example, a combination of multiple FPGAs, or a combination of a CPU and an FPGA). Alternatively, the hardware resource that performs a specific process may consist of a single processor.

[0835] Examples of configurations using a single processor include, firstly, a configuration in which one or more CPUs and software are combined to form a single processor, and this processor functions as a hardware resource that performs a specific process. Secondly, there is a configuration using a processor that realizes the functions of the entire system, including multiple hardware resources that perform a specific process, on a single IC chip, as exemplified by SoCs (System-on-a-chip). In this way, a specific process is realized using one or more of the above types of processors as hardware resources.

[0836] Furthermore, the hardware structure of these various processors can more specifically utilize electrical circuits that combine circuit elements such as semiconductor devices. Also, the specific processing described above is merely an example. Therefore, it goes without saying that unnecessary steps can be deleted, new steps added, or the processing order rearranged, as long as it does not deviate from the main purpose.

[0837] The descriptions and illustrations presented above are detailed explanations of the technical aspects of this disclosure and are merely examples of the technical aspects. For example, the above descriptions of the structure, function, operation, and effect are examples of the structure, function, operation, and effect of the technical aspects of this disclosure. Therefore, it goes without saying that you may delete unnecessary parts, add new elements, or replace elements in the descriptions and illustrations presented above, as long as you do not deviate from the essence of the technical aspects of this disclosure. Furthermore, in order to avoid confusion and facilitate understanding of the technical aspects of this disclosure, explanations of common technical knowledge and the like that do not require special explanation to enable the implementation of the technical aspects of this disclosure have been omitted from the descriptions and illustrations presented above.

[0838] All documents, patent applications, and technical standards described herein are incorporated by reference to the same extent as if each individual document, patent application, and technical standard were specifically and individually noted to be incorporated by reference.

[0839] The following is further disclosed regarding the embodiments described above.

[0840] (Claim 1)

[0841] A terminal means for acquiring video data,

[0842] A server means that receives the aforementioned video data, analyzes the data, and evaluates the actions of the subject,

[0843] A server means for generating data to provide feedback to the subject based on the aforementioned evaluation,

[0844] A terminal means for presenting the generated data to the target person,

[0845] A system that includes this.

[0846] (Claim 2)

[0847] The system according to claim 1, characterized in that the feedback is provided as visual data and audio data.

[0848] (Claim 3)

[0849] The system according to claim 1, characterized in that the server means generates customized feedback based on the subject's profile information.

[0850] "Example 1"

[0851] (Claim 1)

[0852] A device for acquiring video,

[0853] A processing device that receives the aforementioned video, analyzes it using a generated AI model, and evaluates the movements of a person,

[0854] A processing device that generates customized feedback based on the aforementioned evaluation and personal information of the person,

[0855] A device that presents the generated feedback as visual and auditory information,

[0856] A system that includes this.

[0857] (Claim 2)

[0858] The system according to claim 1, characterized in that the generated feedback is provided in real time visually and audibly using prompt statements.

[0859] (Claim 3)

[0860] The system according to claim 1, characterized in that the processing device outputs immediate visual and audio instructions in real time based on motion analysis to support the correction of a person's movements.

[0861] "Application Example 1"

[0862] (Claim 1)

[0863] A device for acquiring video information,

[0864] A computing device that receives the aforementioned video information, analyzes the information, and evaluates the actions of the subject,

[0865] A computing device that generates information to provide feedback to the subject based on the aforementioned evaluation,

[0866] A device that presents the generated information to the target person,

[0867] The aforementioned device is a system for supporting the improvement of the subject's movements.

[0868] (Claim 2)

[0869] The system according to claim 1, characterized in that the feedback is provided as visual information and audio information.

[0870] (Claim 3)

[0871] The system according to claim 1, characterized in that the computing device generates customized feedback based on the attribute information of the subject.

[0872] "Example 2 of combining an emotion engine"

[0873] (Claim 1)

[0874] A terminal means for acquiring video data and audio data,

[0875] A server means that uses deep learning to analyze the aforementioned video data and identify key points and forms of the subject's movements,

[0876] A server means including an emotion analysis engine for analyzing the emotional state of the subject based on the aforementioned audio data,

[0877] A server means that generates feedback including areas for improvement based on the aforementioned actions and emotional states,

[0878] A terminal means for visually and audibly presenting the generated feedback,

[0879] A system that includes this.

[0880] (Claim 2)

[0881] The system according to claim 1, characterized in that the feedback is adjusted according to the real-time mental state of the subject.

[0882] (Claim 3)

[0883] The system according to claim 1, characterized in that the server means integrates the subject's motion data and emotional data and generates individually appropriate feedback.

[0884] "Application example 2 when combining with an emotional engine"

[0885] (Claim 1)

[0886] A device for acquiring video information,

[0887] Processing means that receives the aforementioned video information, analyzes the information, and evaluates the actions of the subject,

[0888] An emotion analysis device means that acquires voice information of a subject and analyzes their emotional state,

[0889] A processing device that generates information for providing feedback to a subject based on performance evaluation and emotional state,

[0890] A device means for presenting the generated information to the target person,

[0891] A system that includes this.

[0892] (Claim 2)

[0893] The system according to claim 1, characterized in that the feedback is provided as visual information and audio information.

[0894] (Claim 3)

[0895] The system according to claim 1, characterized in that the processing device generates customized feedback based on the subject's personal information, thereby supporting improved work efficiency and health management. [Explanation of Symbols]

[0896] 10, 210, 310, 410 Data Processing Systems 12 Data Processing Devices 14 Smart Devices 214 Smart Glasses 314 Headset-type terminal 414 Robots< / url:> < / url:> < / url:> < / url:>

Claims

1. A device for acquiring video information, A device that receives the aforementioned video information, analyzes the information, and evaluates the actions of the subject, Based on the aforementioned evaluation, a device for generating information to provide feedback to the subject, A device that presents the generated information to the target person, A system that includes measures to support the improvement of the subject's movements.

2. The system according to claim 1, characterized in that the feedback is provided as visual information and audio information.

3. The system according to claim 1, characterized in that it generates customized feedback based on the attribute information of the subject.

Citation Information

Patent Citations

  • Persona chatbot control method and system

    JP2022180282A