system
The system automates athlete feedback through video analysis and generative models, addressing the inefficiencies of conventional methods by providing immediate and detailed guidance for improved performance.
Patent Information
- Authority / Receiving Office
- JP · JP
- Patent Type
- Applications
- Current Assignee / Owner
- SOFTBANK GROUP CORP
- Filing Date
- 2024-10-18
- Publication Date
- 2026-05-01
AI Technical Summary
Conventional methods struggle to provide detailed and efficient feedback to athletes, particularly at the primary education stage, leading to increased coach burden and reduced quality of individual guidance.
A system utilizing a video acquisition device, generative model, and communication network to analyze athlete movements, generate precise feedback, and transmit it to external terminals for immediate improvement guidance.
Enables rapid and effective feedback, reducing the need for coach intervention and improving athlete performance by automating detailed technical and strategic analysis.
Smart Images

Figure 2026073348000001_ABST
Abstract
Description
Technical Field
[0001] The technology of the present disclosure relates to a system.
Background Art
[0002] Patent Document 1 discloses a method for controlling a persona chatbot performed by at least one processor, the method including steps of receiving a user utterance, adding the user utterance to a prompt including an instruction sentence related to an explanation of a character of the chatbot, encoding the prompt, and inputting the encoded prompt into a language model to generate a chatbot utterance as a response to the user utterance.
Prior Art Documents
Patent Documents
[0003]
Patent Document 1
Summary of the Invention
Problems to be Solved by the Invention
[0004] In order for athletes to efficiently improve their skills within limited practice time, precise feedback is necessary. However, there is a problem in that it is difficult to provide detailed feedback by a coach to all athletes by conventional methods. This problem is particularly prominent in athletes at the primary education stage. Also, since the burden on the coach increases and the quality of individual guidance may decline, an efficient and effective guidance support system is required.
Means for Solving the Problems
[0005] This invention provides a method for automatically analyzing the actions of athletes by using a generation model to analyze video footage of a competition acquired by a video acquisition device. Based on the analysis results, the method evaluates the athletes' techniques and strategies and generates specific feedback. Furthermore, by transmitting this feedback to an external terminal via a communication network, athletes and coaches can quickly and easily identify areas for improvement. This significantly reduces the need for coaches to provide individual feedback, enabling higher quality instruction.
[0006] A "video acquisition device" is a device used to collect video footage of matches or competitions in real time or as a record, and includes cameras, smartphones, and other similar devices.
[0007] A "generative model" is an artificial intelligence model used to analyze input data and automatically recognize specific patterns or features, and includes models that incorporate deep learning.
[0008] "Movement" refers to a series of physical activities performed by an athlete during a competition, such as body movement, positioning, and shot swing.
[0009] "Analysis" refers to the process of extracting specific information from input video data and evaluating or analyzing that information according to a specific purpose.
[0010] "Feedback" refers to specific advice and information provided to athletes for improvement, based on the results obtained from the analysis.
[0011] A "communication network" refers to a connection system designed to transmit information, including the internet, which utilizes wired or wireless technologies.
[0012] "External devices" refer to devices that can receive analysis results and feedback, and include smartphones, tablets, and personal computers. [Brief explanation of the drawing]
[0013] [Figure 1] This is a conceptual diagram showing an example of the configuration of a data processing system according to the first embodiment. [Figure 2] This is a conceptual diagram showing an example of the essential functions of a data processing device and a smart device according to the first embodiment. [Figure 3] This is a conceptual diagram showing an example of the configuration of a data processing system according to the second embodiment. [Figure 4] This is a conceptual diagram showing an example of the main functions of a data processing device and smart glasses according to the second embodiment. [Figure 5] This is a conceptual diagram showing an example of the configuration of a data processing system according to the third embodiment. [Figure 6] This is a conceptual diagram showing an example of the main functions of a data processing device and a headset-type terminal according to the third embodiment. [Figure 7] This is a conceptual diagram showing an example of the configuration of a data processing system according to the fourth embodiment. [Figure 8] This is a conceptual diagram showing an example of the main functions of a data processing device and a robot according to the fourth embodiment. [Figure 9] This shows an emotion map where multiple emotions are mapped. [Figure 10] This shows an emotion map where multiple emotions are mapped. [Figure 11] This is a sequence diagram showing the processing flow of the data processing system in Example 1. [Figure 12] This is a sequence diagram showing the processing flow of the data processing system in Application Example 1. [Figure 13] This is a sequence diagram showing the processing flow of the data processing system in Example 2, which incorporates an emotion engine. [Figure 14] This is a sequence diagram showing the processing flow of the data processing system in Application Example 2, which combines an emotion engine. [Modes for carrying out the invention]
[0014] Hereinafter, an example of an embodiment of a system according to the technology of the present disclosure will be described with reference to the accompanying drawings.
[0015] First, the terms used in the following description will be explained.
[0016] In the following embodiments, the numbered processor (hereinafter simply referred to as "processor") may be one arithmetic unit or a combination of multiple arithmetic units. Also, the processor may be one type of arithmetic unit or a combination of multiple types of arithmetic units. Examples of arithmetic units include a CPU (Central Processing Unit), a GPU (Graphics Processing Unit), a GPGPU (General-Purpose computing on Graphics Processing Units), an APU (Accelerated Processing Unit), and the like.
[0017] In the following embodiments, the numbered RAM (Random Access Memory) is a memory in which information is temporarily stored and is used as a work memory by the processor.
[0018] In the following embodiments, the numbered storage is one or more non-volatile storage devices that store various programs and various parameters, etc. Examples of non-volatile storage devices include flash memory (SSD (Solid State Drive)), magnetic disks (e.g., hard disks), or magnetic tapes, etc.
[0019] In the following embodiments, the signed communication interface (I / F) is an interface that includes a communication processor and an antenna, etc. The communication interface manages communication between multiple computers. Examples of communication standards applicable to the communication interface include wireless communication standards such as 5G (5th Generation Mobile Communication System), Wi-Fi (registered trademark), or Bluetooth (registered trademark).
[0020] In the following embodiments, "A and / or B" is synonymous with "at least one of A and B." That is, "A and / or B" means that it may be A alone, or B alone, or a combination of A and B. Furthermore, in this specification, the same concept as "A and / or B" applies when expressing three or more things linked by "and / or."
[0021] [First Embodiment]
[0022] Figure 1 shows an example of the configuration of the data processing system 10 according to the first embodiment.
[0023] As shown in Figure 1, the data processing system 10 includes a data processing device 12 and a smart device 14. An example of the data processing device 12 is a server.
[0024] The data processing device 12 comprises a computer 22, a database 24, and a communication interface 26. The computer 22 is an example of a "computer" related to the technology of this disclosure. The computer 22 comprises a processor 28, RAM 30, and storage 32. The processor 28, RAM 30, and storage 32 are connected to a bus 34. The database 24 and the communication interface 26 are also connected to the bus 34. The communication interface 26 is connected to a network 54. An example of the network 54 is a WAN (Wide Area Network) and / or a LAN (Local Area Network).
[0025] The smart device 14 comprises a computer 36, a reception device 38, an output device 40, a camera 42, and a communication interface 44. The computer 36 comprises a processor 46, RAM 48, and storage 50. The processor 46, RAM 48, and storage 50 are connected to a bus 52. The reception device 38, output device 40, and camera 42 are also connected to the bus 52.
[0026] The reception device 38 is equipped with a touch panel 38A and a microphone 38B, etc., and receives user input. The touch panel 38A receives user input by detecting contact with an object (e.g., a pen or finger). The microphone 38B receives user input by detecting the user's voice. The control unit 46A transmits data indicating the user input received by the touch panel 38A and microphone 38B to the data processing device 12. In the data processing device 12, the specific processing unit 290 acquires the data indicating the user input.
[0027] The output device 40 includes a display 40A and a speaker 40B, and presents data to the user 20 by outputting the data in a form perceptible to the user 20 (e.g., audio and / or text). The display 40A displays visible information such as text and images according to instructions from the processor 46. The speaker 40B outputs audio according to instructions from the processor 46. The camera 42 is a small digital camera equipped with an optical system such as a lens, aperture, and shutter, and an image sensor such as a CMOS (Complementary Metal-Oxide-Semiconductor) image sensor or a CCD (Charge Coupled Device) image sensor.
[0028] Communication interface 44 is connected to network 54. Communication interfaces 44 and 26 are responsible for the exchange of various types of information between processor 46 and processor 28 via network 54.
[0029] Figure 2 shows an example of the main functions of the data processing device 12 and the smart device 14.
[0030] As shown in Figure 2, in the data processing device 12, a specific processing is performed by the processor 28. A specific processing program 56 is stored in the storage 32. The specific processing program 56 is an example of a "program" related to the technology of this disclosure. The processor 28 reads the specific processing program 56 from the storage 32 and executes the read specific processing program 56 on the RAM 30. The specific processing is realized by the processor 28 operating as a specific processing unit 290 according to the specific processing program 56 executed on the RAM 30.
[0031] The storage 32 stores the data generation model 58 and the emotion identification model 59. The data generation model 58 and the emotion identification model 59 are used by the identification processing unit 290.
[0032] In the smart device 14, the processor 46 performs the reception output processing. The storage 50 stores the reception output program 60. The reception output program 60 is used in conjunction with a specific processing program 56 by the data processing system 10. The processor 46 reads the reception output program 60 from the storage 50 and executes the read reception output program 60 on the RAM 48. The reception output processing is realized by the processor 46 operating as a control unit 46A according to the reception output program 60 executed on the RAM 48.
[0033] Next, the specific processing performed by the specific processing unit 290 of the data processing device 12 will be described. In the following description, the data processing device 12 will be referred to as the "server" and the smart device 14 as the "terminal".
[0034] This invention relates to an automated analysis system for efficiently improving the technical and strategic capabilities of athletes. This system utilizes the functions of a video acquisition device, a generative model, a communication network, and an external terminal to provide immediate and effective feedback to athletes.
[0035] System Configuration
[0036] 1. Video acquisition device
[0037] Terminal: Using smartphones or dedicated cameras, video data from matches is acquired in real time. This video acquisition device is installed in stadiums and training grounds to record the movements of athletes in detail.
[0038] 2. Transmission of video data
[0039] Terminal: Acquired video data is transmitted to the server using the 5G communication network. Even high-capacity data is transferred quickly, enabling real-time analysis.
[0040] 3. Conduct motion analysis
[0041] Server: The transmitted video data is analyzed by a generative model. The model tracks the athlete's body position and movements, and evaluates their actions and technical elements during the competition. For example, it detects the speed and angle of the athlete's racket swing and compares it to an ideal movement pattern.
[0042] 4. Feedback generation
[0043] Server: Based on the analysis results, it generates individual feedback for each athlete. This feedback evaluates areas for improvement and successful techniques for each athlete, and provides them as goals for further practice.
[0044] 5. Communication and display of feedback
[0045] Device: The generated feedback is sent via the communication network to external devices used by athletes and coaches. Users can review this feedback on their smartphones or tablets and use it to improve their next training session.
[0046] Specific examples of use
[0047] User: If a badminton player's smash swing is incorrect during a match, the system instantly generates feedback on the angle of that swing. The coach, upon receiving this feedback, can then focus their instruction on improving specific forms during the next practice session.
[0048] This invention realizes a system that provides precise technical and strategic feedback through clear image analysis, thereby promoting efficient performance improvement in competitions.
[0049] The following describes the processing flow.
[0050] Step 1:
[0051] Devices: Before the start of the competition, smartphones or dedicated cameras are set up courtside to begin recording the match footage in real time. These devices continuously record the players' movements clearly in high resolution.
[0052] Step 2:
[0053] Terminal: Captured video data is streamed to the server using the 5G network. This allows large amounts of video data to reach the server quickly.
[0054] Step 3:
[0055] Server: The server divides the received video data into frames and performs preprocessing to extract the players' movements and positions from each frame. This process also includes filtering to reduce background noise and improve analysis accuracy.
[0056] Step 4:
[0057] Server: Uses a generative model to analyze the player's technical actions in each frame (e.g., racket swing, player positioning). At this stage, the player's actions are compared to a specific set of rules and criteria to calculate technical evaluation metrics.
[0058] Step 5:
[0059] Server: Based on technical evaluation, it assesses the strategic approach of players. It analyzes the strategies players employ during matches, detecting, for example, delays in defensive movement or positioning during attacks.
[0060] Step 6:
[0061] Server: Based on the analysis results, create specific feedback on areas that need improvement and skills that should be strengthened. This feedback will include technical advice and strategic suggestions.
[0062] Step 7:
[0063] Terminal: Generated feedback is sent to an external terminal managed by the athlete or coach. Users can check this in real time or immediately after the match.
[0064] Step 8:
[0065] User: Based on the feedback, the coach understands areas where the athlete needs improvement and determines the focus points for the next training session. This feedback allows for efficient guidance to improve the athlete's skills and strategies.
[0066] (Example 1)
[0067] Next, we will describe Example 1. In the following description, the data processing device 12 will be referred to as the "server," and the smart device 14 will be referred to as the "terminal."
[0068] The present invention aims to provide a system that offers immediate, highly accurate feedback to efficiently improve athletes' technical skills and strategies. Conventional technologies have made real-time motion analysis and immediate identification of specific areas for improvement difficult, preventing rapid reflection of improvements in practice and competition performance. Therefore, there is a need for a system that automates the analysis of athletes' movements and efficiently provides feedback, including specific improvement suggestions.
[0069] The identification process performed by the identification processing unit 290 of the data processing device 12 in Example 1 is realized by the following means.
[0070] In this invention, the server includes means for taking competition video acquired by a video acquisition device as input and analyzing the athlete's posture from the video using a generative algorithm for analyzing the video; means for evaluating the athlete's technical and strategic elements based on the analysis results and generating feedback including specific improvement suggestions; and means for transmitting the feedback to an external device via a communication path. This allows athletes to immediately grasp areas for improvement in their movements and use that information to improve their technical skills and strategies in their next practice or match.
[0071] "Video acquisition equipment" refers to devices that can capture the actions of athletes and the details of a match in real time, and includes smartphones and dedicated cameras.
[0072] A "generative algorithm" is a mathematical model or method for analyzing input video data, particularly for tracking and evaluating the physical movements of athletes.
[0073] A "communication path" refers to the network used when sending and receiving data, and includes 5G networks and other technologies that enable high-speed and stable communication.
[0074] "External devices" refer to devices used by athletes and coaches to receive feedback, and include terminals such as smartphones and tablets.
[0075] "Feedback" refers to data and advice provided to improve athletes' skills and strategies, based on results analyzed by generative algorithms.
[0076] "Analyzing posture" is the process of tracking an athlete's body position and movements, and using that data to evaluate specific technical and strategic elements.
[0077] "Technical elements" are indicators that show the accuracy and efficiency of individual movements and techniques in a competition, and are used to numerically evaluate an athlete's performance.
[0078] "Strategic elements" are criteria used to evaluate the strategies and tactics that athletes choose during a match, particularly analyzing their reactions to the opponent's movements and their play choices.
[0079] This invention provides a system that effectively analyzes the movements of athletes and provides immediate, specific feedback. This system is primarily implemented using video acquisition equipment, a generative AI model, a communication path, and external devices. Details and specific examples of each element are provided below.
[0080] Video acquisition equipment is a device used to capture the posture and movements of athletes in real time while they are playing. Specifically, this includes smartphones and dedicated cameras. These devices are installed in stadiums and practice fields to record detailed video data.
[0081] The server receives the acquired video data and analyzes it using a generative AI model. The generative AI model tracks the athletes' body positions and movements within the video data and evaluates the technical and strategic elements during the competition. Specifically, it identifies the joint positions and movement speeds of the athletes' bodies and compares them to ideal movement patterns.
[0082] The server then automatically generates feedback based on the analysis results. This feedback includes guidance on areas for improvement in the athlete's movements and specific strategies. For example, it might suggest detailed improvements such as, "The angle of your racket during a smash is 5 degrees shallower than ideal."
[0083] The terminal transmits feedback generated by the server to an external device via a communication path. The user receives and confirms the feedback through an external device such as a smartphone or tablet. This makes it possible to immediately implement specific improvement measures during the next practice session.
[0084] Based on the feedback provided, users can improve their technology and adjust their strategies. Through this process, competitors can quickly improve their performance.
[0085] Example of a prompt:
[0086] "Analyze the angle of the player's smash swing and generate feedback showing the difference from the ideal angle."
[0087] In this way, the system automates a series of processes including video acquisition, data analysis, feedback generation, transmission, and reception, providing strong support for the technical and strategic improvement of athletes.
[0088] The flow of the specific processing in Example 1 will be explained using Figure 11.
[0089] Step 1:
[0090] The terminal uses video acquisition equipment to capture the movements of athletes. Smartphones or dedicated cameras are set up to film movements in real time at the stadium or training ground. The input is video data including the movements of athletes, which is recorded in digital format in preparation for transmission to the server. Specifically, the movements of athletes, the operation of equipment, and interactions with the environment are captured on video.
[0091] Step 2:
[0092] The terminal transfers the captured video data to the server via the communication path. High-capacity video data is transmitted rapidly using a 5G network. The input is the video data acquired in step 1, and the output is the compressed and decompressed video data arriving at the server. At this stage, communication optimization is performed to minimize delays during data transfer.
[0093] Step 3:
[0094] The server inputs the received video data into the generating AI model and begins analyzing the data. The input is the video data sent in step 2, and the output is the evaluation results of the athlete's body position and movement. Specifically, the process involves estimating the athlete's joint positions, calculating the speed and direction of movement, and comparing it with ideal movement for each frame.
[0095] Step 4:
[0096] The server generates feedback for the athlete based on the analysis results. The input is the evaluation results obtained in step 3, and the output is feedback data that includes specific improvement suggestions and quantified metrics. For example, it may include detailed analysis results regarding the athlete's swing angle and speed.
[0097] Step 5:
[0098] The terminal transmits the generated feedback to an external device via a communication path. The input is the feedback data received from the server, and the output is the information presented in a format that the user can verify. At this stage, the feedback is transmitted to the user via a smartphone or tablet.
[0099] Step 6:
[0100] Users utilize the feedback they receive to improve their movements and tactics in competition. Specifically, they adjust their training content or try new strategies based on the feedback. This improves the athletes' technical skills and strategic thinking.
[0101] (Application Example 1)
[0102] Next, we will explain Application Example 1. In the following explanation, the data processing device 12 will be referred to as the "server," and the smart device 14 will be referred to as the "terminal."
[0103] Improving the operational efficiency of various machines is a crucial challenge in modern manufacturing environments. However, there are limited means to constantly monitor machine operation and provide appropriate feedback in real time. Conventional analysis methods make it difficult to quickly and accurately analyze machine operation data and generate appropriate improvement suggestions. Furthermore, even in situations where immediate implementation of improvement suggestions is required, the lack of appropriate information and communication means hinders efficient improvement.
[0104] The specific processing performed by the specific processing unit 290 of the data processing device 12 in Application Example 1 is realized by the following means.
[0105] In this invention, the server includes means for acquiring video footage of the machine's operation using a video acquisition mechanism, means for analyzing the operation using a generation algorithm and evaluating the machine's operation, and means for transmitting improvement suggestions to an external device via an information communication channel. This makes it possible to analyze the machine's operation in real time and quickly provide suggestions for efficient operation improvements.
[0106] A "video acquisition mechanism" is a device or system used to acquire video footage of a machine's operation, and plays a role in capturing the details of its movement.
[0107] A "generative algorithm" is a computational method or program used to analyze the movement of a machine from video footage and evaluate its position and activity.
[0108] "Means for analyzing motion" refers to processes and methods for analyzing video footage of machine movements and evaluating efficient operation.
[0109] An "improvement suggestion" is information that, based on the analysis results, presents specific actions to improve the machine's operation and increase efficiency.
[0110] "Information communication channels" refer to communication paths and network technologies used to transmit improvement suggestions to external devices.
[0111] An "external device" is a computer or mobile device used to receive improvement suggestions sent from the server.
[0112] The system that realizes this invention consists of a combination of an image acquisition mechanism, a generation algorithm, an information communication channel, and external devices. This makes it possible to efficiently monitor and analyze the operation of machinery in a factory and provide improvement suggestions.
[0113] The server uses dedicated cameras to capture video footage of the machinery in operation within the factory. High-resolution cameras can record various machine movements in detail. This allows for precise analysis without any missing information.
[0114] After acquiring video footage, the server uses a generation algorithm to analyze the captured motion video. This algorithm tracks the machine's position and evaluates the efficiency and precision of its movements. For example, it can identify problems such as suboptimal machine movement or inappropriate work speed. TENSORFLOW® and PyTorch can be used as generation algorithms.
[0115] Based on the analysis results, the server generates improvement suggestions. These include areas for improvement in machine operation and action plans for more efficient operation. These suggestions are output in a format that is easy for the user to understand.
[0116] External devices such as smart glasses and tablet terminals are used. Generated improvement suggestions are transmitted to these external devices via a communication channel, allowing the user to view them in real time. As a result, the user can quickly work on improving the machine's operation.
[0117] As a concrete example, suppose a camera installed on a factory robot arm captures and analyzes its movements, revealing that its trajectory is not optimal. The generated AI model then generates specific trajectory correction suggestions to improve the efficiency of the movement, and presents them as a prompt message such as, "Analyze the current movement of the robot arm and compare it to the most efficient movement pattern. Generate feedback on areas for improvement and suggested improvements, and present them to the worker."
[0118] In this way, the machine's operation can always be kept in an optimal state, leading to increased productivity.
[0119] The flow of a specific process in Application Example 1 will be explained using Figure 12.
[0120] Step 1:
[0121] The server uses dedicated cameras to capture video footage of machinery in operation within the factory. By using high-resolution cameras, even subtle changes in movement are captured in detail. The input is real-time video data, and the output is saved video files for analysis.
[0122] Step 2:
[0123] The server transmits the acquired video data to the cloud server via the 5G communication network. This enables rapid data transfer, allowing for immediate transition to video analysis without delay. The input is the stored video file, and the output is the prepared analysis data on the cloud server.
[0124] Step 3:
[0125] The server analyzes the transmitted video data using a generative AI model located in the cloud. The input is analysis data on the cloud server, and data correction and evaluation are performed to track the machine's operating position, speed, and motion trajectory. The output is the analysis result showing the efficiency and accuracy of the operation.
[0126] Step 4:
[0127] The server generates suggestions for improving machine operation based on the analysis results. Using the generated AI model, if the operation is determined to be inefficient, it generates specific suggestions and action plans for optimization. The input is the analysis results, and the output is information including improvement suggestions.
[0128] Step 5:
[0129] The server transmits the generated improvement suggestions to an external terminal via an information communication channel. Smart glasses or tablet devices receive this information and present it to the user in real time. The input is information including improvement suggestions, and the output is the specific improvement suggestions displayed on the external terminal. Based on the received suggestions, the user can adjust the machine and improve operational efficiency.
[0130] Furthermore, an emotion engine that estimates the user's emotions may be incorporated. That is, the identification processing unit 290 may use the emotion identification model 59 to estimate the user's emotions and perform identification processing using the user's emotions.
[0131] The present invention aims to provide more personalized feedback by incorporating an emotion engine that evaluates the user's psychological state into a system that supports the improvement of athletes' technical and strategic abilities.
[0132] System Configuration
[0133] 1. Video acquisition and analysis
[0134] Device: Smartphones or dedicated cameras are used to capture video footage of athletes during matches. This allows for real-time recording of the competition. The match footage is transmitted to a server via a 5G network.
[0135] 2. Analysis of motion and technique
[0136] Server: The server analyzes the received video using a generative model to analyze the athlete's movements and technical elements in detail. This analysis evaluates the athlete's form, position, strategic actions, etc., and identifies areas for improvement.
[0137] 3. Conducting an emotional assessment
[0138] Server: The emotion engine analyzes the competitor's facial expressions and audio data from the video to evaluate their psychological state. The emotion engine detects emotions such as happiness, concentration, and tension in real time.
[0139] 4. Customize feedback
[0140] Server: Based on the results of motion analysis and emotional assessment, it creates feedback tailored to the athlete's specific situation. This includes specific advice that takes psychological state into consideration, and technical improvement suggestions that are adjusted according to emotions.
[0141] 5. Providing feedback
[0142] Device: The generated feedback can be viewed by athletes and coaches on external devices. Users can check this on their smartphones or tablets during or immediately after a match and use it in their next training session.
[0143] Specific examples of use
[0144] User: For example, if a player is experiencing negative emotions during a match, the system can detect this and incorporate positive messages and emotionally uplifting advice into the feedback. The coach can then use this feedback to provide appropriate encouragement to help the player regain their composure.
[0145] This invention enables not only improvement in technical skills but also provides feedback that addresses the psychological aspects of athletes, thereby realizing more comprehensive training.
[0146] The following describes the processing flow.
[0147] Step 1:
[0148] Device: Before the start of the competition, smartphones or dedicated cameras are set up courtside to record the match footage. This video acquisition device records the movements and facial expressions of the athletes during the competition in high resolution.
[0149] Step 2:
[0150] Terminal: Recorded match footage is transmitted to the server via the 5G network. This transmission is done in real time, minimizing latency and enabling immediate analysis.
[0151] Step 3:
[0152] Server: Activates a generative model to analyze the received video. The video is divided frame by frame, allowing for detailed monitoring of the competitor's movements, position, and technical characteristics.
[0153] Step 4:
[0154] Server: Based on the video analysis, the server evaluates the athlete's physical movements, technical form, and strategic positioning. This includes generating technical metrics such as swing angle and motion speed.
[0155] Step 5:
[0156] Server: Using an emotion engine, the system analyzes the facial expressions and vocal patterns of athletes from video footage to assess their emotional state. This assessment aims to understand the psychological aspects of the athletes, measuring things like concentration and tension.
[0157] Step 6:
[0158] Server: Integrates the results of technical analysis and emotional assessment to generate feedback for the athlete. The feedback is optimized according to the athlete's emotional state. For example, if the athlete is anxious, advice for relaxation will be added.
[0159] Step 7:
[0160] Device: The generated feedback is sent to the athlete's or coach's smartphone or tablet. This allows the user to review the feedback and incorporate it into their next training plan.
[0161] Step 8:
[0162] User: Coaches use feedback to provide guidance to athletes, addressing both mental and technical aspects. Users can create plans to promote overall growth, including emotional well-being.
[0163] (Example 2)
[0164] Next, we will describe Example 2. In the following description, the data processing device 12 will be referred to as the "server" and the smart device 14 as the "terminal".
[0165] Traditional athlete support systems have a weakness: they focus on technical feedback and lack feedback that considers the psychological aspects of athletes. This can hinder the overall improvement of athletes' abilities.
[0166] The identification process performed by the identification processing unit 290 of the data processing device 12 in Example 2 is realized by the following means.
[0167] In this invention, the server includes means for receiving match video acquired by video acquisition means as input, analyzing the actions of athletes from the video using a generative model that analyzes the video, generating customized feedback based on the actions and psychological state, and transmitting the feedback to an information processing device via a communication network. This enables comprehensive support that takes into account not only the technical abilities of athletes but also their psychological aspects.
[0168] "Video acquisition means" refers to devices or equipment used to collect video footage of athletes during matches, and plays a role in preparing the footage for later analysis.
[0169] A "generative model" is a program or algorithm that analyzes the movements and technical elements of athletes from input video data, enabling detailed analysis of athletes.
[0170] A "communication network" is a medium for transferring data and serves as a pathway for sending feedback from a server to an information processing device.
[0171] An "information processing device" is a device that receives data and feedback transmitted from a server, and displays or processes it. It refers to a terminal used by a user to check feedback.
[0172] An "emotional engine" refers to software or a system that analyzes an athlete's psychological state from their facial expressions and tone of voice to evaluate their emotional state.
[0173] "Feedback" refers to advice and guidance generated based on an athlete's actions and psychological state, and is information aimed at improving the athlete's technical and strategic abilities.
[0174] This system is designed to support the improvement of athletes' technical and psychological abilities. The main components of the system are video acquisition means, generative models, emotion engines, information processing devices, and communication networks.
[0175] The server receives video footage capturing the athletes' movements during a match. This video is acquired by terminals using smartphones or dedicated cameras and transmitted to the server via a communication network. The server analyzes this video data using a generative model to analyze the athletes' movements and technical elements in detail. The generative model uses machine learning algorithms to extract and quantify the characteristics of the movements. For example, it can measure the athlete's form, the speed of their actions, and the angles of their movements.
[0176] Next, the server uses an emotion engine to analyze the competitor's facial expressions and audio data from the video to evaluate their psychological state. This emotion engine is implemented by combining facial recognition technology and audio analysis technology, and it determines feelings of happiness, tension, concentration, etc., in real time.
[0177] Based on these analysis results, the server generates personalized feedback for the competitor. This feedback includes areas for technical improvement and strategic advice tailored to their psychological state. The feedback is transmitted via the communication network to an information processing device, i.e., the user's terminal.
[0178] The device displays received feedback and is designed to be easily understood by the user. Users can utilize the feedback to improve their next training session. For example, if a negative emotional state is detected during a match, positive messages and suggestions for improvement are provided, enabling coaching that takes the athlete's mental state into consideration.
[0179] Example of a prompt
[0180] "Analyze video footage of athletes during matches to assess areas for improvement in their movements and emotional state, and generate personalized feedback."
[0181] This system enables training support that integrates technology and emotion.
[0182] The flow of the specific processing in Example 2 will be explained using Figure 13.
[0183] Step 1:
[0184] The device acquires video footage of athletes during matches. It uses smartphones or dedicated cameras as input. The output is real-time match video data. This video data is streamed to a server using a 5G network.
[0185] Step 2:
[0186] The server receives video data transmitted from the terminal. The input is high-resolution video data from the terminal. Data preprocessing is performed to convert this data into a format that can be analyzed frame by frame. The output is video data that has been prepared in an analysis-ready format.
[0187] Step 3:
[0188] The server uses a generative AI model to analyze the movements of athletes from pre-processed video data. The input is formatted video data. The generative AI model performs data calculations to extract the athletes' movement patterns and technical characteristics from the video frames. The output is the motion analysis results obtained as numerical data.
[0189] Step 4:
[0190] The server uses an emotion engine to analyze the athlete's facial expressions and voice tone from video data, thereby evaluating their psychological state. The input is video data that has already undergone motion analysis. The emotion engine performs image recognition and voice analysis to quantify the emotional state. The output is data representing the athlete's real-time psychological state.
[0191] Step 5:
[0192] The server integrates motion analysis results and psychological state assessments to generate personalized feedback. The input consists of motion and psychological state data. Based on this, it generates feedback that includes specific advice and technical improvement suggestions for the athlete. The output is customized feedback.
[0193] Step 6:
[0194] The terminal receives feedback sent from the server and displays it in a format viewable by the user. The input is feedback data from the server. The terminal presents this data in a user-friendly interface. The output is feedback information that the user can visually confirm.
[0195] Users can utilize this feedback to improve their next training session.
[0196] (Application Example 2)
[0197] Next, we will explain application example 2. In the following explanation, the data processing device 12 will be referred to as a "server" and the smart device 14 as a "terminal".
[0198] There is a challenge in understanding in real time what customers are feeling in a store and providing optimal customer service tailored to those emotions. Traditional methods rely on sales staff subjectively judging customers' expressions and attitudes, lacking objective means to make appropriate approaches. This makes it difficult to improve customer satisfaction and increase sales efficiency.
[0199] The specific processing performed by the specific processing unit 290 of the data processing device 12 in Application Example 2 is realized by the following means.
[0200] In this invention, the server includes means for taking customer video acquired by a video acquisition device as input and analyzing the customer's facial expressions using a generative model that analyzes the video; means for evaluating the customer's emotional state and generating feedback; and means for transmitting the feedback to the sales staff's terminal via a communication network. This enables optimal customer service based on the customer's emotions.
[0201] "Image acquisition equipment" refers to devices used to capture customers' facial expressions and movements, and includes cameras, smart devices, and other similar equipment.
[0202] A "generative model" is an algorithm that analyzes input video data to estimate the emotional state of a customer, and it utilizes machine learning technology.
[0203] "Methods of analysis" refers to the process of analyzing video data using generative models to identify emotions and psychological states from customers' facial expressions and actions.
[0204] A "communication network" is the infrastructure used to send and receive analyzed information between a server and the terminals of sales staff, and it utilizes the internet or wireless communication.
[0205] "Feedback" refers to information provided to sales staff as advice on customer service and product recommendations based on the customer's emotional state.
[0206] To implement this invention, cameras or smart devices must be installed in the store as video acquisition devices to capture customers' facial expressions and movements. This makes it possible to acquire video data in real time. This video data is transmitted to a server via a communication network such as 5G.
[0207] The server analyzes the received video data using a generative AI model. This model utilizes machine learning algorithms to estimate the customer's emotional state from their facial expressions and movements. During the analysis, the video data is broken down frame by frame, and the customer's facial expressions in each frame are analyzed. Existing video analysis software, such as Google Cloud Vision API, can be used for this process. If audio data is present, customer speech is collected using a noise-canceling microphone, and this content is also incorporated into the emotion analysis.
[0208] The analysis results are generated as feedback, notifying sales staff of the customer's emotional state on their terminals. This feedback is provided in real time or near real time, allowing sales staff to provide optimal customer service based on it. Specifically, the system displays prompts on the sales staff's terminals such as, "Generate advice suggesting an appropriate response based on the customer's current emotional state. If the customer appears confused, suggest popular products for first-time customers." This kind of feedback allows sales staff to immediately take actions that lead to increased customer satisfaction.
[0209] The flow of a specific process in Application Example 2 will be explained using Figure 14.
[0210] Step 1:
[0211] The user (store camera) acquires video data when a customer enters the store. The camera captures the customer's movements and facial expressions inside the store, acquiring video data in a format that can be analyzed by a generative AI model. The input is a real-time video stream, and the output is analyzable digital data.
[0212] Step 2:
[0213] The terminal (camera or smart device) transmits video data to the server via a 5G communication network. Due to the high bandwidth network, data is transmitted in near real-time. The input is digital data from the video acquisition device, and the output is the video data transmitted to the server.
[0214] Step 3:
[0215] The server uses a generated AI model to analyze the received video data. The server extracts the customer's face frame by frame from the video data and performs facial expression analysis. In this analysis process, video analysis algorithms such as the Google Cloud Vision API are used. The input is the video data sent to the server, and the output is data related to the customer's emotional state corresponding to their facial expressions.
[0216] Step 4:
[0217] The server uses the analysis results to generate feedback and create appropriate customer service advice based on the customer's emotional state. The generated feedback includes specific suggestions to help sales staff provide customer service. The input is emotional state data obtained through analysis, and the output is feedback information for sales staff.
[0218] Step 5:
[0219] The server sends the feedback generated in the previous step to the sales staff's terminal via the communication network. The staff checks this information on their smartphone or tablet and uses it to assist customers. The input is the feedback information on the server, and the output is the advice displayed on the sales staff's terminal.
[0220] Step 6:
[0221] The terminal (the sales staff's smart device) checks the received feedback and modifies its approach to the customer. For example, if it detects that the customer is confused, the staff will provide service according to the prompt message, "Suggest popular products to first-time customers." The input is the feedback displayed on the terminal, and the output is the specific action taken by the sales staff based on that feedback.
[0222] The specific processing unit 290 transmits the result of the specific processing to the smart device 14. In the smart device 14, the control unit 46A causes the output device 40 to output the result of the specific processing. The microphone 38B acquires audio indicating user input for the result of the specific processing. The control unit 46A transmits the audio data indicating user input acquired by the microphone 38B to the data processing device 12. In the data processing device 12, the specific processing unit 290 acquires the audio data.
[0223] Data generation model 58 is a so-called generative AI (Artificial Intelligence). An example of data generation model 58 is ChatGPT (registered trademark) (Internet search).<URL: https: / / openai.com / blog / chatgpt> ), Gemini (registered trademark) (Internet search) <url: https: gemini.google.com ?hl="ja">Examples of generative AI include the following. The data generation model 58 is obtained by performing deep learning on a neural network. The data generation model 58 is input with prompts containing instructions, and with inference data such as audio data representing speech, text data representing text, and image data representing images. The data generation model 58 infers from the input inference data according to the instructions indicated by the prompts, and outputs the inference results in data formats such as audio data and text data. Here, inference refers to, for example, analysis, classification, prediction, and / or summarization.
[0224] In the above embodiment, an example was given in which specific processing is performed by the data processing device 12, but the technology of this disclosure is not limited thereto, and the specific processing may also be performed by the smart device 14.
[0225] [Second Embodiment]
[0226] Figure 3 shows an example of the configuration of the data processing system 210 according to the second embodiment.
[0227] As shown in Figure 3, the data processing system 210 includes a data processing device 12 and smart glasses 214. An example of the data processing device 12 is a server.
[0228] The data processing device 12 comprises a computer 22, a database 24, and a communication interface 26. The computer 22 is an example of a "computer" related to the technology of this disclosure. The computer 22 comprises a processor 28, RAM 30, and storage 32. The processor 28, RAM 30, and storage 32 are connected to a bus 34. The database 24 and the communication interface 26 are also connected to the bus 34. The communication interface 26 is connected to a network 54. An example of the network 54 is a WAN (Wide Area Network) and / or a LAN (Local Area Network).
[0229] The smart glasses 214 include a computer 36, a microphone 238, a speaker 240, a camera 42, and a communication interface 44. The computer 36 includes a processor 46, RAM 48, and storage 50. The processor 46, RAM 48, and storage 50 are connected to a bus 52. The microphone 238, speaker 240, and camera 42 are also connected to the bus 52.
[0230] The microphone 238 receives voice signals from the user 20 and receives instructions from the user 20. The microphone 238 captures the voice signals from the user 20, converts the captured voice into audio data, and outputs it to the processor 46. The speaker 240 outputs audio according to the instructions from the processor 46.
[0231] Camera 42 is a small digital camera equipped with an optical system including a lens, aperture, and shutter, and an image sensor such as a CMOS (Complementary Metal-Oxide-Semiconductor) image sensor or a CCD (Charge Coupled Device) image sensor, and captures images of the area around the user 20 (for example, an imaging range defined by a field of view equivalent to the width of a typical healthy person's field of vision).
[0232] Communication interface 44 is connected to network 54. Communication interfaces 44 and 26 are responsible for the exchange of various information between processor 46 and processor 28 via network 54. The exchange of various information between processor 46 and processor 28 using communication interfaces 44 and 26 is performed in a secure manner.
[0233] Figure 4 shows an example of the main functions of the data processing device 12 and the smart glasses 214. As shown in Figure 4, the data processing device 12 performs specific processing using the processor 28. The storage 32 stores the specific processing program 56.
[0234] The specific processing program 56 is an example of a "program" relating to the technology of this disclosure. The processor 28 reads the specific processing program 56 from the storage 32 and executes the read specific processing program 56 on the RAM 30. The specific processing is realized by the processor 28 operating as a specific processing unit 290 in accordance with the specific processing program 56 executed on the RAM 30.
[0235] The storage 32 stores the data generation model 58 and the emotion identification model 59. The data generation model 58 and the emotion identification model 59 are used by the identification processing unit 290.
[0236] In the smart glasses 214, the processor 46 performs the reception output processing. The storage 50 stores the reception output program 60. The processor 46 reads the reception output program 60 from the storage 50 and executes the read reception output program 60 on the RAM 48. The reception output processing is realized by the processor 46 operating as a control unit 46A according to the reception output program 60 executed on the RAM 48.
[0237] Next, the identification processing performed by the identification processing unit 290 of the data processing device 12 will be described. In the following description, the data processing device 12 will be referred to as the "server" and the smart glasses 214 will be referred to as the "terminal".
[0238] This invention relates to an automated analysis system for efficiently improving the technical and strategic capabilities of athletes. This system utilizes the functions of a video acquisition device, a generative model, a communication network, and an external terminal to provide immediate and effective feedback to athletes.
[0239] System Configuration
[0240] 1. Video acquisition device
[0241] Terminal: Using smartphones or dedicated cameras, video data from matches is acquired in real time. This video acquisition device is installed in stadiums and training grounds to record the movements of athletes in detail.
[0242] 2. Transmission of video data
[0243] Terminal: Acquired video data is transmitted to the server using the 5G communication network. Even high-capacity data is transferred quickly, enabling real-time analysis.
[0244] 3. Conduct motion analysis
[0245] Server: The transmitted video data is analyzed by a generative model. The model tracks the athlete's body position and movements, and evaluates their actions and technical elements during the competition. For example, it detects the speed and angle of the athlete's racket swing and compares it to an ideal movement pattern.
[0246] 4. Feedback generation
[0247] Server: Based on the analysis results, it generates individual feedback for each athlete. This feedback evaluates areas for improvement and successful techniques for each athlete, and provides them as goals for further practice.
[0248] 5. Communication and display of feedback
[0249] Device: The generated feedback is sent via the communication network to external devices used by athletes and coaches. Users can review this feedback on their smartphones or tablets and use it to improve their next training session.
[0250] Specific examples of use
[0251] User: If a badminton player's smash swing is incorrect during a match, the system instantly generates feedback on the angle of that swing. The coach, upon receiving this feedback, can then focus their instruction on improving specific forms during the next practice session.
[0252] This invention realizes a system that provides precise technical and strategic feedback through clear image analysis, thereby promoting efficient performance improvement in competitions.
[0253] The following describes the processing flow.
[0254] Step 1:
[0255] Devices: Before the start of the competition, smartphones or dedicated cameras are set up courtside to begin recording the match footage in real time. These devices continuously record the players' movements clearly in high resolution.
[0256] Step 2:
[0257] Terminal: Captured video data is streamed to the server using the 5G network. This allows large amounts of video data to reach the server quickly.
[0258] Step 3:
[0259] Server: The server divides the received video data into frames and performs preprocessing to extract the players' movements and positions from each frame. This process also includes filtering to reduce background noise and improve analysis accuracy.
[0260] Step 4:
[0261] Server: Uses a generative model to analyze the technical actions of the players in each frame (e.g., racket swing, player positioning). At this stage, the players' actions are compared to a specific set of rules and criteria to calculate technical evaluation metrics.
[0262] Step 5:
[0263] Server: Based on technical evaluation, it assesses the strategic approach of players. It analyzes the strategies players employ during matches, detecting, for example, delays in defensive movement or positioning during attacks.
[0264] Step 6:
[0265] Server: Based on the analysis results, create specific feedback on areas that need improvement and skills that should be strengthened. This feedback will include technical advice and strategic suggestions.
[0266] Step 7:
[0267] Terminal: Generated feedback is sent to an external terminal managed by the athlete or coach. Users can check this in real time or immediately after the match.
[0268] Step 8:
[0269] User: Based on the feedback, the coach understands areas where the athlete needs improvement and determines the focus points for the next training session. This feedback allows for efficient guidance to improve the athlete's skills and strategies.
[0270] (Example 1)
[0271] Next, we will describe Example 1. In the following description, the data processing device 12 will be referred to as the "server," and the smart glasses 214 will be referred to as the "terminal."
[0272] The present invention aims to provide a system that offers immediate, highly accurate feedback to efficiently improve athletes' technical skills and strategies. Conventional technologies have made real-time motion analysis and immediate identification of specific areas for improvement difficult, preventing rapid reflection of improvements in practice and competition performance. Therefore, there is a need for a system that automates the analysis of athletes' movements and efficiently provides feedback, including specific improvement suggestions.
[0273] The identification process performed by the identification processing unit 290 of the data processing device 12 in Example 1 is realized by the following means.
[0274] In this invention, the server includes means for taking competition video acquired by a video acquisition device as input and analyzing the athlete's posture from the video using a generative algorithm for analyzing the video; means for evaluating the athlete's technical and strategic elements based on the analysis results and generating feedback including specific improvement suggestions; and means for transmitting the feedback to an external device via a communication path. This allows athletes to immediately grasp areas for improvement in their movements and use that information to improve their technical skills and strategies in their next practice or match.
[0275] "Video acquisition equipment" refers to devices that can capture the actions of athletes and the details of a match in real time, and includes smartphones and dedicated cameras.
[0276] A "generative algorithm" is a mathematical model or method for analyzing input video data, particularly for tracking and evaluating the physical movements of athletes.
[0277] A "communication path" refers to the network used when sending and receiving data, and includes 5G networks and other technologies that enable high-speed and stable communication.
[0278] "External devices" refer to devices used by athletes and coaches to receive feedback, and include terminals such as smartphones and tablets.
[0279] "Feedback" refers to data or advice provided to improve a competitor's skills and strategies based on the results analyzed by a generative algorithm.
[0280] "Analyzing the posture" is a process of tracking the position and movement of a competitor's body and evaluating specific technical and strategic elements based on that data.
[0281] "Technical elements" are indicators showing the accuracy and efficiency of individual movements and techniques in a competition, and are used to numerically evaluate a player's performance.
[0282] "Strategic elements" are criteria for evaluating the strategies and tactics selected by a competitor during a game, and particularly analyze the responses to the movements of opponents and the selection of plays.
[0283] The present invention realizes a system that effectively analyzes the movements of a competitor and provides immediate and specific feedback. This system is mainly implemented using a video acquisition device, a generative AI model, a communication path, and an external device. Details and specific examples of each element are shown below.
[0284] The video acquisition device is a device for capturing the posture and movements of a competitor in real time during play. Specifically, smartphones and dedicated cameras fall under this category. These devices are installed in arenas and practice fields to record detailed video data.
[0285] The server is responsible for receiving the acquired video data and analyzing it using the generative AI model. The generative AI model tracks the body position and movements of the competitor in the video data and evaluates the technical and strategic elements during the competition. Specifically, it identifies the joint positions and movement speeds of the player's body and compares them with ideal movement patterns.
[0286] The server further automatically generates feedback based on the analysis results. This feedback includes points for improving the player's actions and guidance information regarding specific strategies. For example, detailed improvement suggestions such as "The angle of the racket during a smash is 5 degrees shallower than ideal" are presented.
[0287] The terminal transmits the feedback generated by the server to an external device via a communication path. The user receives and checks the feedback through an external device such as a smartphone or a tablet. This enables the user to immediately implement specific improvement measures during the next practice session.
[0288] Based on the provided feedback, the user can improve techniques and adjust strategies. Through this process, the player can quickly enhance their performance.
[0289] Example of a prompt sentence:
[0290] "Analyze the angle of the player's smash swing and generate the difference from the ideal angle as feedback."
[0291] In this way, this system automates a series of processes including video acquisition, data analysis, feedback generation, transmission, and reception, and strongly supports the technical and strategic improvement of the player.
[0292] The flow of the specific process in Example 1 will be described using FIG. 11.
[0293] Step 1:
[0294] The terminal uses a video acquisition device to capture the actions of the player. A smartphone or a dedicated camera is installed to capture the movements in real time at the arena or practice field. The input is video data including the actions of the player, which is recorded in digital format in preparation for transfer to the server. Specific actions such as the player's movement, operation of props, and interaction with the environment are captured in the video.
[0295] Step 2:
[0296] The terminal transfers the captured video data to the server via the communication path. High-capacity video data is transmitted rapidly using a 5G network. The input is the video data acquired in step 1, and the output is the compressed and decompressed video data arriving at the server. At this stage, communication optimization is performed to minimize delays during data transfer.
[0297] Step 3:
[0298] The server inputs the received video data into the generating AI model and begins analyzing the data. The input is the video data sent in step 2, and the output is the evaluation results of the athlete's body position and movement. Specifically, the process involves estimating the athlete's joint positions, calculating the speed and direction of movement, and comparing it with ideal movement for each frame.
[0299] Step 4:
[0300] The server generates feedback for the athlete based on the analysis results. The input is the evaluation results obtained in step 3, and the output is feedback data that includes specific improvement suggestions and quantified metrics. For example, it may include detailed analysis results regarding the athlete's swing angle and speed.
[0301] Step 5:
[0302] The terminal transmits the generated feedback to an external device via a communication path. The input is the feedback data received from the server, and the output is the information presented in a format that the user can verify. At this stage, the feedback is transmitted to the user via a smartphone or tablet.
[0303] Step 6:
[0304] The user utilizes the received feedback to improve their actions and tactics in the competition. Specifically, they adjust their training content based on the feedback and try out new strategies. This improves the technical ability and strategic thinking of the competitor.
[0305] (Application Example 1)
[0306] Next, Application Example 1 will be described. In the following description, the data processing device 12 is referred to as the "server", and the smart glasses 214 are referred to as the "terminal".
[0307] In a modern manufacturing environment, improving the operating efficiency of various machines is an important issue. However, there are limited means to constantly monitor the operation of machines and provide appropriate feedback in real time. With conventional analysis methods, it is difficult to quickly and accurately analyze the operation data of machines and create appropriate improvement proposals. Also, in situations where immediate implementation of improvement proposals is required, there is a problem that efficient improvement cannot be achieved due to the lack of appropriate information communication means.
[0308] The specific processing by the specific processing unit 290 of the data processing device 12 in Application Example 1 is realized by the following respective means.
[0309] In this invention, the server includes means for acquiring the operation video of the machine by a video acquisition mechanism, means for analyzing the operation using a generation algorithm and evaluating the operation of the machine, and means for transmitting improvement proposals to an external device via an information communication path. Thereby, it becomes possible to analyze the operation of the machine in real time and quickly propose efficient operation improvements.
[0310] The "video acquisition mechanism" is a device or system used to acquire the operation video of the machine and plays a role in capturing the details of the operation.
[0311] The "generation algorithm" is a calculation method or program used to analyze the operation of the machine from the video and evaluate the operation position and activities.
[0312] "Means for analyzing motion" refers to processes and methods for analyzing video footage of machine movements and evaluating efficient operation.
[0313] An "improvement suggestion" is information that, based on the analysis results, presents specific actions to improve the machine's operation and increase efficiency.
[0314] "Information communication channels" refer to communication paths and network technologies used to transmit improvement suggestions to external devices.
[0315] An "external device" is a computer or mobile device used to receive improvement suggestions sent from the server.
[0316] The system that realizes this invention consists of a combination of an image acquisition mechanism, a generation algorithm, an information communication channel, and external devices. This makes it possible to efficiently monitor and analyze the operation of machinery in a factory and provide improvement suggestions.
[0317] The server uses dedicated cameras to capture video footage of the machinery in operation within the factory. High-resolution cameras can record various machine movements in detail. This allows for precise analysis without any missing information.
[0318] After acquiring video footage, the server uses a generation algorithm to analyze the captured motion video. This algorithm tracks the machine's position and evaluates the efficiency and precision of its movements. For example, it can identify problems such as suboptimal machine movement or inappropriate work speed. TensorFlow or PyTorch can be used as the generation algorithm.
[0319] Based on the analysis results, the server generates improvement suggestions. These include areas for improvement in machine operation and action plans for more efficient operation. These suggestions are output in a format that is easy for the user to understand.
[0320] External devices such as smart glasses and tablet terminals are used. Generated improvement suggestions are transmitted to these external devices via a communication channel, allowing the user to view them in real time. As a result, the user can quickly work on improving the machine's operation.
[0321] As a concrete example, suppose a camera installed on a factory robot arm captures and analyzes its movements, revealing that its trajectory is not optimal. The generated AI model then generates specific trajectory correction suggestions to improve the efficiency of the movement, and presents them as a prompt message such as, "Analyze the current movement of the robot arm and compare it to the most efficient movement pattern. Generate feedback on areas for improvement and suggested improvements, and present them to the worker."
[0322] In this way, the machine's operation can always be kept in an optimal state, leading to increased productivity.
[0323] The flow of a specific process in Application Example 1 will be explained using Figure 12.
[0324] Step 1:
[0325] The server uses dedicated cameras to capture video footage of machinery in operation within the factory. By using high-resolution cameras, even subtle changes in movement are captured in detail. The input is real-time video data, and the output is saved video files for analysis.
[0326] Step 2:
[0327] The server transmits the acquired video data to the cloud server via the 5G communication network. This enables rapid data transfer, allowing for immediate transition to video analysis without delay. The input is the stored video file, and the output is the prepared analysis data on the cloud server.
[0328] Step 3:
[0329] The server analyzes the transmitted video data using a generative AI model located in the cloud. The input is analysis data on the cloud server, and data correction and evaluation are performed to track the machine's operating position, speed, and motion trajectory. The output is the analysis result showing the efficiency and accuracy of the operation.
[0330] Step 4:
[0331] The server generates suggestions for improving machine operation based on the analysis results. Using the generated AI model, if the operation is determined to be inefficient, it generates specific suggestions and action plans for optimization. The input is the analysis results, and the output is information including improvement suggestions.
[0332] Step 5:
[0333] The server transmits the generated improvement suggestions to an external terminal via an information communication channel. Smart glasses or tablet devices receive this information and present it to the user in real time. The input is information including improvement suggestions, and the output is the specific improvement suggestions displayed on the external terminal. Based on the received suggestions, the user can adjust the machine and improve operational efficiency.
[0334] Furthermore, an emotion engine that estimates the user's emotions may be incorporated. That is, the identification processing unit 290 may use the emotion identification model 59 to estimate the user's emotions and perform identification processing using the user's emotions.
[0335] The present invention aims to provide more personalized feedback by incorporating an emotion engine that evaluates the user's psychological state into a system that supports the improvement of athletes' technical and strategic abilities.
[0336] System Configuration
[0337] 1. Video acquisition and analysis
[0338] Device: Smartphones or dedicated cameras are used to capture video footage of athletes during matches. This allows for real-time recording of the competition. The match footage is transmitted to a server via a 5G network.
[0339] 2. Analysis of motion and technique
[0340] Server: The server analyzes the received video using a generative model to analyze the athlete's movements and technical elements in detail. This analysis evaluates the athlete's form, position, strategic actions, etc., and identifies areas for improvement.
[0341] 3. Conducting an emotional assessment
[0342] Server: The emotion engine analyzes the competitor's facial expressions and audio data from the video to evaluate their psychological state. The emotion engine detects emotions such as happiness, concentration, and tension in real time.
[0343] 4. Customize feedback
[0344] Server: Based on the results of motion analysis and emotional assessment, it creates feedback tailored to the athlete's specific situation. This includes specific advice that takes psychological state into consideration, and technical improvement suggestions that are adjusted according to emotions.
[0345] 5. Providing feedback
[0346] Device: The generated feedback can be viewed by athletes and coaches on external devices. Users can check this on their smartphones or tablets during or immediately after a match and use it in their next training session.
[0347] Specific examples of use
[0348] User: For example, if a player is experiencing negative emotions during a match, the system can detect this and incorporate positive messages and emotionally uplifting advice into the feedback. The coach can then use this feedback to provide appropriate encouragement to help the player regain their composure.
[0349] This invention enables not only improvement in technical skills but also provides feedback that addresses the psychological aspects of athletes, thereby realizing more comprehensive training.
[0350] The following describes the processing flow.
[0351] Step 1:
[0352] Device: Before the start of the competition, smartphones or dedicated cameras are set up courtside to record the match footage. This video acquisition device records the movements and facial expressions of the athletes during the competition in high resolution.
[0353] Step 2:
[0354] Terminal: Recorded match footage is transmitted to the server via the 5G network. This transmission is done in real time, minimizing latency and enabling immediate analysis.
[0355] Step 3:
[0356] Server: Activates a generative model to analyze the received video. The video is divided frame by frame, allowing for detailed monitoring of the competitor's movements, position, and technical characteristics.
[0357] Step 4:
[0358] Server: Based on the video analysis, the server evaluates the athlete's physical movements, technical form, and strategic positioning. This includes generating technical metrics such as swing angle and motion speed.
[0359] Step 5:
[0360] Server: Using an emotion engine, the system analyzes the facial expressions and vocal patterns of athletes from video footage to assess their emotional state. This assessment aims to understand the psychological aspects of the athletes, measuring things like concentration and tension.
[0361] Step 6:
[0362] Server: Integrates the results of technical analysis and emotional assessment to generate feedback for the athlete. The feedback is optimized according to the athlete's emotional state. For example, if the athlete is anxious, advice for relaxation will be added.
[0363] Step 7:
[0364] Device: The generated feedback is sent to the athlete's or coach's smartphone or tablet. This allows the user to review the feedback and incorporate it into their next training plan.
[0365] Step 8:
[0366] User: Coaches use feedback to provide guidance to athletes, addressing both mental and technical aspects. Users can create plans to promote overall growth, including emotional well-being.
[0367] (Example 2)
[0368] Next, we will describe Example 2. In the following description, the data processing device 12 will be referred to as the "server" and the smart glasses 214 will be referred to as the "terminal".
[0369] Traditional athlete support systems have a weakness: they focus on technical feedback and lack feedback that considers the psychological aspects of athletes. This can hinder the overall improvement of athletes' abilities.
[0370] The identification process performed by the identification processing unit 290 of the data processing device 12 in Example 2 is realized by the following means.
[0371] In this invention, the server includes means for receiving match video acquired by video acquisition means as input, analyzing the actions of athletes from the video using a generative model that analyzes the video, generating customized feedback based on the actions and psychological state, and transmitting the feedback to an information processing device via a communication network. This enables comprehensive support that takes into account not only the technical abilities of athletes but also their psychological aspects.
[0372] "Video acquisition means" refers to devices or equipment used to collect video footage of athletes during matches, and plays a role in preparing the footage for later analysis.
[0373] A "generative model" is a program or algorithm that analyzes the movements and technical elements of athletes from input video data, enabling detailed analysis of athletes.
[0374] A "communication network" is a medium for transferring data and serves as a pathway for sending feedback from a server to an information processing device.
[0375] An "information processing device" is a device that receives data and feedback transmitted from a server, and displays or processes it. It refers to a terminal used by a user to check feedback.
[0376] An "emotional engine" refers to software or a system that analyzes an athlete's psychological state from their facial expressions and tone of voice to evaluate their emotional state.
[0377] "Feedback" refers to advice and guidance generated based on an athlete's actions and psychological state, and is information aimed at improving the athlete's technical and strategic abilities.
[0378] This system is designed to support the improvement of athletes' technical and psychological abilities. The main components of the system are video acquisition means, generative models, emotion engines, information processing devices, and communication networks.
[0379] The server receives video footage capturing the athletes' movements during a match. This video is acquired by terminals using smartphones or dedicated cameras and transmitted to the server via a communication network. The server analyzes this video data using a generative model to analyze the athletes' movements and technical elements in detail. The generative model uses machine learning algorithms to extract and quantify the characteristics of the movements. For example, it can measure the athlete's form, the speed of their actions, and the angles of their movements.
[0380] Next, the server uses an emotion engine to analyze the competitor's facial expressions and audio data from the video to evaluate their psychological state. This emotion engine is implemented by combining facial recognition technology and audio analysis technology, and it determines feelings of happiness, tension, concentration, etc., in real time.
[0381] Based on these analysis results, the server generates personalized feedback for the competitor. This feedback includes areas for technical improvement and strategic advice tailored to their psychological state. The feedback is transmitted via the communication network to an information processing device, i.e., the user's terminal.
[0382] The device displays received feedback and is designed to be easily understood by the user. Users can utilize the feedback to improve their next training session. For example, if a negative emotional state is detected during a match, positive messages and suggestions for improvement are provided, enabling coaching that takes the athlete's mental state into consideration.
[0383] Example of a prompt
[0384] "Analyze video footage of athletes during matches to assess areas for improvement in their movements and emotional state, and generate personalized feedback."
[0385] This system enables training support that integrates technology and emotion.
[0386] The flow of the specific processing in Example 2 will be explained using Figure 13.
[0387] Step 1:
[0388] The device acquires video footage of athletes during matches. It uses smartphones or dedicated cameras as input. The output is real-time match video data. This video data is streamed to a server using a 5G network.
[0389] Step 2:
[0390] The server receives video data transmitted from the terminal. The input is high-resolution video data from the terminal. Data preprocessing is performed to convert this data into a format that can be analyzed frame by frame. The output is video data that has been prepared in an analysis-ready format.
[0391] Step 3:
[0392] The server uses a generative AI model to analyze the movements of athletes from pre-processed video data. The input is formatted video data. The generative AI model performs data calculations to extract the athletes' movement patterns and technical characteristics from the video frames. The output is the motion analysis results obtained as numerical data.
[0393] Step 4:
[0394] The server uses an emotion engine to analyze the athlete's facial expressions and voice tone from video data, thereby evaluating their psychological state. The input is video data that has already undergone motion analysis. The emotion engine performs image recognition and voice analysis to quantify the emotional state. The output is data representing the athlete's real-time psychological state.
[0395] Step 5:
[0396] The server integrates motion analysis results and psychological state assessments to generate personalized feedback. The input consists of motion and psychological state data. Based on this, it generates feedback that includes specific advice and technical improvement suggestions for the athlete. The output is customized feedback.
[0397] Step 6:
[0398] The terminal receives feedback sent from the server and displays it in a format viewable by the user. The input is feedback data from the server. The terminal presents this data in a user-friendly interface. The output is feedback information that the user can visually confirm.
[0399] Users can utilize this feedback to improve their next training session.
[0400] (Application Example 2)
[0401] Next, we will explain application example 2. In the following explanation, the data processing device 12 will be referred to as the "server," and the smart glasses 214 will be referred to as the "terminal."
[0402] There is a challenge in understanding in real time what customers are feeling in a store and providing optimal customer service tailored to those emotions. Traditional methods rely on sales staff subjectively judging customers' expressions and attitudes, lacking objective means to make appropriate approaches. This makes it difficult to improve customer satisfaction and increase sales efficiency.
[0403] The specific processing performed by the specific processing unit 290 of the data processing device 12 in Application Example 2 is realized by the following means.
[0404] In this invention, the server includes means for taking customer video acquired by a video acquisition device as input and analyzing the customer's facial expressions using a generative model that analyzes the video; means for evaluating the customer's emotional state and generating feedback; and means for transmitting the feedback to the sales staff's terminal via a communication network. This enables optimal customer service based on the customer's emotions.
[0405] "Image acquisition equipment" refers to devices used to capture customers' facial expressions and movements, and includes cameras, smart devices, and other similar equipment.
[0406] A "generative model" is an algorithm that analyzes input video data to estimate the emotional state of a customer, and it utilizes machine learning technology.
[0407] "Methods of analysis" refers to the process of analyzing video data using generative models to identify emotions and psychological states from customers' facial expressions and actions.
[0408] A "communication network" is the infrastructure used to send and receive analyzed information between a server and the terminals of sales staff, and it utilizes the internet or wireless communication.
[0409] "Feedback" refers to information provided to sales staff as advice on customer service and product recommendations based on the customer's emotional state.
[0410] To implement this invention, cameras or smart devices must be installed in the store as video acquisition devices to capture customers' facial expressions and movements. This makes it possible to acquire video data in real time. This video data is transmitted to a server via a communication network such as 5G.
[0411] The server analyzes the received video data using a generative AI model. This model utilizes machine learning algorithms to estimate the customer's emotional state from their facial expressions and movements. During the analysis, the video data is broken down frame by frame, and the customer's facial expressions in each frame are analyzed. Existing video analysis software, such as the Google Cloud Vision API, can be used for this process. If audio data is present, customer speech is collected using a noise-canceling microphone, and this content is also incorporated into the emotion analysis.
[0412] The analysis results are generated as feedback, notifying sales staff of the customer's emotional state on their terminals. This feedback is provided in real time or near real time, allowing sales staff to provide optimal customer service based on it. Specifically, the system displays prompts on the sales staff's terminals such as, "Generate advice suggesting an appropriate response based on the customer's current emotional state. If the customer appears confused, suggest popular products for first-time customers." This kind of feedback allows sales staff to immediately take actions that lead to increased customer satisfaction.
[0413] The flow of a specific process in Application Example 2 will be explained using Figure 14.
[0414] Step 1:
[0415] The user (store camera) acquires video data when a customer enters the store. The camera captures the customer's movements and facial expressions inside the store, acquiring video data in a format that can be analyzed by a generative AI model. The input is a real-time video stream, and the output is analyzable digital data.
[0416] Step 2:
[0417] The terminal (camera or smart device) transmits video data to the server via a 5G communication network. Due to the high bandwidth network, data is transmitted in near real-time. The input is digital data from the video acquisition device, and the output is the video data transmitted to the server.
[0418] Step 3:
[0419] The server uses a generated AI model to analyze the received video data. The server extracts the customer's face frame by frame from the video data and performs facial expression analysis. In this analysis process, video analysis algorithms such as the Google Cloud Vision API are used. The input is the video data sent to the server, and the output is data related to the customer's emotional state corresponding to their facial expressions.
[0420] Step 4:
[0421] The server uses the analysis results to generate feedback and create appropriate customer service advice based on the customer's emotional state. The generated feedback includes specific suggestions to help sales staff provide customer service. The input is emotional state data obtained through analysis, and the output is feedback information for sales staff.
[0422] Step 5:
[0423] The server sends the feedback generated in the previous step to the sales staff's terminal via the communication network. The staff checks this information on their smartphone or tablet and uses it to assist customers. The input is the feedback information on the server, and the output is the advice displayed on the sales staff's terminal.
[0424] Step 6:
[0425] The terminal (the sales staff's smart device) checks the received feedback and modifies its approach to the customer. For example, if it detects that the customer is confused, the staff will provide service according to the prompt message, "Suggest popular products to first-time customers." The input is the feedback displayed on the terminal, and the output is the specific action taken by the sales staff based on that feedback.
[0426] The specific processing unit 290 transmits the result of the specific processing to the smart glasses 214. In the smart glasses 214, the control unit 46A causes the speaker 240 to output the result of the specific processing. The microphone 238 acquires audio indicating user input for the result of the specific processing. The control unit 46A transmits the audio data indicating user input acquired by the microphone 238 to the data processing unit 12. In the data processing unit 12, the specific processing unit 290 acquires the audio data.
[0427] Data generation model 58 is a type of so-called generative AI (Artificial Intelligence). One example of data generation model 58 is ChatGPT (Internet search<URL: https: / / openai.com / blog / chatgpt> ), Gemini (Internet search) <url: https: gemini.google.com ?hl="ja">Examples of generative AI include the following. The data generation model 58 is obtained by performing deep learning on a neural network. The data generation model 58 is input with prompts containing instructions, and with inference data such as audio data representing speech, text data representing text, and image data representing images. The data generation model 58 infers from the input inference data according to the instructions indicated by the prompts, and outputs the inference results in data formats such as audio data and text data. Here, inference refers to, for example, analysis, classification, prediction, and / or summarization.
[0428] In the above embodiment, an example was given in which specific processing is performed by the data processing device 12, but the technology of this disclosure is not limited thereto, and the specific processing may also be performed by the smart glasses 214.
[0429] [Third Embodiment]
[0430] Figure 5 shows an example of the configuration of the data processing system 310 according to the third embodiment.
[0431] As shown in Figure 5, the data processing system 310 includes a data processing device 12 and a headset terminal 314. An example of the data processing device 12 is a server.
[0432] The data processing device 12 comprises a computer 22, a database 24, and a communication interface 26. The computer 22 is an example of a "computer" related to the technology of this disclosure. The computer 22 comprises a processor 28, RAM 30, and storage 32. The processor 28, RAM 30, and storage 32 are connected to a bus 34. The database 24 and the communication interface 26 are also connected to the bus 34. The communication interface 26 is connected to a network 54. An example of the network 54 is a WAN (Wide Area Network) and / or a LAN (Local Area Network).
[0433] The headset terminal 314 includes a computer 36, a microphone 238, a speaker 240, a camera 42, a communication interface 44, and a display 343. The computer 36 includes a processor 46, RAM 48, and storage 50. The processor 46, RAM 48, and storage 50 are connected to a bus 52. The microphone 238, speaker 240, camera 42, and display 343 are also connected to the bus 52.
[0434] The microphone 238 receives voice signals from the user 20 and receives instructions from the user 20. The microphone 238 captures the voice signals from the user 20, converts the captured voice into audio data, and outputs it to the processor 46. The speaker 240 outputs audio according to the instructions from the processor 46.
[0435] Camera 42 is a small digital camera equipped with an optical system including a lens, aperture, and shutter, and an image sensor such as a CMOS (Complementary Metal-Oxide-Semiconductor) image sensor or a CCD (Charge Coupled Device) image sensor, and captures images of the area around the user 20 (for example, an imaging range defined by a field of view equivalent to the width of a typical healthy person's field of vision).
[0436] Communication interface 44 is connected to network 54. Communication interfaces 44 and 26 are responsible for the exchange of various information between processor 46 and processor 28 via network 54. The exchange of various information between processor 46 and processor 28 using communication interfaces 44 and 26 is performed in a secure manner.
[0437] Figure 6 shows an example of the main functions of the data processing device 12 and the headset terminal 314. As shown in Figure 6, the data processing device 12 performs specific processing using the processor 28. The storage 32 stores the specific processing program 56.
[0438] The specific processing program 56 is an example of a "program" relating to the technology of this disclosure. The processor 28 reads the specific processing program 56 from the storage 32 and executes the read specific processing program 56 on the RAM 30. The specific processing is realized by the processor 28 operating as a specific processing unit 290 in accordance with the specific processing program 56 executed on the RAM 30.
[0439] The storage 32 stores the data generation model 58 and the emotion identification model 59. The data generation model 58 and the emotion identification model 59 are used by the identification processing unit 290.
[0440] In the headset terminal 314, the processor 46 performs the reception output processing. The storage 50 stores the reception output program 60. The processor 46 reads the reception output program 60 from the storage 50 and executes the read reception output program 60 on the RAM 48. The reception output processing is realized by the processor 46 operating as a control unit 46A according to the reception output program 60 executed on the RAM 48.
[0441] Next, the specific processing performed by the specific processing unit 290 of the data processing device 12 will be described. In the following description, the data processing device 12 will be referred to as the "server" and the headset terminal 314 will be referred to as the "terminal".
[0442] This invention relates to an automated analysis system for efficiently improving the technical and strategic capabilities of athletes. This system utilizes the functions of a video acquisition device, a generative model, a communication network, and an external terminal to provide immediate and effective feedback to athletes.
[0443] System Configuration
[0444] 1. Video acquisition device
[0445] Terminal: Using smartphones or dedicated cameras, video data from matches is acquired in real time. This video acquisition device is installed in stadiums and training grounds to record the movements of athletes in detail.
[0446] 2. Transmission of video data
[0447] Terminal: Acquired video data is transmitted to the server using the 5G communication network. Even high-capacity data is transferred quickly, enabling real-time analysis.
[0448] 3. Conduct motion analysis
[0449] Server: The transmitted video data is analyzed by a generative model. The model tracks the athlete's body position and movements, and evaluates their actions and technical elements during the competition. For example, it detects the speed and angle of the athlete's racket swing and compares it to an ideal movement pattern.
[0450] 4. Feedback generation
[0451] Server: Based on the analysis results, it generates individual feedback for each athlete. This feedback evaluates areas for improvement and successful techniques for each athlete, and provides them as goals for further practice.
[0452] 5. Communication and display of feedback
[0453] Device: The generated feedback is sent via the communication network to external devices used by athletes and coaches. Users can review this feedback on their smartphones or tablets and use it to improve their next training session.
[0454] Specific examples of use
[0455] User: If a badminton player's smash swing is incorrect during a match, the system instantly generates feedback on the angle of that swing. The coach, upon receiving this feedback, can then focus their instruction on improving specific forms during the next practice session.
[0456] This invention realizes a system that provides precise technical and strategic feedback through clear image analysis, thereby promoting efficient performance improvement in competitions.
[0457] The following describes the processing flow.
[0458] Step 1:
[0459] Devices: Before the start of the competition, smartphones or dedicated cameras are set up courtside to begin recording the match footage in real time. These devices continuously record the players' movements clearly in high resolution.
[0460] Step 2:
[0461] Terminal: Captured video data is streamed to the server using the 5G network. This allows large amounts of video data to reach the server quickly.
[0462] Step 3:
[0463] Server: The server divides the received video data into frames and performs preprocessing to extract the players' movements and positions from each frame. This process also includes filtering to reduce background noise and improve analysis accuracy.
[0464] Step 4:
[0465] Server: Uses a generative model to analyze the technical actions of the players in each frame (e.g., racket swing, player positioning). At this stage, the players' actions are compared to a specific set of rules and criteria to calculate technical evaluation metrics.
[0466] Step 5:
[0467] Server: Based on technical evaluation, it assesses the strategic approach of players. It analyzes the strategies players employ during matches, detecting, for example, delays in defensive movement or positioning during attacks.
[0468] Step 6:
[0469] Server: Based on the analysis results, create specific feedback on areas that need improvement and skills that should be strengthened. This feedback will include technical advice and strategic suggestions.
[0470] Step 7:
[0471] Terminal: Generated feedback is sent to an external terminal managed by the athlete or coach. Users can check this in real time or immediately after the match.
[0472] Step 8:
[0473] User: Based on the feedback, the coach understands areas where the athlete needs improvement and determines the focus points for the next training session. This feedback allows for efficient guidance to improve the athlete's skills and strategies.
[0474] (Example 1)
[0475] Next, we will describe Example 1. In the following description, the data processing device 12 will be referred to as the "server," and the headset-type terminal 314 will be referred to as the "terminal."
[0476] The present invention aims to provide a system that offers immediate, highly accurate feedback to efficiently improve athletes' technical skills and strategies. Conventional technologies have made real-time motion analysis and immediate identification of specific areas for improvement difficult, preventing rapid reflection of improvements in practice and competition performance. Therefore, there is a need for a system that automates the analysis of athletes' movements and efficiently provides feedback, including specific improvement suggestions.
[0477] The identification process performed by the identification processing unit 290 of the data processing device 12 in Example 1 is realized by the following means.
[0478] In this invention, the server includes means for taking competition video acquired by a video acquisition device as input and analyzing the athlete's posture from the video using a generative algorithm for analyzing the video; means for evaluating the athlete's technical and strategic elements based on the analysis results and generating feedback including specific improvement suggestions; and means for transmitting the feedback to an external device via a communication path. This allows athletes to immediately grasp areas for improvement in their movements and use that information to improve their technical skills and strategies in their next practice or match.
[0479] "Video acquisition equipment" refers to devices that can capture the actions of athletes and the details of a match in real time, and includes smartphones and dedicated cameras.
[0480] A "generative algorithm" is a mathematical model or method for analyzing input video data, particularly for tracking and evaluating the physical movements of athletes.
[0481] A "communication path" refers to the network used when sending and receiving data, and includes 5G networks and other technologies that enable high-speed and stable communication.
[0482] "External devices" refer to devices used by athletes and coaches to receive feedback, and include terminals such as smartphones and tablets.
[0483] "Feedback" refers to data and advice provided to improve athletes' skills and strategies, based on the results of analysis using generative algorithms.
[0484] "Analyzing posture" is the process of tracking an athlete's body position and movements, and using that data to evaluate specific technical and strategic elements.
[0485] "Technical elements" are indicators that show the accuracy and efficiency of individual movements and techniques in a competition, and are used to numerically evaluate an athlete's performance.
[0486] "Strategic elements" are criteria used to evaluate the strategies and tactics that athletes choose during a match, particularly analyzing their reactions to the opponent's movements and their choice of plays.
[0487] This invention provides a system that effectively analyzes the movements of athletes and provides immediate, specific feedback. This system is primarily implemented using video acquisition equipment, a generative AI model, a communication path, and external devices. Details and specific examples of each element are provided below.
[0488] Video acquisition equipment is a device used to capture the posture and movements of athletes in real time while they are playing. Specifically, this includes smartphones and dedicated cameras. These devices are installed in stadiums and practice fields to record detailed video data.
[0489] The server receives the acquired video data and analyzes it using a generative AI model. The generative AI model tracks the athletes' body positions and movements within the video data and evaluates the technical and strategic elements during the competition. Specifically, it identifies the joint positions and movement speeds of the athletes' bodies and compares them to ideal movement patterns.
[0490] The server then automatically generates feedback based on the analysis results. This feedback includes guidance on areas for improvement in the athlete's movements and specific strategies. For example, it might suggest detailed improvements such as, "The angle of your racket during a smash is 5 degrees shallower than ideal."
[0491] The terminal transmits feedback generated by the server to an external device via a communication path. The user receives and confirms the feedback through an external device such as a smartphone or tablet. This makes it possible to immediately implement specific improvement measures during the next practice session.
[0492] Based on the feedback provided, users can improve their technology and adjust their strategies. Through this process, competitors can quickly improve their performance.
[0493] Example of a prompt:
[0494] "Analyze the angle of the player's smash swing and generate feedback showing the difference from the ideal angle."
[0495] In this way, the system automates a series of processes including video acquisition, data analysis, feedback generation, transmission, and reception, providing strong support for the technical and strategic improvement of athletes.
[0496] The flow of the specific processing in Example 1 will be explained using Figure 11.
[0497] Step 1:
[0498] The terminal uses video acquisition equipment to capture the movements of athletes. Smartphones or dedicated cameras are set up to film movements in real time at the stadium or training ground. The input is video data including the movements of athletes, which is recorded in digital format in preparation for transmission to the server. Specifically, the movements of athletes, the operation of equipment, and interactions with the environment are captured on video.
[0499] Step 2:
[0500] The terminal transfers the captured video data to the server via the communication path. High-capacity video data is transmitted rapidly using a 5G network. The input is the video data acquired in step 1, and the output is the compressed and decompressed video data arriving at the server. At this stage, communication optimization is performed to minimize delays during data transfer.
[0501] Step 3:
[0502] The server inputs the received video data into the generating AI model and begins analyzing the data. The input is the video data sent in step 2, and the output is the evaluation results of the athlete's body position and movement. Specifically, the process involves estimating the athlete's joint positions, calculating the speed and direction of movement, and comparing it with ideal movement for each frame.
[0503] Step 4:
[0504] The server generates feedback for the athlete based on the analysis results. The input is the evaluation results obtained in step 3, and the output is feedback data that includes specific improvement suggestions and quantified metrics. For example, it may include detailed analysis results regarding the athlete's swing angle and speed.
[0505] Step 5:
[0506] The terminal transmits the generated feedback to an external device via a communication path. The input is the feedback data received from the server, and the output is the information presented in a format that the user can verify. At this stage, the feedback is transmitted to the user via a smartphone or tablet.
[0507] Step 6:
[0508] Users utilize the feedback they receive to improve their movements and tactics in competition. Specifically, they adjust their training content or try new strategies based on the feedback. This improves the athletes' technical skills and strategic thinking.
[0509] (Application Example 1)
[0510] Next, we will explain Application Example 1. In the following explanation, the data processing device 12 will be referred to as the "server," and the headset-type terminal 314 will be referred to as the "terminal."
[0511] Improving the operational efficiency of various machines is a crucial challenge in modern manufacturing environments. However, there are limited means to constantly monitor machine operation and provide appropriate feedback in real time. Conventional analysis methods make it difficult to quickly and accurately analyze machine operation data and generate appropriate improvement suggestions. Furthermore, even in situations where immediate implementation of improvement suggestions is required, the lack of appropriate information and communication means hinders efficient improvement.
[0512] The specific processing performed by the specific processing unit 290 of the data processing device 12 in Application Example 1 is realized by the following means.
[0513] In this invention, the server includes means for acquiring video footage of the machine's operation using a video acquisition mechanism, means for analyzing the operation using a generation algorithm and evaluating the machine's operation, and means for transmitting improvement suggestions to an external device via an information communication channel. This makes it possible to analyze the machine's operation in real time and quickly provide suggestions for efficient operation improvements.
[0514] A "video acquisition mechanism" is a device or system used to acquire video footage of a machine's operation, and plays a role in capturing the details of its movement.
[0515] A "generative algorithm" is a computational method or program used to analyze the movement of a machine from video footage and evaluate its position and activity.
[0516] "Means for analyzing motion" refers to processes and methods for analyzing video footage of machine movements and evaluating efficient operation.
[0517] An "improvement suggestion" is information that, based on the analysis results, presents specific actions to improve the machine's operation and increase efficiency.
[0518] "Information communication channels" refer to communication paths and network technologies used to transmit improvement suggestions to external devices.
[0519] An "external device" is a computer or mobile device used to receive improvement suggestions sent from the server.
[0520] The system that realizes this invention consists of a combination of an image acquisition mechanism, a generation algorithm, an information communication channel, and external devices. This makes it possible to efficiently monitor and analyze the operation of machinery in a factory and provide improvement suggestions.
[0521] The server uses dedicated cameras to capture video footage of the machinery in operation within the factory. High-resolution cameras can record various machine movements in detail. This allows for precise analysis without any missing information.
[0522] After acquiring video footage, the server uses a generation algorithm to analyze the captured motion video. This algorithm tracks the machine's position and evaluates the efficiency and precision of its movements. For example, it can identify problems such as suboptimal machine movement or inappropriate work speed. TensorFlow or PyTorch can be used as the generation algorithm.
[0523] Based on the analysis results, the server generates improvement suggestions. These include areas for improvement in machine operation and action plans for more efficient operation. These suggestions are output in a format that is easy for the user to understand.
[0524] External devices such as smart glasses and tablet terminals are used. The generated improvement suggestions are transmitted to the external devices via the information communication channel, and the user can view them in real time. As a result, the user can quickly work on improving the machine's operation.
[0525] As a concrete example, suppose a camera installed on a factory robot arm captures and analyzes its movements, revealing that its trajectory is not optimal. The generated AI model then generates specific trajectory correction suggestions to improve the efficiency of the movement, and presents them as a prompt message such as, "Analyze the current movement of the robot arm and compare it to the most efficient movement pattern. Generate feedback on areas for improvement and suggested improvements, and present them to the worker."
[0526] In this way, the machine's operation can always be kept in an optimal state, leading to increased productivity.
[0527] The flow of a specific process in Application Example 1 will be explained using Figure 12.
[0528] Step 1:
[0529] The server uses dedicated cameras to capture video footage of machinery in operation within the factory. By using high-resolution cameras, even subtle changes in movement are captured in detail. The input is real-time video data, and the output is saved video files for analysis.
[0530] Step 2:
[0531] The server transmits the acquired video data to the cloud server via the 5G communication network. This enables rapid data transfer, allowing for immediate transition to video analysis without delay. The input is the stored video file, and the output is the prepared analysis data on the cloud server.
[0532] Step 3:
[0533] The server analyzes the transmitted video data using a generative AI model located in the cloud. The input is analysis data on the cloud server, and data correction and evaluation are performed to track the machine's operating position, speed, and motion trajectory. The output is the analysis result showing the efficiency and accuracy of the operation.
[0534] Step 4:
[0535] The server generates suggestions for improving machine operation based on the analysis results. Using the generated AI model, if the operation is determined to be inefficient, it generates specific suggestions and action plans for optimization. The input is the analysis results, and the output is information including improvement suggestions.
[0536] Step 5:
[0537] The server transmits the generated improvement suggestions to an external terminal via an information communication channel. Smart glasses or tablet devices receive this information and present it to the user in real time. The input is information including improvement suggestions, and the output is the specific improvement suggestions displayed on the external terminal. Based on the received suggestions, the user can adjust the machine and improve operational efficiency.
[0538] Furthermore, an emotion engine that estimates the user's emotions may be incorporated. That is, the identification processing unit 290 may use the emotion identification model 59 to estimate the user's emotions and perform identification processing using the user's emotions.
[0539] The present invention aims to provide more personalized feedback by incorporating an emotion engine that evaluates the user's psychological state into a system that supports the improvement of athletes' technical and strategic abilities.
[0540] System Configuration
[0541] 1. Video acquisition and analysis
[0542] Device: Smartphones or dedicated cameras are used to capture video footage of athletes during matches. This allows for real-time recording of the competition. The match footage is transmitted to a server via a 5G network.
[0543] 2. Analysis of motion and technique
[0544] Server: The server analyzes the received video using a generative model to analyze the athlete's movements and technical elements in detail. This analysis evaluates the athlete's form, position, strategic actions, etc., and identifies areas for improvement.
[0545] 3. Conducting an emotional assessment
[0546] Server: The emotion engine analyzes the competitor's facial expressions and audio data from the video to evaluate their psychological state. The emotion engine detects emotions such as happiness, concentration, and tension in real time.
[0547] 4. Customize feedback
[0548] Server: Based on the results of motion analysis and emotional assessment, it creates feedback tailored to the athlete's specific situation. This includes specific advice that takes psychological state into consideration, and technical improvement suggestions that are adjusted according to emotions.
[0549] 5. Providing feedback
[0550] Device: The generated feedback can be viewed by athletes and coaches on external devices. Users can check this on their smartphones or tablets during or immediately after a match and use it in their next training session.
[0551] Specific examples of use
[0552] User: For example, if a player is experiencing negative emotions during a match, the system can detect this and incorporate positive messages and emotionally uplifting advice into the feedback. The coach can then use this feedback to provide appropriate encouragement to help the player regain their composure.
[0553] This invention enables not only improvement in technical skills but also provides feedback that addresses the psychological aspects of athletes, thereby realizing more comprehensive training.
[0554] The following describes the processing flow.
[0555] Step 1:
[0556] Device: Before the start of the competition, smartphones or dedicated cameras are set up courtside to record the match footage. This video acquisition device records the movements and facial expressions of the athletes during the competition in high resolution.
[0557] Step 2:
[0558] Terminal: Recorded match footage is transmitted to the server via the 5G network. This transmission is done in real time, minimizing latency and enabling immediate analysis.
[0559] Step 3:
[0560] Server: Activates a generative model to analyze the received video. The video is divided frame by frame, allowing for detailed monitoring of the competitor's movements, position, and technical characteristics.
[0561] Step 4:
[0562] Server: Based on the video analysis, the server evaluates the athlete's physical movements, technical form, and strategic positioning. This includes generating technical metrics such as swing angle and motion speed.
[0563] Step 5:
[0564] Server: Using an emotion engine, the system analyzes the facial expressions and vocal patterns of athletes from video footage to assess their emotional state. This assessment aims to understand the psychological aspects of the athletes, measuring things like concentration and tension.
[0565] Step 6:
[0566] Server: Integrates the results of technical analysis and emotional assessment to generate feedback for the athlete. The feedback is optimized according to the athlete's emotional state. For example, if the athlete is anxious, advice for relaxation will be added.
[0567] Step 7:
[0568] Device: The generated feedback is sent to the athlete's or coach's smartphone or tablet. This allows the user to review the feedback and incorporate it into their next training plan.
[0569] Step 8:
[0570] User: Coaches use feedback to provide guidance to athletes, addressing both mental and technical aspects. Users can create plans to promote overall growth, including emotional well-being.
[0571] (Example 2)
[0572] Next, we will describe Example 2. In the following description, the data processing device 12 will be referred to as the "server," and the headset-type terminal 314 will be referred to as the "terminal."
[0573] Traditional athlete support systems have a weakness: they focus on technical feedback and lack feedback that considers the psychological aspects of athletes. This can hinder the overall improvement of athletes' abilities.
[0574] The identification process performed by the identification processing unit 290 of the data processing device 12 in Example 2 is realized by the following means.
[0575] In this invention, the server includes means for receiving match video acquired by video acquisition means as input, analyzing the actions of athletes from the video using a generative model that analyzes the video, generating customized feedback based on the actions and psychological state, and transmitting the feedback to an information processing device via a communication network. This enables comprehensive support that takes into account not only the technical abilities of athletes but also their psychological aspects.
[0576] "Video acquisition means" refers to devices or equipment used to collect video footage of athletes during matches, and plays a role in preparing the footage for later analysis.
[0577] A "generative model" is a program or algorithm used to analyze an athlete's movements and technical elements from input video data, enabling detailed analysis of the athlete.
[0578] A "communication network" is a medium for transferring data and serves as a pathway for sending feedback from a server to an information processing device.
[0579] An "information processing device" is a device that receives data and feedback transmitted from a server, and displays or processes it. It refers to a terminal used by a user to check feedback.
[0580] An "emotional engine" refers to software or a system that analyzes an athlete's psychological state from their facial expressions and tone of voice to evaluate their emotional state.
[0581] "Feedback" refers to advice and guidance generated based on an athlete's actions and psychological state, and is information aimed at improving the athlete's technical and strategic abilities.
[0582] This system is designed to support the improvement of athletes' technical and psychological abilities. The main components of the system are video acquisition means, generative models, emotion engines, information processing devices, and communication networks.
[0583] The server receives video footage capturing the athletes' movements during a match. This video is acquired by terminals using smartphones or dedicated cameras and transmitted to the server via a communication network. The server analyzes this video data using a generative model to analyze the athletes' movements and technical elements in detail. The generative model uses machine learning algorithms to extract and quantify the characteristics of the movements. For example, it can measure the athlete's form, the speed of their actions, and the angles of their movements.
[0584] Next, the server uses an emotion engine to analyze the competitor's facial expressions and audio data from the video to evaluate their psychological state. This emotion engine is implemented by combining facial recognition technology and audio analysis technology, and it determines feelings of happiness, tension, concentration, etc., in real time.
[0585] Based on these analysis results, the server generates personalized feedback for the competitor. This feedback includes areas for technical improvement and strategic advice tailored to their psychological state. The feedback is transmitted via the communication network to an information processing device, i.e., the user's terminal.
[0586] The device displays received feedback and is designed to be easily understood by the user. Users can utilize the feedback to improve their next training session. For example, if a negative emotional state is detected during a match, positive messages and suggestions for improvement are provided, enabling coaching that takes the athlete's mental state into consideration.
[0587] Example of a prompt
[0588] "Analyze video footage of athletes during matches to assess areas for improvement in their movements and emotional state, and generate personalized feedback."
[0589] This system enables training support that integrates technology and emotion.
[0590] The flow of the specific processing in Example 2 will be explained using Figure 13.
[0591] Step 1:
[0592] The device acquires video footage of athletes during matches. It uses smartphones or dedicated cameras as input. The output is real-time match video data. This video data is streamed to a server using a 5G network.
[0593] Step 2:
[0594] The server receives video data transmitted from the terminal. The input is high-resolution video data from the terminal. Data preprocessing is performed to convert this data into a format that can be analyzed frame by frame. The output is video data that has been prepared in an analysis-ready format.
[0595] Step 3:
[0596] The server uses a generative AI model to analyze the movements of athletes from pre-processed video data. The input is formatted video data. The generative AI model performs data calculations to extract the athletes' movement patterns and technical characteristics from the video frames. The output is the motion analysis results obtained as numerical data.
[0597] Step 4:
[0598] The server uses an emotion engine to analyze the athlete's facial expressions and voice tone from video data, thereby evaluating their psychological state. The input is video data that has already undergone motion analysis. The emotion engine performs image recognition and voice analysis to quantify the emotional state. The output is data representing the athlete's real-time psychological state.
[0599] Step 5:
[0600] The server integrates motion analysis results and psychological state assessments to generate personalized feedback. The input consists of motion and psychological state data. Based on this, it generates feedback that includes specific advice and technical improvement suggestions for the athlete. The output is customized feedback.
[0601] Step 6:
[0602] The terminal receives feedback sent from the server and displays it in a format viewable by the user. The input is feedback data from the server. The terminal presents this data in a user-friendly interface. The output is feedback information that the user can visually confirm.
[0603] Users can utilize this feedback to improve their next training session.
[0604] (Application Example 2)
[0605] Next, we will explain Application Example 2. In the following explanation, the data processing device 12 will be referred to as the "server," and the headset-type terminal 314 will be referred to as the "terminal."
[0606] There is a challenge in understanding in real time what customers are feeling in a store and providing optimal customer service tailored to those emotions. Traditional methods rely on sales staff subjectively judging customers' expressions and attitudes, lacking objective means to make appropriate approaches. This makes it difficult to improve customer satisfaction and increase sales efficiency.
[0607] The specific processing performed by the specific processing unit 290 of the data processing device 12 in Application Example 2 is realized by the following means.
[0608] In this invention, the server includes means for taking customer video acquired by a video acquisition device as input and analyzing the customer's facial expressions using a generative model that analyzes the video; means for evaluating the customer's emotional state and generating feedback; and means for transmitting the feedback to the sales staff's terminal via a communication network. This enables optimal customer service based on the customer's emotions.
[0609] "Image acquisition equipment" refers to devices used to capture customers' facial expressions and movements, and includes cameras, smart devices, and other similar devices.
[0610] A "generative model" is an algorithm that analyzes input video data to estimate the emotional state of a customer, and it utilizes machine learning technology.
[0611] "Methods of analysis" refers to the process of analyzing video data using generative models to identify emotions and psychological states from customers' facial expressions and actions.
[0612] A "communication network" is the infrastructure used to send and receive analyzed information between a server and the terminals of sales staff, and it utilizes the internet or wireless communication.
[0613] "Feedback" refers to information provided to sales staff as advice on customer service and product recommendations based on the customer's emotional state.
[0614] To implement this invention, cameras or smart devices must be installed in the store as video acquisition devices to capture customers' facial expressions and movements. This makes it possible to acquire video data in real time. This video data is transmitted to a server via a communication network such as 5G.
[0615] The server analyzes the received video data using a generative AI model. This model utilizes machine learning algorithms to estimate the customer's emotional state from their facial expressions and movements. During the analysis, the video data is broken down frame by frame, and the customer's facial expressions in each frame are analyzed. Existing video analysis software, such as the Google Cloud Vision API, can be used for this process. If audio data is present, customer speech is collected using a noise-canceling microphone, and this content is also incorporated into the emotion analysis.
[0616] The analysis results are generated as feedback, notifying sales staff of the customer's emotional state on their terminals. This feedback is provided in real time or near real time, allowing sales staff to provide optimal customer service based on it. Specifically, the system displays prompts on the sales staff's terminals such as, "Generate advice suggesting an appropriate response based on the customer's current emotional state. If the customer appears confused, suggest popular products for first-time customers." This kind of feedback allows sales staff to immediately take actions that lead to increased customer satisfaction.
[0617] The flow of a specific process in Application Example 2 will be explained using Figure 14.
[0618] Step 1:
[0619] The user (store camera) acquires video data when a customer enters the store. The camera captures the customer's movements and facial expressions inside the store, acquiring video data in a format that can be analyzed by a generative AI model. The input is a real-time video stream, and the output is analyzable digital data.
[0620] Step 2:
[0621] The terminal (camera or smart device) transmits video data to the server via a 5G communication network. Due to the high bandwidth network, data is transmitted in near real-time. The input is digital data from the video acquisition device, and the output is the video data transmitted to the server.
[0622] Step 3:
[0623] The server uses a generated AI model to analyze the received video data. The server extracts the customer's face frame by frame from the video data and performs facial expression analysis. In this analysis process, video analysis algorithms such as the Google Cloud Vision API are used. The input is the video data sent to the server, and the output is data related to the customer's emotional state corresponding to their facial expressions.
[0624] Step 4:
[0625] The server uses the analysis results to generate feedback and create appropriate customer service advice based on the customer's emotional state. The generated feedback includes specific suggestions to help sales staff provide customer service. The input is emotional state data obtained through analysis, and the output is feedback information for sales staff.
[0626] Step 5:
[0627] The server sends the feedback generated in the previous step to the sales staff's terminal via the communication network. The staff checks this information on their smartphone or tablet and uses it to assist customers. The input is the feedback information on the server, and the output is the advice displayed on the sales staff's terminal.
[0628] Step 6:
[0629] The terminal (the sales staff's smart device) checks the received feedback and modifies its approach to the customer. For example, if it detects that the customer is confused, the staff will provide service according to the prompt message, "Suggest popular products to first-time customers." The input is the feedback displayed on the terminal, and the output is the specific action taken by the sales staff based on that feedback.
[0630] The specific processing unit 290 transmits the result of the specific processing to the headset terminal 314. In the headset terminal 314, the control unit 46A causes the speaker 240 and display 343 to output the result of the specific processing. The microphone 238 acquires audio indicating user input for the result of the specific processing. The control unit 46A transmits the audio data indicating user input acquired by the microphone 238 to the data processing unit 12. In the data processing unit 12, the specific processing unit 290 acquires the audio data.
[0631] Data generation model 58 is a type of so-called generative AI (Artificial Intelligence). One example of data generation model 58 is ChatGPT (Internet search<URL: https: / / openai.com / blog / chatgpt> ), Gemini (Internet search) <url: https: gemini.google.com ?hl="ja">Examples of generative AI include the following. The data generation model 58 is obtained by performing deep learning on a neural network. The data generation model 58 is input with prompts containing instructions, and with inference data such as audio data representing speech, text data representing text, and image data representing images. The data generation model 58 infers from the input inference data according to the instructions indicated by the prompts, and outputs the inference results in data formats such as audio data and text data. Here, inference refers to, for example, analysis, classification, prediction, and / or summarization.
[0632] In the above embodiment, an example was given in which specific processing is performed by the data processing device 12, but the technology of this disclosure is not limited thereto, and specific processing may also be performed by the headset terminal 314.
[0633] [Fourth Embodiment]
[0634] Figure 7 shows an example of the configuration of the data processing system 410 according to the fourth embodiment.
[0635] As shown in Figure 7, the data processing system 410 includes a data processing device 12 and a robot 414. An example of the data processing device 12 is a server.
[0636] The data processing device 12 comprises a computer 22, a database 24, and a communication interface 26. The computer 22 is an example of a "computer" related to the technology of this disclosure. The computer 22 comprises a processor 28, RAM 30, and storage 32. The processor 28, RAM 30, and storage 32 are connected to a bus 34. The database 24 and the communication interface 26 are also connected to the bus 34. The communication interface 26 is connected to a network 54. An example of the network 54 is a WAN (Wide Area Network) and / or a LAN (Local Area Network).
[0637] The robot 414 includes a computer 36, a microphone 238, a speaker 240, a camera 42, a communication interface 44, and a controlled object 443. The computer 36 includes a processor 46, RAM 48, and storage 50. The processor 46, RAM 48, and storage 50 are connected to a bus 52. The microphone 238, speaker 240, camera 42, and controlled object 443 are also connected to the bus 52.
[0638] The microphone 238 receives voice signals from the user 20 and receives instructions from the user 20. The microphone 238 captures the voice signals from the user 20, converts the captured voice into audio data, and outputs it to the processor 46. The speaker 240 outputs audio according to the instructions from the processor 46.
[0639] Camera 42 is a small digital camera equipped with an optical system including a lens, aperture, and shutter, and an image sensor such as a CMOS (Complementary Metal-Oxide-Semiconductor) image sensor or a CCD (Charge Coupled Device) image sensor, and captures images of the area around the user 20 (for example, an imaging range defined by a field of view equivalent to the width of a typical healthy person's field of vision).
[0640] Communication interface 44 is connected to network 54. Communication interfaces 44 and 26 are responsible for the exchange of various information between processor 46 and processor 28 via network 54. The exchange of various information between processor 46 and processor 28 using communication interfaces 44 and 26 is performed in a secure manner.
[0641] The controlled object 443 includes a display device, LEDs in the eyes, and motors that drive the arms, hands, and feet. The posture and gestures of the robot 414 are controlled by controlling the motors of the arms, hands, and feet. Some of the robot 414's emotions can be expressed by controlling these motors. Furthermore, the robot 414's facial expressions can also be expressed by controlling the illumination state of the LEDs in its eyes.
[0642] Figure 8 shows an example of the main functions of the data processing device 12 and the robot 414. As shown in Figure 8, the data processing device 12 performs specific processing using the processor 28. The storage 32 stores the specific processing program 56.
[0643] The specific processing program 56 is an example of a "program" relating to the technology of this disclosure. The processor 28 reads the specific processing program 56 from the storage 32 and executes the read specific processing program 56 on the RAM 30. The specific processing is realized by the processor 28 operating as a specific processing unit 290 in accordance with the specific processing program 56 executed on the RAM 30.
[0644] The storage 32 stores the data generation model 58 and the emotion identification model 59. The data generation model 58 and the emotion identification model 59 are used by the identification processing unit 290.
[0645] In robot 414, the processor 46 performs the reception output processing. The storage 50 stores the reception output program 60. The processor 46 reads the reception output program 60 from the storage 50 and executes the read reception output program 60 on the RAM 48. The reception output processing is realized by the processor 46 operating as a control unit 46A according to the reception output program 60 executed on the RAM 48.
[0646] Next, the specific processing performed by the specific processing unit 290 of the data processing device 12 will be described. In the following description, the data processing device 12 will be referred to as the "server" and the robot 414 as the "terminal".
[0647] This invention relates to an automated analysis system for efficiently improving the technical and strategic capabilities of athletes. This system utilizes the functions of a video acquisition device, a generative model, a communication network, and an external terminal to provide immediate and effective feedback to athletes.
[0648] System Configuration
[0649] 1. Video acquisition device
[0650] Terminal: Using smartphones or dedicated cameras, video data from matches is acquired in real time. This video acquisition device is installed in stadiums and training grounds to record the movements of athletes in detail.
[0651] 2. Transmission of video data
[0652] Terminal: Acquired video data is transmitted to the server using the 5G communication network. Even high-capacity data is transferred quickly, enabling real-time analysis.
[0653] 3. Conduct motion analysis
[0654] Server: The transmitted video data is analyzed by a generative model. The model tracks the athlete's body position and movements, and evaluates their actions and technical elements during the competition. For example, it detects the speed and angle of the athlete's racket swing and compares it to an ideal movement pattern.
[0655] 4. Feedback generation
[0656] Server: Based on the analysis results, it generates individual feedback for each athlete. This feedback evaluates areas for improvement and successful techniques for each athlete, and provides them as goals for further practice.
[0657] 5. Communication and display of feedback
[0658] Device: The generated feedback is sent via the communication network to external devices used by athletes and coaches. Users can review this feedback on their smartphones or tablets and use it to improve their next training session.
[0659] Specific examples of use
[0660] User: If a badminton player's smash swing is incorrect during a match, the system instantly generates feedback on the angle of that swing. The coach, upon receiving this feedback, can then focus their instruction on improving specific forms during the next practice session.
[0661] This invention realizes a system that provides precise technical and strategic feedback through clear image analysis, thereby promoting efficient performance improvement in competitions.
[0662] The following describes the processing flow.
[0663] Step 1:
[0664] Devices: Before the start of the competition, smartphones or dedicated cameras are set up courtside to begin recording the match footage in real time. These devices continuously record the players' movements clearly in high resolution.
[0665] Step 2:
[0666] Terminal: Captured video data is streamed to the server using the 5G network. This allows large amounts of video data to reach the server quickly.
[0667] Step 3:
[0668] Server: The server divides the received video data into frames and performs preprocessing to extract the players' movements and positions from each frame. This process also includes filtering to reduce background noise and improve analysis accuracy.
[0669] Step 4:
[0670] Server: Uses a generative model to analyze the technical actions of the players in each frame (e.g., racket swing, player positioning). At this stage, the players' actions are compared to a specific set of rules and criteria to calculate technical evaluation metrics.
[0671] Step 5:
[0672] Server: Based on technical evaluation, it assesses the strategic approach of players. It analyzes the strategies players employ during matches, detecting, for example, delays in defensive movement or positioning during attacks.
[0673] Step 6:
[0674] Server: Based on the analysis results, create specific feedback on areas that need improvement and skills that should be strengthened. This feedback will include technical advice and strategic suggestions.
[0675] Step 7:
[0676] Terminal: Generated feedback is sent to an external terminal managed by the athlete or coach. Users can check this in real time or immediately after the match.
[0677] Step 8:
[0678] User: Based on the feedback, the coach understands areas where the athlete needs improvement and determines the focus points for the next training session. This feedback allows for efficient guidance to improve the athlete's skills and strategies.
[0679] (Example 1)
[0680] Next, we will describe Example 1. In the following description, the data processing device 12 will be referred to as the "server" and the robot 414 as the "terminal".
[0681] The present invention aims to provide a system that offers immediate, highly accurate feedback to efficiently improve athletes' technical skills and strategies. Conventional technologies have made real-time motion analysis and immediate identification of specific areas for improvement difficult, preventing rapid reflection of improvements in practice and competition performance. Therefore, there is a need for a system that automates the analysis of athletes' movements and efficiently provides feedback, including specific improvement suggestions.
[0682] The identification process performed by the identification processing unit 290 of the data processing device 12 in Example 1 is realized by the following means.
[0683] In this invention, the server includes means for taking competition video acquired by a video acquisition device as input and analyzing the athlete's posture from the video using a generative algorithm for analyzing the video; means for evaluating the athlete's technical and strategic elements based on the analysis results and generating feedback including specific improvement suggestions; and means for transmitting the feedback to an external device via a communication path. This allows athletes to immediately grasp areas for improvement in their movements and use that information to improve their technical skills and strategies in their next practice or match.
[0684] "Video acquisition equipment" refers to devices that can capture the actions of athletes and the details of a match in real time, and includes smartphones and dedicated cameras.
[0685] A "generative algorithm" is a mathematical model or method for analyzing input video data, particularly for tracking and evaluating the physical movements of athletes.
[0686] A "communication path" refers to the network used when sending and receiving data, and includes 5G networks and other technologies that enable high-speed and stable communication.
[0687] "External devices" refer to devices used by athletes and coaches to receive feedback, and include terminals such as smartphones and tablets.
[0688] "Feedback" refers to data and advice provided to improve athletes' skills and strategies, based on the results of analysis using generative algorithms.
[0689] "Analyzing posture" is the process of tracking an athlete's body position and movements, and using that data to evaluate specific technical and strategic elements.
[0690] "Technical elements" are indicators that show the accuracy and efficiency of individual movements and techniques in a competition, and are used to numerically evaluate an athlete's performance.
[0691] "Strategic elements" are criteria used to evaluate the strategies and tactics that athletes choose during a match, particularly analyzing their reactions to the opponent's movements and their choice of plays.
[0692] This invention provides a system that effectively analyzes the movements of athletes and provides immediate, specific feedback. This system is primarily implemented using video acquisition equipment, a generative AI model, a communication path, and external devices. Details and specific examples of each element are provided below.
[0693] Video acquisition equipment is a device used to capture the posture and movements of athletes in real time while they are playing. Specifically, this includes smartphones and dedicated cameras. These devices are installed in stadiums and practice fields to record detailed video data.
[0694] The server receives the acquired video data and analyzes it using a generative AI model. The generative AI model tracks the athletes' body positions and movements within the video data and evaluates the technical and strategic elements during the competition. Specifically, it identifies the joint positions and movement speeds of the athletes' bodies and compares them to ideal movement patterns.
[0695] The server then automatically generates feedback based on the analysis results. This feedback includes guidance on areas for improvement in the athlete's movements and specific strategies. For example, it might suggest detailed improvements such as, "The angle of your racket during a smash is 5 degrees shallower than ideal."
[0696] The terminal transmits feedback generated by the server to an external device via a communication path. The user receives and confirms the feedback through an external device such as a smartphone or tablet. This makes it possible to immediately implement specific improvement measures during the next practice session.
[0697] Based on the feedback provided, users can improve their technology and adjust their strategies. Through this process, competitors can quickly improve their performance.
[0698] Example of a prompt:
[0699] "Analyze the angle of the player's smash swing and generate feedback showing the difference from the ideal angle."
[0700] In this way, the system automates a series of processes including video acquisition, data analysis, feedback generation, transmission, and reception, providing strong support for the technical and strategic improvement of athletes.
[0701] The flow of the specific processing in Example 1 will be explained using Figure 11.
[0702] Step 1:
[0703] The terminal uses video acquisition equipment to capture the movements of athletes. Smartphones or dedicated cameras are set up to film movements in real time at the stadium or training ground. The input is video data including the movements of athletes, which is recorded in digital format in preparation for transmission to the server. Specifically, the movements of athletes, the operation of equipment, and interactions with the environment are captured on video.
[0704] Step 2:
[0705] The terminal transfers the captured video data to the server via the communication path. High-capacity video data is transmitted rapidly using a 5G network. The input is the video data acquired in step 1, and the output is the compressed and decompressed video data arriving at the server. At this stage, communication optimization is performed to minimize delays during data transfer.
[0706] Step 3:
[0707] The server inputs the received video data into the generating AI model and begins analyzing the data. The input is the video data sent in step 2, and the output is the evaluation results of the athlete's body position and movement. Specifically, the process involves estimating the athlete's joint positions, calculating the speed and direction of movement, and comparing it with ideal movement for each frame.
[0708] Step 4:
[0709] The server generates feedback for the athlete based on the analysis results. The input is the evaluation results obtained in step 3, and the output is feedback data that includes specific improvement suggestions and quantified metrics. For example, it may include detailed analysis results regarding the athlete's swing angle and speed.
[0710] Step 5:
[0711] The terminal transmits the generated feedback to an external device via a communication path. The input is the feedback data received from the server, and the output is the information presented in a format that the user can verify. At this stage, the feedback is transmitted to the user via a smartphone or tablet.
[0712] Step 6:
[0713] Users utilize the feedback they receive to improve their movements and tactics in competition. Specifically, they adjust their training content or try new strategies based on the feedback. This improves the athletes' technical skills and strategic thinking.
[0714] (Application Example 1)
[0715] Next, we will explain Application Example 1. In the following explanation, the data processing device 12 will be referred to as the "server" and the robot 414 as the "terminal".
[0716] Improving the operational efficiency of various machines is a crucial challenge in modern manufacturing environments. However, there are limited means to constantly monitor machine operation and provide appropriate feedback in real time. Conventional analysis methods make it difficult to quickly and accurately analyze machine operation data and generate appropriate improvement suggestions. Furthermore, even in situations where immediate implementation of improvement suggestions is required, the lack of appropriate information and communication means hinders efficient improvement.
[0717] The specific processing performed by the specific processing unit 290 of the data processing device 12 in Application Example 1 is realized by the following means.
[0718] In this invention, the server includes means for acquiring video footage of the machine's operation using a video acquisition mechanism, means for analyzing the operation using a generation algorithm and evaluating the machine's operation, and means for transmitting improvement suggestions to an external device via an information communication channel. This makes it possible to analyze the machine's operation in real time and quickly provide suggestions for efficient operation improvements.
[0719] A "video acquisition mechanism" is a device or system used to acquire video footage of a machine's operation, and plays a role in capturing the details of its movement.
[0720] A "generative algorithm" is a computational method or program used to analyze the movement of a machine from video footage and evaluate its position and activity.
[0721] "Means for analyzing motion" refers to processes and methods for analyzing video footage of machine movements and evaluating efficient operation.
[0722] An "improvement suggestion" is information that, based on the analysis results, presents specific actions to improve the machine's operation and increase efficiency.
[0723] "Information communication channels" refer to communication paths and network technologies used to transmit improvement suggestions to external devices.
[0724] An "external device" is a computer or mobile device used to receive improvement suggestions sent from the server.
[0725] The system that realizes this invention consists of a combination of an image acquisition mechanism, a generation algorithm, an information communication channel, and external devices. This makes it possible to efficiently monitor and analyze the operation of machinery in a factory and provide improvement suggestions.
[0726] The server uses dedicated cameras to capture video footage of the machinery in operation within the factory. High-resolution cameras can record various machine movements in detail. This allows for precise analysis without any missing information.
[0727] After acquiring video footage, the server uses a generation algorithm to analyze the captured motion video. This algorithm tracks the machine's position and evaluates the efficiency and precision of its movements. For example, it can identify problems such as suboptimal machine movement or inappropriate work speed. TensorFlow or PyTorch can be used as the generation algorithm.
[0728] Based on the analysis results, the server generates improvement suggestions. These include areas for improvement in machine operation and action plans for more efficient operation. These suggestions are output in a format that is easy for the user to understand.
[0729] External devices such as smart glasses and tablet terminals are used. The generated improvement suggestions are transmitted to the external devices via the information communication channel, and the user can view them in real time. As a result, the user can quickly work on improving the machine's operation.
[0730] As a concrete example, suppose a camera installed on a factory robot arm captures and analyzes its movements, revealing that its trajectory is not optimal. The generated AI model then generates specific trajectory correction suggestions to improve the efficiency of the movement, and presents them as a prompt message such as, "Analyze the current movement of the robot arm and compare it to the most efficient movement pattern. Generate feedback on areas for improvement and suggested improvements, and present them to the worker."
[0731] In this way, the machine's operation can always be kept in an optimal state, leading to increased productivity.
[0732] The flow of a specific process in Application Example 1 will be explained using Figure 12.
[0733] Step 1:
[0734] The server uses dedicated cameras to capture video footage of machinery in operation within the factory. By using high-resolution cameras, even subtle changes in movement are captured in detail. The input is real-time video data, and the output is saved video files for analysis.
[0735] Step 2:
[0736] The server transmits the acquired video data to the cloud server via the 5G communication network. This enables rapid data transfer, allowing for immediate transition to video analysis without delay. The input is the stored video file, and the output is the prepared analysis data on the cloud server.
[0737] Step 3:
[0738] The server analyzes the transmitted video data using a generative AI model located in the cloud. The input is analysis data on the cloud server, and data correction and evaluation are performed to track the machine's operating position, speed, and motion trajectory. The output is the analysis result showing the efficiency and accuracy of the operation.
[0739] Step 4:
[0740] The server generates suggestions for improving machine operation based on the analysis results. Using the generated AI model, if the operation is determined to be inefficient, it generates specific suggestions and action plans for optimization. The input is the analysis results, and the output is information including improvement suggestions.
[0741] Step 5:
[0742] The server transmits the generated improvement suggestions to an external terminal via an information communication channel. Smart glasses or tablet devices receive this information and present it to the user in real time. The input is information including improvement suggestions, and the output is the specific improvement suggestions displayed on the external terminal. Based on the received suggestions, the user can adjust the machine and improve operational efficiency.
[0743] Furthermore, an emotion engine that estimates the user's emotions may be incorporated. That is, the identification processing unit 290 may use the emotion identification model 59 to estimate the user's emotions and perform identification processing using the user's emotions.
[0744] The present invention aims to provide more personalized feedback by incorporating an emotion engine that evaluates the user's psychological state into a system that supports the improvement of athletes' technical and strategic abilities.
[0745] System Configuration
[0746] 1. Video acquisition and analysis
[0747] Device: Smartphones or dedicated cameras are used to capture video footage of athletes during matches. This allows for real-time recording of the competition. The match footage is transmitted to a server via a 5G network.
[0748] 2. Analysis of motion and technique
[0749] Server: The server analyzes the received video using a generative model to analyze the athlete's movements and technical elements in detail. This analysis evaluates the athlete's form, position, strategic actions, etc., and identifies areas for improvement.
[0750] 3. Conducting an emotional assessment
[0751] Server: The emotion engine analyzes the competitor's facial expressions and audio data from the video to evaluate their psychological state. The emotion engine detects emotions such as happiness, concentration, and tension in real time.
[0752] 4. Customize feedback
[0753] Server: Based on the results of motion analysis and emotional assessment, it creates feedback tailored to the athlete's specific situation. This includes specific advice that takes psychological state into consideration, and technical improvement suggestions that are adjusted according to emotions.
[0754] 5. Providing feedback
[0755] Device: The generated feedback can be viewed by athletes and coaches on external devices. Users can check this on their smartphones or tablets during or immediately after a match and use it in their next training session.
[0756] Specific examples of use
[0757] User: For example, if a player is experiencing negative emotions during a match, the system can detect this and incorporate positive messages and emotionally uplifting advice into the feedback. The coach can then use this feedback to provide appropriate encouragement to help the player regain their composure.
[0758] This invention enables not only improvement in technical skills but also provides feedback that addresses the psychological aspects of athletes, thereby realizing more comprehensive training.
[0759] The following describes the processing flow.
[0760] Step 1:
[0761] Device: Before the start of the competition, smartphones or dedicated cameras are set up courtside to record the match footage. This video acquisition device records the movements and facial expressions of the athletes during the competition in high resolution.
[0762] Step 2:
[0763] Terminal: Recorded match footage is transmitted to the server via the 5G network. This transmission is done in real time, minimizing latency and enabling immediate analysis.
[0764] Step 3:
[0765] Server: Activates a generative model to analyze the received video. The video is divided frame by frame, allowing for detailed monitoring of the competitor's movements, position, and technical characteristics.
[0766] Step 4:
[0767] Server: Based on the video analysis, the server evaluates the athlete's physical movements, technical form, and strategic positioning. This includes generating technical metrics such as swing angle and motion speed.
[0768] Step 5:
[0769] Server: Using an emotion engine, the system analyzes the facial expressions and vocal patterns of athletes from video footage to assess their emotional state. This assessment aims to understand the psychological aspects of the athletes, measuring things like concentration and tension.
[0770] Step 6:
[0771] Server: Integrates the results of technical analysis and emotional assessment to generate feedback for the athlete. The feedback is optimized according to the athlete's emotional state. For example, if the athlete is anxious, advice for relaxation will be added.
[0772] Step 7:
[0773] Device: The generated feedback is sent to the athlete's or coach's smartphone or tablet. This allows the user to review the feedback and incorporate it into their next training plan.
[0774] Step 8:
[0775] User: Coaches use feedback to provide guidance to athletes, addressing both mental and technical aspects. Users can create plans to promote overall growth, including emotional well-being.
[0776] (Example 2)
[0777] Next, we will describe Example 2. In the following description, the data processing device 12 will be referred to as the "server" and the robot 414 as the "terminal".
[0778] Traditional athlete support systems have a weakness: they focus on technical feedback and lack feedback that considers the psychological aspects of athletes. This can hinder the overall improvement of athletes' abilities.
[0779] The identification process performed by the identification processing unit 290 of the data processing device 12 in Example 2 is realized by the following means.
[0780] In this invention, the server includes means for receiving match video acquired by video acquisition means as input, analyzing the actions of athletes from the video using a generative model that analyzes the video, generating customized feedback based on the actions and psychological state, and transmitting the feedback to an information processing device via a communication network. This enables comprehensive support that takes into account not only the technical abilities of athletes but also their psychological aspects.
[0781] "Video acquisition means" refers to devices or equipment used to collect video footage of athletes during matches, and plays a role in preparing the footage for later analysis.
[0782] A "generative model" is a program or algorithm used to analyze an athlete's movements and technical elements from input video data, enabling detailed analysis of the athlete.
[0783] A "communication network" is a medium for transferring data and serves as a pathway for sending feedback from a server to an information processing device.
[0784] An "information processing device" is a device that receives data and feedback transmitted from a server, and displays or processes it. It refers to a terminal used by a user to check feedback.
[0785] An "emotional engine" refers to software or a system that analyzes an athlete's psychological state from their facial expressions and tone of voice to evaluate their emotional state.
[0786] "Feedback" refers to advice and guidance generated based on an athlete's actions and psychological state, and is information aimed at improving the athlete's technical and strategic abilities.
[0787] This system is designed to support the improvement of athletes' technical and psychological abilities. The main components of the system are video acquisition means, generative models, emotion engines, information processing devices, and communication networks.
[0788] The server receives video footage capturing the athletes' movements during a match. This video is acquired by terminals using smartphones or dedicated cameras and transmitted to the server via a communication network. The server analyzes this video data using a generative model to analyze the athletes' movements and technical elements in detail. The generative model uses machine learning algorithms to extract and quantify the characteristics of the movements. For example, it can measure the athlete's form, the speed of their actions, and the angles of their movements.
[0789] Next, the server uses an emotion engine to analyze the competitor's facial expressions and audio data from the video to evaluate their psychological state. This emotion engine is implemented by combining facial recognition technology and audio analysis technology, and it determines feelings of happiness, tension, concentration, etc., in real time.
[0790] Based on these analysis results, the server generates personalized feedback for the competitor. This feedback includes areas for technical improvement and strategic advice tailored to their psychological state. The feedback is transmitted via the communication network to an information processing device, i.e., the user's terminal.
[0791] The device displays received feedback and is designed to be easily understood by the user. Users can utilize the feedback to improve their next training session. For example, if a negative emotional state is detected during a match, positive messages and suggestions for improvement are provided, enabling coaching that takes the athlete's mental state into consideration.
[0792] Example of a prompt
[0793] "Analyze video footage of athletes during matches to assess areas for improvement in their movements and emotional state, and generate personalized feedback."
[0794] This system enables training support that integrates technology and emotion.
[0795] The flow of the specific processing in Example 2 will be explained using Figure 13.
[0796] Step 1:
[0797] The device acquires video footage of athletes during matches. It uses smartphones or dedicated cameras as input. The output is real-time match video data. This video data is streamed to a server using a 5G network.
[0798] Step 2:
[0799] The server receives video data transmitted from the terminal. The input is high-resolution video data from the terminal. Data preprocessing is performed to convert this data into a format that can be analyzed frame by frame. The output is video data that has been prepared in an analysis-ready format.
[0800] Step 3:
[0801] The server uses a generative AI model to analyze the movements of athletes from pre-processed video data. The input is formatted video data. The generative AI model performs data calculations to extract the athletes' movement patterns and technical characteristics from the video frames. The output is the motion analysis results obtained as numerical data.
[0802] Step 4:
[0803] The server uses an emotion engine to analyze the athlete's facial expressions and voice tone from video data, thereby evaluating their psychological state. The input is video data that has already undergone motion analysis. The emotion engine performs image recognition and voice analysis to quantify the emotional state. The output is data representing the athlete's real-time psychological state.
[0804] Step 5:
[0805] The server integrates motion analysis results and psychological state assessments to generate personalized feedback. The input consists of motion and psychological state data. Based on this, it generates feedback that includes specific advice and technical improvement suggestions for the athlete. The output is customized feedback.
[0806] Step 6:
[0807] The terminal receives feedback sent from the server and displays it in a format viewable by the user. The input is feedback data from the server. The terminal presents this data in a user-friendly interface. The output is feedback information that the user can visually confirm.
[0808] Users can utilize this feedback to improve their next training session.
[0809] (Application Example 2)
[0810] Next, we will explain application example 2. In the following explanation, the data processing device 12 will be referred to as the "server" and the robot 414 as the "terminal".
[0811] There is a challenge in understanding in real time what customers are feeling in a store and providing optimal customer service tailored to those emotions. Traditional methods rely on sales staff subjectively judging customers' expressions and attitudes, lacking objective means to make appropriate approaches. This makes it difficult to improve customer satisfaction and increase sales efficiency.
[0812] The specific processing performed by the specific processing unit 290 of the data processing device 12 in Application Example 2 is realized by the following means.
[0813] In this invention, the server includes means for taking customer video acquired by a video acquisition device as input and analyzing the customer's facial expressions using a generative model that analyzes the video; means for evaluating the customer's emotional state and generating feedback; and means for transmitting the feedback to the sales staff's terminal via a communication network. This enables optimal customer service based on the customer's emotions.
[0814] "Image acquisition equipment" refers to devices used to capture customers' facial expressions and movements, and includes cameras, smart devices, and other similar devices.
[0815] A "generative model" is an algorithm that analyzes input video data to estimate the emotional state of a customer, and it utilizes machine learning technology.
[0816] "Methods of analysis" refers to the process of analyzing video data using generative models to identify emotions and psychological states from customers' facial expressions and actions.
[0817] A "communication network" is the infrastructure used to send and receive analyzed information between a server and the terminals of sales staff, and it utilizes the internet or wireless communication.
[0818] "Feedback" refers to information provided to sales staff as advice on customer service and product recommendations based on the customer's emotional state.
[0819] To implement this invention, cameras or smart devices must be installed in the store as video acquisition devices to capture customers' facial expressions and movements. This makes it possible to acquire video data in real time. This video data is transmitted to a server via a communication network such as 5G.
[0820] The server analyzes the received video data using a generative AI model. This model utilizes machine learning algorithms to estimate the customer's emotional state from their facial expressions and movements. During the analysis, the video data is broken down frame by frame, and the customer's facial expressions in each frame are analyzed. Existing video analysis software, such as the Google Cloud Vision API, can be used for this process. If audio data is present, customer speech is collected using a noise-canceling microphone, and this content is also incorporated into the emotion analysis.
[0821] The analysis results are generated as feedback, notifying sales staff of the customer's emotional state on their terminals. This feedback is provided in real time or near real time, allowing sales staff to provide optimal customer service based on it. Specifically, the system displays prompts on the sales staff's terminals such as, "Generate advice suggesting an appropriate response based on the customer's current emotional state. If the customer appears confused, suggest popular products for first-time customers." This kind of feedback allows sales staff to immediately take actions that lead to increased customer satisfaction.
[0822] The flow of a specific process in Application Example 2 will be explained using Figure 14.
[0823] Step 1:
[0824] The user (store camera) acquires video data when a customer enters the store. The camera captures the customer's movements and facial expressions inside the store, acquiring video data in a format that can be analyzed by a generative AI model. The input is a real-time video stream, and the output is analyzable digital data.
[0825] Step 2:
[0826] The terminal (camera or smart device) transmits video data to the server via a 5G communication network. Due to the high bandwidth network, data is transmitted in near real-time. The input is digital data from the video acquisition device, and the output is the video data transmitted to the server.
[0827] Step 3:
[0828] The server uses a generated AI model to analyze the received video data. The server extracts the customer's face frame by frame from the video data and performs facial expression analysis. In this analysis process, video analysis algorithms such as the Google Cloud Vision API are used. The input is the video data sent to the server, and the output is data related to the customer's emotional state corresponding to their facial expressions.
[0829] Step 4:
[0830] The server uses the analysis results to generate feedback and create appropriate customer service advice based on the customer's emotional state. The generated feedback includes specific suggestions to help sales staff provide customer service. The input is emotional state data obtained through analysis, and the output is feedback information for sales staff.
[0831] Step 5:
[0832] The server sends the feedback generated in the previous step to the sales staff's terminal via the communication network. The staff checks this information on their smartphone or tablet and uses it to assist customers. The input is the feedback information on the server, and the output is the advice displayed on the sales staff's terminal.
[0833] Step 6:
[0834] The terminal (the sales staff's smart device) checks the received feedback and modifies its approach to the customer. For example, if it detects that the customer is confused, the staff will provide service according to the prompt message, "Suggest popular products to first-time customers." The input is the feedback displayed on the terminal, and the output is the specific action taken by the sales staff based on that feedback.
[0835] The specific processing unit 290 transmits the result of the specific processing to the robot 414. In the robot 414, the control unit 46A causes the speaker 240 and the controlled object 443 to output the result of the specific processing. The microphone 238 acquires audio indicating user input for the result of the specific processing. The control unit 46A transmits the audio data indicating user input acquired by the microphone 238 to the data processing unit 12. In the data processing unit 12, the specific processing unit 290 acquires the audio data.
[0836] Data generation model 58 is a type of so-called generative AI (Artificial Intelligence). One example of data generation model 58 is ChatGPT (Internet search<URL: https: / / openai.com / blog / chatgpt> ), Gemini (Internet search) <url: https: gemini.google.com ?hl="ja">Examples of generative AI include the following. The data generation model 58 is obtained by performing deep learning on a neural network. The data generation model 58 is input with prompts containing instructions, and with inference data such as audio data representing speech, text data representing text, and image data representing images. The data generation model 58 infers from the input inference data according to the instructions indicated by the prompts, and outputs the inference results in data formats such as audio data and text data. Here, inference refers to, for example, analysis, classification, prediction, and / or summarization.
[0837] In the above embodiment, an example was given in which the specific processing is performed by the data processing device 12, but the technology of this disclosure is not limited thereto, and the specific processing may also be performed by the robot 414.
[0838] Furthermore, the emotion identification model 59, acting as an emotion engine, may determine the user's emotion according to a specific mapping. Specifically, the emotion identification model 59 may determine the user's emotion according to a specific mapping, which is an emotion map (see Figure 9). Similarly, the emotion identification model 59 may also determine the robot's emotion, and the identification processing unit 290 may perform identification processing using the robot's emotion.
[0839] Figure 9 shows an emotion map 400 in which multiple emotions are mapped. In the emotion map 400, emotions are arranged in concentric circles radiating from the center. The closer to the center of the concentric circles, the more primitive the emotions are located. Further out of the concentric circles, emotions representing states and actions arising from mental states are located. Emotion is a concept that includes feelings and mental states. On the left side of the concentric circles, emotions that are generally generated from reactions occurring in the brain are located. On the right side of the concentric circles, emotions that are generally induced by situational judgment are located. Above and below the concentric circles, emotions that are generally generated from reactions occurring in the brain and induced by situational judgment are located. In addition, the emotion of "pleasure" is located on the upper side of the concentric circles, and the emotion of "displeasure" is located on the lower side. Thus, in the emotion map 400, multiple emotions are mapped based on the structure in which emotions arise, and emotions that are likely to occur simultaneously are mapped close together.
[0840] These emotions are distributed at the 3 o'clock position on the Emotion Map 400, and usually fluctuate between feelings of security and anxiety. In the right half of the Emotion Map 400, situational awareness takes precedence over internal feelings, resulting in a calm impression.
[0841] The inside of the Emotion Map 400 represents inner thoughts, while the outside represents actions. Therefore, the further you go from the outside of the Emotion Map 400, the more visible (expressed in actions) your emotions become.
[0842] Here, human emotions are based on various balances, such as posture and blood sugar levels. When these balances deviate from the ideal, it results in discomfort, and when they approach the ideal, it results in pleasure. Similarly, in robots, cars, motorcycles, etc., emotions can be created based on various balances, such as posture and battery level. When these balances deviate from the ideal, it results in discomfort, and when they approach the ideal, it results in pleasure. The emotion map can be generated, for example, based on Dr. Mitsuyoshi's emotion map (Research on a system for analyzing brain physiological signals of speech emotion recognition and emotion, Tokushima University, doctoral dissertation: https: / / ci.nii.ac.jp / naid / 500000375379). The left half of the emotion map contains emotions belonging to a region called "response," where sensation is dominant. The right half of the emotion map contains emotions belonging to a region called "situation," where situational awareness is dominant.
[0843] The emotion map defines two emotions that promote learning. One is the emotion around the middle of the negative "repentance" and "reflection" on the situation side. In other words, it is when the robot experiences negative emotions such as "I never want to feel this way again" or "I don't want to be scolded again." The other is the emotion around the positive "desire" on the reaction side. In other words, it is when the robot has positive feelings such as "I want more" or "I want to know more."
[0844] The emotion identification model 59 inputs user input into a pre-trained neural network, obtains emotion values representing each emotion shown in the emotion map 400, and determines the user's emotion. This neural network is pre-trained based on multiple training data sets, which are combinations of user input and emotion values representing each emotion shown in the emotion map 400. Furthermore, this neural network is trained so that emotions located close together have similar values, as shown in the emotion map 900 in Figure 10. Figure 10 shows an example where multiple emotions such as "reassured," "calm," and "confident" have similar emotion values.
[0845] The above description primarily focuses on the functions of the data processing device 12 in relation to this disclosure. However, the system related to this disclosure is not necessarily implemented on a server. The system related to this disclosure may be implemented as a general information processing system. This disclosure may be implemented, for example, as a software program that runs on a personal computer or as an application that runs on a smartphone. The method related to this disclosure may be provided to users in SaaS (Software as a Service) format.
[0846] In the above embodiment, an example was given in which a specific process is performed by a single computer 22. However, the technology of this disclosure is not limited thereto, and a distributed processing of the specific process may be performed by multiple computers, including computer 22. For example, a data generation model 58 may be provided in an external device of the data processing device 12, and the external device may generate data according to the input data.
[0847] In the above embodiment, an example was given in which the specific processing program 56 is stored in the storage 32, but the technology of this disclosure is not limited thereto. For example, the specific processing program 56 may be stored in a portable, computer-readable, non-temporary storage medium such as a USB (Universal Serial Bus) memory. The specific processing program 56 stored in the non-temporary storage medium is installed in the computer 22 of the data processing device 12. The processor 28 executes specific processing according to the specific processing program 56.
[0848] Alternatively, the specific processing program 56 may be stored in a storage device such as a server connected to the data processing device 12 via the network 54, and the specific processing program 56 may be downloaded and installed on the computer 22 in response to a request from the data processing device 12.
[0849] Furthermore, it is not necessary to store the entirety of the specific processing program 56 in a storage device such as a server connected to the data processing device 12 via the network 54, or to store the entirety of the specific processing program 56 in the storage 32; it is acceptable to store only a portion of the specific processing program 56.
[0850] The following types of processors can be used as hardware resources to perform specific processing. Examples of processors include a CPU, a general-purpose processor that functions as a hardware resource to perform specific processing by executing software, i.e., a program. Other examples of processors include dedicated electrical circuits, such as FPGAs (Field-Programmable Gate Arrays), PLDs (Programmable Logic Devices), or ASICs (Application Specific Integrated Circuits), which have circuit configurations specifically designed to perform specific processing. All of these processors have built-in or connected memory, and all of them perform specific processing by using memory.
[0851] The hardware resource that performs a specific process may consist of one of these various processors, or it may consist of a combination of two or more processors of the same or different types (for example, a combination of multiple FPGAs, or a combination of a CPU and an FPGA). Alternatively, the hardware resource that performs a specific process may consist of a single processor.
[0852] Examples of configurations using a single processor include, firstly, a configuration in which one or more CPUs and software are combined to form a single processor, and this processor functions as a hardware resource that performs a specific process. Secondly, there is a configuration using a processor that realizes the functions of the entire system, including multiple hardware resources that perform a specific process, on a single IC chip, as exemplified by SoCs (System-on-a-chip). In this way, a specific process is realized using one or more of the above types of processors as hardware resources.
[0853] Furthermore, the hardware structure of these various processors can more specifically utilize electrical circuits that combine circuit elements such as semiconductor devices. Also, the specific processing described above is merely an example. Therefore, it goes without saying that unnecessary steps can be deleted, new steps added, or the processing order rearranged, as long as it does not deviate from the main purpose.
[0854] The descriptions and illustrations presented above are detailed explanations of the technical aspects of this disclosure and are merely examples of the technical aspects. For example, the above descriptions of the structure, function, operation, and effect are examples of the structure, function, operation, and effect of the technical aspects of this disclosure. Therefore, it goes without saying that you may delete unnecessary parts, add new elements, or replace elements in the descriptions and illustrations presented above, as long as you do not deviate from the essence of the technical aspects of this disclosure. Furthermore, in order to avoid confusion and facilitate understanding of the technical aspects of this disclosure, explanations of common technical knowledge and the like that do not require special explanation to enable the implementation of the technical aspects of this disclosure have been omitted from the descriptions and illustrations presented above.
[0855] All documents, patent applications, and technical standards described herein are incorporated by reference to the same extent as if each individual document, patent application, and technical standard were specifically and individually noted to be incorporated by reference.
[0856] The following is further disclosed regarding the embodiments described above.
[0857] (Claim 1)
[0858] A means for analyzing the movements of athletes from match footage acquired by a video acquisition device, using a generative model that analyzes the said footage,
[0859] Based on the aforementioned analysis results, a means for evaluating the technical and strategic elements of an athlete and generating feedback,
[0860] A means for transmitting the aforementioned feedback to an external terminal via a communication network,
[0861] A system that includes this.
[0862] (Claim 2)
[0863] The system according to claim 1, wherein the generation model tracks the body position and movements of the athlete and evaluates their posture during play.
[0864] (Claim 3)
[0865] The system according to claim 1, wherein the feedback is designed to include advice on areas for improvement in the athlete's movements and strategies.
[0866] "Example 1"
[0867] (Claim 1)
[0868] A means for analyzing the posture of athletes from video footage, using a generative algorithm that analyzes the video footage, with the input being video footage of the competition acquired by video acquisition equipment.
[0869] Based on the aforementioned analysis results, a means for evaluating the technical and strategic elements of an athlete and generating feedback including specific improvement suggestions,
[0870] Means for transmitting the aforementioned feedback to an external device via a communication path,
[0871] A system that includes this.
[0872] (Claim 2)
[0873] The system according to claim 1, wherein the generation algorithm tracks the body position and movements of the athlete and evaluates their posture and behavior during the competition.
[0874] (Claim 3)
[0875] The system according to claim 1, wherein the feedback is designed to include guidance information regarding quantified areas for improvement in performance and specific strategies.
[0876] "Application Example 1"
[0877] (Claim 1)
[0878] A means for analyzing the operation of a machine from motion video acquired by a video acquisition mechanism, using a generation algorithm that analyzes the motion,
[0879] Based on the aforementioned analysis results, a means for evaluating the functional and efficient elements of the machine and generating improvement suggestions,
[0880] Means for transmitting the aforementioned improvement proposal to an external device via an information communication channel,
[0881] A system that includes this.
[0882] (Claim 2)
[0883] The system according to claim 1, wherein the generation algorithm tracks the operating position and activity of the machine and evaluates its operation during work.
[0884] (Claim 3)
[0885] The system according to claim 1, wherein the improvement suggestions are designed to include suggestions for improvements or efficiencies in the operation of the machine.
[0886] "Example 2 of combining an emotion engine"
[0887] (Claim 1)
[0888] A means for analyzing the movements of athletes from video footage, using a generative model that analyzes the video footage, with the video footage acquired by the video acquisition means as input.
[0889] A means for evaluating the technical and strategic elements of an athlete based on the aforementioned analysis results, and for generating customized feedback based on their movements and psychological state,
[0890] Means for transmitting the aforementioned feedback to an information processing device via a communication network,
[0891] A system that includes this.
[0892] (Claim 2)
[0893] The system according to claim 1, wherein the generative model tracks the body position and movements of the athlete and uses an emotion engine to evaluate the athlete's emotional state during the competition.
[0894] (Claim 3)
[0895] The system according to claim 1, wherein the feedback is designed to include advice on areas for improvement in the athlete's movements and strategies tailored to their psychological state.
[0896] "Application example 2 when combining with an emotional engine"
[0897] (Claim 1)
[0898] A means for analyzing customer facial expressions from video using a generative model that analyzes video, with customer video acquired by a video acquisition device as input.
[0899] Based on the aforementioned analysis results, means for evaluating the customer's emotional state and generating feedback,
[0900] A means for transmitting the aforementioned feedback to the sales staff's terminal via a communication network,
[0901] A system that includes this.
[0902] (Claim 2)
[0903] The system according to claim 1, wherein the generative model tracks the customer's facial expressions and voice and evaluates the situation inside the store.
[0904] (Claim 3)
[0905] The system according to claim 1, wherein the feedback is designed to include advice on customer service and product suggestions that are tailored to the customer's emotions. [Explanation of symbols]
[0906] 10, 210, 310, 410 Data Processing Systems 12 Data Processing Devices 14 Smart Devices 214 Smart Glasses 314 Headset-type terminal 414 Robots< / url:> < / url:> < / url:> < / url:>
Claims
1. A means for analyzing the movements of athletes from match footage acquired by a video acquisition device, using a generative model that analyzes the said footage, Based on the aforementioned analysis results, a means for evaluating the technical and strategic elements of an athlete and generating feedback, A means for transmitting the aforementioned feedback to an external terminal via a communication network, A system that includes this.
2. The system according to claim 1, wherein the generation model tracks the body position and movements of the athlete and evaluates their posture during play.
3. The system according to claim 1, wherein the feedback is designed to include advice on areas for improvement in the athlete's movements and strategies.
Citation Information
Patent Citations
Persona chatbot control method and system
JP2022180282A