system

The system provides real-time analysis and feedback using a video camera and server with deep learning models to enhance billiards skills by identifying strengths and weaknesses, maintaining motivation, and supporting long-term skill improvement.

JP2026071558APending Publication Date: 2026-04-30SOFTBANK GROUP CORP
View PDF 1 Cites 0 Cited by

Patent Information

Authority / Receiving Office
JP · JP
Patent Type
Applications
Current Assignee / Owner
SOFTBANK GROUP CORP
Filing Date
2024-10-17
Publication Date
2026-04-30

AI Technical Summary

Technical Problem

Billiard players face challenges in receiving appropriate feedback, objectively analyzing their play, and maintaining motivation due to limited opportunities for improvement, which hinders skill development and enjoyment of the sport.

Method used

A system utilizing a video camera device and a server to analyze real-time video with deep learning models, evaluate player movements and ball motion, provide precise feedback, and archive data for personalized practice plans, enhancing skill improvement and motivation.

Benefits of technology

Enables accurate real-time feedback and personalized practice plans, helping players identify strengths and weaknesses, maintain motivation, and track long-term progress, thereby improving athletic performance and enjoyment.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure 2026071558000001_ABST
    Figure 2026071558000001_ABST
Patent Text Reader

Abstract

We provide the system. [Solution] A system comprising: means for receiving real-time video data from a video camera device and detecting the player's movements and the ball's motion using a deep learning model for analyzing the video data; means for evaluating the player's technical characteristics based on their shots and identifying their strengths and weaknesses in shots; means for generating areas for improvement and shot strategies in natural language and providing feedback to the player based on these technical characteristics; means for presenting this feedback visually and audibly on a display device; means for archiving the video data and technical characteristics in a database and managing a history for analyzing the player's progress; and means for generating an individual practice plan based on the player's performance data.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The technology of the present disclosure relates to a system.

Background Art

[0002] Patent Document 1 discloses a method for controlling a persona chatbot, which is performed by at least one processor, including steps of receiving a user utterance, adding the user utterance to a prompt including an instruction sentence related to an explanation of a chatbot character, encoding the prompt, and inputting the encoded prompt into a language model to generate a chatbot utterance in response to the user utterance.

Prior Art Documents

Patent Documents

[0003]

Patent Document 1

Summary of the Invention

Problems to be Solved by the Invention

[0004] Billiard players need to spend a great deal of time and effort on improving their skills, and there is a problem that they have few opportunities to receive appropriate feedback. Also, it is difficult to objectively analyze their own play from a third-person perspective and identify areas for improvement. Furthermore, players find it difficult to maintain motivation, especially when consecutive failures can exacerbate the situation. As a result, the improvement of billiards technology is hindered, and the enjoyment and competitiveness of the entire sport decline.

Means for Solving the Problems

[0005] This invention solves these problems by providing a means for analyzing real-time video acquired from a video camera device to detect the player's movements and the ball's motion. Furthermore, by evaluating technical characteristics using a deep learning model and identifying the player's strengths and weaknesses in shots, it provides the player with precise areas for improvement and strategies. This allows the player to reflect on their own skills and improve them efficiently. In addition, if the player experiences repeated failures, it maintains motivation by suggesting their strengths in shots to promote successful experiences. The player's data is archived and made reviewable at a later date, allowing for the visualization of long-term progress and the provision of individualized practice plans. Through these configurations, the aim is to improve athletic performance and promote enjoyment of billiards.

[0006] A "video camera device" is an electronic device for capturing video and audio data in real time.

[0007] "Real-time video data" refers to data that can be processed and analyzed continuously with virtually no delay after being captured.

[0008] A "deep learning model" is a type of artificial intelligence technology that uses a large amount of data to extract features and make predictions; it is a multi-layered neural network.

[0009] "Player actions" refers to the body movements and posture required for a shot in billiards.

[0010] "Ball motion" refers to the movement of a ball on a billiard table, including its trajectory and physical properties.

[0011] "Technical characteristics" refer to various indicators used to identify the skills and characteristics of a billiards player's shots.

[0012] A "favorite shot" is a shot shape or technique that a player has a high success rate with and can execute relatively easily.

[0013] A "difficult shot" is a shot shape or technique that a player has a low success rate with and needs improvement.

[0014] "Generating in natural language" refers to the process by which a machine generates information in a language format that is easy for humans to understand.

[0015] "Feedback" refers to a form of improvement and strategic advice provided to encourage players to improve their skills.

[0016] A "display device" is a device used to convey given information to a user visually or audibly.

[0017] "Archiving in a database" is the process of storing acquired data in a digital format that facilitates information storage and management.

[0018] An "individualized training plan" refers to a training program and schedule designed for each player, aimed at improving their skills. [Brief explanation of the drawing]

[0019] [Figure 1] This is a conceptual diagram showing an example of the configuration of a data processing system according to the first embodiment. [Figure 2] This is a conceptual diagram showing an example of the essential functions of a data processing device and a smart device according to the first embodiment. [Figure 3] This is a conceptual diagram showing an example of the configuration of a data processing system according to the second embodiment. [Figure 4] This is a conceptual diagram showing an example of the main functions of a data processing device and smart glasses according to the second embodiment. [Figure 5] This is a conceptual diagram showing an example of the configuration of a data processing system according to the third embodiment. [Figure 6]It is a conceptual diagram showing an example of the main functions of a data processing device and a headset-type terminal according to the third embodiment. [Figure 7] It is a conceptual diagram showing an example of the configuration of a data processing system according to the fourth embodiment. [Figure 8] It is a conceptual diagram showing an example of the main functions of a data processing device and a robot according to the fourth embodiment. [Figure 9] It shows an emotion map to which multiple emotions are mapped. [Figure 10] It shows an emotion map to which multiple emotions are mapped. [Figure 11] It is a sequence diagram showing the processing flow of the data processing system in Example 1. [Figure 12] It is a sequence diagram showing the processing flow of the data processing system in Application Example 1. [Figure 13] It is a sequence diagram showing the processing flow of the data processing system in Example 2 when an emotion engine is combined. [Figure 14] It is a sequence diagram showing the processing flow of the data processing system in Application Example 2 when an emotion engine is combined.

Embodiments for Carrying Out the Invention

[0020] Hereinafter, an example of an embodiment of a system according to the technology of the present disclosure will be described according to the accompanying drawings.

[0021] First, the language used in the following description will be explained.

[0022] In the following embodiments, the signed processor (hereinafter simply referred to as "processor") may be a single arithmetic unit or a combination of multiple arithmetic units. Furthermore, the processor may be a single type of arithmetic unit or a combination of multiple types of arithmetic units. Examples of arithmetic units include CPU (Central Processing Unit), GPU (Graphics Processing Unit), GPGPU (General-Purpose computing on Graphics Processing Units), and APU (Accelerated Processing Unit).

[0023] In the following embodiments, signed RAM (Random Access Memory) is a memory that temporarily stores information and is used as work memory by the processor.

[0024] In the following embodiments, the signed storage is one or more non-volatile storage devices that store various programs and various parameters. Examples of non-volatile storage devices include flash memory (SSD (Solid State Drive)), magnetic disks (e.g., hard disks), or magnetic tapes.

[0025] In the following embodiments, the signed communication interface (I / F) is an interface that includes a communication processor and an antenna, etc. The communication interface manages communication between multiple computers. Examples of communication standards applicable to the communication interface include wireless communication standards such as 5G (5th Generation Mobile Communication System), Wi-Fi (registered trademark), or Bluetooth (registered trademark).

[0026] In the following embodiments, "A and / or B" is synonymous with "at least one of A and B." That is, "A and / or B" means that it may be A alone, or B alone, or a combination of A and B. Furthermore, in this specification, the same concept as "A and / or B" applies when expressing three or more things linked by "and / or."

[0027] [First Embodiment]

[0028] Figure 1 shows an example of the configuration of the data processing system 10 according to the first embodiment.

[0029] As shown in Figure 1, the data processing system 10 includes a data processing device 12 and a smart device 14. An example of the data processing device 12 is a server.

[0030] The data processing device 12 comprises a computer 22, a database 24, and a communication interface 26. The computer 22 is an example of a "computer" related to the technology of this disclosure. The computer 22 comprises a processor 28, RAM 30, and storage 32. The processor 28, RAM 30, and storage 32 are connected to a bus 34. The database 24 and the communication interface 26 are also connected to the bus 34. The communication interface 26 is connected to a network 54. An example of the network 54 is a WAN (Wide Area Network) and / or a LAN (Local Area Network).

[0031] The smart device 14 comprises a computer 36, a reception device 38, an output device 40, a camera 42, and a communication interface 44. The computer 36 comprises a processor 46, RAM 48, and storage 50. The processor 46, RAM 48, and storage 50 are connected to a bus 52. The reception device 38, output device 40, and camera 42 are also connected to the bus 52.

[0032] The reception device 38 is equipped with a touch panel 38A and a microphone 38B, etc., and receives user input. The touch panel 38A receives user input by detecting contact with an object (e.g., a pen or finger). The microphone 38B receives user input by detecting the user's voice. The control unit 46A transmits data indicating the user input received by the touch panel 38A and microphone 38B to the data processing device 12. In the data processing device 12, the specific processing unit 290 acquires the data indicating the user input.

[0033] The output device 40 includes a display 40A and a speaker 40B, and presents data to the user 20 by outputting the data in a form perceptible to the user 20 (e.g., audio and / or text). The display 40A displays visible information such as text and images according to instructions from the processor 46. The speaker 40B outputs audio according to instructions from the processor 46. The camera 42 is a small digital camera equipped with an optical system such as a lens, aperture, and shutter, and an image sensor such as a CMOS (Complementary Metal-Oxide-Semiconductor) image sensor or a CCD (Charge Coupled Device) image sensor.

[0034] Communication interface 44 is connected to network 54. Communication interfaces 44 and 26 are responsible for the exchange of various types of information between processor 46 and processor 28 via network 54.

[0035] Figure 2 shows an example of the main functions of the data processing device 12 and the smart device 14.

[0036] As shown in Figure 2, in the data processing device 12, a specific processing is performed by the processor 28. A specific processing program 56 is stored in the storage 32. The specific processing program 56 is an example of a "program" related to the technology of this disclosure. The processor 28 reads the specific processing program 56 from the storage 32 and executes the read specific processing program 56 on the RAM 30. The specific processing is realized by the processor 28 operating as a specific processing unit 290 according to the specific processing program 56 executed on the RAM 30.

[0037] The storage 32 stores the data generation model 58 and the emotion identification model 59. The data generation model 58 and the emotion identification model 59 are used by the identification processing unit 290.

[0038] In the smart device 14, the processor 46 performs the reception output processing. The storage 50 stores the reception output program 60. The reception output program 60 is used in conjunction with a specific processing program 56 by the data processing system 10. The processor 46 reads the reception output program 60 from the storage 50 and executes the read reception output program 60 on the RAM 48. The reception output processing is realized by the processor 46 operating as a control unit 46A according to the reception output program 60 executed on the RAM 48.

[0039] Next, the specific processing performed by the specific processing unit 290 of the data processing device 12 will be described. In the following description, the data processing device 12 will be referred to as the "server" and the smart device 14 as the "terminal".

[0040] This invention is a system using a video camera device and a server to help billiard players improve their skills. The role and operation of each component are described below.

[0041] Acquisition of video data

[0042] The device uses a video camera to capture real-time footage of the billiards game. This ensures that shots and ball movements are accurately recorded.

[0043] Analysis of real-time video

[0044] The server receives video data sent from the terminal and uses a deep learning model to analyze the player's movements and the ball's motion. The model processes each captured frame and extracts data for analyzing techniques such as shot speed, angle, and ball rotation.

[0045] Player skill evaluation

[0046] The server evaluates the player's technical characteristics based on the analyzed data. The computer automatically identifies the player's strong and weak shots, and pinpoints the success rate of those shots and areas for improvement.

[0047] Generating and providing feedback

[0048] Based on the results of the technical evaluation, the server generates improvement suggestions and strategic advice for the player in natural language. The generated feedback is sent to the device in real time. The device then communicates this to the user through visual displays or audio notifications.

[0049] Data storage and practice plan generation

[0050] The server archives data from each play session in a database. This saved data is used by players to review their gameplay at a later date. The server also generates personalized practice plans based on the accumulated performance data, helping users improve their skills more effectively.

[0051] For example, if a player repeatedly misses difficult shots, the server analyzes past success data and suggests alternative shots that have a higher success rate for that player. This provides a sense of accomplishment and helps prevent a decline in motivation.

[0052] The following describes the processing flow.

[0053] Step 1:

[0054] The terminal captures real-time video from the billiard table via a video camera device and sends the data to the server. The terminal continues to acquire video at the appropriate resolution and frame rate.

[0055] Step 2:

[0056] The server acquires the received real-time video data and begins analysis frame by frame. Using a deep learning model, it detects the player's movements and the ball's motion. Specifically, it identifies the angle and force of the player's shot, as well as the ball's collision point and trajectory.

[0057] Step 3:

[0058] The server evaluates the player's shots based on the analyzed data. By comparing technical data, it identifies the player's strengths and weaknesses in shots. By comparing this data with past play data, it analyzes the player's consistency and areas for improvement.

[0059] Step 4:

[0060] The server generates feedback for the player based on the evaluation results. Using natural language processing technology, it provides intuitive and effective improvement advice and shot strategies. The feedback is customized according to the player's skill level and the situation.

[0061] Step 5:

[0062] The server sends the generated feedback to the terminal. The terminal then communicates this to the user either by displaying it on the screen or announcing it through an audio device. The user receives real-time feedback and can immediately use it to improve their gameplay.

[0063] Step 6:

[0064] The server stores data from each play session in a database. This ensures that the underlying data supports the long-term improvement of players' skills. The data is used to generate individual practice plans and is utilized when users plan their future gameplay.

[0065] (Example 1)

[0066] Next, we will describe Example 1. In the following description, the data processing device 12 will be referred to as the "server," and the smart device 14 will be referred to as the "terminal."

[0067] Conventional skill improvement support systems have struggled to analyze user actions in real time and provide immediate improvement suggestions based on individual skill characteristics. Furthermore, they lacked the ability to automatically generate practice plans based on past performance data, preventing users from efficiently improving their skills. This created a need for an integrated skill improvement support system that includes rapid and accurate feedback.

[0068] The identification process performed by the identification processing unit 290 of the data processing device 12 in Example 1 is realized by the following means.

[0069] In this invention, the server includes means for acquiring real-time visual information from a visual sensor and detecting user operations and object movements using a machine learning model for analyzing the visual information; means for evaluating technical characteristics based on the user's operations and identifying strengths and operations that need improvement; and means for generating and providing information to the user in natural language regarding areas for improvement and operation plans based on the technical characteristics. This enables the user to receive detailed feedback in real time and obtain accurate advice and practice plans for improving their skills.

[0070] A "visual sensor" is a device that captures the movement of an object and converts the visual information into digital data.

[0071] "Visual information" refers to digital data of images and videos acquired from visual sensors, and this data is the subject of analysis.

[0072] A "machine learning model" is a collection of algorithms that analyze large amounts of data and automatically learn specific patterns and features.

[0073] "Users" refer to individuals or organizations that wish to use the system to improve specific technologies.

[0074] "Operation" refers to physical or digital actions performed by the user, including actions detected by visual sensors.

[0075] "Object movement" refers to changes in the position of physical objects detected within visual information.

[0076] "Technical characteristics" refer to the specific features of performance and capabilities related to user operation.

[0077] "Operations that need improvement" are specific actions or procedures that users should focus on modifying or practicing in order to improve their skills.

[0078] "Natural language" refers to the form of language that humans use in everyday life, and is used when computers generate and provide information to users.

[0079] "Information provision" refers to the act of conveying analyzed data and suggestions to users, and this can be done visually or audibly.

[0080] This invention is a system designed to support the improvement of users' skills and comprises a visual sensor, a server, and a terminal. Specifically, a general-purpose video acquisition device is used as the visual sensor to capture the actions of physical sports such as billiards in real time. This video data is transmitted to the server via the terminal.

[0081] The server analyzes the received visual information using a machine learning model. This process utilizes commonly used deep learning frameworks (e.g., TENSORFLOW®) to evaluate the user's actions and object movements frame by frame. The server then extracts numerical data for the speed, angle, and ball rotation of each shot. Based on this data, the server assesses the user's technical characteristics and identifies their strengths and areas for improvement.

[0082] The server uses a generative AI model (e.g., GPT-3®) based on the evaluation results to generate improvement suggestions and technical advice for the user in natural language. An example of a prompt used for this purpose would be, "Please generate improvement advice based on this player's recent shot data."

[0083] The generated feedback is sent to the device in real time. The device uses visual display devices and audio output devices to convey this to the user. This allows the user to continuously review their gameplay and effectively improve their skills.

[0084] Furthermore, the server stores visual information and analysis results in memory. This data is used as a history for users to review past gameplay data later, and is also utilized to create personalized practice plans for continuous skill improvement. Based on the accumulated data, the server performs statistical analysis and provides practice plans optimized for the user. Through this entire process, the system can provide comprehensive support for users to improve their skills.

[0085] The flow of the specific processing in Example 1 will be explained using Figure 11.

[0086] Step 1:

[0087] The device uses a visual sensor to capture the user's athletic performance in real time. The input is the physical athletic scene, and the output is digital video data. This ensures that the user's movements and the movement of objects are always recorded in their most up-to-date state. Specifically, the sensor continuously captures images at regular intervals according to the frame rate.

[0088] Step 2:

[0089] The terminal sends the captured video data to the server. The input is the video data acquired in step 1, and the output is streaming data for the server to receive. The data is transmitted in real time, and compression technology is used to reduce network load while maintaining data integrity.

[0090] Step 3:

[0091] The server analyzes the received video data using a machine learning model. The input is video data sent from the terminal, and the output is numerical data related to the user's movements and the position of objects. This analysis involves feature point extraction and motion vector calculation for each frame. Specifically, the server runs a deep learning framework and applies algorithms to extract important physical information from each frame.

[0092] Step 4:

[0093] The server evaluates the user's technical characteristics based on the analysis results. The input is the output data from step 3, and the output is scoring data related to the technical evaluation. The server automatically identifies the user's strengths and areas for improvement through a script and uses a technical evaluation model. Specifically, it applies statistical analysis and scores each operation based on evaluation criteria.

[0094] Step 5:

[0095] The server uses a generative AI model to generate improvement suggestions in natural language based on technical evaluations. The input is the technical evaluation data from step 4, and the output is advice in natural language provided to the user. Specifically, prompt sentences are used and applied to the generative AI model to form concrete advice that helps improve the user's skills.

[0096] Step 6:

[0097] The terminal receives advice from the server and presents it to the user in real time. Input is feedback data from the server, and output is advice expressed visually or audibly. The terminal uses a display and speakers to inform the user of specific improvement measures. This allows the user to directly and immediately see the improvement measures.

[0098] Step 7:

[0099] The server stores video data and technical evaluations in its memory. The input is past pre-session data, and the output is historical data stored in the memory. This storage process is performed using database management functions. This enables long-term growth analysis of users.

[0100] Step 8:

[0101] The server generates individual practice plans based on stored data. The input is performance data stored in memory, and the output is a practice plan optimized for the user. The server combines statistical analysis and machine learning algorithms to construct effective practice suggestions.

[0102] (Application Example 1)

[0103] Next, we will explain Application Example 1. In the following explanation, the data processing device 12 will be referred to as the "server," and the smart device 14 will be referred to as the "terminal."

[0104] For billiards players to efficiently improve their skills, real-time motion analysis and feedback are essential. Furthermore, being able to receive this information instantly in a real playing environment greatly enhances the effectiveness of practice. However, conventional technology has not been able to create a system that accurately analyzes a player's movements and provides real-time, visually and audibly effective feedback.

[0105] The specific processing performed by the specific processing unit 290 of the data processing device 12 in Application Example 1 is realized by the following means.

[0106] In this invention, the server includes means for receiving real-time video information from a video camera device and detecting the player's body movements and the movement of a sphere using a machine learning model for analyzing the video information; means for generating and providing feedback to the player in natural language regarding areas for improvement and shot strategies based on these technical characteristics; and means for overlaying the feedback onto the real environment using smart glasses. This enables the player to obtain detailed and practical analytical information and feedback in real time, allowing for more effective technical improvement.

[0107] A "video camera device" is a device used to capture video information in real time.

[0108] "Real-time video information" refers to information that provides video data occurring at the present moment immediately.

[0109] A "machine learning model" is a model that uses algorithms to automatically learn patterns and rules from data.

[0110] "Physical movement" refers to the movement and posture of the player's body, and analyzing this is a means of extracting technical characteristics.

[0111] "Sphere motion" refers to the kinetic characteristics of a billiard ball, specifically its direction of movement, speed, and rotation.

[0112] "Technical characteristics" refer to a set of parameters used to evaluate the accuracy and movement characteristics of a player's shots.

[0113] "Natural language" refers to the linguistic expressions that people use on a daily basis, and is a means of conveying information to players in an easily understandable way as feedback.

[0114] "Smart glasses" are wearable devices that have the function of displaying visual information as augmented reality.

[0115] "Overlaying information onto the real world" is a technology that displays additional information superimposed on the visual information of the real world.

[0116] This invention provides a system to support the skill improvement of billiard players. This system operates in combination with a video camera device, smart glasses, and a server.

[0117] The server receives real-time video data from the video camera device. This data is processed using machine learning models to analyze in detail the player's body movements and the ball's motion during the shot. Specifically, it utilizes libraries such as TensorFlow in Python to recognize the ball's speed, rotation, and shot angle for each video frame.

[0118] Based on the analyzed data, the system evaluates the player's technical characteristics and identifies their strengths and weaknesses in shots. This allows the server to generate natural language feedback, providing players with areas for improvement and strategic advice. This feedback is overlaid onto the real-world gameplay environment via smart glasses. The system is also designed to allow users to access information in real time using devices such as Google Glass®.

[0119] Accumulated video and technical data are archived in a database on the server. This allows users to analyze past gameplay and track their progress. Player performance data is managed using database software such as Sarvapi and PostgreSQL. Furthermore, individual practice plans are generated, and practice menus tailored to each player are suggested.

[0120] For example, if a player repeatedly misses difficult shots, this system analyzes past success data and suggests alternative shots with a higher success rate. In this way, it helps the player gain a sense of accomplishment. An example of a prompt to input into the generative AI model might be, "Please tell me a temporary strategy for successfully executing a great mid-range shot while playing billiards."

[0121] The flow of a specific process in Application Example 1 will be explained using Figure 12.

[0122] Step 1:

[0123] The terminal captures the billiards game using a video camera device. The input is real-time video data, and the output is the same video data sent directly to the server. In this process, the camera operates at a high frame rate to capture the fast movement of the ball without delay.

[0124] Step 2:

[0125] The server receives video data, which is then used as input for a machine learning model. The input consists of captured image frames, and the output includes the ball's position and motion vector, as well as the player's movement pattern. TensorFlow is used to analyze the ball's speed and spin, and the angle of the shot, generating this numerical data.

[0126] Step 3:

[0127] The server evaluates the player's technical characteristics based on the generated analysis data. The input is the analysis data obtained in step 2, and the output includes strong and weak shots, success rates, etc. Based on this evaluation information, the server inputs suggestions for improvement and strategies for the next shot as prompts into the generating AI model.

[0128] Step 4:

[0129] The server generates feedback from the generative AI model in natural language and sends it to the smart glasses. The input is the suggestion information from the generative AI model, and the output is a natural language feedback message. In this step, the suggestions are summarized in a way that is easy for the user to understand.

[0130] Step 5:

[0131] The user's smart glasses receive feedback from the server and overlay it onto the real environment. The input is the feedback message from the server, and the output is the information visually displayed through the glasses. This allows the user to receive improvement advice in real time.

[0132] Step 6:

[0133] The server archives video data and analysis results from play sessions into a database. Inputs are the original video data and technical evaluation data, while outputs are accumulated historical data. This data serves as a basis for players to later review their progress.

[0134] Step 7:

[0135] The server uses accumulated historical data to generate individual practice plans. The input is the player's past performance data, and the output is the suggested practice plan. This step is performed automatically to support the player's skill improvement.

[0136] Furthermore, an emotion engine that estimates the user's emotions may be incorporated. That is, the identification processing unit 290 may use the emotion identification model 59 to estimate the user's emotions and perform identification processing using the user's emotions.

[0137] This invention combines a system designed to improve the skills of billiard players with an emotion engine that recognizes the user's emotions and provides appropriate feedback. Its configuration and operation are described in detail below.

[0138] Acquisition and analysis of video data

[0139] The terminal uses a video camera device to capture real-time video of a billiards game and sends it to the server. The server uses a deep learning model to analyze the player's movements and the ball's motion. This provides data for evaluating the technical aspects of the shots.

[0140] Technical evaluation and feedback generation

[0141] The server evaluates the player's technical characteristics based on the analyzed data. It identifies strengths and weaknesses in shots and generates improvement suggestions and shot strategies in natural language. The generated feedback is provided to the user visually and audibly through the terminal.

[0142] Emotional recognition and feedback regulation

[0143] The emotion engine analyzes the user's facial expressions and tone of voice to recognize their emotional state. The server then adjusts the feedback based on this emotional data, providing optimal feedback tailored to the user's motivation level. For example, if a player is feeling frustrated, the emotion engine generates positive reinforcement messages to encourage further attempts.

[0144] Data storage and practice plan generation

[0145] The server stores all play data and emotion recognition results in a database, enabling the visualization of long-term progress. This allows for the creation of personalized practice plans for each player, supporting skill improvement. Users can review their play history and emotion change trends and utilize their individual practice plans to efficiently improve their skills.

[0146] For example, if a player repeatedly misses a particular shot, the server uses an emotion engine to recognize the player's frustration and suggests shots they are good at while also providing encouraging messages. This helps the player regain confidence and continue practicing.

[0147] The following describes the processing flow.

[0148] Step 1:

[0149] The terminal uses a video camera device to capture video footage of the billiards game. The video data is transmitted to the server in real time. Camera settings are adjusted to ensure high-quality video and enable accurate motion analysis.

[0150] Step 2:

[0151] The server inputs the received video data into a deep learning model to analyze the player's movements and the ball's motion. The analysis includes shot speed, angle, and spin rate, and the player's current technical characteristics are extracted as numerical data.

[0152] Step 3:

[0153] The server evaluates the player's skills based on the analysis results, identifying their strengths and weaknesses in shots. By comparing this with past gameplay data, it highlights areas for improvement to enhance their skills.

[0154] Step 4:

[0155] The emotion engine recognizes the user's emotional state from their facial expressions, tone of voice, and body movements. The server then incorporates this emotional information into the feedback. For example, if the player is focused, it provides detailed technical analysis; on the other hand, if they are feeling frustrated, it generates encouraging messages.

[0156] Step 5:

[0157] The server generates feedback for the user using natural language processing, based on technical evaluation results and emotional states. The feedback is created as a customized message and includes strategic advice and suggestions for improvement tailored to the player's current situation.

[0158] Step 6:

[0159] The device presents the generated feedback to the user visually and audibly. Users can receive real-time advice and gain a deeper understanding of their own shots.

[0160] Step 7:

[0161] The server stores technical and emotional data for each session in a database. This allows users to have a history that they can refer to in later sessions and use to create practice plans.

[0162] Step 8:

[0163] Based on accumulated data, the server analyzes players' technical and emotional tendencies and generates personalized practice plans. These plans provide users with specific steps to achieve their goals and support long-term skill improvement.

[0164] (Example 2)

[0165] Next, we will describe Example 2. In the following description, the data processing device 12 will be referred to as the "server" and the smart device 14 as the "terminal".

[0166] It is difficult for billiards players to objectively evaluate their own skills and clearly identify areas for improvement. Furthermore, providing feedback that takes into account the player's emotional state is needed to boost their motivation for practice. A system is required to address these challenges and provide effective practice plans tailored to each individual.

[0167] The identification process performed by the identification processing unit 290 of the data processing device 12 in Example 2 is realized by the following means.

[0168] In this invention, the server includes means for receiving real-time image information from a video recording device and confirming the player's movements and the ball's movement using a deep learning model for analyzing the image information; means for recording the image information and technical elements in an information storage medium and managing a history for analyzing the player's progress; and means for analyzing the user's emotional state, adjusting the response content based on the emotional information, and providing an optimal response. This enables billiard players to effectively improve their skills and receive feedback tailored to their emotional state.

[0169] "Video recording equipment" is a general term for devices used to record video and audio and to later play them back or transmit them.

[0170] "Real-time image information" refers to data used to acquire and process images from the real world in real time.

[0171] A "deep learning model" is a type of algorithm that uses a multi-layer neuron network to learn the features of data and perform analysis and prediction.

[0172] "Player movements" in billiards refers to all physical actions performed by a player, including shots and movement.

[0173] "Ball movement" in billiards refers to the motion of the ball as it rolls, spins, or collides in response to the force exerted by the player.

[0174] "Technical elements" refer to skill-based characteristics such as force and precision, which are evaluated as parameters of shots and play in billiards.

[0175] "Information storage medium" is a general term for hardware or digital storage used to record, store, and manage data.

[0176] "Emotional state" refers to the psychological and emotional state of a user, measured based on their behavior, facial expressions, voice, etc.

[0177] "Response content" refers to all feedback messages provided to the user, such as criticisms, instructions, or encouragement.

[0178] This invention is a system that operates by combining multiple components to support the skill improvement of billiard players. Specific embodiments of this system are shown below.

[0179] The terminal uses a video recording device to capture real-time image information of a billiards game and transmits the data to a server. The terminal is positioned optimally to smoothly capture high-resolution video. The video data is efficiently transmitted to the server via the network.

[0180] The server inputs the received real-time image information into a deep learning model. Specifically, it uses software frameworks such as TensorFlow and PyTorch to analyze the player's motion and the ball's movement. The server then extracts technical elements such as shot speed, angle, and contact point, thereby evaluating the player's skill.

[0181] Based on the evaluated technical elements, the server generates suggestions for improvement and shortcuts for the user in natural language. This process utilizes a natural language generation AI model, and the generated feedback is provided to the user visually and audibly through the device. The device presents the feedback using both screen display and audio output.

[0182] Furthermore, the emotion engine analyzes the user's emotional state. The device obtains emotional data by recording the user's facial expressions and collecting their voice tone. The server uses this data to understand the user's emotional state and adjust the feedback accordingly. For example, if the user is irritated, a positive message may be added to boost their motivation.

[0183] Finally, the server saves all play data and emotion recognition results to a data storage medium. This allows users to refer to past performance data and emotion fluctuations to generate individually customized practice plans.

[0184] For example, if a player repeatedly misses a particular shot, the system recognizes the player's frustration and offers words of encouragement while suggesting alternative shots the player is good at. This allows the player to move on to the next step with confidence.

[0185] An example of a prompt is: "Analyze the player's movements and the ball's trajectory during a billiards match, evaluate their technical characteristics, and determine their strengths and weaknesses. Furthermore, create appropriate feedback in natural language, taking into account the player's emotional state."

[0186] The flow of the specific processing in Example 2 will be explained using Figure 13.

[0187] Step 1:

[0188] The terminal activates a video recording device and captures real-time image information of a billiards game. The input is a real billiards game in action, and the output is high-resolution video data. The quality of this data is ensured by adjusting the camera position to accurately capture the players' movements and the ball's trajectory.

[0189] Step 2:

[0190] The terminal transmits captured image information to the server. The input is video data obtained from a video recording device, and the output is data transfer to the server. In this process, data compression technology is used to efficiently utilize network bandwidth.

[0191] Step 3:

[0192] The server inputs the received image information into a deep learning model. The input is video data sent from the terminal, and the output is the analysis results of the player's motion and the ball's movement. As part of the data processing, the model analyzes the image frame by frame and extracts technical elements such as the motion trajectory and velocity of each object.

[0193] Step 4:

[0194] The server evaluates the player's technical characteristics based on the analysis data. The input is the analysis results obtained in step 3, and the output is the evaluated technical characteristics as well as identification data of strengths and weaknesses. This makes it possible to evaluate shot accuracy and performance.

[0195] Step 5:

[0196] The server generates improvement suggestions and shortcuts in natural language based on the evaluated technical features. The input is evaluation data for the technical features, and the output is feedback messages in natural language. A generative AI model is used, which creates sentences that are easy for humans to understand.

[0197] Step 6:

[0198] The device displays and provides generated feedback to the user via audio. Input is feedback messages sent from the server, and output is a visual and audible presentation of specific advice to the user. Information is delivered to the user using the screen and speaker.

[0199] Step 7:

[0200] The emotion engine analyzes the user's facial expressions and voice tone to recognize their emotional state. The input is image and audio data obtained from the user, and the output is the result of the emotional state analysis. Facial recognition and voice analysis algorithms are used for data processing.

[0201] Step 8:

[0202] The server adjusts the feedback content according to the user's emotional state to maintain or improve their motivation. The input consists of emotional state analysis data and feedback data generated by the server, while the output is the adjusted feedback message. This enables personalized responses that take the user's emotions into consideration.

[0203] Step 9:

[0204] The server stores all gameplay data and emotion recognition results in an information storage medium. Input is general data, and output is historical information stored in a database. This allows for tracking and analysis of players' long-term progress.

[0205] (Application Example 2)

[0206] Next, we will explain application example 2. In the following explanation, the data processing device 12 will be referred to as a "server" and the smart device 14 as a "terminal".

[0207] For players who want to improve their skills in billiards and other sports, there is a need to provide not only technical feedback but also emotional support to boost their motivation and acquire skills more effectively. However, conventional technical support systems lack mechanisms to adjust feedback according to the player's emotional state, which hinders skill improvement.

[0208] The specific processing performed by the specific processing unit 290 of the data processing device 12 in Application Example 2 is realized by the following means.

[0209] In this invention, the server includes means for receiving image information captured without delay from a video recording device and detecting the user's actions and the movement of objects using a machine learning model for analyzing the image information; means for evaluating the technical characteristics based on the user's actions and identifying actions in which the user is good and bad; and means for analyzing the user's emotional state using an emotion analysis engine and adjusting the feedback content based on that state. This makes it possible to improve both the technical and emotional aspects of the user by providing technical feedback along with support tailored to their individual emotions.

[0210] A "video recording device" is a device for acquiring image information in real time without any time lag.

[0211] A "machine learning model" is an algorithm that analyzes data, recognizes patterns, and predicts the future.

[0212] "Users" refers to individuals or organizations that use the system for the purpose of improving their skills.

[0213] "Action and motion of objects" refers to physical movements, including the position, posture, and velocity of the user or object.

[0214] "Technical characteristics" refer to the unique technical operating patterns and characteristics of the user or object.

[0215] An "emotion analysis engine" is an analytical device and program that determines a user's emotional state from their facial expressions and tone of voice.

[0216] "Feedback" refers to information provided to users, either technical or emotional, to encourage improvement and support.

[0217] As an example of the application of this invention, the details of a system for supporting the skill improvement of users in a billiard hall are shown below.

[0218] First, a video recording device is used to capture the user's movements and the movement of the billiard ball in real time. The terminal sends this video data to a server. On the server, a machine learning model is applied to analyze the video data, and the user's shots and the movement of the ball are analyzed to extract technical features.

[0219] Next, the server identifies the user's strengths and weaknesses based on the extracted technical characteristics and generates feedback in natural language based on the results. This feedback is provided to the user visually and audibly via a visual display device.

[0220] Furthermore, an emotion analysis engine is used to analyze the user's facial expressions and tone of voice to recognize their emotional state. Based on this information, the content of the feedback is adjusted according to the user's emotional state to provide optimal support.

[0221] For example, if a user repeatedly misses a particular shot, the server uses an emotion analysis engine to recognize the user's frustration, suggest shots that the user excels at, and provide an encouraging message. In this process, a generative AI model is used for natural language processing, and a prompt message such as, "Please provide the latest billiard shot analysis and feedback. Please also consider the player's emotional data and create an encouraging message," is used.

[0222] In this way, users can efficiently improve their billiards skills while receiving support from both a technical and emotional perspective.

[0223] The flow of a specific process in Application Example 2 will be explained using Figure 14.

[0224] Step 1:

[0225] The terminal captures the user's actions and the movement of the billiard ball in real time via a video recording device and transmits the video data to the server. The input is video data of the user's actions and the ball's movement, and the output is the data captured without any time lag and sent to the server.

[0226] Step 2:

[0227] The server uses a machine learning model to analyze the received video data. The input is video data, and the output is the extraction of technical features related to the user's actions and the ball's movement. This allows the data to be analyzed through the model, identifying specific patterns and trends.

[0228] Step 3:

[0229] The server evaluates the acquired technical characteristics and identifies the user's strengths and weaknesses in shots. The input is data on technical characteristics, and the output is information indicating strengths and weaknesses in shots. Based on this, the server prepares the basic data for generating feedback.

[0230] Step 4:

[0231] Using a generative AI model, the server generates suggestions for improvement and shot strategies for the user in natural language. The input is information about a specific shot, and the output is a detailed feedback message. The server uses prompts to construct the feedback naturally.

[0232] Step 5:

[0233] The server activates an emotion analysis engine to recognize the user's emotional state from their facial expressions and voice data. The input is the user's facial expressions and voice tone information, and the output is an evaluation of their emotional state. This data allows the server to understand how the user is feeling.

[0234] Step 6:

[0235] The server adjusts the feedback content to reflect the emotional state and delivers it to the user visually and audibly via a visual display device. Input is an assessment of the emotional state and pre-generated feedback, while output is the adjusted visual and audible feedback. The goal is to deliver the most appropriate feedback to the user.

[0236] Step 7:

[0237] The server stores technical characteristics and emotional state data in an information aggregation device to manage user progress over the long term. The input is the entire analyzed data, and the output is the stored historical data. This ensures that the history for later use is always up-to-date.

[0238] The specific processing unit 290 transmits the result of the specific processing to the smart device 14. In the smart device 14, the control unit 46A causes the output device 40 to output the result of the specific processing. The microphone 38B acquires audio indicating user input for the result of the specific processing. The control unit 46A transmits the audio data indicating user input acquired by the microphone 38B to the data processing device 12. In the data processing device 12, the specific processing unit 290 acquires the audio data.

[0239] Data generation model 58 is a so-called generative AI (Artificial Intelligence). An example of data generation model 58 is ChatGPT (registered trademark) (Internet search).<URL: https: / / openai.com / blog / chatgpt> ), Gemini (registered trademark) (Internet search) <url: https: gemini.google.com ?hl="ja">Examples of generative AI include the following. The data generation model 58 is obtained by performing deep learning on a neural network. The data generation model 58 is input with prompts containing instructions, and with inference data such as audio data representing speech, text data representing text, and image data representing images. The data generation model 58 infers from the input inference data according to the instructions indicated by the prompts, and outputs the inference results in data formats such as audio data and text data. Here, inference refers to, for example, analysis, classification, prediction, and / or summarization.

[0240] In the above embodiment, an example was given in which specific processing is performed by the data processing device 12, but the technology of this disclosure is not limited thereto, and the specific processing may also be performed by the smart device 14.

[0241] [Second Embodiment]

[0242] Figure 3 shows an example of the configuration of the data processing system 210 according to the second embodiment.

[0243] As shown in Figure 3, the data processing system 210 includes a data processing device 12 and smart glasses 214. An example of the data processing device 12 is a server.

[0244] The data processing device 12 comprises a computer 22, a database 24, and a communication interface 26. The computer 22 is an example of a "computer" related to the technology of this disclosure. The computer 22 comprises a processor 28, RAM 30, and storage 32. The processor 28, RAM 30, and storage 32 are connected to a bus 34. The database 24 and the communication interface 26 are also connected to the bus 34. The communication interface 26 is connected to a network 54. An example of the network 54 is a WAN (Wide Area Network) and / or a LAN (Local Area Network).

[0245] The smart glasses 214 include a computer 36, a microphone 238, a speaker 240, a camera 42, and a communication interface 44. The computer 36 includes a processor 46, RAM 48, and storage 50. The processor 46, RAM 48, and storage 50 are connected to a bus 52. The microphone 238, speaker 240, and camera 42 are also connected to the bus 52.

[0246] The microphone 238 receives voice signals from the user 20 and receives instructions from the user 20. The microphone 238 captures the voice signals from the user 20, converts the captured voice into audio data, and outputs it to the processor 46. The speaker 240 outputs audio according to the instructions from the processor 46.

[0247] Camera 42 is a small digital camera equipped with an optical system including a lens, aperture, and shutter, and an image sensor such as a CMOS (Complementary Metal-Oxide-Semiconductor) image sensor or a CCD (Charge Coupled Device) image sensor, and captures images of the area around the user 20 (for example, an imaging range defined by a field of view equivalent to the width of a typical healthy person's field of vision).

[0248] Communication interface 44 is connected to network 54. Communication interfaces 44 and 26 are responsible for the exchange of various information between processor 46 and processor 28 via network 54. The exchange of various information between processor 46 and processor 28 using communication interfaces 44 and 26 is performed in a secure manner.

[0249] Figure 4 shows an example of the main functions of the data processing device 12 and the smart glasses 214. As shown in Figure 4, the data processing device 12 performs specific processing using the processor 28. The storage 32 stores the specific processing program 56.

[0250] The specific processing program 56 is an example of a "program" relating to the technology of this disclosure. The processor 28 reads the specific processing program 56 from the storage 32 and executes the read specific processing program 56 on the RAM 30. The specific processing is realized by the processor 28 operating as a specific processing unit 290 in accordance with the specific processing program 56 executed on the RAM 30.

[0251] The storage 32 stores the data generation model 58 and the emotion identification model 59. The data generation model 58 and the emotion identification model 59 are used by the identification processing unit 290.

[0252] In the smart glasses 214, the processor 46 performs the reception output processing. The storage 50 stores the reception output program 60. The processor 46 reads the reception output program 60 from the storage 50 and executes the read reception output program 60 on the RAM 48. The reception output processing is realized by the processor 46 operating as a control unit 46A according to the reception output program 60 executed on the RAM 48.

[0253] Next, the identification processing performed by the identification processing unit 290 of the data processing device 12 will be described. In the following description, the data processing device 12 will be referred to as the "server" and the smart glasses 214 will be referred to as the "terminal".

[0254] This invention is a system using a video camera device and a server to help billiard players improve their skills. The role and operation of each component are described below.

[0255] Acquisition of video data

[0256] The device uses a video camera to capture real-time footage of the billiards game. This ensures that shots and ball movements are accurately recorded.

[0257] Analysis of real-time video

[0258] The server receives video data sent from the terminal and uses a deep learning model to analyze the player's movements and the ball's motion. The model processes each captured frame and extracts data for analyzing techniques such as shot speed, angle, and ball rotation.

[0259] Player skill evaluation

[0260] The server evaluates the player's technical characteristics based on the analyzed data. The computer automatically identifies the player's strong and weak shots, and pinpoints the success rate of those shots and areas for improvement.

[0261] Generating and providing feedback

[0262] Based on the results of the technical evaluation, the server generates improvement suggestions and strategic advice for the player in natural language. The generated feedback is sent to the device in real time. The device then communicates this to the user through visual displays or audio notifications.

[0263] Data storage and practice plan generation

[0264] The server archives data from each play session in a database. This saved data is used by players to review their gameplay at a later date. The server also generates personalized practice plans based on the accumulated performance data, helping users improve their skills more effectively.

[0265] For example, if a player repeatedly misses difficult shots, the server analyzes past success data and suggests alternative shots that have a higher success rate for that player. This provides a sense of accomplishment and helps prevent a decline in motivation.

[0266] The following describes the processing flow.

[0267] Step 1:

[0268] The terminal captures real-time video from the billiard table via a video camera device and sends the data to the server. The terminal continues to acquire video at the appropriate resolution and frame rate.

[0269] Step 2:

[0270] The server acquires the received real-time video data and begins analysis frame by frame. Using a deep learning model, it detects the player's movements and the ball's motion. Specifically, it identifies the angle and force of the player's shot, as well as the ball's collision point and trajectory.

[0271] Step 3:

[0272] The server evaluates the player's shots based on the analyzed data. By comparing technical data, it identifies the player's strengths and weaknesses in shots. By comparing this data with past play data, it analyzes the player's consistency and areas for improvement.

[0273] Step 4:

[0274] The server generates feedback for the player based on the evaluation results. Using natural language processing technology, it provides intuitive and effective improvement advice and shot strategies. The feedback is customized according to the player's skill level and the situation.

[0275] Step 5:

[0276] The server sends the generated feedback to the terminal. The terminal then communicates this to the user either by displaying it on the screen or announcing it through an audio device. The user receives real-time feedback and can immediately use it to improve their gameplay.

[0277] Step 6:

[0278] The server stores data from each play session in a database. This ensures that the underlying data supports the long-term improvement of players' skills. The data is used to generate individual practice plans and is utilized when users plan their future gameplay.

[0279] (Example 1)

[0280] Next, we will describe Example 1. In the following description, the data processing device 12 will be referred to as the "server" and the smart glasses 214 will be referred to as the "terminal".

[0281] Conventional skill improvement support systems have struggled to analyze user actions in real time and provide immediate improvement suggestions based on individual skill characteristics. Furthermore, they lacked the ability to automatically generate practice plans based on past performance data, preventing users from efficiently improving their skills. This created a need for an integrated skill improvement support system that includes rapid and accurate feedback.

[0282] The identification process performed by the identification processing unit 290 of the data processing device 12 in Example 1 is realized by the following means.

[0283] In this invention, the server includes means for acquiring real-time visual information from a visual sensor and detecting user operations and object movements using a machine learning model for analyzing the visual information; means for evaluating technical characteristics based on the user's operations and identifying strengths and operations that need improvement; and means for generating and providing information to the user in natural language regarding areas for improvement and operation plans based on the technical characteristics. This enables the user to receive detailed feedback in real time and obtain accurate advice and practice plans for improving their skills.

[0284] A "visual sensor" is a device that captures the movement of an object and converts the visual information into digital data.

[0285] "Visual information" refers to the digital data of videos and images obtained from visual sensors, and this data is the object of analysis.

[0286] "Machine learning model" refers to a set of algorithms for analyzing a large amount of data and automatically learning specific patterns and features.

[0287] "User" refers to an individual or group who wants to improve specific technologies using the system.

[0288] "Operation" refers to physical or digital actions by the user, including actions detected by visual sensors.

[0289] "Movement of an object" refers to the change in the position of a physical object detected within visual information.

[0290] "Technical characteristics" refer to specific features of performance and capabilities related to the operations of the user.

[0291] "Operations to be improved" refer to specific actions or procedures that the user should focus on modifying or practicing for technology improvement.

[0292] "Natural language" is the form of words that humans use in daily life and is used when the computer generates and provides information to the user.

[0293] "Information provision" refers to the act of conveying the analyzed data and proposals to the user, which is implemented visually or audibly.

[0294] This invention is a system for assisting users in improving their technology, and has a configuration including a visual sensor, a server, and a terminal. Specifically, a general video acquisition device is used as the visual sensor to capture the actions of physical competitions such as billiards in real time. This video data is transmitted to the server via the terminal.

[0295] The server analyzes the received visual information using a machine learning model. This process utilizes commonly used deep learning frameworks (e.g., TensorFlow) to evaluate the user's actions and object movements frame by frame. The server then extracts numerical data for the speed, angle, and ball rotation of each shot. Based on this data, the server assesses the user's technical characteristics and identifies their strengths and areas for improvement.

[0296] The server uses a generative AI model (e.g., GPT-3) based on the evaluation results to generate improvement suggestions and technical advice for the user in natural language. An example of a prompt used for this purpose would be, "Generate improvement advice based on this player's recent shot data."

[0297] The generated feedback is sent to the device in real time. The device uses visual display devices and audio output devices to convey this to the user. This allows the user to continuously review their gameplay and effectively improve their skills.

[0298] Furthermore, the server stores visual information and analysis results in memory. This data is used as a history for users to review past gameplay data later, and is also utilized to create personalized practice plans for continuous skill improvement. Based on the accumulated data, the server performs statistical analysis and provides practice plans optimized for the user. Through this entire process, the system can provide comprehensive support for users to improve their skills.

[0299] The flow of the specific processing in Example 1 will be explained using Figure 11.

[0300] Step 1:

[0301] The terminal uses a visual sensor to capture the user's competitive scene in real time. The input is the physical competitive landscape, and digital video data is generated as the output. As a result, the user's actions and the movement of objects are always recorded in the latest state. As a specific action, the sensor continuously takes pictures at regular intervals according to the frame rate.

[0302] Step 2:

[0303] The terminal sends the captured video data to the server. The input is the video data obtained in Step 1, and the output is the streaming data for the server to receive. The data is sent in real time, and at this time, compression technology is used to reduce the network load while maintaining the integrity of the data.

[0304] Step 3:

[0305] The server analyzes the received video data using a machine learning model. The input is the video data sent from the terminal, and the output is numerical data related to the user's actions and the position information of objects. In this analysis, feature point extraction and calculation of motion vectors are performed for each frame. Specifically, the server executes a deep learning framework and applies an algorithm to extract important physical information from each frame.

[0306] Step 4:

[0307] Based on the analysis results, the server evaluates the user's technical characteristics. The input is the output data of Step 3, and scoring data related to technical evaluation is generated as the output. The server automatically identifies the user's proficient operations and improvement points through a script and uses a technical evaluation model. As a specific action, statistical analysis is applied to score each operation based on the evaluation criteria.

[0308] Step 5:

[0309] The server uses a generative AI model to generate improvement suggestions in natural language based on technical evaluations. The input is the technical evaluation data from step 4, and the output is advice in natural language provided to the user. Specifically, prompt sentences are used and applied to the generative AI model to form concrete advice that helps improve the user's skills.

[0310] Step 6:

[0311] The terminal receives advice from the server and presents it to the user in real time. Input is feedback data from the server, and output is advice expressed visually or audibly. The terminal uses a display and speakers to inform the user of specific improvement measures. This allows the user to directly and immediately see the improvement measures.

[0312] Step 7:

[0313] The server stores video data and technical evaluations in its memory. The input is past pre-session data, and the output is historical data stored in the memory. This storage process is performed using database management functions. This enables long-term growth analysis of users.

[0314] Step 8:

[0315] The server generates individual practice plans based on stored data. The input is performance data stored in memory, and the output is a practice plan optimized for the user. The server combines statistical analysis and machine learning algorithms to construct effective practice suggestions.

[0316] (Application Example 1)

[0317] Next, we will explain Application Example 1. In the following explanation, the data processing device 12 will be referred to as the "server," and the smart glasses 214 will be referred to as the "terminal."

[0318] For billiards players to efficiently improve their skills, real-time motion analysis and feedback are essential. Furthermore, being able to receive this information instantly in a real playing environment greatly enhances the effectiveness of practice. However, conventional technology has not been able to create a system that accurately analyzes a player's movements and provides real-time, visually and audibly effective feedback.

[0319] The specific processing performed by the specific processing unit 290 of the data processing device 12 in Application Example 1 is realized by the following means.

[0320] In this invention, the server includes means for receiving real-time video information from a video camera device and detecting the player's body movements and the movement of a sphere using a machine learning model for analyzing the video information; means for generating and providing feedback to the player in natural language regarding areas for improvement and shot strategies based on these technical characteristics; and means for overlaying the feedback onto the real environment using smart glasses. This enables the player to obtain detailed and practical analytical information and feedback in real time, allowing for more effective technical improvement.

[0321] A "video camera device" is a device used to capture video information in real time.

[0322] "Real-time video information" refers to information that provides video data occurring at the present moment immediately.

[0323] A "machine learning model" is a model that uses algorithms to automatically learn patterns and rules from data.

[0324] "Physical movement" refers to the movement and posture of the player's body, and analyzing this is a means of extracting technical characteristics.

[0325] "Sphere motion" refers to the kinetic characteristics of a billiard ball, specifically its direction of movement, speed, and rotation.

[0326] "Technical characteristics" refer to a set of parameters used to evaluate the accuracy and movement characteristics of a player's shots.

[0327] "Natural language" refers to the linguistic expressions that people use on a daily basis, and is a means of conveying information to players in an easily understandable way as feedback.

[0328] "Smart glasses" are wearable devices that have the function of displaying visual information as augmented reality.

[0329] "Overlaying information onto the real world" is a technology that displays additional information superimposed on the visual information of the real world.

[0330] This invention provides a system to support the skill improvement of billiard players. This system operates in combination with a video camera device, smart glasses, and a server.

[0331] The server receives real-time video data from the video camera device. This data is processed using machine learning models to analyze in detail the player's body movements and the ball's motion during the shot. Specifically, it utilizes libraries such as TensorFlow in Python to recognize the ball's speed, rotation, and shot angle for each video frame.

[0332] Based on the analyzed data, the system evaluates the player's technical characteristics and identifies their strengths and weaknesses in shots. The server then generates natural language feedback, providing players with areas for improvement and strategic advice. This feedback is overlaid onto the real-world gameplay environment via smart glasses. The system is also designed to allow users to access information in real time using devices such as Google Glass.

[0333] Accumulated video and technical data are archived in a database on the server. This allows users to analyze past gameplay and track their progress. Player performance data is managed using database software such as Sarvapi and PostgreSQL. Furthermore, individual practice plans are generated, and practice menus tailored to each player are suggested.

[0334] For example, if a player repeatedly misses difficult shots, this system analyzes past success data and suggests alternative shots with a higher success rate. In this way, it helps the player gain a sense of accomplishment. An example of a prompt to input into the generative AI model might be, "Please tell me a temporary strategy for successfully executing a great mid-range shot while playing billiards."

[0335] The flow of a specific process in Application Example 1 will be explained using Figure 12.

[0336] Step 1:

[0337] The terminal captures the billiards game using a video camera device. The input is real-time video data, and the output is the same video data sent directly to the server. In this process, the camera operates at a high frame rate to capture the fast movement of the ball without delay.

[0338] Step 2:

[0339] The server receives video data, which is then used as input for a machine learning model. The input consists of captured image frames, and the output includes the ball's position and motion vector, as well as the player's movement pattern. TensorFlow is used to analyze the ball's speed and spin, and the angle of the shot, generating this numerical data.

[0340] Step 3:

[0341] The server evaluates the player's technical characteristics based on the generated analysis data. The input is the analysis data obtained in step 2, and the output includes strong and weak shots, success rates, etc. Based on this evaluation information, the server inputs suggestions for improvement and strategies for the next shot as prompts into the generating AI model.

[0342] Step 4:

[0343] The server generates feedback from the generative AI model in natural language and sends it to the smart glasses. The input is the suggestion information from the generative AI model, and the output is a natural language feedback message. In this step, the suggestions are summarized in a way that is easy for the user to understand.

[0344] Step 5:

[0345] The user's smart glasses receive feedback from the server and overlay it onto the real environment. The input is the feedback message from the server, and the output is the information visually displayed through the glasses. This allows the user to receive improvement advice in real time.

[0346] Step 6:

[0347] The server archives video data and analysis results from play sessions into a database. Inputs are the original video data and technical evaluation data, while outputs are accumulated historical data. This data serves as a basis for players to later review their progress.

[0348] Step 7:

[0349] The server uses accumulated historical data to generate individual practice plans. The input is the player's past performance data, and the output is the suggested practice plan. This step is performed automatically to support the player's skill improvement.

[0350] Furthermore, an emotion engine that estimates the user's emotions may be incorporated. That is, the identification processing unit 290 may use the emotion identification model 59 to estimate the user's emotions and perform identification processing using the user's emotions.

[0351] This invention combines a system designed to improve the skills of billiard players with an emotion engine that recognizes the user's emotions and provides appropriate feedback. Its configuration and operation are described in detail below.

[0352] Acquisition and analysis of video data

[0353] The terminal uses a video camera device to capture real-time video of a billiards game and sends it to the server. The server uses a deep learning model to analyze the player's movements and the ball's motion. This provides data for evaluating the technical aspects of the shots.

[0354] Technical evaluation and feedback generation

[0355] The server evaluates the player's technical characteristics based on the analyzed data. It identifies strengths and weaknesses in shots and generates improvement suggestions and shot strategies in natural language. The generated feedback is provided to the user visually and audibly through the terminal.

[0356] Emotional recognition and feedback regulation

[0357] The emotion engine analyzes the user's facial expressions and tone of voice to recognize their emotional state. The server then adjusts the feedback based on this emotional data, providing optimal feedback tailored to the user's motivation level. For example, if a player is feeling frustrated, the emotion engine generates positive reinforcement messages to encourage further attempts.

[0358] Data storage and practice plan generation

[0359] The server stores all play data and emotion recognition results in a database, enabling the visualization of long-term progress. This allows for the creation of personalized practice plans for each player, supporting skill improvement. Users can review their play history and emotion change trends and utilize their individual practice plans to efficiently improve their skills.

[0360] For example, if a player repeatedly misses a particular shot, the server uses an emotion engine to recognize the player's frustration and suggests shots they are good at while also providing encouraging messages. This helps the player regain confidence and continue practicing.

[0361] The following describes the processing flow.

[0362] Step 1:

[0363] The terminal uses a video camera device to capture video footage of the billiards game. The video data is transmitted to the server in real time. Camera settings are adjusted to ensure high-quality video and enable accurate motion analysis.

[0364] Step 2:

[0365] The server inputs the received video data into a deep learning model to analyze the player's movements and the ball's motion. The analysis includes shot speed, angle, and spin rate, and the player's current technical characteristics are extracted as numerical data.

[0366] Step 3:

[0367] The server evaluates the player's skills based on the analysis results, identifying their strengths and weaknesses in shots. By comparing this with past gameplay data, it highlights areas for improvement to enhance their skills.

[0368] Step 4:

[0369] The emotion engine recognizes the user's emotional state from their facial expressions, tone of voice, and body movements. The server then incorporates this emotional information into the feedback. For example, if the player is focused, it provides detailed technical analysis; on the other hand, if they are feeling frustrated, it generates encouraging messages.

[0370] Step 5:

[0371] The server generates feedback for the user using natural language processing, based on technical evaluation results and emotional states. The feedback is created as a customized message and includes strategic advice and suggestions for improvement tailored to the player's current situation.

[0372] Step 6:

[0373] The device presents the generated feedback to the user visually and audibly. Users can receive real-time advice and gain a deeper understanding of their own shots.

[0374] Step 7:

[0375] The server stores technical and emotional data for each session in a database. This allows users to have a history that they can refer to in later sessions and use to create practice plans.

[0376] Step 8:

[0377] Based on accumulated data, the server analyzes players' technical and emotional tendencies and generates personalized practice plans. These plans provide users with specific steps to achieve their goals and support long-term skill improvement.

[0378] (Example 2)

[0379] Next, we will describe Example 2. In the following description, the data processing device 12 will be referred to as the "server" and the smart glasses 214 will be referred to as the "terminal".

[0380] It is difficult for billiards players to objectively evaluate their own skills and clearly identify areas for improvement. Furthermore, providing feedback that takes into account the player's emotional state is needed to boost their motivation for practice. A system is required to address these challenges and provide effective practice plans tailored to each individual.

[0381] The identification process performed by the identification processing unit 290 of the data processing device 12 in Example 2 is realized by the following means.

[0382] In this invention, the server includes means for receiving real-time image information from a video recording device and confirming the player's movements and the ball's movement using a deep learning model for analyzing the image information; means for recording the image information and technical elements in an information storage medium and managing a history for analyzing the player's progress; and means for analyzing the user's emotional state, adjusting the response content based on the emotional information, and providing an optimal response. This enables billiard players to effectively improve their skills and receive feedback tailored to their emotional state.

[0383] "Video recording equipment" is a general term for devices used to record video and audio and to later play them back or transmit them.

[0384] "Real-time image information" refers to data used to acquire and process images from the real world in real time.

[0385] A "deep learning model" is a type of algorithm that uses a multi-layer neuron network to learn the features of data and perform analysis and prediction.

[0386] "Player movements" in billiards refers to all physical actions performed by a player, including shots and movement.

[0387] "Ball movement" in billiards refers to the motion of the ball as it rolls, spins, or collides in response to the force exerted by the player.

[0388] "Technical elements" refer to skill-based characteristics such as force and precision, which are evaluated as parameters of shots and play in billiards.

[0389] "Information storage medium" is a general term for hardware or digital storage used to record, store, and manage data.

[0390] "Emotional state" refers to the psychological and emotional state of a user, measured based on their behavior, facial expressions, voice, etc.

[0391] "Response content" refers to all feedback messages provided to the user, such as criticisms, instructions, or encouragement.

[0392] This invention is a system that operates by combining multiple components to support the skill improvement of billiard players. Specific embodiments of this system are shown below.

[0393] The terminal uses a video recording device to capture real-time image information of a billiards game and transmits the data to a server. The terminal is positioned optimally to smoothly capture high-resolution video. The video data is efficiently transmitted to the server via the network.

[0394] The server inputs the received real-time image information into a deep learning model. Specifically, it uses software frameworks such as TensorFlow and PyTorch to analyze the player's motion and the ball's movement. The server then extracts technical elements such as shot speed, angle, and contact point, thereby evaluating the player's skill.

[0395] Based on the evaluated technical elements, the server generates suggestions for improvement and shortcuts for the user in natural language. This process utilizes a natural language generation AI model, and the generated feedback is provided to the user visually and audibly through the device. The device presents the feedback using both screen display and audio output.

[0396] Furthermore, the emotion engine analyzes the user's emotional state. The device obtains emotional data by recording the user's facial expressions and collecting their voice tone. The server uses this data to understand the user's emotional state and adjust the feedback accordingly. For example, if the user is irritated, a positive message may be added to boost their motivation.

[0397] Finally, the server saves all play data and emotion recognition results to a data storage medium. This allows users to refer to past performance data and emotion fluctuations to generate individually customized practice plans.

[0398] For example, if a player repeatedly misses a particular shot, the system recognizes the player's frustration and offers words of encouragement while suggesting alternative shots the player is good at. This allows the player to move on to the next step with confidence.

[0399] An example of a prompt is: "Analyze the player's movements and the ball's trajectory during a billiards match, evaluate their technical characteristics, and determine their strengths and weaknesses. Furthermore, create appropriate feedback in natural language, taking into account the player's emotional state."

[0400] The flow of the specific processing in Example 2 will be explained using Figure 13.

[0401] Step 1:

[0402] The terminal activates a video recording device and captures real-time image information of a billiards game. The input is a real billiards game in action, and the output is high-resolution video data. The quality of this data is ensured by adjusting the camera position to accurately capture the players' movements and the ball's trajectory.

[0403] Step 2:

[0404] The terminal transmits captured image information to the server. The input is video data obtained from a video recording device, and the output is data transfer to the server. In this process, data compression technology is used to efficiently utilize network bandwidth.

[0405] Step 3:

[0406] The server inputs the received image information into a deep learning model. The input is video data sent from the terminal, and the output is the analysis results of the player's motion and the ball's movement. As part of the data processing, the model analyzes the image frame by frame and extracts technical elements such as the motion trajectory and velocity of each object.

[0407] Step 4:

[0408] The server evaluates the player's technical characteristics based on the analysis data. The input is the analysis results obtained in step 3, and the output is the evaluated technical characteristics as well as identification data of strengths and weaknesses. This makes it possible to evaluate shot accuracy and performance.

[0409] Step 5:

[0410] The server generates improvement suggestions and shortcuts in natural language based on the evaluated technical features. The input is evaluation data for the technical features, and the output is feedback messages in natural language. A generative AI model is used, which creates sentences that are easy for humans to understand.

[0411] Step 6:

[0412] The device displays and provides generated feedback to the user via audio. Input is feedback messages sent from the server, and output is a visual and audible presentation of specific advice to the user. Information is delivered to the user using the screen and speaker.

[0413] Step 7:

[0414] The emotion engine analyzes the user's facial expressions and voice tone to recognize their emotional state. The input is image and audio data obtained from the user, and the output is the result of the emotional state analysis. Facial recognition and voice analysis algorithms are used for data processing.

[0415] Step 8:

[0416] The server adjusts the feedback content according to the user's emotional state to maintain or improve their motivation. The input consists of emotional state analysis data and feedback data generated by the server, while the output is the adjusted feedback message. This enables personalized responses that take the user's emotions into consideration.

[0417] Step 9:

[0418] The server stores all gameplay data and emotion recognition results in an information storage medium. Input is general data, and output is historical information stored in a database. This allows for tracking and analysis of players' long-term progress.

[0419] (Application Example 2)

[0420] Next, we will explain application example 2. In the following explanation, the data processing device 12 will be referred to as the "server," and the smart glasses 214 will be referred to as the "terminal."

[0421] For players who want to improve their skills in billiards and other sports, there is a need to provide not only technical feedback but also emotional support to boost their motivation and acquire skills more effectively. However, conventional technical support systems lack mechanisms to adjust feedback according to the player's emotional state, which hinders skill improvement.

[0422] The specific processing performed by the specific processing unit 290 of the data processing device 12 in Application Example 2 is realized by the following means.

[0423] In this invention, the server includes means for receiving image information captured without delay from a video recording device and detecting the user's actions and the movement of objects using a machine learning model for analyzing the image information; means for evaluating the technical characteristics based on the user's actions and identifying actions in which the user is good and bad; and means for analyzing the user's emotional state using an emotion analysis engine and adjusting the feedback content based on that state. This makes it possible to improve both the technical and emotional aspects of the user by providing technical feedback along with support tailored to their individual emotions.

[0424] A "video recording device" is a device for acquiring image information in real time without any time lag.

[0425] A "machine learning model" is an algorithm that analyzes data, recognizes patterns, and predicts the future.

[0426] "Users" refers to individuals or organizations that use the system for the purpose of improving their skills.

[0427] "Action and motion of objects" refers to physical movements, including the position, posture, and velocity of the user or object.

[0428] "Technical characteristics" refer to the unique technical operating patterns and characteristics of the user or object.

[0429] An "emotion analysis engine" is an analytical device and program that determines a user's emotional state from their facial expressions and tone of voice.

[0430] "Feedback" refers to information provided to users, either technical or emotional, to encourage improvement and support.

[0431] As an example of the application of this invention, the details of a system for supporting the skill improvement of users in a billiard hall are shown below.

[0432] First, a video recording device is used to capture the user's movements and the movement of the billiard ball in real time. The terminal sends this video data to a server. On the server, a machine learning model is applied to analyze the video data, and the user's shots and the movement of the ball are analyzed to extract technical features.

[0433] Next, the server identifies the user's strengths and weaknesses based on the extracted technical characteristics and generates feedback in natural language based on the results. This feedback is provided to the user visually and audibly via a visual display device.

[0434] Furthermore, an emotion analysis engine is used to analyze the user's facial expressions and tone of voice to recognize their emotional state. Based on this information, the content of the feedback is adjusted according to the user's emotional state to provide optimal support.

[0435] For example, if a user repeatedly misses a particular shot, the server uses an emotion analysis engine to recognize the user's frustration, suggest shots that the user excels at, and provide an encouraging message. In this process, a generative AI model is used for natural language processing, and a prompt message such as, "Please provide the latest billiard shot analysis and feedback. Please also consider the player's emotional data and create an encouraging message," is used.

[0436] In this way, users can efficiently improve their billiards skills while receiving support from both a technical and emotional perspective.

[0437] The flow of a specific process in Application Example 2 will be explained using Figure 14.

[0438] Step 1:

[0439] The terminal captures the user's actions and the movement of the billiard ball in real time via a video recording device and transmits the video data to the server. The input is video data of the user's actions and the ball's movement, and the output is the data captured without any time lag and sent to the server.

[0440] Step 2:

[0441] The server uses a machine learning model to analyze the received video data. The input is video data, and the output is the extraction of technical features related to the user's actions and the ball's movement. This allows the data to be analyzed through the model, identifying specific patterns and trends.

[0442] Step 3:

[0443] The server evaluates the acquired technical characteristics and identifies the user's strengths and weaknesses in shots. The input is data on technical characteristics, and the output is information indicating strengths and weaknesses in shots. Based on this, the server prepares the basic data for generating feedback.

[0444] Step 4:

[0445] Using a generative AI model, the server generates suggestions for improvement and shot strategies for the user in natural language. The input is information about a specific shot, and the output is a detailed feedback message. The server uses prompts to construct the feedback naturally.

[0446] Step 5:

[0447] The server activates an emotion analysis engine to recognize the user's emotional state from their facial expressions and voice data. The input is the user's facial expressions and voice tone information, and the output is an evaluation of their emotional state. This data allows the server to understand how the user is feeling.

[0448] Step 6:

[0449] The server adjusts the feedback content to reflect the emotional state and delivers it to the user visually and audibly via a visual display device. Input is an assessment of the emotional state and pre-generated feedback, while output is the adjusted visual and audible feedback. The goal is to deliver the most appropriate feedback to the user.

[0450] Step 7:

[0451] The server stores technical characteristics and emotional state data in an information aggregation device to manage user progress over the long term. The input is the entire analyzed data, and the output is the stored historical data. This ensures that the history for later use is always up-to-date.

[0452] The specific processing unit 290 transmits the result of the specific processing to the smart glasses 214. In the smart glasses 214, the control unit 46A causes the speaker 240 to output the result of the specific processing. The microphone 238 acquires audio indicating user input for the result of the specific processing. The control unit 46A transmits the audio data indicating user input acquired by the microphone 238 to the data processing unit 12. In the data processing unit 12, the specific processing unit 290 acquires the audio data.

[0453] Data generation model 58 is a type of so-called generative AI (Artificial Intelligence). One example of data generation model 58 is ChatGPT (Internet Search).<URL: https: / / openai.com / blog / chatgpt> ), Gemini (Internet search) <url: https: gemini.google.com ?hl="ja">Examples of generative AI include the following. The data generation model 58 is obtained by performing deep learning on a neural network. The data generation model 58 is input with prompts containing instructions, and with inference data such as audio data representing speech, text data representing text, and image data representing images. The data generation model 58 infers from the input inference data according to the instructions indicated by the prompts, and outputs the inference results in data formats such as audio data and text data. Here, inference refers to, for example, analysis, classification, prediction, and / or summarization.

[0454] In the above embodiment, an example was given in which specific processing is performed by the data processing device 12, but the technology of this disclosure is not limited thereto, and the specific processing may also be performed by the smart glasses 214.

[0455] [Third Embodiment]

[0456] Figure 5 shows an example of the configuration of the data processing system 310 according to the third embodiment.

[0457] As shown in Figure 5, the data processing system 310 includes a data processing device 12 and a headset terminal 314. An example of the data processing device 12 is a server.

[0458] The data processing device 12 comprises a computer 22, a database 24, and a communication interface 26. The computer 22 is an example of a "computer" related to the technology of this disclosure. The computer 22 comprises a processor 28, RAM 30, and storage 32. The processor 28, RAM 30, and storage 32 are connected to a bus 34. The database 24 and the communication interface 26 are also connected to the bus 34. The communication interface 26 is connected to a network 54. An example of the network 54 is a WAN (Wide Area Network) and / or a LAN (Local Area Network).

[0459] The headset terminal 314 includes a computer 36, a microphone 238, a speaker 240, a camera 42, a communication interface 44, and a display 343. The computer 36 includes a processor 46, RAM 48, and storage 50. The processor 46, RAM 48, and storage 50 are connected to a bus 52. The microphone 238, speaker 240, camera 42, and display 343 are also connected to the bus 52.

[0460] The microphone 238 receives voice signals from the user 20 and receives instructions from the user 20. The microphone 238 captures the voice signals from the user 20, converts the captured voice into audio data, and outputs it to the processor 46. The speaker 240 outputs audio according to the instructions from the processor 46.

[0461] Camera 42 is a small digital camera equipped with an optical system including a lens, aperture, and shutter, and an image sensor such as a CMOS (Complementary Metal-Oxide-Semiconductor) image sensor or a CCD (Charge Coupled Device) image sensor, and captures images of the area around the user 20 (for example, an imaging range defined by a field of view equivalent to the width of a typical healthy person's field of vision).

[0462] Communication interface 44 is connected to network 54. Communication interfaces 44 and 26 are responsible for the exchange of various information between processor 46 and processor 28 via network 54. The exchange of various information between processor 46 and processor 28 using communication interfaces 44 and 26 is performed in a secure manner.

[0463] Figure 6 shows an example of the main functions of the data processing device 12 and the headset terminal 314. As shown in Figure 6, the data processing device 12 performs specific processing using the processor 28. The storage 32 stores the specific processing program 56.

[0464] The specific processing program 56 is an example of a "program" relating to the technology of this disclosure. The processor 28 reads the specific processing program 56 from the storage 32 and executes the read specific processing program 56 on the RAM 30. The specific processing is realized by the processor 28 operating as a specific processing unit 290 in accordance with the specific processing program 56 executed on the RAM 30.

[0465] The storage 32 stores the data generation model 58 and the emotion identification model 59. The data generation model 58 and the emotion identification model 59 are used by the identification processing unit 290.

[0466] In the headset terminal 314, the processor 46 performs the reception output processing. The storage 50 stores the reception output program 60. The processor 46 reads the reception output program 60 from the storage 50 and executes the read reception output program 60 on the RAM 48. The reception output processing is realized by the processor 46 operating as a control unit 46A according to the reception output program 60 executed on the RAM 48.

[0467] Next, the specific processing performed by the specific processing unit 290 of the data processing device 12 will be described. In the following description, the data processing device 12 will be referred to as the "server" and the headset terminal 314 will be referred to as the "terminal".

[0468] This invention is a system using a video camera device and a server to help billiard players improve their skills. The role and operation of each component are described below.

[0469] Acquisition of video data

[0470] The device uses a video camera to capture real-time footage of the billiards game. This ensures that shots and ball movements are accurately recorded.

[0471] Analysis of real-time video

[0472] The server receives video data sent from the terminal and uses a deep learning model to analyze the player's movements and the ball's motion. The model processes each captured frame and extracts data for analyzing techniques such as shot speed, angle, and ball rotation.

[0473] Player skill evaluation

[0474] The server evaluates the player's technical characteristics based on the analyzed data. The computer automatically identifies the player's strong and weak shots, and pinpoints the success rate of those shots and areas for improvement.

[0475] Generating and providing feedback

[0476] Based on the results of the technical evaluation, the server generates improvement suggestions and strategic advice for the player in natural language. The generated feedback is sent to the device in real time. The device then communicates this to the user through visual displays or audio notifications.

[0477] Data storage and practice plan generation

[0478] The server archives data from each play session in a database. This saved data is used by players to review their gameplay at a later date. The server also generates personalized practice plans based on the accumulated performance data, helping users improve their skills more effectively.

[0479] For example, if a player repeatedly misses difficult shots, the server analyzes past success data and suggests alternative shots that have a higher success rate for that player. This provides a sense of accomplishment and helps prevent a decline in motivation.

[0480] The following describes the processing flow.

[0481] Step 1:

[0482] The terminal captures real-time video from the billiard table via a video camera device and sends the data to the server. The terminal continues to acquire video at the appropriate resolution and frame rate.

[0483] Step 2:

[0484] The server acquires the received real-time video data and begins analysis frame by frame. Using a deep learning model, it detects the player's movements and the ball's motion. Specifically, it identifies the angle and force of the player's shot, as well as the ball's collision point and trajectory.

[0485] Step 3:

[0486] The server evaluates the player's shots based on the analyzed data. By comparing technical data, it identifies the player's strengths and weaknesses in shots. By comparing this data with past play data, it analyzes the player's consistency and areas for improvement.

[0487] Step 4:

[0488] The server generates feedback for the player based on the evaluation results. Using natural language processing technology, it provides intuitive and effective improvement advice and shot strategies. The feedback is customized according to the player's skill level and the situation.

[0489] Step 5:

[0490] The server sends the generated feedback to the terminal. The terminal then communicates this to the user either by displaying it on the screen or announcing it through an audio device. The user receives real-time feedback and can immediately use it to improve their gameplay.

[0491] Step 6:

[0492] The server stores data from each play session in a database. This ensures that the underlying data supports the long-term improvement of players' skills. The data is used to generate individual practice plans and is utilized when users plan their future gameplay.

[0493] (Example 1)

[0494] Next, we will describe Example 1. In the following description, the data processing device 12 will be referred to as the "server," and the headset-type terminal 314 will be referred to as the "terminal."

[0495] Conventional skill improvement support systems have struggled to analyze user actions in real time and provide immediate improvement suggestions based on individual skill characteristics. Furthermore, they lacked the ability to automatically generate practice plans based on past performance data, preventing users from efficiently improving their skills. This created a need for an integrated skill improvement support system that includes rapid and accurate feedback.

[0496] The identification process performed by the identification processing unit 290 of the data processing device 12 in Example 1 is realized by the following means.

[0497] In this invention, the server includes means for acquiring real-time visual information from a visual sensor and detecting user operations and object movements using a machine learning model for analyzing the visual information; means for evaluating technical characteristics based on the user's operations and identifying strengths and operations that need improvement; and means for generating and providing information to the user in natural language regarding areas for improvement and operation plans based on the technical characteristics. This enables the user to receive detailed feedback in real time and obtain accurate advice and practice plans for improving their skills.

[0498] A "visual sensor" is a device that captures the movement of an object and converts the visual information into digital data.

[0499] "Visual information" refers to digital data of images and videos acquired from visual sensors, and this data is the subject of analysis.

[0500] A "machine learning model" is a collection of algorithms that analyze large amounts of data and automatically learn specific patterns and features.

[0501] "Users" refer to individuals or organizations that wish to use the system to improve specific technologies.

[0502] "Operation" refers to physical or digital actions performed by the user, including actions detected by visual sensors.

[0503] "Object movement" refers to changes in the position of physical objects detected within visual information.

[0504] "Technical characteristics" refer to the specific features of performance and capabilities related to user operation.

[0505] "Operations that need improvement" are specific actions or procedures that users should focus on modifying or practicing in order to improve their skills.

[0506] "Natural language" refers to the form of language that humans use in everyday life, and is used when computers generate and provide information to users.

[0507] "Information provision" refers to the act of conveying analyzed data and suggestions to users, and this can be done visually or audibly.

[0508] This invention is a system designed to support the improvement of users' skills and comprises a visual sensor, a server, and a terminal. Specifically, a general-purpose video acquisition device is used as the visual sensor to capture the actions of physical sports such as billiards in real time. This video data is transmitted to the server via the terminal.

[0509] The server analyzes the received visual information using a machine learning model. This process utilizes commonly used deep learning frameworks (e.g., TensorFlow) to evaluate the user's actions and object movements frame by frame. The server then extracts numerical data for the speed, angle, and ball rotation of each shot. Based on this data, the server assesses the user's technical characteristics and identifies their strengths and areas for improvement.

[0510] The server uses a generative AI model (e.g., GPT-3) based on the evaluation results to generate improvement suggestions and technical advice for the user in natural language. An example of a prompt used for this purpose would be, "Generate improvement advice based on this player's recent shot data."

[0511] The generated feedback is sent to the device in real time. The device uses visual display devices and audio output devices to convey this to the user. This allows the user to continuously review their gameplay and effectively improve their skills.

[0512] Furthermore, the server stores visual information and analysis results in memory. This data is used as a history for users to review past gameplay data later, and is also utilized to create personalized practice plans for continuous skill improvement. Based on the accumulated data, the server performs statistical analysis and provides practice plans optimized for the user. Through this entire process, the system can provide comprehensive support for users to improve their skills.

[0513] The flow of the specific processing in Example 1 will be explained using Figure 11.

[0514] Step 1:

[0515] The device uses a visual sensor to capture the user's athletic performance in real time. The input is the physical athletic scene, and the output is digital video data. This ensures that the user's movements and the movement of objects are always recorded in their most up-to-date state. Specifically, the sensor continuously captures images at regular intervals according to the frame rate.

[0516] Step 2:

[0517] The terminal sends the captured video data to the server. The input is the video data acquired in step 1, and the output is streaming data for the server to receive. The data is transmitted in real time, and compression technology is used to reduce network load while maintaining data integrity.

[0518] Step 3:

[0519] The server analyzes the received video data using a machine learning model. The input is video data sent from the terminal, and the output is numerical data related to the user's movements and the position of objects. This analysis involves feature point extraction and motion vector calculation for each frame. Specifically, the server runs a deep learning framework and applies algorithms to extract important physical information from each frame.

[0520] Step 4:

[0521] The server evaluates the user's technical characteristics based on the analysis results. The input is the output data from step 3, and the output is scoring data related to the technical evaluation. The server automatically identifies the user's strengths and areas for improvement through a script and uses a technical evaluation model. Specifically, it applies statistical analysis and scores each operation based on evaluation criteria.

[0522] Step 5:

[0523] The server uses a generative AI model to generate improvement suggestions in natural language based on technical evaluations. The input is the technical evaluation data from step 4, and the output is advice in natural language provided to the user. Specifically, prompt sentences are used and applied to the generative AI model to form concrete advice that helps improve the user's skills.

[0524] Step 6:

[0525] The terminal receives advice from the server and presents it to the user in real time. Input is feedback data from the server, and output is advice expressed visually or audibly. The terminal uses a display and speakers to inform the user of specific improvement measures. This allows the user to directly and immediately see the improvement measures.

[0526] Step 7:

[0527] The server stores video data and technical evaluations in its memory. The input is past pre-session data, and the output is historical data stored in the memory. This storage process is performed using database management functions. This enables long-term growth analysis of users.

[0528] Step 8:

[0529] The server generates individual practice plans based on stored data. The input is performance data stored in memory, and the output is a practice plan optimized for the user. The server combines statistical analysis and machine learning algorithms to construct effective practice suggestions.

[0530] (Application Example 1)

[0531] Next, we will explain Application Example 1. In the following explanation, the data processing device 12 will be referred to as the "server," and the headset-type terminal 314 will be referred to as the "terminal."

[0532] For billiards players to efficiently improve their skills, real-time motion analysis and feedback are essential. Furthermore, being able to receive this information instantly in a real playing environment greatly enhances the effectiveness of practice. However, conventional technology has not been able to create a system that accurately analyzes a player's movements and provides real-time, visually and audibly effective feedback.

[0533] The specific processing performed by the specific processing unit 290 of the data processing device 12 in Application Example 1 is realized by the following means.

[0534] In this invention, the server includes means for receiving real-time video information from a video camera device and detecting the player's body movements and the movement of a sphere using a machine learning model for analyzing the video information; means for generating and providing feedback to the player in natural language regarding areas for improvement and shot strategies based on these technical characteristics; and means for overlaying the feedback onto the real environment using smart glasses. This enables the player to obtain detailed and practical analytical information and feedback in real time, allowing for more effective technical improvement.

[0535] A "video camera device" is a device used to capture video information in real time.

[0536] "Real-time video information" refers to information that provides video data occurring at the present moment immediately.

[0537] A "machine learning model" is a model that uses algorithms to automatically learn patterns and rules from data.

[0538] "Physical movement" refers to the movement and posture of the player's body, and analyzing this is a means of extracting technical characteristics.

[0539] "Sphere motion" refers to the kinetic characteristics of a billiard ball, specifically its direction of movement, speed, and rotation.

[0540] "Technical characteristics" refer to a set of parameters used to evaluate the accuracy and movement characteristics of a player's shots.

[0541] "Natural language" refers to the linguistic expressions that people use on a daily basis, and is a means of conveying information to players in an easily understandable way as feedback.

[0542] "Smart glasses" are wearable devices that have the function of displaying visual information as augmented reality.

[0543] "Overlaying information onto the real world" is a technology that displays additional information superimposed on the visual information of the real world.

[0544] This invention provides a system to support the skill improvement of billiard players. This system operates in combination with a video camera device, smart glasses, and a server.

[0545] The server receives real-time video data from the video camera device. This data is processed using machine learning models to analyze in detail the player's body movements and the ball's motion during the shot. Specifically, it utilizes libraries such as TensorFlow in Python to recognize the ball's speed, rotation, and shot angle for each video frame.

[0546] Based on the analyzed data, the system evaluates the player's technical characteristics and identifies their strengths and weaknesses in shots. The server then generates natural language feedback, providing players with areas for improvement and strategic advice. This feedback is overlaid onto the real-world gameplay environment via smart glasses. The system is also designed to allow users to access information in real time using devices such as Google Glass.

[0547] Accumulated video and technical data are archived in a database on the server. This allows users to analyze past gameplay and track their progress. Player performance data is managed using database software such as Sarvapi and PostgreSQL. Furthermore, individual practice plans are generated, and practice menus tailored to each player are suggested.

[0548] For example, if a player repeatedly misses difficult shots, this system analyzes past success data and suggests alternative shots with a higher success rate. In this way, it helps the player gain a sense of accomplishment. An example of a prompt to input into the generative AI model might be, "Please tell me a temporary strategy for successfully executing a great mid-range shot while playing billiards."

[0549] The flow of a specific process in Application Example 1 will be explained using Figure 12.

[0550] Step 1:

[0551] The terminal captures the billiards game using a video camera device. The input is real-time video data, and the output is the same video data sent directly to the server. In this process, the camera operates at a high frame rate to capture the fast movement of the ball without delay.

[0552] Step 2:

[0553] The server receives video data, which is then used as input for a machine learning model. The input consists of captured image frames, and the output includes the ball's position and motion vector, as well as the player's movement pattern. TensorFlow is used to analyze the ball's speed and spin, and the angle of the shot, generating this numerical data.

[0554] Step 3:

[0555] The server evaluates the player's technical characteristics based on the generated analysis data. The input is the analysis data obtained in step 2, and the output includes strong and weak shots, success rates, etc. Based on this evaluation information, the server inputs suggestions for improvement and strategies for the next shot as prompts into the generating AI model.

[0556] Step 4:

[0557] The server generates feedback from the generative AI model in natural language and sends it to the smart glasses. The input is the suggestion information from the generative AI model, and the output is a natural language feedback message. In this step, the suggestions are summarized in a way that is easy for the user to understand.

[0558] Step 5:

[0559] The user's smart glasses receive feedback from the server and overlay it onto the real environment. The input is the feedback message from the server, and the output is the information visually displayed through the glasses. This allows the user to receive improvement advice in real time.

[0560] Step 6:

[0561] The server archives video data and analysis results from play sessions into a database. Inputs are the original video data and technical evaluation data, while outputs are accumulated historical data. This data serves as a basis for players to later review their progress.

[0562] Step 7:

[0563] The server uses accumulated historical data to generate individual practice plans. The input is the player's past performance data, and the output is the suggested practice plan. This step is performed automatically to support the player's skill improvement.

[0564] Furthermore, an emotion engine that estimates the user's emotions may be incorporated. That is, the identification processing unit 290 may use the emotion identification model 59 to estimate the user's emotions and perform identification processing using the user's emotions.

[0565] This invention combines a system designed to improve the skills of billiard players with an emotion engine that recognizes the user's emotions and provides appropriate feedback. Its configuration and operation are described in detail below.

[0566] Acquisition and analysis of video data

[0567] The terminal uses a video camera device to capture real-time video of a billiards game and sends it to the server. The server uses a deep learning model to analyze the player's movements and the ball's motion. This provides data for evaluating the technical aspects of the shots.

[0568] Technical evaluation and feedback generation

[0569] The server evaluates the player's technical characteristics based on the analyzed data. It identifies strengths and weaknesses in shots and generates improvement suggestions and shot strategies in natural language. The generated feedback is provided to the user visually and audibly through the terminal.

[0570] Emotional recognition and feedback regulation

[0571] The emotion engine analyzes the user's facial expressions and tone of voice to recognize their emotional state. The server then adjusts the feedback based on this emotional data, providing optimal feedback tailored to the user's motivation level. For example, if a player is feeling frustrated, the emotion engine generates positive reinforcement messages to encourage further attempts.

[0572] Data storage and practice plan generation

[0573] The server stores all play data and emotion recognition results in a database, enabling the visualization of long-term progress. This allows for the creation of personalized practice plans for each player, supporting skill improvement. Users can review their play history and emotion change trends and utilize their individual practice plans to efficiently improve their skills.

[0574] For example, if a player repeatedly misses a particular shot, the server uses an emotion engine to recognize the player's frustration and suggests shots they are good at while also providing encouraging messages. This helps the player regain confidence and continue practicing.

[0575] The following describes the processing flow.

[0576] Step 1:

[0577] The terminal uses a video camera device to capture video footage of the billiards game. The video data is transmitted to the server in real time. Camera settings are adjusted to ensure high-quality video and enable accurate motion analysis.

[0578] Step 2:

[0579] The server inputs the received video data into a deep learning model to analyze the player's movements and the ball's motion. The analysis includes shot speed, angle, and spin rate, and the player's current technical characteristics are extracted as numerical data.

[0580] Step 3:

[0581] The server evaluates the player's skills based on the analysis results, identifying their strengths and weaknesses in shots. By comparing this with past gameplay data, it highlights areas for improvement to enhance their skills.

[0582] Step 4:

[0583] The emotion engine recognizes the user's emotional state from their facial expressions, tone of voice, and body movements. The server then incorporates this emotional information into the feedback. For example, if the player is focused, it provides detailed technical analysis; on the other hand, if they are feeling frustrated, it generates encouraging messages.

[0584] Step 5:

[0585] The server generates feedback for the user using natural language processing, based on technical evaluation results and emotional states. The feedback is created as a customized message and includes strategic advice and suggestions for improvement tailored to the player's current situation.

[0586] Step 6:

[0587] The device presents the generated feedback to the user visually and audibly. Users can receive real-time advice and gain a deeper understanding of their own shots.

[0588] Step 7:

[0589] The server stores technical and emotional data for each session in a database. This allows users to have a history that they can refer to in later sessions and use to create practice plans.

[0590] Step 8:

[0591] Based on accumulated data, the server analyzes players' technical and emotional tendencies and generates personalized practice plans. These plans provide users with specific steps to achieve their goals and support long-term skill improvement.

[0592] (Example 2)

[0593] Next, we will describe Example 2. In the following description, the data processing device 12 will be referred to as the "server," and the headset-type terminal 314 will be referred to as the "terminal."

[0594] It is difficult for billiards players to objectively evaluate their own skills and clearly identify areas for improvement. Furthermore, providing feedback that takes into account the player's emotional state is needed to boost their motivation for practice. A system is required to address these challenges and provide effective practice plans tailored to each individual.

[0595] The identification process performed by the identification processing unit 290 of the data processing device 12 in Example 2 is realized by the following means.

[0596] In this invention, the server includes means for receiving real-time image information from a video recording device and confirming the player's movements and the ball's movement using a deep learning model for analyzing the image information; means for recording the image information and technical elements in an information storage medium and managing a history for analyzing the player's progress; and means for analyzing the user's emotional state, adjusting the response content based on the emotional information, and providing an optimal response. This enables billiard players to effectively improve their skills and receive feedback tailored to their emotional state.

[0597] "Video recording equipment" is a general term for devices used to record video and audio and to later play them back or transmit them.

[0598] "Real-time image information" refers to data used to acquire and process images from the real world in real time.

[0599] A "deep learning model" is a type of algorithm that uses a multi-layer neuron network to learn the features of data and perform analysis and prediction.

[0600] "Player movements" in billiards refers to all physical actions performed by a player, including shots and movement.

[0601] "Ball movement" in billiards refers to the motion of the ball as it rolls, spins, or collides in response to the force exerted by the player.

[0602] "Technical elements" refer to skill-based characteristics such as force and precision, which are evaluated as parameters of shots and play in billiards.

[0603] "Information storage medium" is a general term for hardware or digital storage used to record, store, and manage data.

[0604] "Emotional state" refers to the psychological and emotional state of a user, measured based on their behavior, facial expressions, voice, etc.

[0605] "Response content" refers to all feedback messages provided to the user, such as criticisms, instructions, or encouragement.

[0606] This invention is a system that operates by combining multiple components to support the skill improvement of billiard players. Specific embodiments of this system are shown below.

[0607] The terminal uses a video recording device to capture real-time image information of a billiards game and transmits the data to a server. The terminal is positioned optimally to smoothly capture high-resolution video. The video data is efficiently transmitted to the server via the network.

[0608] The server inputs the received real-time image information into a deep learning model. Specifically, it uses software frameworks such as TensorFlow and PyTorch to analyze the player's motion and the ball's movement. The server then extracts technical elements such as shot speed, angle, and contact point, thereby evaluating the player's skill.

[0609] Based on the evaluated technical elements, the server generates suggestions for improvement and shortcuts for the user in natural language. This process utilizes a natural language generation AI model, and the generated feedback is provided to the user visually and audibly through the device. The device presents the feedback using both screen display and audio output.

[0610] Furthermore, the emotion engine analyzes the user's emotional state. The device obtains emotional data by recording the user's facial expressions and collecting their voice tone. The server uses this data to understand the user's emotional state and adjust the feedback accordingly. For example, if the user is irritated, a positive message may be added to boost their motivation.

[0611] Finally, the server saves all play data and emotion recognition results to a data storage medium. This allows users to refer to past performance data and emotion fluctuations to generate individually customized practice plans.

[0612] For example, if a player repeatedly misses a particular shot, the system recognizes the player's frustration and offers words of encouragement while suggesting alternative shots the player is good at. This allows the player to move on to the next step with confidence.

[0613] An example of a prompt is: "Analyze the player's movements and the ball's trajectory during a billiards match, evaluate their technical characteristics, and determine their strengths and weaknesses. Furthermore, create appropriate feedback in natural language, taking into account the player's emotional state."

[0614] The flow of the specific processing in Example 2 will be explained using Figure 13.

[0615] Step 1:

[0616] The terminal activates a video recording device and captures real-time image information of a billiards game. The input is a real billiards game in action, and the output is high-resolution video data. The quality of this data is ensured by adjusting the camera position to accurately capture the players' movements and the ball's trajectory.

[0617] Step 2:

[0618] The terminal transmits captured image information to the server. The input is video data obtained from a video recording device, and the output is data transfer to the server. In this process, data compression technology is used to efficiently utilize network bandwidth.

[0619] Step 3:

[0620] The server inputs the received image information into a deep learning model. The input is video data sent from the terminal, and the output is the analysis results of the player's motion and the ball's movement. As part of the data processing, the model analyzes the image frame by frame and extracts technical elements such as the motion trajectory and velocity of each object.

[0621] Step 4:

[0622] The server evaluates the player's technical characteristics based on the analysis data. The input is the analysis results obtained in step 3, and the output is the evaluated technical characteristics as well as identification data of strengths and weaknesses. This makes it possible to evaluate shot accuracy and performance.

[0623] Step 5:

[0624] The server generates improvement suggestions and shortcuts in natural language based on the evaluated technical features. The input is evaluation data for the technical features, and the output is feedback messages in natural language. A generative AI model is used, which creates sentences that are easy for humans to understand.

[0625] Step 6:

[0626] The device displays and provides generated feedback to the user via audio. Input is feedback messages sent from the server, and output is a visual and audible presentation of specific advice to the user. Information is delivered to the user using the screen and speaker.

[0627] Step 7:

[0628] The emotion engine analyzes the user's facial expressions and voice tone to recognize their emotional state. The input is image and audio data obtained from the user, and the output is the result of the emotional state analysis. Facial recognition and voice analysis algorithms are used for data processing.

[0629] Step 8:

[0630] The server adjusts the feedback content according to the user's emotional state to maintain or improve their motivation. The input consists of emotional state analysis data and feedback data generated by the server, while the output is the adjusted feedback message. This enables personalized responses that take the user's emotions into consideration.

[0631] Step 9:

[0632] The server stores all gameplay data and emotion recognition results in an information storage medium. Input is general data, and output is historical information stored in a database. This allows for tracking and analysis of players' long-term progress.

[0633] (Application Example 2)

[0634] Next, we will explain application example 2. In the following explanation, the data processing device 12 will be referred to as the "server," and the headset-type terminal 314 will be referred to as the "terminal."

[0635] For players who want to improve their skills in billiards and other sports, there is a need to provide not only technical feedback but also emotional support to boost their motivation and acquire skills more effectively. However, conventional technical support systems lack mechanisms to adjust feedback according to the player's emotional state, which hinders skill improvement.

[0636] The specific processing performed by the specific processing unit 290 of the data processing device 12 in Application Example 2 is realized by the following means.

[0637] In this invention, the server includes means for receiving image information captured without delay from a video recording device and detecting the user's actions and the movement of objects using a machine learning model for analyzing the image information; means for evaluating the technical characteristics based on the user's actions and identifying actions in which the user is good and bad; and means for analyzing the user's emotional state using an emotion analysis engine and adjusting the feedback content based on that state. This makes it possible to improve both the technical and emotional aspects of the user by providing technical feedback along with support tailored to their individual emotions.

[0638] A "video recording device" is a device for acquiring image information in real time without any time lag.

[0639] A "machine learning model" is an algorithm that analyzes data, recognizes patterns, and predicts the future.

[0640] "Users" refers to individuals or organizations that use the system for the purpose of improving their skills.

[0641] "Action and motion of objects" refers to physical movements, including the position, posture, and velocity of the user or object.

[0642] "Technical characteristics" refer to the unique technical operating patterns and characteristics of the user or object.

[0643] An "emotion analysis engine" is an analytical device and program that determines a user's emotional state from their facial expressions and tone of voice.

[0644] "Feedback" refers to information provided to users, either technical or emotional, to encourage improvement and support.

[0645] As an example of the application of this invention, the details of a system for supporting the skill improvement of users in a billiard hall are shown below.

[0646] First, a video recording device is used to capture the user's movements and the movement of the billiard ball in real time. The terminal sends this video data to a server. On the server, a machine learning model is applied to analyze the video data, and the user's shots and the movement of the ball are analyzed to extract technical features.

[0647] Next, the server identifies the user's strengths and weaknesses based on the extracted technical characteristics and generates feedback in natural language based on the results. This feedback is provided to the user visually and audibly via a visual display device.

[0648] Furthermore, an emotion analysis engine is used to analyze the user's facial expressions and tone of voice to recognize their emotional state. Based on this information, the content of the feedback is adjusted according to the user's emotional state to provide optimal support.

[0649] For example, if a user repeatedly misses a particular shot, the server uses an emotion analysis engine to recognize the user's frustration, suggest shots that the user excels at, and provide an encouraging message. In this process, a generative AI model is used for natural language processing, and a prompt message such as, "Please provide the latest billiard shot analysis and feedback. Please also consider the player's emotional data and create an encouraging message," is used.

[0650] In this way, users can efficiently improve their billiards skills while receiving support from both a technical and emotional perspective.

[0651] The flow of a specific process in Application Example 2 will be explained using Figure 14.

[0652] Step 1:

[0653] The terminal captures the user's actions and the movement of the billiard ball in real time via a video recording device and transmits the video data to the server. The input is video data of the user's actions and the ball's movement, and the output is the data captured without any time lag and sent to the server.

[0654] Step 2:

[0655] The server uses a machine learning model to analyze the received video data. The input is video data, and the output is the extraction of technical features related to the user's actions and the ball's movement. This allows the data to be analyzed through the model, identifying specific patterns and trends.

[0656] Step 3:

[0657] The server evaluates the acquired technical characteristics and identifies the user's strengths and weaknesses in shots. The input is data on technical characteristics, and the output is information indicating strengths and weaknesses in shots. Based on this, the server prepares the basic data for generating feedback.

[0658] Step 4:

[0659] Using a generative AI model, the server generates suggestions for improvement and shot strategies for the user in natural language. The input is information about a specific shot, and the output is a detailed feedback message. The server uses prompts to construct the feedback naturally.

[0660] Step 5:

[0661] The server activates an emotion analysis engine to recognize the user's emotional state from their facial expressions and voice data. The input is the user's facial expressions and voice tone information, and the output is an evaluation of their emotional state. This data allows the server to understand how the user is feeling.

[0662] Step 6:

[0663] The server adjusts the feedback content to reflect the emotional state and delivers it to the user visually and audibly via a visual display device. Input is an assessment of the emotional state and pre-generated feedback, while output is the adjusted visual and audible feedback. The goal is to deliver the most appropriate feedback to the user.

[0664] Step 7:

[0665] The server stores technical characteristics and emotional state data in an information aggregation device to manage user progress over the long term. The input is the entire analyzed data, and the output is the stored historical data. This ensures that the history for later use is always up-to-date.

[0666] The specific processing unit 290 transmits the result of the specific processing to the headset terminal 314. In the headset terminal 314, the control unit 46A causes the speaker 240 and display 343 to output the result of the specific processing. The microphone 238 acquires audio indicating user input for the result of the specific processing. The control unit 46A transmits the audio data indicating user input acquired by the microphone 238 to the data processing unit 12. In the data processing unit 12, the specific processing unit 290 acquires the audio data.

[0667] Data generation model 58 is a type of so-called generative AI (Artificial Intelligence). One example of data generation model 58 is ChatGPT (Internet Search).<URL: https: / / openai.com / blog / chatgpt> ), Gemini (Internet search) <url: https: gemini.google.com ?hl="ja">Examples of generative AI include the following. The data generation model 58 is obtained by performing deep learning on a neural network. The data generation model 58 is input with prompts containing instructions, and with inference data such as audio data representing speech, text data representing text, and image data representing images. The data generation model 58 infers from the input inference data according to the instructions indicated by the prompts, and outputs the inference results in data formats such as audio data and text data. Here, inference refers to, for example, analysis, classification, prediction, and / or summarization.

[0668] In the above embodiment, an example was given in which specific processing is performed by the data processing device 12, but the technology of this disclosure is not limited thereto, and specific processing may also be performed by the headset terminal 314.

[0669] [Fourth Embodiment]

[0670] Figure 7 shows an example of the configuration of the data processing system 410 according to the fourth embodiment.

[0671] As shown in Figure 7, the data processing system 410 includes a data processing device 12 and a robot 414. An example of the data processing device 12 is a server.

[0672] The data processing device 12 comprises a computer 22, a database 24, and a communication interface 26. The computer 22 is an example of a "computer" related to the technology of this disclosure. The computer 22 comprises a processor 28, RAM 30, and storage 32. The processor 28, RAM 30, and storage 32 are connected to a bus 34. The database 24 and the communication interface 26 are also connected to the bus 34. The communication interface 26 is connected to a network 54. An example of the network 54 is a WAN (Wide Area Network) and / or a LAN (Local Area Network).

[0673] The robot 414 includes a computer 36, a microphone 238, a speaker 240, a camera 42, a communication interface 44, and a controlled object 443. The computer 36 includes a processor 46, RAM 48, and storage 50. The processor 46, RAM 48, and storage 50 are connected to a bus 52. The microphone 238, speaker 240, camera 42, and controlled object 443 are also connected to the bus 52.

[0674] The microphone 238 receives voice signals from the user 20 and receives instructions from the user 20. The microphone 238 captures the voice signals from the user 20, converts the captured voice into audio data, and outputs it to the processor 46. The speaker 240 outputs audio according to the instructions from the processor 46.

[0675] Camera 42 is a small digital camera equipped with an optical system including a lens, aperture, and shutter, and an image sensor such as a CMOS (Complementary Metal-Oxide-Semiconductor) image sensor or a CCD (Charge Coupled Device) image sensor, and captures images of the area around the user 20 (for example, an imaging range defined by a field of view equivalent to the width of a typical healthy person's field of vision).

[0676] Communication interface 44 is connected to network 54. Communication interfaces 44 and 26 are responsible for the exchange of various information between processor 46 and processor 28 via network 54. The exchange of various information between processor 46 and processor 28 using communication interfaces 44 and 26 is performed in a secure manner.

[0677] The controlled object 443 includes a display device, LEDs in the eyes, and motors that drive the arms, hands, and feet. The posture and gestures of the robot 414 are controlled by controlling the motors of the arms, hands, and feet. Some of the robot 414's emotions can be expressed by controlling these motors. Furthermore, the robot 414's facial expressions can also be expressed by controlling the illumination state of the LEDs in its eyes.

[0678] Figure 8 shows an example of the main functions of the data processing device 12 and the robot 414. As shown in Figure 8, the data processing device 12 performs specific processing using the processor 28. The storage 32 stores the specific processing program 56.

[0679] The specific processing program 56 is an example of a "program" relating to the technology of this disclosure. The processor 28 reads the specific processing program 56 from the storage 32 and executes the read specific processing program 56 on the RAM 30. The specific processing is realized by the processor 28 operating as a specific processing unit 290 in accordance with the specific processing program 56 executed on the RAM 30.

[0680] The storage 32 stores the data generation model 58 and the emotion identification model 59. The data generation model 58 and the emotion identification model 59 are used by the identification processing unit 290.

[0681] In robot 414, the processor 46 performs the reception output processing. The storage 50 stores the reception output program 60. The processor 46 reads the reception output program 60 from the storage 50 and executes the read reception output program 60 on the RAM 48. The reception output processing is realized by the processor 46 operating as a control unit 46A according to the reception output program 60 executed on the RAM 48.

[0682] Next, the specific processing performed by the specific processing unit 290 of the data processing device 12 will be described. In the following description, the data processing device 12 will be referred to as the "server" and the robot 414 as the "terminal".

[0683] This invention is a system using a video camera device and a server to help billiard players improve their skills. The role and operation of each component are described below.

[0684] Acquisition of video data

[0685] The device uses a video camera to capture real-time footage of the billiards game. This ensures that shots and ball movements are accurately recorded.

[0686] Analysis of real-time video

[0687] The server receives video data sent from the terminal and uses a deep learning model to analyze the player's movements and the ball's motion. The model processes each captured frame and extracts data for analyzing techniques such as shot speed, angle, and ball rotation.

[0688] Player skill evaluation

[0689] The server evaluates the player's technical characteristics based on the analyzed data. The computer automatically identifies the player's strong and weak shots, and pinpoints the success rate of those shots and areas for improvement.

[0690] Generating and providing feedback

[0691] Based on the results of the technical evaluation, the server generates improvement suggestions and strategic advice for the player in natural language. The generated feedback is sent to the device in real time. The device then communicates this to the user through visual displays or audio notifications.

[0692] Data storage and practice plan generation

[0693] The server archives data from each play session in a database. This saved data is used by players to review their gameplay at a later date. The server also generates personalized practice plans based on the accumulated performance data, helping users improve their skills more effectively.

[0694] For example, if a player repeatedly misses difficult shots, the server analyzes past success data and suggests alternative shots that have a higher success rate for that player. This provides a sense of accomplishment and helps prevent a decline in motivation.

[0695] The following describes the processing flow.

[0696] Step 1:

[0697] The terminal captures real-time video from the billiard table via a video camera device and sends the data to the server. The terminal continues to acquire video at the appropriate resolution and frame rate.

[0698] Step 2:

[0699] The server acquires the received real-time video data and begins analysis frame by frame. Using a deep learning model, it detects the player's movements and the ball's motion. Specifically, it identifies the angle and force of the player's shot, as well as the ball's collision point and trajectory.

[0700] Step 3:

[0701] The server evaluates the player's shots based on the analyzed data. By comparing technical data, it identifies the player's strengths and weaknesses in shots. By comparing this data with past play data, it analyzes the player's consistency and areas for improvement.

[0702] Step 4:

[0703] The server generates feedback for the player based on the evaluation results. Using natural language processing technology, it provides intuitive and effective improvement advice and shot strategies. The feedback is customized according to the player's skill level and the situation.

[0704] Step 5:

[0705] The server sends the generated feedback to the terminal. The terminal then communicates this to the user either by displaying it on the screen or announcing it through an audio device. The user receives real-time feedback and can immediately use it to improve their gameplay.

[0706] Step 6:

[0707] The server stores data from each play session in a database. This ensures that the underlying data supports the long-term improvement of players' skills. The data is used to generate individual practice plans and is utilized when users plan their future gameplay.

[0708] (Example 1)

[0709] Next, we will describe Example 1. In the following description, the data processing device 12 will be referred to as the "server" and the robot 414 as the "terminal".

[0710] Conventional skill improvement support systems have struggled to analyze user actions in real time and provide immediate improvement suggestions based on individual skill characteristics. Furthermore, they lacked the ability to automatically generate practice plans based on past performance data, preventing users from efficiently improving their skills. This created a need for an integrated skill improvement support system that includes rapid and accurate feedback.

[0711] The identification process performed by the identification processing unit 290 of the data processing device 12 in Example 1 is realized by the following means.

[0712] In this invention, the server includes means for acquiring real-time visual information from a visual sensor and detecting user operations and object movements using a machine learning model for analyzing the visual information; means for evaluating technical characteristics based on the user's operations and identifying strengths and operations that need improvement; and means for generating and providing information to the user in natural language regarding areas for improvement and operation plans based on the technical characteristics. This enables the user to receive detailed feedback in real time and obtain accurate advice and practice plans for improving their skills.

[0713] A "visual sensor" is a device that captures the movement of an object and converts the visual information into digital data.

[0714] "Visual information" refers to digital data of images and videos acquired from visual sensors, and this data is the subject of analysis.

[0715] A "machine learning model" is a collection of algorithms that analyze large amounts of data and automatically learn specific patterns and features.

[0716] "Users" refer to individuals or organizations that wish to use the system to improve specific technologies.

[0717] "Operation" refers to physical or digital actions performed by the user, including actions detected by visual sensors.

[0718] "Object movement" refers to changes in the position of physical objects detected within visual information.

[0719] "Technical characteristics" refer to the specific features of performance and capabilities related to user operation.

[0720] "Operations that need improvement" are specific actions or procedures that users should focus on modifying or practicing in order to improve their skills.

[0721] "Natural language" refers to the form of language that humans use in everyday life, and is used when computers generate and provide information to users.

[0722] "Information provision" refers to the act of conveying analyzed data and suggestions to users, and this can be done visually or audibly.

[0723] This invention is a system designed to support the improvement of users' skills and comprises a visual sensor, a server, and a terminal. Specifically, a general-purpose video acquisition device is used as the visual sensor to capture the actions of physical sports such as billiards in real time. This video data is transmitted to the server via the terminal.

[0724] The server analyzes the received visual information using a machine learning model. This process utilizes commonly used deep learning frameworks (e.g., TensorFlow) to evaluate the user's actions and object movements frame by frame. The server then extracts numerical data for the speed, angle, and ball rotation of each shot. Based on this data, the server assesses the user's technical characteristics and identifies their strengths and areas for improvement.

[0725] The server uses a generative AI model (e.g., GPT-3) based on the evaluation results to generate improvement suggestions and technical advice for the user in natural language. An example of a prompt used for this purpose would be, "Generate improvement advice based on this player's recent shot data."

[0726] The generated feedback is sent to the device in real time. The device uses visual display devices and audio output devices to convey this to the user. This allows the user to continuously review their gameplay and effectively improve their skills.

[0727] Furthermore, the server stores visual information and analysis results in memory. This data is used as a history for users to review past gameplay data later, and is also utilized to create personalized practice plans for continuous skill improvement. Based on the accumulated data, the server performs statistical analysis and provides practice plans optimized for the user. Through this entire process, the system can provide comprehensive support for users to improve their skills.

[0728] The flow of the specific processing in Example 1 will be explained using Figure 11.

[0729] Step 1:

[0730] The device uses a visual sensor to capture the user's athletic performance in real time. The input is the physical athletic scene, and the output is digital video data. This ensures that the user's movements and the movement of objects are always recorded in their most up-to-date state. Specifically, the sensor continuously captures images at regular intervals according to the frame rate.

[0731] Step 2:

[0732] The terminal sends the captured video data to the server. The input is the video data acquired in step 1, and the output is streaming data for the server to receive. The data is transmitted in real time, and compression technology is used to reduce network load while maintaining data integrity.

[0733] Step 3:

[0734] The server analyzes the received video data using a machine learning model. The input is video data sent from the terminal, and the output is numerical data related to the user's movements and the position of objects. This analysis involves feature point extraction and motion vector calculation for each frame. Specifically, the server runs a deep learning framework and applies algorithms to extract important physical information from each frame.

[0735] Step 4:

[0736] The server evaluates the user's technical characteristics based on the analysis results. The input is the output data from step 3, and the output is scoring data related to the technical evaluation. The server automatically identifies the user's strengths and areas for improvement through a script and uses a technical evaluation model. Specifically, it applies statistical analysis and scores each operation based on evaluation criteria.

[0737] Step 5:

[0738] The server uses a generative AI model to generate improvement suggestions in natural language based on technical evaluations. The input is the technical evaluation data from step 4, and the output is advice in natural language provided to the user. Specifically, prompt sentences are used and applied to the generative AI model to form concrete advice that helps improve the user's skills.

[0739] Step 6:

[0740] The terminal receives advice from the server and presents it to the user in real time. Input is feedback data from the server, and output is advice expressed visually or audibly. The terminal uses a display and speakers to inform the user of specific improvement measures. This allows the user to directly and immediately see the improvement measures.

[0741] Step 7:

[0742] The server stores video data and technical evaluations in its memory. The input is past pre-session data, and the output is historical data stored in the memory. This storage process is performed using database management functions. This enables long-term growth analysis of users.

[0743] Step 8:

[0744] The server generates individual practice plans based on stored data. The input is performance data stored in memory, and the output is a practice plan optimized for the user. The server combines statistical analysis and machine learning algorithms to construct effective practice suggestions.

[0745] (Application Example 1)

[0746] Next, we will explain Application Example 1. In the following explanation, the data processing device 12 will be referred to as the "server" and the robot 414 as the "terminal".

[0747] For billiards players to efficiently improve their skills, real-time motion analysis and feedback are essential. Furthermore, being able to receive this information instantly in a real playing environment greatly enhances the effectiveness of practice. However, conventional technology has not been able to create a system that accurately analyzes a player's movements and provides real-time, visually and audibly effective feedback.

[0748] The specific processing performed by the specific processing unit 290 of the data processing device 12 in Application Example 1 is realized by the following means.

[0749] In this invention, the server includes means for receiving real-time video information from a video camera device and detecting the player's body movements and the movement of a sphere using a machine learning model for analyzing the video information; means for generating and providing feedback to the player in natural language regarding areas for improvement and shot strategies based on these technical characteristics; and means for overlaying the feedback onto the real environment using smart glasses. This enables the player to obtain detailed and practical analytical information and feedback in real time, allowing for more effective technical improvement.

[0750] A "video camera device" is a device used to capture video information in real time.

[0751] "Real-time video information" refers to information that provides video data occurring at the present moment immediately.

[0752] A "machine learning model" is a model that uses algorithms to automatically learn patterns and rules from data.

[0753] "Physical movement" refers to the movement and posture of the player's body, and analyzing this is a means of extracting technical characteristics.

[0754] "Sphere motion" refers to the kinetic characteristics of a billiard ball, specifically its direction of movement, speed, and rotation.

[0755] "Technical characteristics" refer to a set of parameters used to evaluate the accuracy and movement characteristics of a player's shots.

[0756] "Natural language" refers to the linguistic expressions that people use on a daily basis, and is a means of conveying information to players in an easily understandable way as feedback.

[0757] "Smart glasses" are wearable devices that have the function of displaying visual information as augmented reality.

[0758] "Overlaying information onto the real world" is a technology that displays additional information superimposed on the visual information of the real world.

[0759] This invention provides a system to support the skill improvement of billiard players. This system operates in combination with a video camera device, smart glasses, and a server.

[0760] The server receives real-time video data from the video camera device. This data is processed using machine learning models to analyze in detail the player's body movements and the ball's motion during the shot. Specifically, it utilizes libraries such as TensorFlow in Python to recognize the ball's speed, rotation, and shot angle for each video frame.

[0761] Based on the analyzed data, the system evaluates the player's technical characteristics and identifies their strengths and weaknesses in shots. The server then generates natural language feedback, providing players with areas for improvement and strategic advice. This feedback is overlaid onto the real-world gameplay environment via smart glasses. The system is also designed to allow users to access information in real time using devices such as Google Glass.

[0762] Accumulated video and technical data are archived in a database on the server. This allows users to analyze past gameplay and track their progress. Player performance data is managed using database software such as Sarvapi and PostgreSQL. Furthermore, individual practice plans are generated, and practice menus tailored to each player are suggested.

[0763] For example, if a player repeatedly misses difficult shots, this system analyzes past success data and suggests alternative shots with a higher success rate. In this way, it helps the player gain a sense of accomplishment. An example of a prompt to input into the generative AI model might be, "Please tell me a temporary strategy for successfully executing a great mid-range shot while playing billiards."

[0764] The flow of a specific process in Application Example 1 will be explained using Figure 12.

[0765] Step 1:

[0766] The terminal captures the billiards game using a video camera device. The input is real-time video data, and the output is the same video data sent directly to the server. In this process, the camera operates at a high frame rate to capture the fast movement of the ball without delay.

[0767] Step 2:

[0768] The server receives video data, which is then used as input for a machine learning model. The input consists of captured image frames, and the output includes the ball's position and motion vector, as well as the player's movement pattern. TensorFlow is used to analyze the ball's speed and spin, and the angle of the shot, generating this numerical data.

[0769] Step 3:

[0770] The server evaluates the player's technical characteristics based on the generated analysis data. The input is the analysis data obtained in step 2, and the output includes strong and weak shots, success rates, etc. Based on this evaluation information, the server inputs suggestions for improvement and strategies for the next shot as prompts into the generating AI model.

[0771] Step 4:

[0772] The server generates feedback from the generative AI model in natural language and sends it to the smart glasses. The input is the suggestion information from the generative AI model, and the output is a natural language feedback message. In this step, the suggestions are summarized in a way that is easy for the user to understand.

[0773] Step 5:

[0774] The user's smart glasses receive feedback from the server and overlay it onto the real environment. The input is the feedback message from the server, and the output is the information visually displayed through the glasses. This allows the user to receive improvement advice in real time.

[0775] Step 6:

[0776] The server archives video data and analysis results from play sessions into a database. Inputs are the original video data and technical evaluation data, while outputs are accumulated historical data. This data serves as a basis for players to later review their progress.

[0777] Step 7:

[0778] The server uses accumulated historical data to generate individual practice plans. The input is the player's past performance data, and the output is the suggested practice plan. This step is performed automatically to support the player's skill improvement.

[0779] Furthermore, an emotion engine that estimates the user's emotions may be incorporated. That is, the identification processing unit 290 may use the emotion identification model 59 to estimate the user's emotions and perform identification processing using the user's emotions.

[0780] This invention combines a system designed to improve the skills of billiard players with an emotion engine that recognizes the user's emotions and provides appropriate feedback. Its configuration and operation are described in detail below.

[0781] Acquisition and analysis of video data

[0782] The terminal uses a video camera device to capture real-time video of a billiards game and sends it to the server. The server uses a deep learning model to analyze the player's movements and the ball's motion. This provides data for evaluating the technical aspects of the shots.

[0783] Technical evaluation and feedback generation

[0784] The server evaluates the player's technical characteristics based on the analyzed data. It identifies strengths and weaknesses in shots and generates improvement suggestions and shot strategies in natural language. The generated feedback is provided to the user visually and audibly through the terminal.

[0785] Emotional recognition and feedback regulation

[0786] The emotion engine analyzes the user's facial expressions and tone of voice to recognize their emotional state. The server then adjusts the feedback based on this emotional data, providing optimal feedback tailored to the user's motivation level. For example, if a player is feeling frustrated, the emotion engine generates positive reinforcement messages to encourage further attempts.

[0787] Data storage and practice plan generation

[0788] The server stores all play data and emotion recognition results in a database, enabling the visualization of long-term progress. This allows for the creation of personalized practice plans for each player, supporting skill improvement. Users can review their play history and emotion change trends and utilize their individual practice plans to efficiently improve their skills.

[0789] For example, if a player repeatedly misses a particular shot, the server uses an emotion engine to recognize the player's frustration and suggests shots they are good at while also providing encouraging messages. This helps the player regain confidence and continue practicing.

[0790] The following describes the processing flow.

[0791] Step 1:

[0792] The terminal uses a video camera device to capture video footage of the billiards game. The video data is transmitted to the server in real time. Camera settings are adjusted to ensure high-quality video and enable accurate motion analysis.

[0793] Step 2:

[0794] The server inputs the received video data into a deep learning model to analyze the player's movements and the ball's motion. The analysis includes shot speed, angle, and spin rate, and the player's current technical characteristics are extracted as numerical data.

[0795] Step 3:

[0796] The server evaluates the player's skills based on the analysis results, identifying their strengths and weaknesses in shots. By comparing this with past gameplay data, it highlights areas for improvement to enhance their skills.

[0797] Step 4:

[0798] The emotion engine recognizes the user's emotional state from their facial expressions, tone of voice, and body movements. The server then incorporates this emotional information into the feedback. For example, if the player is focused, it provides detailed technical analysis; on the other hand, if they are feeling frustrated, it generates encouraging messages.

[0799] Step 5:

[0800] The server generates feedback for the user using natural language processing, based on technical evaluation results and emotional states. The feedback is created as a customized message and includes strategic advice and suggestions for improvement tailored to the player's current situation.

[0801] Step 6:

[0802] The device presents the generated feedback to the user visually and audibly. Users can receive real-time advice and gain a deeper understanding of their own shots.

[0803] Step 7:

[0804] The server stores technical and emotional data for each session in a database. This allows users to have a history that they can refer to in later sessions and use to create practice plans.

[0805] Step 8:

[0806] Based on accumulated data, the server analyzes players' technical and emotional tendencies and generates personalized practice plans. These plans provide users with specific steps to achieve their goals and support long-term skill improvement.

[0807] (Example 2)

[0808] Next, we will describe Example 2. In the following description, the data processing device 12 will be referred to as the "server" and the robot 414 as the "terminal".

[0809] It is difficult for billiards players to objectively evaluate their own skills and clearly identify areas for improvement. Furthermore, providing feedback that takes into account the player's emotional state is needed to boost their motivation for practice. A system is required to address these challenges and provide effective practice plans tailored to each individual.

[0810] The identification process performed by the identification processing unit 290 of the data processing device 12 in Example 2 is realized by the following means.

[0811] In this invention, the server includes means for receiving real-time image information from a video recording device and confirming the player's movements and the ball's movement using a deep learning model for analyzing the image information; means for recording the image information and technical elements in an information storage medium and managing a history for analyzing the player's progress; and means for analyzing the user's emotional state, adjusting the response content based on the emotional information, and providing an optimal response. This enables billiard players to effectively improve their skills and receive feedback tailored to their emotional state.

[0812] "Video recording equipment" is a general term for devices used to record video and audio and to later play them back or transmit them.

[0813] "Real-time image information" refers to data used to acquire and process images from the real world in real time.

[0814] A "deep learning model" is a type of algorithm that uses a multi-layer neuron network to learn the features of data and perform analysis and prediction.

[0815] "Player movements" in billiards refers to all physical actions performed by a player, including shots and movement.

[0816] "Ball movement" in billiards refers to the motion of the ball as it rolls, spins, or collides in response to the force exerted by the player.

[0817] "Technical elements" refer to skill-based characteristics such as force and precision, which are evaluated as parameters of shots and play in billiards.

[0818] "Information storage medium" is a general term for hardware or digital storage used to record, store, and manage data.

[0819] "Emotional state" refers to the psychological and emotional state of a user, measured based on their behavior, facial expressions, voice, etc.

[0820] "Response content" refers to all feedback messages provided to the user, such as criticisms, instructions, or encouragement.

[0821] This invention is a system that operates by combining multiple components to support the skill improvement of billiard players. Specific embodiments of this system are shown below.

[0822] The terminal uses a video recording device to capture real-time image information of a billiards game and transmits the data to a server. The terminal is positioned optimally to smoothly capture high-resolution video. The video data is efficiently transmitted to the server via the network.

[0823] The server inputs the received real-time image information into a deep learning model. Specifically, it uses software frameworks such as TensorFlow and PyTorch to analyze the player's motion and the ball's movement. The server then extracts technical elements such as shot speed, angle, and contact point, thereby evaluating the player's skill.

[0824] Based on the evaluated technical elements, the server generates suggestions for improvement and shortcuts for the user in natural language. This process utilizes a natural language generation AI model, and the generated feedback is provided to the user visually and audibly through the device. The device presents the feedback using both screen display and audio output.

[0825] Furthermore, the emotion engine analyzes the user's emotional state. The device obtains emotional data by recording the user's facial expressions and collecting their voice tone. The server uses this data to understand the user's emotional state and adjust the feedback accordingly. For example, if the user is irritated, a positive message may be added to boost their motivation.

[0826] Finally, the server saves all play data and emotion recognition results to a data storage medium. This allows users to refer to past performance data and emotion fluctuations to generate individually customized practice plans.

[0827] For example, if a player repeatedly misses a particular shot, the system recognizes the player's frustration and offers words of encouragement while suggesting alternative shots the player is good at. This allows the player to move on to the next step with confidence.

[0828] An example of a prompt is: "Analyze the player's movements and the ball's trajectory during a billiards match, evaluate their technical characteristics, and determine their strengths and weaknesses. Furthermore, create appropriate feedback in natural language, taking into account the player's emotional state."

[0829] The flow of the specific processing in Example 2 will be explained using Figure 13.

[0830] Step 1:

[0831] The terminal activates a video recording device and captures real-time image information of a billiards game. The input is a real billiards game in action, and the output is high-resolution video data. The quality of this data is ensured by adjusting the camera position to accurately capture the players' movements and the ball's trajectory.

[0832] Step 2:

[0833] The terminal transmits captured image information to the server. The input is video data obtained from a video recording device, and the output is data transfer to the server. In this process, data compression technology is used to efficiently utilize network bandwidth.

[0834] Step 3:

[0835] The server inputs the received image information into a deep learning model. The input is video data sent from the terminal, and the output is the analysis results of the player's motion and the ball's movement. As part of the data processing, the model analyzes the image frame by frame and extracts technical elements such as the motion trajectory and velocity of each object.

[0836] Step 4:

[0837] The server evaluates the player's technical characteristics based on the analysis data. The input is the analysis results obtained in step 3, and the output is the evaluated technical characteristics as well as identification data of strengths and weaknesses. This makes it possible to evaluate shot accuracy and performance.

[0838] Step 5:

[0839] The server generates improvement suggestions and shortcuts in natural language based on the evaluated technical features. The input is evaluation data for the technical features, and the output is feedback messages in natural language. A generative AI model is used, which creates sentences that are easy for humans to understand.

[0840] Step 6:

[0841] The device displays and provides generated feedback to the user via audio. Input is feedback messages sent from the server, and output is a visual and audible presentation of specific advice to the user. Information is delivered to the user using the screen and speaker.

[0842] Step 7:

[0843] The emotion engine analyzes the user's facial expressions and voice tone to recognize their emotional state. The input is image and audio data obtained from the user, and the output is the result of the emotional state analysis. Facial recognition and voice analysis algorithms are used for data processing.

[0844] Step 8:

[0845] The server adjusts the feedback content according to the user's emotional state to maintain or improve their motivation. The input consists of emotional state analysis data and feedback data generated by the server, while the output is the adjusted feedback message. This enables personalized responses that take the user's emotions into consideration.

[0846] Step 9:

[0847] The server stores all gameplay data and emotion recognition results in an information storage medium. Input is general data, and output is historical information stored in a database. This allows for tracking and analysis of players' long-term progress.

[0848] (Application Example 2)

[0849] Next, we will explain application example 2. In the following explanation, the data processing device 12 will be referred to as the "server" and the robot 414 as the "terminal".

[0850] For players who want to improve their skills in billiards and other sports, there is a need to provide not only technical feedback but also emotional support to boost their motivation and acquire skills more effectively. However, conventional technical support systems lack mechanisms to adjust feedback according to the player's emotional state, which hinders skill improvement.

[0851] The specific processing performed by the specific processing unit 290 of the data processing device 12 in Application Example 2 is realized by the following means.

[0852] In this invention, the server includes means for receiving image information captured without delay from a video recording device and detecting the user's actions and the movement of objects using a machine learning model for analyzing the image information; means for evaluating the technical characteristics based on the user's actions and identifying actions in which the user is good and bad; and means for analyzing the user's emotional state using an emotion analysis engine and adjusting the feedback content based on that state. This makes it possible to improve both the technical and emotional aspects of the user by providing technical feedback along with support tailored to their individual emotions.

[0853] A "video recording device" is a device for acquiring image information in real time without any time lag.

[0854] A "machine learning model" is an algorithm that analyzes data, recognizes patterns, and predicts the future.

[0855] "Users" refers to individuals or organizations that use the system for the purpose of improving their skills.

[0856] "Action and motion of objects" refers to physical movements, including the position, posture, and velocity of the user or object.

[0857] "Technical characteristics" refer to the unique technical operating patterns and characteristics of the user or object.

[0858] An "emotion analysis engine" is an analytical device and program that determines a user's emotional state from their facial expressions and tone of voice.

[0859] "Feedback" refers to information provided to users, either technical or emotional, to encourage improvement and support.

[0860] As an example of the application of this invention, the details of a system for supporting the skill improvement of users in a billiard hall are shown below.

[0861] First, a video recording device is used to capture the user's movements and the movement of the billiard ball in real time. The terminal sends this video data to a server. On the server, a machine learning model is applied to analyze the video data, and the user's shots and the movement of the ball are analyzed to extract technical features.

[0862] Next, the server identifies the user's strengths and weaknesses based on the extracted technical characteristics and generates feedback in natural language based on the results. This feedback is provided to the user visually and audibly via a visual display device.

[0863] Furthermore, an emotion analysis engine is used to analyze the user's facial expressions and tone of voice to recognize their emotional state. Based on this information, the content of the feedback is adjusted according to the user's emotional state to provide optimal support.

[0864] For example, if a user repeatedly misses a particular shot, the server uses an emotion analysis engine to recognize the user's frustration, suggest shots that the user excels at, and provide an encouraging message. In this process, a generative AI model is used for natural language processing, and a prompt message such as, "Please provide the latest billiard shot analysis and feedback. Please also consider the player's emotional data and create an encouraging message," is used.

[0865] In this way, users can efficiently improve their billiards skills while receiving support from both a technical and emotional perspective.

[0866] The flow of a specific process in Application Example 2 will be explained using Figure 14.

[0867] Step 1:

[0868] The terminal captures the user's actions and the movement of the billiard ball in real time via a video recording device and transmits the video data to the server. The input is video data of the user's actions and the ball's movement, and the output is the data captured without any time lag and sent to the server.

[0869] Step 2:

[0870] The server uses a machine learning model to analyze the received video data. The input is video data, and the output is the extraction of technical features related to the user's actions and the ball's movement. This allows the data to be analyzed through the model, identifying specific patterns and trends.

[0871] Step 3:

[0872] The server evaluates the acquired technical characteristics and identifies the user's strengths and weaknesses in shots. The input is data on technical characteristics, and the output is information indicating strengths and weaknesses in shots. Based on this, the server prepares the basic data for generating feedback.

[0873] Step 4:

[0874] Using a generative AI model, the server generates suggestions for improvement and shot strategies for the user in natural language. The input is information about a specific shot, and the output is a detailed feedback message. The server uses prompts to construct the feedback naturally.

[0875] Step 5:

[0876] The server activates an emotion analysis engine to recognize the user's emotional state from their facial expressions and voice data. The input is the user's facial expressions and voice tone information, and the output is an evaluation of their emotional state. This data allows the server to understand how the user is feeling.

[0877] Step 6:

[0878] The server adjusts the feedback content to reflect the emotional state and delivers it to the user visually and audibly via a visual display device. Input is an assessment of the emotional state and pre-generated feedback, while output is the adjusted visual and audible feedback. The goal is to deliver the most appropriate feedback to the user.

[0879] Step 7:

[0880] The server stores technical characteristics and emotional state data in an information aggregation device to manage user progress over the long term. The input is the entire analyzed data, and the output is the stored historical data. This ensures that the history for later use is always up-to-date.

[0881] The specific processing unit 290 transmits the result of the specific processing to the robot 414. In the robot 414, the control unit 46A causes the speaker 240 and the controlled object 443 to output the result of the specific processing. The microphone 238 acquires audio indicating user input for the result of the specific processing. The control unit 46A transmits the audio data indicating user input acquired by the microphone 238 to the data processing unit 12. In the data processing unit 12, the specific processing unit 290 acquires the audio data.

[0882] Data generation model 58 is a type of so-called generative AI (Artificial Intelligence). One example of data generation model 58 is ChatGPT (Internet Search).<URL: https: / / openai.com / blog / chatgpt> ), Gemini (Internet search) <url: https: gemini.google.com ?hl="ja">Examples of generative AI include the following. The data generation model 58 is obtained by performing deep learning on a neural network. The data generation model 58 is input with prompts containing instructions, and with inference data such as audio data representing speech, text data representing text, and image data representing images. The data generation model 58 infers from the input inference data according to the instructions indicated by the prompts, and outputs the inference results in data formats such as audio data and text data. Here, inference refers to, for example, analysis, classification, prediction, and / or summarization.

[0883] In the above embodiment, an example was given in which the specific processing is performed by the data processing device 12, but the technology of this disclosure is not limited thereto, and the specific processing may also be performed by the robot 414.

[0884] Furthermore, the emotion identification model 59, acting as an emotion engine, may determine the user's emotion according to a specific mapping. Specifically, the emotion identification model 59 may determine the user's emotion according to a specific mapping, which is an emotion map (see Figure 9). Similarly, the emotion identification model 59 may also determine the robot's emotion, and the identification processing unit 290 may perform identification processing using the robot's emotion.

[0885] Figure 9 shows an emotion map 400 in which multiple emotions are mapped. In the emotion map 400, emotions are arranged in concentric circles radiating from the center. The closer to the center of the concentric circles, the more primitive the emotions are located. Further out of the concentric circles, emotions representing states and actions arising from mental states are located. Emotion is a concept that includes feelings and mental states. On the left side of the concentric circles, emotions that are generally generated from reactions occurring in the brain are located. On the right side of the concentric circles, emotions that are generally induced by situational judgment are located. Above and below the concentric circles, emotions that are generally generated from reactions occurring in the brain and induced by situational judgment are located. In addition, the emotion of "pleasure" is located on the upper side of the concentric circles, and the emotion of "displeasure" is located on the lower side. Thus, in the emotion map 400, multiple emotions are mapped based on the structure in which emotions arise, and emotions that are likely to occur simultaneously are mapped close together.

[0886] These emotions are distributed at the 3 o'clock position on the Emotion Map 400, and usually fluctuate between feelings of security and anxiety. In the right half of the Emotion Map 400, situational awareness takes precedence over internal feelings, resulting in a calm impression.

[0887] The inside of the Emotion Map 400 represents inner thoughts, while the outside represents actions. Therefore, the further you go from the outside of the Emotion Map 400, the more visible (expressed in actions) your emotions become.

[0888] Here, human emotions are based on various balances, such as posture and blood sugar levels. When these balances deviate from the ideal, it results in discomfort, and when they approach the ideal, it results in pleasure. Similarly, in robots, cars, motorcycles, etc., emotions can be created based on various balances, such as posture and battery level. When these balances deviate from the ideal, it results in discomfort, and when they approach the ideal, it results in pleasure. The emotion map can be generated, for example, based on Dr. Mitsuyoshi's emotion map (Research on a system for analyzing brain physiological signals of speech emotion recognition and emotion, Tokushima University, doctoral dissertation: https: / / ci.nii.ac.jp / naid / 500000375379). The left half of the emotion map contains emotions belonging to a region called "response," where sensation is dominant. The right half of the emotion map contains emotions belonging to a region called "situation," where situational awareness is dominant.

[0889] The emotion map defines two emotions that promote learning. One is the emotion around the middle of the negative "repentance" and "reflection" on the situation side. In other words, it is when the robot experiences negative emotions such as "I never want to feel this way again" or "I don't want to be scolded again." The other is the emotion around the positive "desire" on the reaction side. In other words, it is when the robot has positive feelings such as "I want more" or "I want to know more."

[0890] The emotion identification model 59 inputs user input into a pre-trained neural network, obtains emotion values ​​representing each emotion shown in the emotion map 400, and determines the user's emotion. This neural network is pre-trained based on multiple training data sets, which are combinations of user input and emotion values ​​representing each emotion shown in the emotion map 400. Furthermore, this neural network is trained so that emotions located close together have similar values, as shown in the emotion map 900 in Figure 10. Figure 10 shows an example where multiple emotions such as "reassured," "calm," and "confident" have similar emotion values.

[0891] The above description primarily focuses on the functions of the data processing device 12 in relation to this disclosure. However, the system related to this disclosure is not necessarily implemented on a server. The system related to this disclosure may be implemented as a general information processing system. This disclosure may be implemented, for example, as a software program that runs on a personal computer or as an application that runs on a smartphone. The method related to this disclosure may be provided to users in SaaS (Software as a Service) format.

[0892] In the above embodiment, an example was given in which a specific process is performed by a single computer 22. However, the technology of this disclosure is not limited thereto, and a distributed processing of the specific process may be performed by multiple computers, including computer 22. For example, a data generation model 58 may be provided in an external device of the data processing device 12, and the external device may generate data according to the input data.

[0893] In the above embodiment, an example was given in which the specific processing program 56 is stored in the storage 32, but the technology of this disclosure is not limited thereto. For example, the specific processing program 56 may be stored in a portable, computer-readable, non-temporary storage medium such as a USB (Universal Serial Bus) memory. The specific processing program 56 stored in the non-temporary storage medium is installed in the computer 22 of the data processing device 12. The processor 28 executes specific processing according to the specific processing program 56.

[0894] Alternatively, the specific processing program 56 may be stored in a storage device such as a server connected to the data processing device 12 via the network 54, and the specific processing program 56 may be downloaded and installed on the computer 22 in response to a request from the data processing device 12.

[0895] Furthermore, it is not necessary to store the entirety of the specific processing program 56 in a storage device such as a server connected to the data processing device 12 via the network 54, or to store the entirety of the specific processing program 56 in the storage 32; it is acceptable to store only a portion of the specific processing program 56.

[0896] The following types of processors can be used as hardware resources to perform specific processing. Examples of processors include a CPU, a general-purpose processor that functions as a hardware resource to perform specific processing by executing software, i.e., a program. Other examples of processors include dedicated electrical circuits, such as FPGAs (Field-Programmable Gate Arrays), PLDs (Programmable Logic Devices), or ASICs (Application Specific Integrated Circuits), which have circuit configurations specifically designed to perform specific processing. All of these processors have built-in or connected memory, and all of them perform specific processing by using memory.

[0897] The hardware resource that performs a specific process may consist of one of these various processors, or it may consist of a combination of two or more processors of the same or different types (for example, a combination of multiple FPGAs, or a combination of a CPU and an FPGA). Alternatively, the hardware resource that performs a specific process may consist of a single processor.

[0898] Examples of configurations using a single processor include, firstly, a configuration in which one or more CPUs and software are combined to form a single processor, and this processor functions as a hardware resource that performs a specific process. Secondly, there is a configuration using a processor that realizes the functions of the entire system, including multiple hardware resources that perform a specific process, on a single IC chip, as exemplified by SoCs (System-on-a-chip). In this way, a specific process is realized using one or more of the above types of processors as hardware resources.

[0899] Furthermore, the hardware structure of these various processors can more specifically utilize electrical circuits that combine circuit elements such as semiconductor devices. Also, the specific processing described above is merely an example. Therefore, it goes without saying that unnecessary steps can be deleted, new steps added, or the processing order rearranged, as long as it does not deviate from the main purpose.

[0900] The descriptions and illustrations presented above are detailed explanations of the technical aspects of this disclosure and are merely examples of the technical aspects. For example, the above descriptions of the structure, function, operation, and effect are examples of the structure, function, operation, and effect of the technical aspects of this disclosure. Therefore, it goes without saying that you may delete unnecessary parts, add new elements, or replace elements in the descriptions and illustrations presented above, as long as you do not deviate from the essence of the technical aspects of this disclosure. Furthermore, in order to avoid confusion and facilitate understanding of the technical aspects of this disclosure, explanations of common technical knowledge and the like that do not require special explanation to enable the implementation of the technical aspects of this disclosure have been omitted from the descriptions and illustrations presented above.

[0901] All documents, patent applications, and technical standards described herein are incorporated by reference to the same extent as if each individual document, patent application, and technical standard were specifically and individually noted as being incorporated by reference.

[0902] The following is further disclosed regarding the embodiments described above.

[0903] (Claim 1)

[0904] A means for receiving real-time video data from a video camera device and detecting the player's movements and the ball's motion using a deep learning model for analyzing the video data,

[0905] A means for evaluating the technical characteristics of the player based on their shots and identifying their strengths and weaknesses in shots,

[0906] A means for generating and providing feedback to players on areas for improvement and shot strategies in natural language based on the technical features,

[0907] Means for presenting the feedback visually and audibly on a display device,

[0908] A means for archiving the video data and technical features in a database and managing a history for analyzing the player's progress,

[0909] A means for generating an individual training plan based on the player's performance data,

[0910] A system that includes this.

[0911] (Claim 2)

[0912] The system according to claim 1, further comprising means for suggesting a shot the player is good at in order to promote the player's sense of success when the player has failed a particular shot multiple times in a row.

[0913] (Claim 3)

[0914] The system according to claim 1, further comprising a module for the user to later review the feedback and performance data using a graphical user interface.

[0915] "Example 1"

[0916] (Claim 1)

[0917] A means for acquiring real-time visual information from a visual sensor and detecting user operations and object movement using a machine learning model for analyzing said visual information,

[0918] A means for evaluating the technical characteristics based on the user's operations and identifying strengths and operations that need improvement,

[0919] A means for generating improvements and operational plans in natural language and providing information to users based on the technical characteristics,

[0920] Means for presenting the information visually and audibly using an output device,

[0921] Means for storing the visual information and technical characteristics in a memory area and managing a history for analyzing the user's growth,

[0922] A means for creating an individual practice plan based on the user's performance data,

[0923] A system that includes this.

[0924] (Claim 2)

[0925] The system according to claim 1, further comprising means for recommending operations in which the user has strengths in order to promote the user's successful experience when the user fails to perform a particular operation consecutively.

[0926] (Claim 3)

[0927] The system according to claim 1, further comprising a module for users to later review the information and performance data using a visual user interface.

[0928] "Application Example 1"

[0929] (Claim 1)

[0930] A means for receiving real-time video information from a video camera device and detecting the player's body movements and the movement of a sphere using a machine learning model for analyzing the video information,

[0931] A means for evaluating the technical characteristics of the player based on their shots and identifying their strengths and weaknesses in shots,

[0932] A means for generating and providing feedback to players on areas for improvement and shot strategies in natural language based on the technical features,

[0933] Means for presenting the feedback visually and audibly on a display device,

[0934] A means for archiving the video information and technical features on an information recording medium and managing a history for analyzing the player's progress,

[0935] Means for generating an individual training plan based on the player's performance information,

[0936] A means for displaying the feedback as an overlay in the real environment using smart glasses,

[0937] A system that includes this.

[0938] (Claim 2)

[0939] The system according to claim 1, further comprising means for suggesting a shot the player is good at in order to promote the player's sense of success when the player has failed a particular shot multiple times in a row.

[0940] (Claim 3)

[0941] The system according to claim 1, further comprising a module for the user to later review the feedback and performance information using a user interface.

[0942] "Example 2 of combining an emotion engine"

[0943] (Claim 1)

[0944] A means for receiving real-time image information from a video recording device and using a deep learning model to analyze the image information to confirm the player's movement and the ball's movement,

[0945] A means for evaluating the technical elements based on the player's actions and identifying skilled and unskilled actions,

[0946] A means for generating improvements and operational strategies in natural language and providing responses to the user based on the said technical elements,

[0947] Means for presenting the reaction visually and audibly using a display device,

[0948] A means for recording the image information and technical elements on an information storage medium and managing a history for analyzing the player's progress,

[0949] Means for generating individual practice plans based on the player's performance data,

[0950] A means for analyzing the user's emotional state, adjusting the response based on the emotional information, and providing the optimal response,

[0951] A system that includes this.

[0952] (Claim 2)

[0953] The system according to claim 1, further comprising means for suggesting skilled operations to facilitate a successful experience for a player if the player repeatedly fails to perform a particular operation.

[0954] (Claim 3)

[0955] The system according to claim 1, comprising a unit for the user to later view the reaction and execution data using a graphic user connection surface.

[0956] "Application example 2 when combining with an emotional engine"

[0957] (Claim 1)

[0958] A means for receiving image information captured without delay from a video recording device and detecting the user's movements and the motion of objects using a machine learning model for analyzing said image information,

[0959] A means for evaluating the technical characteristics based on the user's behavior and identifying the user's strengths and weaknesses in performing certain actions,

[0960] A means for generating improvements and action strategies in natural language based on the technical features and providing feedback to the user,

[0961] Means for presenting the feedback visually and audibly using a visual display device,

[0962] A means for archiving the image information and technical features in an information accumulating device and managing a history for analyzing the user's progress,

[0963] A means for generating an individual training plan based on the user's performance data,

[0964] A means of incorporating an emotion analysis engine to analyze the user's emotional state and adjust the feedback content based on that state,

[0965] A system that includes this.

[0966] (Claim 2)

[0967] The system according to claim 1, further comprising means for suggesting actions the user is good at in order to promote the user's successful experiences when the user repeatedly fails at a particular action.

[0968] (Claim 3)

[0969] The system according to claim 1, further comprising a module for the user to visually review the feedback and performance data at a later date. [Explanation of Symbols]

[0970] 10, 210, 310, 410 Data Processing Systems 12 Data Processing Devices 14 Smart Devices 214 Smart Glasses 314 Headset-type terminal 414 Robots< / url:> < / url:> < / url:> < / url:>

Claims

1. A means for receiving real-time video data from a video camera device and detecting the player's movements and the ball's motion using a deep learning model for analyzing the video data, A means for evaluating the technical characteristics of the player based on their shots and identifying their strengths and weaknesses in shots, A means for generating and providing feedback to players on areas for improvement and shot strategies in natural language based on the technical features, Means for presenting the feedback visually and audibly on a display device, A means for archiving the video data and technical features in a database and managing a history for analyzing the player's progress, A means for generating an individual training plan based on the player's performance data, A system that includes this.

2. The system according to claim 1, further comprising means for suggesting a shot the player is good at in order to promote the player's sense of success when the player has failed a particular shot multiple times in a row.

3. The system according to claim 1, further comprising a module for the user to later review the feedback and performance data using a graphical user interface.

Citation Information

Patent Citations

  • Persona chatbot control method and system

    JP2022180282A