system

A system for analyzing golf swings through uploaded videos or images provides efficient, objective, and cost-effective swing analysis with visual feedback, addressing the need for specialized equipment and expertise.

JP2026062298APending Publication Date: 2026-04-09SOFTBANK GROUP CORP
View PDF 1 Cites 0 Cited by

Patent Information

Authority / Receiving Office
JP · JP
Patent Type
Applications
Current Assignee / Owner
Filing Date
2024-09-30
Publication Date
2026-04-09

AI Technical Summary

Technical Problem

Conventional golf swing analysis requires expensive equipment and specialized knowledge, necessitating expert visual evaluation, making it difficult for ordinary users to perform efficient and objective evaluations and grasp swing improvement points.

Method used

A system that allows users to upload videos or images of their swing, extracts key scenes, executes a diagnostic algorithm for analysis, and provides diagnostic results with visual feedback, enabling users to analyze and improve their swing form without specialized equipment or expertise.

Benefits of technology

Enables efficient, objective, and cost-effective golf swing analysis with intuitive visual feedback, allowing users to identify and address specific improvement areas.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure 2026062298000001_ABST
    Figure 2026062298000001_ABST
Patent Text Reader

Abstract

We provide the system. [Solution] A means for the user to upload a video or image of their swing, A method for extracting each scene from an uploaded video, A means for analyzing extracted scenes and executing a diagnostic algorithm to evaluate swing form, A system that includes means for generating and displaying diagnostic results to the user.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The technology of the present disclosure relates to a system.

Background Art

[0002] Patent Document 1 discloses a method for controlling a persona chatbot, which is performed by at least one processor, and includes steps of receiving a user utterance, adding the user utterance to a prompt including an instruction sentence related to an explanation of a character of the chatbot, encoding the prompt, and inputting the encoded prompt into a language model to generate a chatbot utterance in response to the user utterance.

Prior Art Documents

Patent Documents

[0003]

Patent Document 1

Summary of the Invention

Problems to be Solved by the Invention

[0004] Conventional golf swing analysis often requires dedicated expensive equipment and specialized knowledge, making it difficult for ordinary users to use easily. Also, there was a problem that an expert's visual evaluation was required for swing analysis, and it was difficult to perform an efficient and objective evaluation. Furthermore, it was difficult to accurately grasp each scene of the swing and appropriately feedback improvement points based on it.

Means for Solving the Problems

[0005] The present invention solves the above problems by providing a system that includes means for a user to upload a video or image of their swing, means for extracting each scene (address, backswing, top, impact, finish) from the uploaded video, means for executing a diagnostic algorithm to analyze the extracted scenes and evaluate the swing form, and means for generating and displaying the diagnostic results to the user. By further including means for directly analyzing the uploaded image and executing a diagnostic algorithm to evaluate the swing form, the system can also handle cases where video analysis is not necessary, thereby improving efficiency. In addition, by including means for adding annotations to the images of each scene to indicate areas for improvement, the system can provide the user with concrete visual feedback and aid in understanding.

[0006] A "user" is an individual or group that uses the swing analysis system.

[0007] A "swing" is a series of actions performed using a golf club to hit a golf ball.

[0008] A "video" is a digital file that displays a series of still images over time to represent movement.

[0009] An "image" is a digital file that holds static visual information.

[0010] "Uploading" refers to the act of a user sending digital files from their device to a server.

[0011] A "scene" refers to a frame that represents a specific moment or state of motion within a swing, including the address, backswing, top, impact, and finish phases.

[0012] "Address" refers to the stance you take in relation to the ball before starting a golf swing.

[0013] "Takeback" refers to the action of pulling the club backward in a golf swing.

[0014] "Top" refers to the moment when the club is held at the highest position in a golf swing.

[0015] "Impact" refers to the moment when the golf club contacts the ball.

[0016] "Finish" refers to the final posture after the swing is completed.

[0017] "Extraction" refers to the process of cutting out a specific scene from a video.

[0018] "Analysis" refers to the process of evaluating the content of a video or an image using digital algorithms.

[0019] "Diagnostic algorithm" refers to the calculation methods and programs used to evaluate the swing form.

[0020] "Generation" refers to the process of creating useful information based on the analysis results.

[0021] "Annotation" refers to the explanations and notes added to an image or a video.

[0022] "Feedback" refers to the evaluation information and improvement suggestions provided to the user based on the analysis results.

[0023] "Terminal" refers to the electronic device used by the user, including smartphones, personal computers, etc.

Brief Description of the Drawings

[0024] [Figure 1] It is a conceptual diagram showing an example of the configuration of a data processing system according to the first embodiment. [Figure 2] It is a conceptual diagram showing an example of the main functions of a data processing device and a smart device according to the first embodiment. [Figure 3] It is a conceptual diagram showing an example of the configuration of a data processing system according to the second embodiment. [Figure 4] It is a conceptual diagram showing an example of the main functions of a data processing device and smart glasses according to the second embodiment. [Figure 5] It is a conceptual diagram showing an example of the configuration of a data processing system according to the third embodiment. [Figure 6] It is a conceptual diagram showing an example of the main functions of a data processing device and a headset-type terminal according to the third embodiment. [Figure 7] It is a conceptual diagram showing an example of the configuration of a data processing system according to the fourth embodiment. [Figure 8] It is a conceptual diagram showing an example of the main functions of a data processing device and a robot according to the fourth embodiment. [Figure 9] It shows an emotion map to which a plurality of emotions are mapped. [Figure 10] It shows an emotion map to which a plurality of emotions are mapped. [Figure 11] It is a sequence diagram showing the processing flow of the data processing system in Example 1. [Figure 12] It is a sequence diagram showing the processing flow of the data processing system in Application Example 1. [Figure 13] It is a sequence diagram showing the processing flow of the data processing system in Example 2 when an emotion engine is combined. [Figure 14] It is a sequence diagram showing the processing flow of the data processing system in Application Example 2 when an emotion engine is combined.

Embodiments for Carrying Out the Invention

[0025] Hereinafter, an example of an embodiment of a system according to the technology of the present disclosure will be described according to the accompanying drawings.

[0026] First, the terms used in the following description will be explained.

[0027]

[0027] In the following embodiments, the signed processor (hereinafter simply referred to as "processor") may be a single arithmetic unit or a combination of multiple arithmetic units. Furthermore, the processor may be a single type of arithmetic unit or a combination of multiple types of arithmetic units. Examples of arithmetic units include CPU (Central Processing Unit), GPU (Graphics Processing Unit), GPGPU (General-Purpose computing on Graphics Processing Units), and APU (Accelerated Processing Unit).

[0028] In the following embodiments, signed RAM (Random Access Memory) is a memory that temporarily stores information and is used as work memory by the processor.

[0029] In the following embodiments, the signed storage is one or more non-volatile storage devices that store various programs and various parameters. Examples of non-volatile storage devices include flash memory (SSD (Solid State Drive)), magnetic disks (e.g., hard disks), or magnetic tapes.

[0030] In the following embodiments, the signed communication interface (I / F) is an interface that includes a communication processor and an antenna, etc. The communication interface manages communication between multiple computers. Examples of communication standards applicable to the communication interface include wireless communication standards such as 5G (5th Generation Mobile Communication System), Wi-Fi (registered trademark), or Bluetooth (registered trademark).

[0031] In the following embodiments, "A and / or B" is synonymous with "at least one of A and B." That is, "A and / or B" means that it may be A alone, or B alone, or a combination of A and B. Furthermore, in this specification, the same concept as "A and / or B" applies when expressing three or more things linked by "and / or."

[0032] [First Embodiment]

[0033] Figure 1 shows an example of the configuration of the data processing system 10 according to the first embodiment.

[0034] As shown in Figure 1, the data processing system 10 includes a data processing device 12 and a smart device 14. An example of the data processing device 12 is a server.

[0035] The data processing device 12 comprises a computer 22, a database 24, and a communication interface 26. The computer 22 is an example of a "computer" related to the technology of this disclosure. The computer 22 comprises a processor 28, RAM 30, and storage 32. The processor 28, RAM 30, and storage 32 are connected to a bus 34. The database 24 and the communication interface 26 are also connected to the bus 34. The communication interface 26 is connected to a network 54. An example of the network 54 is a WAN (Wide Area Network) and / or a LAN (Local Area Network).

[0036] The smart device 14 comprises a computer 36, a reception device 38, an output device 40, a camera 42, and a communication interface 44. The computer 36 comprises a processor 46, RAM 48, and storage 50. The processor 46, RAM 48, and storage 50 are connected to a bus 52. The reception device 38, output device 40, and camera 42 are also connected to the bus 52.

[0037] The reception device 38 is equipped with a touch panel 38A and a microphone 38B, etc., and receives user input. The touch panel 38A receives user input by detecting contact with an object (e.g., a pen or finger). The microphone 38B receives user input by detecting the user's voice. The control unit 46A transmits data indicating the user input received by the touch panel 38A and microphone 38B to the data processing device 12. In the data processing device 12, the specific processing unit 290 acquires the data indicating the user input.

[0038] The output device 40 includes a display 40A and a speaker 40B, and presents data to the user 20 by outputting the data in a form perceptible to the user 20 (e.g., audio and / or text). The display 40A displays visible information such as text and images according to instructions from the processor 46. The speaker 40B outputs audio according to instructions from the processor 46. The camera 42 is a small digital camera equipped with an optical system such as a lens, aperture, and shutter, and an image sensor such as a CMOS (Complementary Metal-Oxide-Semiconductor) image sensor or a CCD (Charge Coupled Device) image sensor.

[0039] Communication interface 44 is connected to network 54. Communication interfaces 44 and 26 are responsible for the exchange of various types of information between processor 46 and processor 28 via network 54.

[0040] Figure 2 shows an example of the main functions of the data processing device 12 and the smart device 14.

[0041] As shown in Figure 2, in the data processing device 12, a specific processing is performed by the processor 28. A specific processing program 56 is stored in the storage 32. The specific processing program 56 is an example of a "program" related to the technology of this disclosure. The processor 28 reads the specific processing program 56 from the storage 32 and executes the read specific processing program 56 on the RAM 30. The specific processing is realized by the processor 28 operating as a specific processing unit 290 according to the specific processing program 56 executed on the RAM 30.

[0042] The storage 32 stores the data generation model 58 and the emotion identification model 59. The data generation model 58 and the emotion identification model 59 are used by the identification processing unit 290.

[0043] In the smart device 14, the processor 46 performs the reception output processing. The storage 50 stores the reception output program 60. The reception output program 60 is used in conjunction with a specific processing program 56 by the data processing system 10. The processor 46 reads the reception output program 60 from the storage 50 and executes the read reception output program 60 on the RAM 48. The reception output processing is realized by the processor 46 operating as a control unit 46A according to the reception output program 60 executed on the RAM 48.

[0044] Next, the specific processing performed by the specific processing unit 290 of the data processing device 12 will be described. In the following description, the data processing device 12 will be referred to as the "server" and the smart device 14 as the "terminal".

[0045] The present invention provides a system for which a user uploads a video or image of their swing, and for which the swing form is analyzed and diagnosed. The operation of each element and the program's processing are described below in natural language.

[0046] System Overview

[0047] This system consists of three main components: users, terminals, and servers.

[0048] User

[0049] Users record videos or images of their golf swing using a device such as a smartphone or camera. After recording, they access a swing analysis application or web interface and select the file.

[0050] terminal

[0051] The terminal's role is to send the user-selected video or image file to the server. Communication methods such as HTTP POST requests are used for transmission. Furthermore, the terminal displays the diagnostic results received from the server on the user interface.

[0052] server

[0053] The server plays a central role in receiving and analyzing videos or images sent from the terminal. If a video is sent, the server analyzes the frames and extracts each scene of the swing (address, backswing, top, impact, finish). It then runs a swing diagnostic algorithm on the extracted scenes or images to generate a diagnostic result. The generated diagnostic result is then sent back to the terminal.

[0054] Specific examples of program processing

[0055] 1. The user films their golf swing with their smartphone.

[0056] Use your smartphone's camera function to record your swing motion in video format.

[0057] 2. The user opens the application, selects the recorded video, and uploads it.

[0058] Use the application's file selection screen to upload the recorded video file to the designated location.

[0059] 3. The device sends the selected video file to the server.

[0060] File transmission is performed using HTTP POST requests or other appropriate communication methods.

[0061] 4. The server receives the video and extracts frames for each scene of the swing.

[0062] The system analyzes the video content and executes an algorithm to extract specific scenes (address, backswing, top of swing, impact, finish).

[0063] 5. The server analyzes the images from each scene and evaluates them based on the swing diagnostic algorithm.

[0064] The system analyzes the swing form for each extracted scene image. The diagnostic algorithm evaluates the accuracy of posture and movement, identifying abnormalities and areas for improvement.

[0065] 6. The server generates diagnostic results and creates feedback such as "Your body is unbalanced at address" and "Your timing of impact is slow."

[0066] Based on the output of the diagnostic algorithm, specific feedback is generated to be provided to the user.

[0067] 7. The server sends the diagnostic results back to the smartphone.

[0068] The generated diagnostic results are sent back to the terminal as an HTTP response.

[0069] 8. The terminal displays the diagnostic results to the user and indicates areas for improvement.

[0070] The diagnostic results are displayed in the user interface, visually and textually presenting areas for improvement to the user. Annotations are added to the images in each scene to clearly indicate specific problems.

[0071] This system allows users to efficiently and objectively analyze their golf swing and visually understand areas for improvement without requiring expensive specialized equipment or expertise.

[0072] The following describes the processing flow.

[0073] Step 1: The user records a video of their golf swing using their smartphone.

[0074] Users record their golf swing as a video using their smartphone's camera function.

[0075] Step 2: The user opens the application and selects the recorded video.

[0076] Display the application's file selection screen and select the video file you recorded.

[0077] Step 3: The user uploads a video file.

[0078] Press the upload button in the application to send the selected video file to the server.

[0079] Step 4: The device sends the video file to the server.

[0080] The device uses an HTTP POST request to send the video file to the server.

[0081] Step 5: The server receives the video file.

[0082] The server retrieves and saves the video file from the received HTTP POST request.

[0083] Step 6: The server analyzes the video and extracts each scene.

[0084] The server analyzes the video frames to identify each scene—address, backswing, top of swing, impact, and finish—and extracts the corresponding frames.

[0085] Step 7: The server preprocesses the frames of each extracted scene.

[0086] The server standardizes the frames of each scene and converts them into a format that allows image analysis algorithms to function properly. For example, it performs image size unification and noise reduction.

[0087] Step 8: The server inputs the pre-processed images into the diagnostic algorithm.

[0088] The server inputs the pre-processed images of each scene into a swing diagnostic algorithm for evaluation.

[0089] Step 9: The server generates a diagnostic result based on the image analysis results.

[0090] Note) The server evaluates the analysis results and identifies anomalies in the swing form for each scene. It generates feedback on the identified anomalies and compiles them into a diagnostic report.

[0091] Step 10: The server sends the diagnostic results back to the terminal.

[0092] The server returns the diagnostic results to the terminal as an HTTP response.

[0093] Step 11: The terminal displays the diagnostic results to the user.

[0094] The terminal displays the received diagnostic results on the user interface, showing the user specific areas for improvement and evaluation of their swing form.

[0095] This processing step allows users to efficiently analyze their golf swing and receive specific feedback.

[0096] (Example 1)

[0097] Next, we will describe Example 1. In the following description, the data processing device 12 will be referred to as the "server," and the smart device 14 will be referred to as the "terminal."

[0098] Conventional technologies required expensive specialized equipment and the assistance of skilled technicians to analyze golf swings, making it difficult for ordinary users to easily evaluate and improve their swing form. Furthermore, there was a lack of methods to provide swing form diagnostic results in an intuitively understandable format. Therefore, there was a need for a technology that could analyze and evaluate golf swings easily and at low cost, allowing users to clearly identify areas for improvement themselves.

[0099] The identification process performed by the identification processing unit 290 of the data processing device 12 in Example 1 is realized by the following means.

[0100] In this invention, the server includes means for a user to upload a video or image of their swing, means for analyzing and extracting each scene (address, backswing, top, impact, finish) frame by frame from the uploaded video, means for executing a diagnostic algorithm to analyze the extracted scenes and evaluate the swing form, means for returning and displaying the generated diagnostic results to the user, means for generating specific feedback (e.g., issues with body balance or timing) based on the diagnostic results, and means for adding annotations to the images of each scene as feedback to indicate areas for improvement. This makes it possible for users to analyze their golf swing in detail and intuitively understand specific areas for improvement without using expensive specialized equipment.

[0101] A "user" refers to a person who uses the system to film their own golf swing and obtain the analysis results.

[0102] "Swing video or image" refers to a video file or still image file recorded by the user of their golf swing motion.

[0103] "Means for uploading" refers to the interface or function that allows users to provide videos or images they have taken to the system.

[0104] "Each scene" refers to the specific phases of a golf swing: address, backswing, top, impact, and finish.

[0105] "Means for analyzing and extracting frame by frame" refers to technologies and devices that divide a video into multiple still images and identify a specific scene from each frame.

[0106] A "diagnostic algorithm" refers to a calculation method or program that analyzes extracted scene images to evaluate golf swing form.

[0107] The means for returning and displaying the "generated diagnostic results" to the user refers to the communication method and display function for returning the analysis results from the diagnostic algorithm and making them viewable by the user.

[0108] "Specific feedback" refers to detailed observations and advice based on the diagnostic results, such as issues with body balance or timing.

[0109] "Annotation" refers to annotations, marks, and visual information added to images in each scene to indicate areas for improvement.

[0110] The present invention provides a system for users to record their golf swings as videos or images, and to analyze and evaluate those swings. This system consists of three main components: the user, the terminal, and the server.

[0111] User operations and terminal functions

[0112] Users record their golf swing as a video or image using a device such as a smartphone or tablet. This is done using the device's camera function. For example, the iPhone® camera app can be used to shoot high-resolution video. Afterward, the user launches a dedicated application such as "Golf Swing Analyzer," selects the recorded video through the application's interface, and uploads it.

[0113] The terminal is responsible for sending the video file selected by the user to the server. Common communication technologies such as HTTP POST requests are used for this transmission. During transmission, the appropriate communication protocol is selected considering the file format and size. Furthermore, the terminal displays the diagnostic results received from the server on the user interface. Therefore, the terminal requires visual and textual display functions that allow the user to intuitively understand the diagnostic results.

[0114] Server analysis and diagnostic functions

[0115] The server plays a central role in receiving and analyzing videos or images sent from the terminal. After receiving the data, the server uses image analysis libraries such as OpenCV to divide the video into frames and extract each scene of the swing (address, backswing, top, impact, finish). A specific algorithm is used to extract each scene.

[0116] For each extracted scene, the server uses a generative AI model such as TENSORFLOW® to evaluate the swing form. This diagnostic algorithm evaluates the accuracy of posture and movement, and identifies abnormalities and areas for improvement. For example, it can evaluate things like "the body is unbalanced at address" or "the timing of impact is late."

[0117] Examples of specific cases and prompt statements

[0118] Further implementation examples include the following: A user uploads a video of their golf swing, the server analyzes it, and sends specific feedback back to the user's device. This feedback is displayed as annotations for each scene in the video, making it easy for the user to understand. This system allows users to obtain analysis and evaluation of their golf swing without needing expensive specialized equipment.

[0119] Examples of prompt messages are as follows:

[0120] Design a system that automatically analyzes the movements in each scene of a golf swing video recorded by a user on their iPhone, and diagnoses the swing form. Describe the specific analysis method and algorithm, and clearly specify the communication method between the server and client.

[0121] This invention allows users to efficiently and accurately analyze their golf swing and identify areas for improvement without requiring specialized knowledge or expensive equipment.

[0122] The flow of the specific processing in Example 1 will be explained using Figure 11.

[0123] Program processing flow

[0124] Step 1:

[0125] The user films their golf swing.

[0126] (Specific actions)

[0127] Users record their golf swings in video format using their smartphones or tablets. They use a camera app to capture the entire swing within the frame.

[0128] (input)

[0129] Videos filmed by users.

[0130] (output)

[0131] A video of a swing saved on a smartphone.

[0132] Step 2:

[0133] The user opens the application, selects the video they have recorded, and uploads it.

[0134] (Specific actions)

[0135] The user launches a dedicated app such as "Golf Swing Analyzer" and navigates to the "file selection screen" within the app's interface. The user selects the recorded video and clicks the "upload button".

[0136] (input)

[0137] A video of a swing saved on a smartphone.

[0138] (output)

[0139] The selected video will be displayed within the application, and an upload request will be generated.

[0140] Step 3:

[0141] The device sends the video file to the server.

[0142] (Specific actions)

[0143] The device sends the video file to the server using an HTTP POST request. During this process, the video file is appropriately split and transmitted according to the communication protocol.

[0144] (input)

[0145] The selected video file within the application, and the upload request.

[0146] (output)

[0147] Video file being sent to the server.

[0148] Step 4:

[0149] The server receives the video and analyzes it frame by frame.

[0150] (Specific actions)

[0151] The server saves the received video file and uses the OpenCV library to split the video into frames. Each frame is sequentially loaded into memory and used for analysis.

[0152] (input)

[0153] The video file sent to the server.

[0154] (output)

[0155] Video data divided into individual frames.

[0156] Step 5:

[0157] The server extracts each scene and applies a swing analysis algorithm.

[0158] (Specific actions)

[0159] The server analyzes each frame and detects specific scenes (address, backswing, top of trajectory, impact, finish). After this, a TensorFlow-based generative AI model is used to evaluate each scene.

[0160] (input)

[0161] Video data divided into individual frames.

[0162] (output)

[0163] Data for each detected scene and its diagnostic results.

[0164] Step 6:

[0165] The server generates diagnostic results and creates feedback for the user.

[0166] (Specific actions)

[0167] The server generates feedback based on the analysis results. For example, it creates diagnostic statements such as "body balance is off at address" or "impact timing is slow," and formats them in JSON format.

[0168] (input)

[0169] Data for each detected scene and its diagnostic results.

[0170] (output)

[0171] User feedback data.

[0172] Step 7:

[0173] The server sends the generated diagnostic results back to the user's terminal.

[0174] (Specific actions)

[0175] The server sends the generated diagnostic results back to the user's terminal as an HTTP response. The returned data is sent in an appropriate format, such as JSON or XML.

[0176] (input)

[0177] User feedback data.

[0178] (output)

[0179] Diagnostic results sent to the user's terminal.

[0180] Step 8:

[0181] The device displays the diagnostic results.

[0182] (Specific actions)

[0183] The terminal analyzes the received diagnostic results and displays them in the user interface. Because the user can visually confirm the diagnostic results, specific annotations are added for each scene.

[0184] (input)

[0185] Diagnostic results sent from the server.

[0186] (output)

[0187] Diagnostic results and annotations displayed in the user interface.

[0188] (Application Example 1)

[0189] Next, we will explain Application Example 1. In the following explanation, the data processing device 12 will be referred to as the "server," and the smart device 14 will be referred to as the "terminal."

[0190] Until now, there had been no system that could efficiently and quickly analyze golf swing form and provide users with useful feedback in real time. As a result, users needed expert advice or expensive equipment to improve their swing form. This meant that many amateur golfers did not have enough opportunities for self-improvement. Furthermore, because the feedback was not in real time, users could not immediately obtain information for improvement, which reduced the efficiency of their practice. In addition, the lack of a feedback system using wearable devices made it difficult for users to check swing improvement measures on the spot.

[0191] The specific processing performed by the specific processing unit 290 of the data processing device 12 in Application Example 1 is realized by the following means.

[0192] In this invention, the server includes means for a user to upload a video or image of their swing, means for extracting each scene (address, backswing, top, impact, finish) from the uploaded video, means for executing a diagnostic algorithm to analyze the extracted scenes and evaluate the swing form, and means for providing real-time feedback of the displayed diagnostic results to a wearable device worn by the user. This enables real-time feedback of golf swing form analysis and diagnostic results, allowing the user to immediately see areas for improvement and increase the efficiency of their practice.

[0193] A "user" is an individual or group that uses the system to upload videos or images of their golf swing and receive analysis and diagnostic results.

[0194] A "video" is a dynamic media file that displays a series of images taken by a user in sequence.

[0195] An "image" is a media file that includes still images taken by the user.

[0196] "Uploading" refers to the act of a user transferring video or image data from their device to a server.

[0197] A "scene" refers to a specific moment or stage in a swing, specifically the address, backswing, top of the swing, impact, and finish.

[0198] "Extraction" refers to taking a specific scene from an uploaded video.

[0199] "Analysis" refers to evaluating the extracted images and videos of scenes using a diagnostic algorithm.

[0200] A "diagnostic algorithm" is a specific set of calculation procedures or analytical models used to evaluate swing form and generate diagnostic results.

[0201] "Generation" refers to creating diagnostic results based on the analyzed data.

[0202] "Display" refers to outputting the generated diagnostic results to the terminal in a format that the user can review.

[0203] A "wearable device" refers to a device that a user can wear, typically in the form of glasses, a head-mounted display, or a wristband.

[0204] "Real-time" refers to processing and feedback being performed instantly without delay.

[0205] "Feedback" is the process of returning the generated diagnostic results and areas for improvement to the user.

[0206] This invention is a system that analyzes videos and images of golf swings to diagnose swing form, and is designed to be easily accessible to users. The main components are the user, a terminal, a server, and a wearable device.

[0207] User

[0208] Users record their golf swing using a device such as a smartphone or smart glasses, capturing it as a video or image. After recording, they upload the video or image using a dedicated application. This allows the system to input the user's swing form.

[0209] terminal

[0210] The device is used to send videos and images taken by the user to the server. HTTP POST requests are typically used for transmission. Furthermore, the device displays the diagnostic results received from the server on the user interface. Examples of devices include smartphones, smart glasses, and head-mounted displays.

[0211] server

[0212] The server plays a central role in analyzing videos and images sent from the terminal and evaluating the swing form. Specifically, it performs the following processes:

[0213] 1. When a video is sent, the server analyzes it frame by frame and extracts each scene of the swing (address, backswing, top, impact, finish).

[0214] 2. A diagnostic algorithm is run on the extracted scenes to evaluate the swing form of each scene. The software used includes OpenCV (image processing), TensorFlow and PyTorch (machine learning algorithms), and Flask (web framework).

[0215] 3. Generate diagnostic results and create feedback to provide to the user. This feedback will include specific areas for improvement and problems, and will be displayed in text and annotation format.

[0216] Wearable devices

[0217] Wearable devices are devices that users wear to receive real-time feedback. Examples include smart glasses and head-mounted displays. Diagnostic results and feedback transmitted from a server are displayed on these wearable devices in real time.

[0218] Specific example

[0219] As a concrete example, imagine a scenario where a user wears smart glasses at a golf driving range and performs a swing. The smart glasses film the user's swing in real time and send the video to a server. The server analyzes the video and diagnoses the swing form. As a result, feedback such as "Your body balance is off at address" or "Your top position is too high" is immediately displayed on the smart glasses' screen. The user can then correct their swing form on the spot, film again, and receive further evaluation.

[0220] Examples of prompts to input into a generative AI model

[0221] "Design an application that allows users to film their golf swing with smart glasses and receive real-time form analysis and diagnostic results. The system would involve filming the frame, uploading the footage to a server, and displaying the analyzed results on the smart glasses. Please provide a detailed explanation, including the necessary hardware and software, and the program's processing steps."

[0222] The flow of a specific process in Application Example 1 will be explained using Figure 12.

[0223] Step 1:

[0224] The user takes a video or image of their swing.

[0225] Input: A video or image of a swing taken using the camera of a smartphone or smart glasses.

[0226] Operation: The user performs a golf swing, and the camera records the motion. In the case of video, a series of movements are recorded continuously; in the case of still images, specific moments are saved as still images.

[0227] Output: The recorded video or image of the golf swing is saved on the device.

[0228] Step 2:

[0229] Users upload videos or images they have taken to the server via a dedicated application.

[0230] Input: Swing video or image taken after shooting.

[0231] Operation: The user opens the dedicated application, selects a video or image from the file selection screen, and presses the upload button. The file is sent to the server using a data transfer protocol such as an HTTP POST request.

[0232] Output: Video or image data is sent to the server.

[0233] Step 3:

[0234] The server analyzes the received video or images and extracts each scene of the swing.

[0235] Input: A swing video or image uploaded to the server.

[0236] Operation: The server divides the video frame by frame and detects important scenes in the swing (address, backswing, top, impact, finish). This analysis and extraction is performed using image processing libraries such as OpenCV.

[0237] Output: Frame images of each extracted scene.

[0238] Step 4:

[0239] The server analyzes each extracted scene and executes a diagnostic algorithm to evaluate the swing form.

[0240] Input: Frame images from each extracted scene.

[0241] Operation: The server uses machine learning models (e.g., TensorFlow, PyTorch) to evaluate the swing form. It analyzes the accuracy of posture and movement, and identifies abnormalities and areas for improvement.

[0242] Output: Swing form evaluation results for each scene.

[0243] Step 5:

[0244] The server generates diagnostic results based on the swing form evaluation and creates feedback to provide to the user.

[0245] Input: Swing form evaluation results.

[0246] Operation: The server generates specific feedback based on the evaluation results. For example, it might create text-based feedback such as "Your body balance is off at address" or "Your top position is too high."

[0247] Output: Generated diagnostic results and feedback.

[0248] Step 6:

[0249] The server sends diagnostic results to the terminal and wearable device, and displays feedback in real time.

[0250] Input: Generated diagnostic results and feedback.

[0251] Operation: The server sends the diagnostic results to the device (smartphone, smart glasses, etc.). The results are returned as an HTTP response, and in the case of wearable devices, the data is transferred via a dedicated API. The device or wearable device displays the received data.

[0252] Output: Diagnostic results and feedback displayed on the user's device or wearable device.

[0253] Step 7:

[0254] We will work on improving the swing based on the feedback displayed to the user.

[0255] Input: Diagnostic results and feedback displayed on a terminal or wearable device.

[0256] Operation: The user reviews the displayed feedback and implements measures to improve their swing form. If necessary, they film their swing again and use the system for further diagnosis.

[0257] Output: Improved swing form, revised diagnostic results.

[0258] Furthermore, an emotion engine that estimates the user's emotions may be incorporated. That is, the identification processing unit 290 may use the emotion identification model 59 to estimate the user's emotions and perform identification processing using the user's emotions.

[0259] The system of the present invention not only allows users to upload videos or images of their swings and analyzes and diagnoses their swing form, but also includes a function to recognize the user's emotions and adjust the diagnosis results based on those emotions. The operation of each element and the processing of the program are described below in natural language.

[0260] System Overview

[0261] This system consists of four main components: the user, the terminal, the server, and the emotion engine.

[0262] User

[0263] Users record videos or images of their golf swing using a device such as a smartphone or camera. After recording, they access a swing analysis application or web interface to select the file. In addition, the user's facial expressions are also recorded during the upload process.

[0264] terminal

[0265] The terminal's role is to send the video or image file selected by the user to the server. Communication methods such as HTTP POST requests are used for transmission. The terminal also displays the diagnostic results received from the server on the user interface.

[0266] server

[0267] The server plays a central role in receiving and analyzing videos or images sent from the terminal. If a video is sent, the server analyzes the frames and extracts each scene of the swing (address, backswing, top, impact, finish). It then runs a swing diagnostic algorithm on the extracted scenes and images to generate a diagnostic result. It also analyzes the user's emotional data and incorporates it into the diagnostic result.

[0268] Emotional Engine

[0269] The emotion engine analyzes the user's facial expressions to identify their emotional state. This allows the system to understand the user's emotions when they upload swing videos or images and adjust the swing analysis results accordingly.

[0270] Specific examples of program processing

[0271] 1. The user simultaneously records a video of their golf swing and their facial expressions using their smartphone.

[0272] The smartphone's camera function is used to simultaneously record the golf swing and facial expressions.

[0273] 2. The user opens the application, selects the recorded video, and uploads it.

[0274] Use the application's file selection screen to upload the recorded video file to the designated location.

[0275] 3. The device sends the selected video file to the server.

[0276] File transmission is performed using HTTP POST requests or other appropriate communication methods.

[0277] 4. The server receives the video file and extracts frames for each scene of the swing.

[0278] The system analyzes the video content and executes an algorithm to extract specific scenes (address, backswing, top of swing, impact, finish).

[0279] 5. The server preprocesses the frames of each extracted scene.

[0280] The frames of each scene are standardized and converted into a format that allows image analysis algorithms to function properly. For example, this may involve unifying image sizes and removing noise.

[0281] 6. The server inputs the pre-processed images into the diagnostic algorithm.

[0282] The images from each scene, after preprocessing is complete, are input into the swing analysis algorithm for evaluation.

[0283] 7. The emotion engine analyzes the user's facial expressions and identifies their emotional state.

[0284] Analyze the facial expression data of the captured face and perform emotion classification such as joy, surprise, sadness, anger, etc.

[0285] [[ID=,4]]8. Based on the output result of the emotion engine, the server adjusts the swing diagnosis result.

[0286] Considering the emotion data, add emotion feedback to the diagnosis result. For example, when the user has a worried expression, provide more positive feedback.

[0287] 9. The server generates the final diagnosis result and returns it to the terminal.

[0288] Return the diagnosis result to the terminal as an HTTP response.

[0289] 10. The terminal displays the diagnosis result to the user.

[0290] The terminal displays the received diagnosis result on the user interface, shows the specific improvement points, evaluation of the swing form, and feedback based on emotions to the user.

[0291] With this system, the user can not only efficiently analyze their own golf swing and obtain specific feedback, but also receive flexible advice according to their emotions. This is expected to contribute to improving the user's motivation and reducing stress.

[0292] The following describes the processing flow.

[0293] Step 1: The user simultaneously captures the golf swing and facial expressions with a smartphone.

[0294] The user uses the camera function of the smartphone to record their own golf swing and facial expressions in video format.

[0295] Step 2: The user opens the application and selects the captured video.

[0296] The user opens the application for swing diagnosis and selects the captured video file.

[0297] Step 3: The user uploads the video file.

[0298] The user presses the upload button in the application and sends the selected video file to the server.

[0299] Step 4: The terminal sends the video file to the server.

[0300] The terminal uses an HTTP POST request to send the video file to the server.

[0301] Step 5: The server receives the video file.

[0302] The server extracts and saves the video file from the received HTTP POST request.

[0303] Step 6: The server analyzes the video and extracts each scene.

[0304] The server analyzes the frames of the video, identifies each scene of address, takeback, top, impact, and finish, and extracts the corresponding frames.

[0305] Step 7: The server preprocesses the frames of each extracted scene.

[0306] The server normalizes the frames of each scene and converts them into a format in which the image analysis algorithm can operate properly. For example, it performs operations such as unifying the image size and removing noise.

[0307] Step 8: The server inputs the preprocessed images into the diagnostic algorithm.

[0308] The server inputs the pre-processed images of each scene into a swing diagnostic algorithm for evaluation.

[0309] Step 9: The emotion engine analyzes the user's facial expressions.

[0310] An emotion engine installed on the server analyzes the user's facial expressions in the video and classifies them into emotions such as joy, surprise, sadness, and anger.

[0311] Step 10: The server adjusts the swing diagnosis results based on the output of the emotion engine.

[0312] The server considers emotional data obtained from facial expressions and includes emotion-based feedback in the generated swing diagnosis results. For example, if the user has an anxious expression, it will provide more positive feedback.

[0313] Step 11: The server generates the final diagnostic results and sends them back to the terminal.

[0314] The server returns the adjusted diagnostic results to the terminal as an HTTP response.

[0315] Step 12: The terminal displays the diagnostic results to the user.

[0316] The device displays the received diagnostic results in the user interface, visually showing specific areas for improvement and evaluation of the swing form, as well as feedback based on the user's emotions.

[0317] These detailed processing steps allow users to efficiently analyze their golf swing and receive specific, emotionally responsive feedback. This provides not only improvement to their swing form but also mental support.

[0318] (Example 2)

[0319] Next, we will describe Example 2. In the following description, the data processing device 12 will be referred to as the "server" and the smart device 14 as the "terminal".

[0320] Conventional exercise form analysis systems fail to consider the user's emotional state during video analysis and evaluation, sometimes leading to user stress or decreased motivation. Furthermore, the lack of specific feedback for each scene of exercise form makes it difficult for users to identify areas for improvement. Therefore, there is a need to develop a system that recognizes the user's emotions and reflects them in the diagnostic results, thereby providing more effective feedback and improving user motivation.

[0321] The identification process performed by the identification processing unit 290 of the data processing device 12 in Example 2 is realized by the following means.

[0322] In this invention, the server includes means for a user to upload a video or image of exercise, means for extracting each scene (start, in motion, end) from the uploaded video, means for executing a diagnostic algorithm to analyze the extracted scenes and evaluate the exercise form, means for analyzing facial expressions to identify the emotional state, means for adjusting the diagnostic results considering the emotional state, and means for generating the diagnostic results and displaying them to the user. This enables more effective and motivating feedback for the user by analyzing the user's emotions and reflecting them in the diagnostic results of the exercise form.

[0323] A "user" refers to someone who uses the system to upload videos or images of their exercise and receives the analysis results.

[0324] "Movement" refers to specific actions such as a golf swing, and is the subject of analysis by the system.

[0325] "Video or image" refers to data in file format in which a user records their exercise movements.

[0326] "Means for uploading" refers to methods and devices for sending videos and images taken by users to a server.

[0327] "Means for extraction" refers to algorithms or methods used to select specific scenes from uploaded videos.

[0328] A "diagnostic algorithm" refers to an algorithm used to analyze extracted scenes and evaluate their movement form.

[0329] "Means for analyzing facial expressions" refers to methods or devices for analyzing a user's facial expressions to identify their emotional state.

[0330] "Means of adjusting diagnostic results by considering emotional state" refers to methods or algorithms that modify or correct diagnostic results based on the analyzed emotional state.

[0331] "Means for generating diagnostic results" refers to methods or devices for compiling and providing users with evaluation results of exercise form.

[0332] "Means for display" refers to methods or devices for visually presenting the generated diagnostic results to the user.

[0333] A "scene" refers to any frame in a video of an exercise in which a specific action takes place.

[0334] "Start," "In Progress," and "End" refer to the scenes in the video that indicate the first, middle, and final stages of the exercise, respectively.

[0335] The system of the present invention not only allows users to upload videos or images of their exercise and analyzes and diagnoses their exercise form, but also includes a function to recognize the user's emotions and adjust the diagnosis results based on those emotions. This system consists of four main components: the user, the terminal, the server, and the emotion engine.

[0336] User behavior

[0337] Users record videos or images of their exercise (e.g., a golf swing) using a device such as a smartphone or camera. After recording, they access an application or web interface for exercise analysis and select the files. In addition, the user's facial expressions are recorded simultaneously during the upload process.

[0338] Terminal operation

[0339] The terminal's role is to send the user-selected video or image file to the server. Communication methods such as HTTP POST requests are used for transmission. The terminal also displays the diagnostic results received from the server on the user interface.

[0340] Server Role

[0341] The server plays a central role in receiving and analyzing videos or images sent from the terminal. Specifically, it performs the following steps:

[0342] 1. Once the video is sent, the server analyzes the frames and extracts each scene of the motion (e.g., start, in motion, end).

[0343] Example of a tool used: OpenCV

[0344] 2. The motion diagnosis algorithm is executed on the extracted scenes and images to generate the diagnosis results.

[0345] 3. Analyze user emotional data and incorporate it into the diagnostic results.

[0346] Examples of tools used: Microsoft® Azure® Cognitive Services

[0347] Functions of the Emotion Engine

[0348] The emotion engine analyzes the user's facial expressions to identify their emotional state. This allows the system to understand the user's emotions when they upload exercise videos or images and adjust the exercise diagnosis results accordingly.

[0349] Specific examples of program processing

[0350] 1. The user simultaneously records a video of their exercise and their facial expressions using their smartphone.

[0351] The smartphone's camera function is used to record exercise and facial expressions simultaneously.

[0352] 2. The user opens the application, selects the recorded video, and uploads it.

[0353] Use the application's file selection screen to upload the recorded video file to the designated location.

[0354] 3. The device sends the selected video file to the server.

[0355] File transmission is performed using HTTP POST requests or other appropriate communication methods.

[0356] 4. The server receives the video file and extracts frames for each scene of the exercise.

[0357] The program analyzes the video content and executes an algorithm to extract specific scenes (start, in progress, end). For example, it uses OpenCV.

[0358] 5. The server preprocesses the frames of each extracted scene.

[0359] The frames of each scene are standardized and converted into a format that allows image analysis algorithms to function properly. For example, this may involve unifying image sizes and removing noise.

[0360] 6. The server inputs the pre-processed images into the diagnostic algorithm.

[0361] The images from each scene, after preprocessing is complete, are input into a motion diagnosis algorithm for evaluation.

[0362] 7. The emotion engine analyzes the user's facial expressions and identifies their emotional state.

[0363] The system analyzes captured facial expression data to classify emotions such as joy, surprise, sadness, and anger. For example, it uses Microsoft Azure Cognitive Services.

[0364] 8. The server adjusts the motor diagnosis results based on the output of the emotion engine.

[0365] The system takes emotional data into account and adds emotional feedback to the diagnostic results. For example, if the user has an anxious expression, it provides more positive feedback.

[0366] 9. The server generates the final diagnostic results and sends them back to the terminal.

[0367] The diagnostic results are sent back to the terminal as an HTTP response.

[0368] 10. The terminal displays the diagnostic results to the user.

[0369] The device displays the received diagnostic results in the user interface, showing the user specific areas for improvement in their exercise form, evaluations, and emotion-based feedback.

[0370] Example of a prompt

[0371] "Please upload a video of your swing. We will analyze your swing form and your emotions, and provide feedback."

[0372] The flow of the specific processing in Example 2 will be explained using Figure 13.

[0373] Step 1:

[0374] The user simultaneously records a video of their exercise and their facial expressions using their smartphone.

[0375] Input: Smartphone camera function

[0376] Instructions: Mount your smartphone on a tripod and position it so that your whole body is visible. Press the record button to record your physical movements and facial expressions simultaneously.

[0377] Output: Exercise video file and facial expression video file

[0378] Step 2:

[0379] The user opens the application for swing analysis, selects the recorded video, and uploads it.

[0380] Input: Recorded exercise video file

[0381] How to do it: Tap the "Select Video" button in the application, choose the video file you've recorded, and then tap "Upload" to send it to the server.

[0382] Output: Exercise video file ready to upload

[0383] Step 3:

[0384] The device sends the selected video file to the server.

[0385] Input: Exercise video file selected by the user and uploaded by pressing the upload button.

[0386] Operation: The device sends a video file to the server using an HTTP POST request and displays the transmission progress.

[0387] Output: Exercise video file sent to the server

[0388] Step 4:

[0389] The server extracts each scene frame by frame from the received video file.

[0390] Input: Exercise video file sent to the server

[0391] Operation: The video is divided into frames using a video analysis library (e.g., OpenCV), and the start, middle, and end scenes are identified and extracted.

[0392] Output: Frame images divided by scene

[0393] Step 5:

[0394] The server preprocesses the extracted frame images.

[0395] Input: Frame images divided by scene

[0396] Operation: Unify image sizes and smooth images by applying a denoising filter. For example, unify the resolution to 640x480 pixels and run a denoising algorithm.

[0397] Output: Preprocessed frame image

[0398] Step 6:

[0399] The server inputs the pre-processed images into the motion diagnosis algorithm.

[0400] Input: Pre-processed frame image

[0401] Operation: Pre-processed frame images are passed to a motion diagnostic algorithm for evaluation scoring and analysis of movement form. For example, swing trajectory and body tilt are quantified.

[0402] Output: Evaluation results and scores of exercise form

[0403] Step 7:

[0404] The emotion engine analyzes the user's facial expressions to identify their emotional state.

[0405] Input: Facial expression video file being sent to the server

[0406] Operation: Uses an emotion analysis model (e.g., Microsoft Azure Cognitive Services) to classify emotions such as joy, surprise, sadness, and anger from facial expressions.

[0407] Output: Identified emotional state

[0408] Step 8:

[0409] The server adjusts the motor diagnosis results based on the output of the emotion engine.

[0410] Input: Evaluation results and scores of exercise form, identified emotional state

[0411] Operation: It takes emotional data into account and adjusts the diagnostic results to be positive or negative. For example, if the expression is anxious, it adds a positive message such as "Stay calm and continue practicing at this pace."

[0412] Output: Adjusted final diagnostic results

[0413] Step 9:

[0414] The server generates the final diagnostic results and sends them back to the terminal.

[0415] Input: Adjusted final diagnostic result

[0416] Function: Organizes the diagnostic results and sends them to the terminal in JSON format or another appropriate format.

[0417] Output: Final diagnostic results sent to the terminal

[0418] Step 10:

[0419] The terminal displays the diagnostic results to the user.

[0420] Input: Final diagnostic results received from the server

[0421] Operation: The user interface visually presents evaluation results of exercise form, areas for improvement, and emotion-based feedback. For example, it might display, "Your movement is generally good. Be careful as your body tends to move forward at the finish."

[0422] Output: Diagnostic results and feedback displayed to the user

[0423] Example of a prompt

[0424] "Please upload a video of your swing. We will analyze your swing form and your emotions, and provide feedback."

[0425] (Application Example 2)

[0426] Next, we will explain application example 2. In the following explanation, the data processing device 12 will be referred to as a "server" and the smart device 14 as a "terminal".

[0427] Traditional swing analysis systems only evaluate the user's swing form and do not consider feedback based on the user's emotions, thus limiting their effectiveness in improving user motivation and reducing stress. Furthermore, they lacked the functionality to provide specific training methods, and therefore could not directly contribute to improving the user's skills.

[0428] The specific processing performed by the specific processing unit 290 of the data processing device 12 in Application Example 2 is realized by the following means.

[0429] In this invention, the server includes means for analyzing the user's facial expressions to identify their emotional state, means for adjusting the swing diagnosis results based on the emotional state, and means for providing specific training methods based on an evaluation of the swing form. This allows the user to receive flexible feedback tailored to their emotional state and specific methods for improvement.

[0430] A "user" is an individual or group that uses the system, uploads videos or images of their swing, and receives the analysis results.

[0431] A "video" is video data composed of a series of image frames, and is used to record the user's swing motion.

[0432] An "image" is visual data represented in a single still image format and is used to capture a specific scene of a swing.

[0433] A "scene" refers to a visual representation of each stage of the swing (address, backswing, top, impact, finish).

[0434] A "diagnostic algorithm" is a set of calculation procedures and rules used to analyze and evaluate swing form, and the system uses them to scrutinize each aspect of the swing.

[0435] "Facial expression" refers to the various muscle movements and arrangements that appear on a user's face, and serves as basic information for inferring their emotional state.

[0436] "Emotional state" refers to the internal psychological state analyzed from the user's facial expressions, and includes classifications such as joy, surprise, sadness, and anger.

[0437] "Training methods" refer to specific practice procedures and exercises aimed at improving swing form and technique.

[0438] In an embodiment for carrying out this invention, the system includes the following components.

[0439] User

[0440] A "user" is an individual or group that uses the system, and is the entity that takes videos or images of their swing using a device such as a smartphone and uploads them. Users can receive not only technical evaluations of their swing form, but also emotional feedback derived from their facial expressions.

[0441] terminal

[0442] A "device" is a device used by a user to upload data, and includes smartphones and tablets. A device utilizes the following hardware and software:

[0443] Hardware: Smartphones (iPhone, Android®)

[0444] Software: Dedicated application (Golf Swing Master)

[0445] Its role is to handle everything from video recording to data transmission, and then receive diagnostic results from the server and display them to the user.

[0446] server

[0447] A "server" is a central processing unit that receives data transmitted from terminals and performs analysis and diagnosis. The server fulfills the following roles:

[0448] Swing Analysis: Frames are extracted from the video, and a diagnostic algorithm is executed to analyze each scene (address, backswing, top, impact, finish).

[0449] Emotion Recognition: Analyzes the user's facial expressions to identify their emotional state and reflect it in the swing analysis results.

[0450] Result Adjustment: The diagnostic results are adjusted based on emotional data to provide users with appropriate feedback.

[0451] Emotional Engine

[0452] The "emotion engine" is a software module that analyzes a user's facial expressions and classifies their emotional state. The emotion engine is used to understand the user's psychological state when uploading swing videos and images.

[0453] Training provided

[0454] The server provides specific training methods based on an evaluation of swing form. This allows users to learn concrete practical techniques that help improve their skills.

[0455] Examples

[0456] For example, a user films their swing at a golf driving range with their smartphone and uploads the video to a dedicated app. The device sends the video data to a server, which analyzes the video to evaluate the swing form and uses an emotion engine to analyze the user's facial expressions to identify their emotional state. As a result, in addition to the technical evaluation, the server provides feedback and training methods tailored to the user's emotional state. This feedback might include something like, "Your swing is good, but try to relax a little more. Refer to this video and practice swinging in a more relaxed state."

[0457] Examples of prompts for generative AI models

[0458] Create an application that analyzes a user's golf swing video and facial expressions, providing a technical evaluation of their swing form and emotionally responsive feedback.

[0459] 1. How to extract frames from a video.

[0460] 2. A method for identifying emotions by analyzing facial expressions.

[0461] 3. How to send swing analysis results and emotional feedback to the server.

[0462] 4. How to receive diagnostic results and how to display them in the user interface.

[0463] The flow of a specific process in Application Example 2 will be explained using Figure 14.

[0464] Step 1:

[0465] The user simultaneously records a video of their golf swing and their facial expressions using their smartphone.

[0466] Input: Video to be recorded and facial expressions

[0467] Output: Video files saved on the smartphone's storage.

[0468] Specific operation: The user uses the smartphone's camera app to record a video, and at the same time, their facial expressions are also recorded using the front camera.

[0469] Step 2:

[0470] The user opens the application, selects the video they have recorded, and uploads it.

[0471] Input: User-recorded video

[0472] Output: Video file uploaded to the application

[0473] Specific steps: Open the file selection screen within the application, select the recorded video file, and press the upload button.

[0474] Step 3:

[0475] The device sends the selected video file to the server.

[0476] Input: Uploaded video file

[0477] Output: Video data sent to the server

[0478] Specific action: Send the video file to the server using an HTTP POST request.

[0479] Step 4:

[0480] The server receives the video file and extracts frames for each scene of the swing.

[0481] Input: Video data sent from the device

[0482] Output: Frame images of each extracted scene

[0483] Specific operation: Use OpenCV to extract frames from the video and separate the address, backswing, top of swing, impact, and finish scenes.

[0484] Step 5:

[0485] The server preprocesses the frames of each extracted scene.

[0486] Input: Extracted frame image

[0487] Output: Preprocessed image data

[0488] Specific actions: Standardize frames, unify image sizes, and remove noise.

[0489] Step 6:

[0490] The server inputs the pre-processed images into the diagnostic algorithm.

[0491] Input: Preprocessed image data

[0492] Output: Swing form evaluation results

[0493] Specific operation: Pre-processed images are input into the diagnostic algorithm, and a technical evaluation of the swing form is performed.

[0494] Step 7:

[0495] The emotion engine analyzes the user's facial expressions to identify their emotional state.

[0496] Input: Facial expression data included in the video

[0497] Output: Classified emotional state

[0498] Specific operation: Facial expression data is input into an emotion recognition model to classify emotions such as joy, surprise, sadness, and anger.

[0499] Step 8:

[0500] The server adjusts the swing analysis results based on the output of the emotion engine.

[0501] Input: Swing form evaluation results and emotional state

[0502] Output: Final diagnostic results including emotional feedback

[0503] Specific actions: Consider emotional data and incorporate emotional feedback into the diagnostic results, generating feedback such as, "Your swing is good, but try to relax a little more."

[0504] Step 9:

[0505] The server generates the final diagnostic results and sends them back to the terminal.

[0506] Input: Final diagnostic results including emotional feedback

[0507] Output: Diagnostic results sent to the terminal

[0508] Specific operation: The final diagnostic results are sent back to the terminal using an HTTP response.

[0509] Step 10:

[0510] The terminal displays the diagnostic results to the user.

[0511] Input: Diagnostic results received from the server

[0512] Output: Diagnostic results displayed in the user interface

[0513] Specific action: Display the received feedback and swing evaluation in the application's user interface.

[0514] The specific processing unit 290 transmits the result of the specific processing to the smart device 14. In the smart device 14, the control unit 46A causes the output device 40 to output the result of the specific processing. The microphone 38B acquires audio indicating user input for the result of the specific processing. The control unit 46A transmits the audio data indicating user input acquired by the microphone 38B to the data processing device 12. In the data processing device 12, the specific processing unit 290 acquires the audio data.

[0515] Data generation model 58 is a so-called generative AI (Artificial Intelligence). An example of data generation model 58 is ChatGPT (registered trademark) (Internet search).<URL: https: / / openai.com / blog / chatgpt> ), Gemini (registered trademark) (Internet search) <url: https: gemini.google.com ?hl="ja">Examples of generative AI include the following. The data generation model 58 is obtained by performing deep learning on a neural network. The data generation model 58 is input with prompts containing instructions, and with inference data such as audio data representing speech, text data representing text, and image data representing images. The data generation model 58 infers from the input inference data according to the instructions indicated by the prompts, and outputs the inference results in data formats such as audio data and text data. Here, inference refers to, for example, analysis, classification, prediction, and / or summarization.

[0516] In the above embodiment, an example was given in which specific processing is performed by the data processing device 12, but the technology of this disclosure is not limited thereto, and the specific processing may also be performed by the smart device 14.

[0517] [Second Embodiment]

[0518] Figure 3 shows an example of the configuration of the data processing system 210 according to the second embodiment.

[0519] As shown in Figure 3, the data processing system 210 includes a data processing device 12 and smart glasses 214. An example of the data processing device 12 is a server.

[0520] The data processing device 12 comprises a computer 22, a database 24, and a communication interface 26. The computer 22 is an example of a "computer" related to the technology of this disclosure. The computer 22 comprises a processor 28, RAM 30, and storage 32. The processor 28, RAM 30, and storage 32 are connected to a bus 34. The database 24 and the communication interface 26 are also connected to the bus 34. The communication interface 26 is connected to a network 54. An example of the network 54 is a WAN (Wide Area Network) and / or a LAN (Local Area Network).

[0521] The smart glasses 214 include a computer 36, a microphone 238, a speaker 240, a camera 42, and a communication interface 44. The computer 36 includes a processor 46, RAM 48, and storage 50. The processor 46, RAM 48, and storage 50 are connected to a bus 52. The microphone 238, speaker 240, and camera 42 are also connected to the bus 52.

[0522] The microphone 238 receives voice signals from the user 20 and receives instructions from the user 20. The microphone 238 captures the voice signals from the user 20, converts the captured voice into audio data, and outputs it to the processor 46. The speaker 240 outputs audio according to the instructions from the processor 46.

[0523] Camera 42 is a small digital camera equipped with an optical system including a lens, aperture, and shutter, and an image sensor such as a CMOS (Complementary Metal-Oxide-Semiconductor) image sensor or a CCD (Charge Coupled Device) image sensor, and captures images of the area around the user 20 (for example, an imaging range defined by a field of view equivalent to the width of a typical healthy person's field of vision).

[0524] Communication interface 44 is connected to network 54. Communication interfaces 44 and 26 are responsible for the exchange of various information between processor 46 and processor 28 via network 54. The exchange of various information between processor 46 and processor 28 using communication interfaces 44 and 26 is performed in a secure manner.

[0525] Figure 4 shows an example of the main functions of the data processing device 12 and the smart glasses 214. As shown in Figure 4, the data processing device 12 performs specific processing using the processor 28. The storage 32 stores the specific processing program 56.

[0526] The specific processing program 56 is an example of a "program" relating to the technology of this disclosure. The processor 28 reads the specific processing program 56 from the storage 32 and executes the read specific processing program 56 on the RAM 30. The specific processing is realized by the processor 28 operating as a specific processing unit 290 in accordance with the specific processing program 56 executed on the RAM 30.

[0527] The storage 32 stores the data generation model 58 and the emotion identification model 59. The data generation model 58 and the emotion identification model 59 are used by the identification processing unit 290.

[0528] In the smart glasses 214, the processor 46 performs the reception output processing. The storage 50 stores the reception output program 60. The processor 46 reads the reception output program 60 from the storage 50 and executes the read reception output program 60 on the RAM 48. The reception output processing is realized by the processor 46 operating as a control unit 46A according to the reception output program 60 executed on the RAM 48.

[0529] Next, the identification processing performed by the identification processing unit 290 of the data processing device 12 will be described. In the following description, the data processing device 12 will be referred to as the "server" and the smart glasses 214 will be referred to as the "terminal".

[0530] The present invention provides a system for which a user uploads a video or image of their swing, and for which the swing form is analyzed and diagnosed. The operation of each element and the program's processing are described below in natural language.

[0531] System Overview

[0532] This system consists of three main components: users, terminals, and servers.

[0533] User

[0534] Users record videos or images of their golf swing using a device such as a smartphone or camera. After recording, they access a swing analysis application or web interface and select the file.

[0535] terminal

[0536] The terminal's role is to send the user-selected video or image file to the server. Communication methods such as HTTP POST requests are used for transmission. Furthermore, the terminal displays the diagnostic results received from the server on the user interface.

[0537] server

[0538] The server plays a central role in receiving and analyzing videos or images sent from the terminal. If a video is sent, the server analyzes the frames and extracts each scene of the swing (address, backswing, top, impact, finish). It then runs a swing diagnostic algorithm on the extracted scenes or images to generate a diagnostic result. The generated diagnostic result is then sent back to the terminal.

[0539] Specific examples of program processing

[0540] 1. The user films their golf swing with their smartphone.

[0541] Use your smartphone's camera function to record your swing motion in video format.

[0542] 2. The user opens the application, selects the recorded video, and uploads it.

[0543] Use the application's file selection screen to upload the recorded video file to the designated location.

[0544] 3. The device sends the selected video file to the server.

[0545] File transmission is performed using HTTP POST requests or other appropriate communication methods.

[0546] 4. The server receives the video and extracts frames for each scene of the swing.

[0547] The system analyzes the video content and executes an algorithm to extract specific scenes (address, backswing, top of swing, impact, finish).

[0548] 5. The server analyzes the images from each scene and evaluates them based on the swing diagnostic algorithm.

[0549] The system analyzes the swing form for each extracted scene image. The diagnostic algorithm evaluates the accuracy of posture and movement, identifying abnormalities and areas for improvement.

[0550] 6. The server generates diagnostic results and creates feedback such as "Your body is unbalanced at address" and "Your timing of impact is slow."

[0551] Based on the output of the diagnostic algorithm, specific feedback is generated to be provided to the user.

[0552] 7. The server sends the diagnostic results back to the smartphone.

[0553] The generated diagnostic results are sent back to the terminal as an HTTP response.

[0554] 8. The terminal displays the diagnostic results to the user and indicates areas for improvement.

[0555] The diagnostic results are displayed in the user interface, visually and textually presenting areas for improvement to the user. Annotations are added to the images in each scene to clearly indicate specific problems.

[0556] This system allows users to efficiently and objectively analyze their golf swing and visually understand areas for improvement without requiring expensive specialized equipment or expertise.

[0557] The following describes the processing flow.

[0558] Step 1: The user records a video of their golf swing using their smartphone.

[0559] Users record their golf swing as a video using their smartphone's camera function.

[0560] Step 2: The user opens the application and selects the recorded video.

[0561] Display the application's file selection screen and select the video file you recorded.

[0562] Step 3: The user uploads a video file.

[0563] Press the upload button in the application to send the selected video file to the server.

[0564] Step 4: The device sends the video file to the server.

[0565] The device uses an HTTP POST request to send the video file to the server.

[0566] Step 5: The server receives the video file.

[0567] The server retrieves and saves the video file from the received HTTP POST request.

[0568] Step 6: The server analyzes the video and extracts each scene.

[0569] The server analyzes the video frames to identify each scene—address, backswing, top of swing, impact, and finish—and extracts the corresponding frames.

[0570] Step 7: The server preprocesses the frames of each extracted scene.

[0571] The server standardizes the frames of each scene and converts them into a format that allows image analysis algorithms to function properly. For example, it performs image size unification and noise reduction.

[0572] Step 8: The server inputs the pre-processed images into the diagnostic algorithm.

[0573] The server inputs the pre-processed images of each scene into a swing diagnostic algorithm for evaluation.

[0574] Step 9: The server generates a diagnostic result based on the image analysis results.

[0575] Note) The server evaluates the analysis results and identifies anomalies in the swing form for each scene. It generates feedback on the identified anomalies and compiles them into a diagnostic report.

[0576] Step 10: The server sends the diagnostic results back to the terminal.

[0577] The server returns the diagnostic results to the terminal as an HTTP response.

[0578] Step 11: The terminal displays the diagnostic results to the user.

[0579] The terminal displays the received diagnostic results on the user interface, showing the user specific areas for improvement and evaluation of their swing form.

[0580] This processing step allows users to efficiently analyze their golf swing and receive specific feedback.

[0581] (Example 1)

[0582] Next, we will describe Example 1. In the following description, the data processing device 12 will be referred to as the "server" and the smart glasses 214 will be referred to as the "terminal".

[0583] Conventional technologies required expensive specialized equipment and the assistance of skilled technicians to analyze golf swings, making it difficult for ordinary users to easily evaluate and improve their swing form. Furthermore, there was a lack of methods to provide swing form diagnostic results in an intuitively understandable format. Therefore, there was a need for a technology that could analyze and evaluate golf swings easily and at low cost, allowing users to clearly identify areas for improvement themselves.

[0584] The identification process performed by the identification processing unit 290 of the data processing device 12 in Example 1 is realized by the following means.

[0585] In this invention, the server includes means for a user to upload a video or image of their swing, means for analyzing and extracting each scene (address, backswing, top, impact, finish) frame by frame from the uploaded video, means for executing a diagnostic algorithm to analyze the extracted scenes and evaluate the swing form, means for returning and displaying the generated diagnostic results to the user, means for generating specific feedback (e.g., issues with body balance or timing) based on the diagnostic results, and means for adding annotations to the images of each scene as feedback to indicate areas for improvement. This makes it possible for users to analyze their golf swing in detail and intuitively understand specific areas for improvement without using expensive specialized equipment.

[0586] A "user" refers to a person who uses the system to film their own golf swing and obtain the analysis results.

[0587] "Swing video or image" refers to a video file or still image file recorded by the user of their golf swing motion.

[0588] "Means for uploading" refers to the interface or function that allows users to provide videos or images they have taken to the system.

[0589] "Each scene" refers to the specific phases of a golf swing: address, backswing, top, impact, and finish.

[0590] "Means for analyzing and extracting frame by frame" refers to technologies and devices that divide a video into multiple still images and identify a specific scene from each frame.

[0591] A "diagnostic algorithm" refers to a calculation method or program that analyzes extracted scene images to evaluate golf swing form.

[0592] The means for returning and displaying the "generated diagnostic results" to the user refers to the communication method and display function for returning the analysis results from the diagnostic algorithm and making them viewable by the user.

[0593] "Specific feedback" refers to detailed observations and advice based on the diagnostic results, such as issues with body balance or timing.

[0594] "Annotation" refers to annotations, marks, and visual information added to images in each scene to indicate areas for improvement.

[0595] The present invention provides a system for users to record their golf swings as videos or images, and to analyze and evaluate those swings. This system consists of three main components: the user, the terminal, and the server.

[0596] User operations and terminal functions

[0597] Users record their golf swing as a video or image using a device such as a smartphone or tablet. They utilize the device's camera function for this purpose. For example, they can use the iPhone's camera app to shoot high-resolution video. Afterward, the user launches a dedicated application such as "Golf Swing Analyzer," selects the recorded video through the application's interface, and uploads it.

[0598] The terminal is responsible for sending the video file selected by the user to the server. Common communication technologies such as HTTP POST requests are used for this transmission. During transmission, the appropriate communication protocol is selected considering the file format and size. Furthermore, the terminal displays the diagnostic results received from the server on the user interface. Therefore, the terminal requires visual and textual display functions that allow the user to intuitively understand the diagnostic results.

[0599] Server analysis and diagnostic functions

[0600] The server plays a central role in receiving and analyzing videos or images sent from the terminal. After receiving the data, the server uses image analysis libraries such as OpenCV to divide the video into frames and extract each scene of the swing (address, backswing, top, impact, finish). A specific algorithm is used to extract each scene.

[0601] For each extracted scene, the server uses a generative AI model such as TensorFlow to evaluate the swing form. This diagnostic algorithm assesses the accuracy of posture and movement, and identifies abnormalities and areas for improvement. For example, it can evaluate things like "the body is unbalanced at address" or "the timing of impact is late."

[0602] Examples of specific cases and prompt statements

[0603] Further implementation examples include the following: A user uploads a video of their golf swing, the server analyzes it, and sends specific feedback back to the user's device. This feedback is displayed as annotations for each scene in the video, making it easy for the user to understand. This system allows users to obtain analysis and evaluation of their golf swing without needing expensive specialized equipment.

[0604] Examples of prompt messages are as follows:

[0605] Design a system that automatically analyzes the movements in each scene of a golf swing video recorded by a user on their iPhone, and diagnoses the swing form. Describe the specific analysis method and algorithm, and clearly specify the communication method between the server and client.

[0606] This invention allows users to efficiently and accurately analyze their golf swing and identify areas for improvement without requiring specialized knowledge or expensive equipment.

[0607] The flow of the specific processing in Example 1 will be explained using Figure 11.

[0608] Program processing flow

[0609] Step 1:

[0610] The user films their golf swing.

[0611] (Specific actions)

[0612] Users record their golf swings in video format using their smartphones or tablets. They use a camera app to capture the entire swing within the frame.

[0613] (input)

[0614] Videos filmed by users.

[0615] (output)

[0616] A video of a swing saved on a smartphone.

[0617] Step 2:

[0618] The user opens the application, selects the video they have recorded, and uploads it.

[0619] (Specific actions)

[0620] The user launches a dedicated app such as "Golf Swing Analyzer" and navigates to the "file selection screen" within the app's interface. The user selects the recorded video and clicks the "upload button".

[0621] (input)

[0622] A video of a swing saved on a smartphone.

[0623] (output)

[0624] The selected video will be displayed within the application, and an upload request will be generated.

[0625] Step 3:

[0626] The device sends the video file to the server.

[0627] (Specific actions)

[0628] The device sends the video file to the server using an HTTP POST request. During this process, the video file is appropriately split and transmitted according to the communication protocol.

[0629] (input)

[0630] The selected video file within the application, and the upload request.

[0631] (output)

[0632] Video file being sent to the server.

[0633] Step 4:

[0634] The server receives the video and analyzes it frame by frame.

[0635] (Specific actions)

[0636] The server saves the received video file and uses the OpenCV library to split the video into frames. Each frame is sequentially loaded into memory and used for analysis.

[0637] (input)

[0638] The video file sent to the server.

[0639] (output)

[0640] Video data divided into individual frames.

[0641] Step 5:

[0642] The server extracts each scene and applies a swing analysis algorithm.

[0643] (Specific actions)

[0644] The server analyzes each frame and detects specific scenes (address, backswing, top of trajectory, impact, finish). After this, a TensorFlow-based generative AI model is used to evaluate each scene.

[0645] (input)

[0646] Video data divided into individual frames.

[0647] (output)

[0648] Data for each detected scene and its diagnostic results.

[0649] Step 6:

[0650] The server generates diagnostic results and creates feedback for the user.

[0651] (Specific actions)

[0652] The server generates feedback based on the analysis results. For example, it creates diagnostic statements such as "body balance is off at address" or "impact timing is slow," and formats them in JSON format.

[0653] (input)

[0654] Data for each detected scene and its diagnostic results.

[0655] (output)

[0656] User feedback data.

[0657] Step 7:

[0658] The server sends the generated diagnostic results back to the user's terminal.

[0659] (Specific actions)

[0660] The server sends the generated diagnostic results back to the user's terminal as an HTTP response. The returned data is sent in an appropriate format, such as JSON or XML.

[0661] (input)

[0662] User feedback data.

[0663] (output)

[0664] Diagnostic results sent to the user's terminal.

[0665] Step 8:

[0666] The device displays the diagnostic results.

[0667] (Specific actions)

[0668] The terminal analyzes the received diagnostic results and displays them in the user interface. Because the user can visually confirm the diagnostic results, specific annotations are added for each scene.

[0669] (input)

[0670] Diagnostic results sent from the server.

[0671] (output)

[0672] Diagnostic results and annotations displayed in the user interface.

[0673] (Application Example 1)

[0674] Next, we will explain Application Example 1. In the following explanation, the data processing device 12 will be referred to as the "server," and the smart glasses 214 will be referred to as the "terminal."

[0675] Until now, there had been no system that could efficiently and quickly analyze golf swing form and provide users with useful feedback in real time. As a result, users needed expert advice or expensive equipment to improve their swing form. This meant that many amateur golfers did not have enough opportunities for self-improvement. Furthermore, because the feedback was not in real time, users could not immediately obtain information for improvement, which reduced the efficiency of their practice. In addition, the lack of a feedback system using wearable devices made it difficult for users to check swing improvement measures on the spot.

[0676] The specific processing performed by the specific processing unit 290 of the data processing device 12 in Application Example 1 is realized by the following means.

[0677] In this invention, the server includes means for a user to upload a video or image of their swing, means for extracting each scene (address, backswing, top, impact, finish) from the uploaded video, means for executing a diagnostic algorithm to analyze the extracted scenes and evaluate the swing form, and means for providing real-time feedback of the displayed diagnostic results to a wearable device worn by the user. This enables real-time feedback of golf swing form analysis and diagnostic results, allowing the user to immediately see areas for improvement and increase the efficiency of their practice.

[0678] A "user" is an individual or group that uses the system to upload videos or images of their golf swing and receive analysis and diagnostic results.

[0679] A "video" is a dynamic media file that displays a series of images taken by a user in sequence.

[0680] An "image" is a media file that includes still images taken by the user.

[0681] "Uploading" refers to the act of a user transferring video or image data from their device to a server.

[0682] A "scene" refers to a specific moment or stage in a swing, specifically the address, backswing, top of the swing, impact, and finish.

[0683] "Extraction" refers to taking a specific scene from an uploaded video.

[0684] "Analysis" refers to evaluating the extracted images and videos of scenes using a diagnostic algorithm.

[0685] A "diagnostic algorithm" is a specific set of calculation procedures or analytical models used to evaluate swing form and generate diagnostic results.

[0686] "Generation" refers to creating diagnostic results based on the analyzed data.

[0687] "Display" refers to outputting the generated diagnostic results to the terminal in a format that the user can review.

[0688] A "wearable device" refers to a device that a user can wear, typically in the form of glasses, a head-mounted display, or a wristband.

[0689] "Real-time" refers to processing and feedback being performed instantly without delay.

[0690] "Feedback" is the process of returning the generated diagnostic results and areas for improvement to the user.

[0691] This invention is a system that analyzes videos and images of golf swings to diagnose swing form, and is designed to be easily accessible to users. The main components are the user, a terminal, a server, and a wearable device.

[0692] User

[0693] Users record their golf swing using a device such as a smartphone or smart glasses, capturing it as a video or image. After recording, they upload the video or image using a dedicated application. This allows the system to input the user's swing form.

[0694] terminal

[0695] The device is used to send videos and images taken by the user to the server. HTTP POST requests are typically used for transmission. Furthermore, the device displays the diagnostic results received from the server on the user interface. Examples of devices include smartphones, smart glasses, and head-mounted displays.

[0696] server

[0697] The server plays a central role in analyzing videos and images sent from the terminal and evaluating the swing form. Specifically, it performs the following processes:

[0698] 1. When a video is sent, the server analyzes it frame by frame and extracts each scene of the swing (address, backswing, top, impact, finish).

[0699] 2. A diagnostic algorithm is run on the extracted scenes to evaluate the swing form of each scene. The software used includes OpenCV (image processing), TensorFlow and PyTorch (machine learning algorithms), and Flask (web framework).

[0700] 3. Generate diagnostic results and create feedback to provide to the user. This feedback will include specific areas for improvement and problems, and will be displayed in text and annotation format.

[0701] Wearable devices

[0702] Wearable devices are devices that users wear to receive real-time feedback. Examples include smart glasses and head-mounted displays. Diagnostic results and feedback transmitted from a server are displayed on these wearable devices in real time.

[0703] Specific example

[0704] As a concrete example, imagine a scenario where a user wears smart glasses at a golf driving range and performs a swing. The smart glasses film the user's swing in real time and send the video to a server. The server analyzes the video and diagnoses the swing form. As a result, feedback such as "Your body balance is off at address" or "Your top position is too high" is immediately displayed on the smart glasses' screen. The user can then correct their swing form on the spot, film again, and receive further evaluation.

[0705] Examples of prompts to input into a generative AI model

[0706] "Design an application that allows users to film their golf swing with smart glasses and receive real-time form analysis and diagnostic results. The system would involve filming the frame, uploading the footage to a server, and displaying the analyzed results on the smart glasses. Please provide a detailed explanation, including the necessary hardware and software, and the program's processing steps."

[0707] The flow of a specific process in Application Example 1 will be explained using Figure 12.

[0708] Step 1:

[0709] The user takes a video or image of their swing.

[0710] Input: A video or image of a swing taken using the camera of a smartphone or smart glasses.

[0711] Operation: The user performs a golf swing, and the camera records the motion. In the case of video, a series of movements are recorded continuously; in the case of still images, specific moments are saved as still images.

[0712] Output: The recorded video or image of the golf swing is saved on the device.

[0713] Step 2:

[0714] Users upload videos or images they have taken to the server via a dedicated application.

[0715] Input: Swing video or image taken after shooting.

[0716] Operation: The user opens the dedicated application, selects a video or image from the file selection screen, and presses the upload button. The file is sent to the server using a data transfer protocol such as an HTTP POST request.

[0717] Output: Video or image data is sent to the server.

[0718] Step 3:

[0719] The server analyzes the received video or images and extracts each scene of the swing.

[0720] Input: A swing video or image uploaded to the server.

[0721] Operation: The server divides the video frame by frame and detects important scenes in the swing (address, backswing, top, impact, finish). This analysis and extraction is performed using image processing libraries such as OpenCV.

[0722] Output: Frame images of each extracted scene.

[0723] Step 4:

[0724] The server analyzes each extracted scene and executes a diagnostic algorithm to evaluate the swing form.

[0725] Input: Frame images from each extracted scene.

[0726] Operation: The server uses machine learning models (e.g., TensorFlow, PyTorch) to evaluate the swing form. It analyzes the accuracy of posture and movement, and identifies abnormalities and areas for improvement.

[0727] Output: Swing form evaluation results for each scene.

[0728] Step 5:

[0729] The server generates diagnostic results based on the swing form evaluation and creates feedback to provide to the user.

[0730] Input: Swing form evaluation results.

[0731] Operation: The server generates specific feedback based on the evaluation results. For example, it might create text-based feedback such as "Your body balance is off at address" or "Your top position is too high."

[0732] Output: Generated diagnostic results and feedback.

[0733] Step 6:

[0734] The server sends diagnostic results to the terminal and wearable device, and displays feedback in real time.

[0735] Input: Generated diagnostic results and feedback.

[0736] Operation: The server sends the diagnostic results to the device (smartphone, smart glasses, etc.). The results are returned as an HTTP response, and in the case of wearable devices, the data is transferred via a dedicated API. The device or wearable device displays the received data.

[0737] Output: Diagnostic results and feedback displayed on the user's device or wearable device.

[0738] Step 7:

[0739] We will work on improving the swing based on the feedback displayed to the user.

[0740] Input: Diagnostic results and feedback displayed on a terminal or wearable device.

[0741] Operation: The user reviews the displayed feedback and implements measures to improve their swing form. If necessary, they film their swing again and use the system for further diagnosis.

[0742] Output: Improved swing form, revised diagnostic results.

[0743] Furthermore, an emotion engine that estimates the user's emotions may be incorporated. That is, the identification processing unit 290 may use the emotion identification model 59 to estimate the user's emotions and perform identification processing using the user's emotions.

[0744] The system of the present invention not only allows users to upload videos or images of their swings and analyzes and diagnoses their swing form, but also includes a function to recognize the user's emotions and adjust the diagnosis results based on those emotions. The operation of each element and the processing of the program are described below in natural language.

[0745] System Overview

[0746] This system consists of four main components: the user, the terminal, the server, and the emotion engine.

[0747] User

[0748] Users record videos or images of their golf swing using a device such as a smartphone or camera. After recording, they access a swing analysis application or web interface to select the file. In addition, the user's facial expressions are also recorded during the upload process.

[0749] terminal

[0750] The terminal's role is to send the video or image file selected by the user to the server. Communication methods such as HTTP POST requests are used for transmission. The terminal also displays the diagnostic results received from the server on the user interface.

[0751] server

[0752] The server plays a central role in receiving and analyzing videos or images sent from the terminal. If a video is sent, the server analyzes the frames and extracts each scene of the swing (address, backswing, top, impact, finish). It then runs a swing diagnostic algorithm on the extracted scenes and images to generate a diagnostic result. It also analyzes the user's emotional data and incorporates it into the diagnostic result.

[0753] Emotional Engine

[0754] The emotion engine analyzes the user's facial expressions to identify their emotional state. This allows the system to understand the user's emotions when they upload swing videos or images and adjust the swing analysis results accordingly.

[0755] Specific examples of program processing

[0756] 1. The user simultaneously records a video of their golf swing and their facial expressions using their smartphone.

[0757] The smartphone's camera function is used to simultaneously record the golf swing and facial expressions.

[0758] 2. The user opens the application, selects the recorded video, and uploads it.

[0759] Use the application's file selection screen to upload the recorded video file to the designated location.

[0760] 3. The device sends the selected video file to the server.

[0761] File transmission is performed using HTTP POST requests or other appropriate communication methods.

[0762] 4. The server receives the video file and extracts frames for each scene of the swing.

[0763] The system analyzes the video content and executes an algorithm to extract specific scenes (address, backswing, top of swing, impact, finish).

[0764] 5. The server preprocesses the frames of each extracted scene.

[0765] The frames of each scene are standardized and converted into a format that allows image analysis algorithms to function properly. For example, this may involve unifying image sizes and removing noise.

[0766] 6. The server inputs the pre-processed images into the diagnostic algorithm.

[0767] The images from each scene, after preprocessing is complete, are input into the swing analysis algorithm for evaluation.

[0768] 7. The emotion engine analyzes the user's facial expressions and identifies their emotional state.

[0769] The system analyzes captured facial expression data to classify emotions such as joy, surprise, sadness, and anger.

[0770] 8. The server adjusts the swing analysis results based on the output of the emotion engine.

[0771] The system takes emotional data into account and adds emotional feedback to the diagnostic results. For example, if the user has an anxious expression, it provides more positive feedback.

[0772] 9. The server generates the final diagnostic results and sends them back to the terminal.

[0773] The diagnostic results are sent back to the terminal as an HTTP response.

[0774] 10. The terminal displays the diagnostic results to the user.

[0775] The device displays the received diagnostic results in the user interface, showing the user specific areas for improvement in their swing form, evaluations, and emotion-based feedback.

[0776] This system allows users to efficiently analyze their golf swing and receive specific feedback, as well as flexible advice tailored to their emotions. This is expected to contribute to increased user motivation and reduced stress.

[0777] The following describes the processing flow.

[0778] Step 1: The user simultaneously films their golf swing and facial expression using their smartphone.

[0779] Users use their smartphone's camera function to record themselves and their facial expressions while performing a golf swing in video format.

[0780] Step 2: The user opens the application and selects the recorded video.

[0781] The user opens the swing analysis application and selects the recorded video file.

[0782] Step 3: The user uploads a video file.

[0783] The user presses the upload button within the application and sends the selected video file to the server.

[0784] Step 4: The device sends the video file to the server.

[0785] The device uses an HTTP POST request to send the video file to the server.

[0786] Step 5: The server receives the video file.

[0787] The server retrieves and saves the video file from the received HTTP POST request.

[0788] Step 6: The server analyzes the video and extracts each scene.

[0789] The server analyzes the video frames to identify each scene—address, backswing, top of swing, impact, and finish—and extracts the corresponding frames.

[0790] Step 7: The server preprocesses the frames of each extracted scene.

[0791] The server standardizes the frames of each scene and converts them into a format that allows image analysis algorithms to function properly. For example, it performs image size unification and noise reduction.

[0792] Step 8: The server inputs the pre-processed images into the diagnostic algorithm.

[0793] The server inputs the pre-processed images of each scene into a swing diagnostic algorithm for evaluation.

[0794] Step 9: The emotion engine analyzes the user's facial expressions.

[0795] An emotion engine installed on the server analyzes the user's facial expressions in the video and classifies them into emotions such as joy, surprise, sadness, and anger.

[0796] Step 10: The server adjusts the swing diagnosis results based on the output of the emotion engine.

[0797] The server considers emotional data obtained from facial expressions and includes emotion-based feedback in the generated swing diagnosis results. For example, if the user has an anxious expression, it will provide more positive feedback.

[0798] Step 11: The server generates the final diagnostic results and sends them back to the terminal.

[0799] The server returns the adjusted diagnostic results to the terminal as an HTTP response.

[0800] Step 12: The terminal displays the diagnostic results to the user.

[0801] The device displays the received diagnostic results in the user interface, visually showing specific areas for improvement and evaluation of the swing form, as well as feedback based on the user's emotions.

[0802] These detailed processing steps allow users to efficiently analyze their golf swing and receive specific, emotionally responsive feedback. This provides not only improvement to their swing form but also mental support.

[0803] (Example 2)

[0804] Next, we will describe Example 2. In the following description, the data processing device 12 will be referred to as the "server" and the smart glasses 214 will be referred to as the "terminal".

[0805] Conventional exercise form analysis systems fail to consider the user's emotional state during video analysis and evaluation, sometimes leading to user stress or decreased motivation. Furthermore, the lack of specific feedback for each scene of exercise form makes it difficult for users to identify areas for improvement. Therefore, there is a need to develop a system that recognizes the user's emotions and reflects them in the diagnostic results, thereby providing more effective feedback and improving user motivation.

[0806] The identification process performed by the identification processing unit 290 of the data processing device 12 in Example 2 is realized by the following means.

[0807] In this invention, the server includes means for a user to upload a video or image of exercise, means for extracting each scene (start, in motion, end) from the uploaded video, means for executing a diagnostic algorithm to analyze the extracted scenes and evaluate the exercise form, means for analyzing facial expressions to identify the emotional state, means for adjusting the diagnostic results considering the emotional state, and means for generating the diagnostic results and displaying them to the user. This enables more effective and motivating feedback for the user by analyzing the user's emotions and reflecting them in the diagnostic results of the exercise form.

[0808] A "user" refers to someone who uses the system to upload videos or images of their exercise and receives the analysis results.

[0809] "Movement" refers to specific actions such as a golf swing, and is the subject of analysis by the system.

[0810] "Video or image" refers to data in file format in which a user records their exercise movements.

[0811] "Means for uploading" refers to methods and devices for sending videos and images taken by users to a server.

[0812] "Means for extraction" refers to algorithms or methods used to select specific scenes from uploaded videos.

[0813] A "diagnostic algorithm" refers to an algorithm used to analyze extracted scenes and evaluate their movement form.

[0814] "Means for analyzing facial expressions" refers to methods or devices for analyzing a user's facial expressions to identify their emotional state.

[0815] "Means of adjusting diagnostic results by considering emotional state" refers to methods or algorithms that modify or correct diagnostic results based on the analyzed emotional state.

[0816] "Means for generating diagnostic results" refers to methods or devices for compiling and providing users with evaluation results of exercise form.

[0817] "Means for display" refers to methods or devices for visually presenting the generated diagnostic results to the user.

[0818] A "scene" refers to any frame in a video of an exercise in which a specific action takes place.

[0819] "Start," "In Progress," and "End" refer to the scenes in the video that indicate the first, middle, and final stages of the exercise, respectively.

[0820] The system of the present invention not only allows users to upload videos or images of their exercise and analyzes and diagnoses their exercise form, but also includes a function to recognize the user's emotions and adjust the diagnosis results based on those emotions. This system consists of four main components: the user, the terminal, the server, and the emotion engine.

[0821] User behavior

[0822] Users record videos or images of their exercise (e.g., a golf swing) using a device such as a smartphone or camera. After recording, they access an application or web interface for exercise analysis and select the files. In addition, the user's facial expressions are recorded simultaneously during the upload process.

[0823] Terminal operation

[0824] The terminal's role is to send the video or image file selected by the user to the server. Communication methods such as HTTP POST requests are used for transmission. The terminal also displays the diagnostic results received from the server on the user interface.

[0825] Server Role

[0826] The server plays a central role in receiving and analyzing videos or images sent from the terminal. Specifically, it performs the following steps:

[0827] 1. Once the video is sent, the server analyzes the frames and extracts each scene of the motion (e.g., start, in motion, end).

[0828] Example of a tool used: OpenCV

[0829] 2. The motion diagnosis algorithm is executed on the extracted scenes and images to generate the diagnosis results.

[0830] 3. Analyze user emotional data and incorporate it into the diagnostic results.

[0831] Example of tools used: Microsoft Azure Cognitive Services

[0832] Functions of the Emotion Engine

[0833] The emotion engine analyzes the user's facial expressions to identify their emotional state. This allows the system to understand the user's emotions when they upload exercise videos or images and adjust the exercise diagnosis results accordingly.

[0834] Specific examples of program processing

[0835] 1. The user simultaneously records a video of their exercise and their facial expressions using their smartphone.

[0836] The smartphone's camera function is used to record exercise and facial expressions simultaneously.

[0837] 2. The user opens the application, selects the recorded video, and uploads it.

[0838] Use the application's file selection screen to upload the recorded video file to the designated location.

[0839] 3. The device sends the selected video file to the server.

[0840] File transmission is performed using HTTP POST requests or other appropriate communication methods.

[0841] 4. The server receives the video file and extracts frames for each scene of the exercise.

[0842] The program analyzes the video content and executes an algorithm to extract specific scenes (start, during, and end). For example, it uses OpenCV.

[0843] 5. The server preprocesses the frames of each extracted scene.

[0844] The frames of each scene are standardized and converted into a format that allows image analysis algorithms to function properly. For example, this may involve unifying image sizes and removing noise.

[0845] 6. The server inputs the pre-processed images into the diagnostic algorithm.

[0846] The images from each scene, after preprocessing is complete, are input into a motion diagnosis algorithm for evaluation.

[0847] 7. The emotion engine analyzes the user's facial expressions and identifies their emotional state.

[0848] The system analyzes captured facial expression data to classify emotions such as joy, surprise, sadness, and anger. For example, it uses Microsoft Azure Cognitive Services.

[0849] 8. The server adjusts the motor diagnosis results based on the output of the emotion engine.

[0850] The system takes emotional data into account and adds emotional feedback to the diagnostic results. For example, if the user has an anxious expression, it provides more positive feedback.

[0851] 9. The server generates the final diagnostic results and sends them back to the terminal.

[0852] The diagnostic results are sent back to the terminal as an HTTP response.

[0853] 10. The terminal displays the diagnostic results to the user.

[0854] The device displays the received diagnostic results in the user interface, showing the user specific areas for improvement in their exercise form, evaluations, and emotion-based feedback.

[0855] Example of a prompt

[0856] "Please upload a video of your swing. We will analyze your swing form and your emotions, and provide feedback."

[0857] The flow of the specific processing in Example 2 will be explained using Figure 13.

[0858] Step 1:

[0859] The user simultaneously records a video of their exercise and their facial expressions using their smartphone.

[0860] Input: Smartphone camera function

[0861] Instructions: Mount your smartphone on a tripod and position it so that your whole body is visible. Press the record button to record your physical movements and facial expressions simultaneously.

[0862] Output: Exercise video file and facial expression video file

[0863] Step 2:

[0864] The user opens the application for swing analysis, selects the recorded video, and uploads it.

[0865] Input: Recorded exercise video file

[0866] How to do it: Tap the "Select Video" button in the application, choose the video file you've recorded, and then tap "Upload" to send it to the server.

[0867] Output: Exercise video file ready to upload

[0868] Step 3:

[0869] The device sends the selected video file to the server.

[0870] Input: Exercise video file selected by the user and uploaded by pressing the upload button.

[0871] Operation: The device sends a video file to the server using an HTTP POST request and displays the transmission progress.

[0872] Output: Exercise video file sent to the server

[0873] Step 4:

[0874] The server extracts each scene frame by frame from the received video file.

[0875] Input: Exercise video file sent to the server

[0876] Operation: The video is divided into frames using a video analysis library (e.g., OpenCV), and the start, middle, and end scenes are identified and extracted.

[0877] Output: Frame images divided by scene

[0878] Step 5:

[0879] The server preprocesses the extracted frame images.

[0880] Input: Frame images divided by scene

[0881] Operation: Unify image sizes and smooth images by applying a denoising filter. For example, unify the resolution to 640x480 pixels and run a denoising algorithm.

[0882] Output: Preprocessed frame image

[0883] Step 6:

[0884] The server inputs the pre-processed images into the motion diagnosis algorithm.

[0885] Input: Pre-processed frame image

[0886] Operation: Pre-processed frame images are passed to a motion diagnostic algorithm for evaluation scoring and analysis of movement form. For example, swing trajectory and body tilt are quantified.

[0887] Output: Evaluation results and scores of exercise form

[0888] Step 7:

[0889] The emotion engine analyzes the user's facial expressions to identify their emotional state.

[0890] Input: Facial expression video file being sent to the server

[0891] Operation: Uses an emotion analysis model (e.g., Microsoft Azure Cognitive Services) to classify emotions such as joy, surprise, sadness, and anger from facial expressions.

[0892] Output: Identified emotional state

[0893] Step 8:

[0894] The server adjusts the motor diagnosis results based on the output of the emotion engine.

[0895] Input: Evaluation results and scores of exercise form, identified emotional state

[0896] Operation: It takes emotional data into account and adjusts the diagnostic results to be positive or negative. For example, if the expression is anxious, it adds a positive message such as "Stay calm and continue practicing at this pace."

[0897] Output: Adjusted final diagnostic results

[0898] Step 9:

[0899] The server generates the final diagnostic results and sends them back to the terminal.

[0900] Input: Adjusted final diagnostic result

[0901] Function: Organizes the diagnostic results and sends them to the terminal in JSON format or another appropriate format.

[0902] Output: Final diagnostic results sent to the terminal

[0903] Step 10:

[0904] The terminal displays the diagnostic results to the user.

[0905] Input: Final diagnostic results received from the server

[0906] Operation: The user interface visually presents evaluation results of exercise form, areas for improvement, and emotion-based feedback. For example, it might display, "Your movement is generally good. Be careful as your body tends to move forward at the finish."

[0907] Output: Diagnostic results and feedback displayed to the user

[0908] Example of a prompt

[0909] "Please upload a video of your swing. We will analyze your swing form and your emotions, and provide feedback."

[0910] (Application Example 2)

[0911] Next, we will explain application example 2. In the following explanation, the data processing device 12 will be referred to as the "server," and the smart glasses 214 will be referred to as the "terminal."

[0912] Traditional swing analysis systems only evaluate the user's swing form and do not consider feedback based on the user's emotions, thus limiting their effectiveness in improving user motivation and reducing stress. Furthermore, they lacked the functionality to provide specific training methods, preventing them from directly contributing to user skill improvement.

[0913] The specific processing performed by the specific processing unit 290 of the data processing device 12 in Application Example 2 is realized by the following means.

[0914] In this invention, the server includes means for analyzing the user's facial expressions to identify their emotional state, means for adjusting the swing diagnosis results based on the emotional state, and means for providing specific training methods based on an evaluation of the swing form. This allows the user to receive flexible feedback tailored to their emotional state and specific methods for improvement.

[0915] A "user" is an individual or group that uses the system, uploads videos or images of their swing, and receives the analysis results.

[0916] A "video" is video data composed of a series of image frames, and is used to record the user's swing motion.

[0917] An "image" is visual data represented in a single still image format and is used to capture a specific scene of a swing.

[0918] A "scene" refers to a visual representation of each stage of the swing (address, backswing, top, impact, finish).

[0919] A "diagnostic algorithm" is a set of calculation procedures and rules used to analyze and evaluate swing form, and the system uses them to scrutinize each aspect of the swing.

[0920] "Facial expression" refers to the various muscle movements and arrangements that appear on a user's face, and serves as basic information for inferring their emotional state.

[0921] "Emotional state" refers to the internal psychological state analyzed from the user's facial expressions, and includes classifications such as joy, surprise, sadness, and anger.

[0922] "Training methods" refer to specific practice procedures and exercises aimed at improving swing form and technique.

[0923] In an embodiment for carrying out this invention, the system includes the following components.

[0924] User

[0925] A "user" is an individual or group that uses the system, and is the entity that takes videos or images of their swing using a device such as a smartphone and uploads them. Users can receive not only technical evaluations of their swing form, but also emotional feedback derived from their facial expressions.

[0926] terminal

[0927] A "device" is a device used by a user to upload data, and includes smartphones and tablets. A device utilizes the following hardware and software:

[0928] Hardware: Smartphones (iPhone, Android)

[0929] Software: Dedicated application (Golf Swing Master)

[0930] Its role is to handle everything from video recording to data transmission, and then receive diagnostic results from the server and display them to the user.

[0931] server

[0932] A "server" is a central processing unit that receives data transmitted from terminals and performs analysis and diagnosis. The server fulfills the following roles:

[0933] Swing Analysis: Frames are extracted from the video, and a diagnostic algorithm is executed to analyze each scene (address, backswing, top, impact, finish).

[0934] Emotion Recognition: Analyzes the user's facial expressions to identify their emotional state and reflect it in the swing analysis results.

[0935] Result Adjustment: The diagnostic results are adjusted based on emotional data to provide users with appropriate feedback.

[0936] Emotional Engine

[0937] The "emotion engine" is a software module that analyzes a user's facial expressions and classifies their emotional state. The emotion engine is used to understand the user's psychological state when uploading swing videos and images.

[0938] Training provided

[0939] The server provides specific training methods based on an evaluation of swing form. This allows users to learn concrete practical techniques that help improve their skills.

[0940] Examples

[0941] For example, a user films their swing at a golf driving range with their smartphone and uploads the video to a dedicated app. The device sends the video data to a server, which analyzes the video to evaluate the swing form and uses an emotion engine to analyze the user's facial expressions to identify their emotional state. As a result, in addition to the technical evaluation, the server provides feedback and training methods tailored to the user's emotional state. This feedback might include something like, "Your swing is good, but try to relax a little more. Refer to this video and practice swinging in a more relaxed state."

[0942] Examples of prompts for generative AI models

[0943] Create an application that analyzes a user's golf swing video and facial expressions to provide a technical evaluation of their swing form and emotionally responsive feedback.

[0944] 1. How to extract frames from a video.

[0945] 2. A method for identifying emotions by analyzing facial expressions.

[0946] 3. How to send swing analysis results and emotional feedback to the server.

[0947] 4. How to receive diagnostic results and how to display them in the user interface.

[0948] The flow of a specific process in Application Example 2 will be explained using Figure 14.

[0949] Step 1:

[0950] The user simultaneously records a video of their golf swing and their facial expressions using their smartphone.

[0951] Input: Video to be recorded and facial expressions

[0952] Output: Video files saved on the smartphone's storage.

[0953] Specific operation: The user uses the smartphone's camera app to record a video, and at the same time, their facial expressions are also recorded using the front camera.

[0954] Step 2:

[0955] The user opens the application, selects the video they have recorded, and uploads it.

[0956] Input: User-recorded video

[0957] Output: Video file uploaded to the application

[0958] Specific steps: Open the file selection screen within the application, select the recorded video file, and press the upload button.

[0959] Step 3:

[0960] The device sends the selected video file to the server.

[0961] Input: Uploaded video file

[0962] Output: Video data sent to the server

[0963] Specific action: Send the video file to the server using an HTTP POST request.

[0964] Step 4:

[0965] The server receives the video file and extracts frames for each scene of the swing.

[0966] Input: Video data sent from the device

[0967] Output: Frame images of each extracted scene

[0968] Specific operation: Use OpenCV to extract frames from the video and separate the address, backswing, top of swing, impact, and finish scenes.

[0969] Step 5:

[0970] The server preprocesses each frame of the extracted scene.

[0971] Input: Extracted frame image

[0972] Output: Preprocessed image data

[0973] Specific actions: Standardize frames, unify image sizes, and remove noise.

[0974] Step 6:

[0975] The server inputs the pre-processed images into the diagnostic algorithm.

[0976] Input: Preprocessed image data

[0977] Output: Swing form evaluation results

[0978] Specific operation: Pre-processed images are input into the diagnostic algorithm, and a technical evaluation of the swing form is performed.

[0979] Step 7:

[0980] The emotion engine analyzes the user's facial expressions to identify their emotional state.

[0981] Input: Facial expression data included in the video

[0982] Output: Classified emotional state

[0983] Specific operation: Facial expression data is input into an emotion recognition model to classify emotions such as joy, surprise, sadness, and anger.

[0984] Step 8:

[0985] The server adjusts the swing analysis results based on the output of the emotion engine.

[0986] Input: Swing form evaluation results and emotional state

[0987] Output: Final diagnostic results including emotional feedback

[0988] Specific actions: Consider emotional data and incorporate emotional feedback into the diagnostic results, generating feedback such as, "Your swing is good, but try to relax a little more."

[0989] Step 9:

[0990] The server generates the final diagnostic results and sends them back to the terminal.

[0991] Input: Final diagnostic results including emotional feedback

[0992] Output: Diagnostic results sent to the terminal

[0993] Specific operation: The final diagnostic results are sent back to the terminal using an HTTP response.

[0994] Step 10:

[0995] The terminal displays the diagnostic results to the user.

[0996] Input: Diagnostic results received from the server

[0997] Output: Diagnostic results displayed in the user interface

[0998] Specific action: Display the received feedback and swing evaluation in the application's user interface.

[0999] The specific processing unit 290 transmits the result of the specific processing to the smart glasses 214. In the smart glasses 214, the control unit 46A causes the speaker 240 to output the result of the specific processing. The microphone 238 acquires audio indicating user input for the result of the specific processing. The control unit 46A transmits the audio data indicating user input acquired by the microphone 238 to the data processing unit 12. In the data processing unit 12, the specific processing unit 290 acquires the audio data.

[1000] Data generation model 58 is a type of so-called generative AI (Artificial Intelligence). One example of data generation model 58 is ChatGPT (Internet search<URL: https: / / openai.com / blog / chatgpt> ), Gemini (Internet search) <url: https: gemini.google.com ?hl="ja">Examples of generative AI include the following. The data generation model 58 is obtained by performing deep learning on a neural network. The data generation model 58 is input with prompts containing instructions, and with inference data such as audio data representing speech, text data representing text, and image data representing images. The data generation model 58 infers from the input inference data according to the instructions indicated by the prompts, and outputs the inference results in data formats such as audio data and text data. Here, inference refers to, for example, analysis, classification, prediction, and / or summarization.

[1001] In the above embodiment, an example was given in which specific processing is performed by the data processing device 12, but the technology of this disclosure is not limited thereto, and the specific processing may also be performed by the smart glasses 214.

[1002] [Third Embodiment]

[1003] Figure 5 shows an example of the configuration of the data processing system 310 according to the third embodiment.

[1004] As shown in Figure 5, the data processing system 310 includes a data processing device 12 and a headset terminal 314. An example of the data processing device 12 is a server.

[1005] The data processing device 12 comprises a computer 22, a database 24, and a communication interface 26. The computer 22 is an example of a "computer" related to the technology of this disclosure. The computer 22 comprises a processor 28, RAM 30, and storage 32. The processor 28, RAM 30, and storage 32 are connected to a bus 34. The database 24 and the communication interface 26 are also connected to the bus 34. The communication interface 26 is connected to a network 54. An example of the network 54 is a WAN (Wide Area Network) and / or a LAN (Local Area Network).

[1006] The headset terminal 314 includes a computer 36, a microphone 238, a speaker 240, a camera 42, a communication interface 44, and a display 343. The computer 36 includes a processor 46, RAM 48, and storage 50. The processor 46, RAM 48, and storage 50 are connected to a bus 52. The microphone 238, speaker 240, camera 42, and display 343 are also connected to the bus 52.

[1007] The microphone 238 receives voice signals from the user 20 and receives instructions from the user 20. The microphone 238 captures the voice signals from the user 20, converts the captured voice into audio data, and outputs it to the processor 46. The speaker 240 outputs audio according to the instructions from the processor 46.

[1008] Camera 42 is a small digital camera equipped with an optical system including a lens, aperture, and shutter, and an image sensor such as a CMOS (Complementary Metal-Oxide-Semiconductor) image sensor or a CCD (Charge Coupled Device) image sensor, and captures images of the area around the user 20 (for example, an imaging range defined by a field of view equivalent to the width of a typical healthy person's field of vision).

[1009] Communication interface 44 is connected to network 54. Communication interfaces 44 and 26 are responsible for the exchange of various information between processor 46 and processor 28 via network 54. The exchange of various information between processor 46 and processor 28 using communication interfaces 44 and 26 is performed in a secure manner.

[1010] Figure 6 shows an example of the main functions of the data processing device 12 and the headset terminal 314. As shown in Figure 6, the data processing device 12 performs specific processing using the processor 28. The storage 32 stores the specific processing program 56.

[1011] The specific processing program 56 is an example of a "program" relating to the technology of this disclosure. The processor 28 reads the specific processing program 56 from the storage 32 and executes the read specific processing program 56 on the RAM 30. The specific processing is realized by the processor 28 operating as a specific processing unit 290 in accordance with the specific processing program 56 executed on the RAM 30.

[1012] The storage 32 stores the data generation model 58 and the emotion identification model 59. The data generation model 58 and the emotion identification model 59 are used by the identification processing unit 290.

[1013] In the headset terminal 314, the processor 46 performs the reception output processing. The storage 50 stores the reception output program 60. The processor 46 reads the reception output program 60 from the storage 50 and executes the read reception output program 60 on the RAM 48. The reception output processing is realized by the processor 46 operating as a control unit 46A according to the reception output program 60 executed on the RAM 48.

[1014] Next, the specific processing performed by the specific processing unit 290 of the data processing device 12 will be described. In the following description, the data processing device 12 will be referred to as the "server" and the headset terminal 314 will be referred to as the "terminal".

[1015] The present invention provides a system for which a user uploads a video or image of their swing, and for which the swing form is analyzed and diagnosed. The operation of each element and the program's processing are described below in natural language.

[1016] System Overview

[1017] This system consists of three main components: users, terminals, and servers.

[1018] User

[1019] Users record videos or images of their golf swing using a device such as a smartphone or camera. After recording, they access a swing analysis application or web interface and select the file.

[1020] terminal

[1021] The terminal's role is to send the user-selected video or image file to the server. Communication methods such as HTTP POST requests are used for transmission. Furthermore, the terminal displays the diagnostic results received from the server on the user interface.

[1022] server

[1023] The server plays a central role in receiving and analyzing videos or images sent from the terminal. If a video is sent, the server analyzes the frames and extracts each scene of the swing (address, backswing, top, impact, finish). It then runs a swing diagnostic algorithm on the extracted scenes or images to generate a diagnostic result. The generated diagnostic result is then sent back to the terminal.

[1024] Specific examples of program processing

[1025] 1. The user films their golf swing with their smartphone.

[1026] Use your smartphone's camera function to record your swing motion in video format.

[1027] 2. The user opens the application, selects the recorded video, and uploads it.

[1028] Use the application's file selection screen to upload the recorded video file to the designated location.

[1029] 3. The device sends the selected video file to the server.

[1030] File transmission is performed using HTTP POST requests or other appropriate communication methods.

[1031] 4. The server receives the video and extracts frames for each scene of the swing.

[1032] The system analyzes the video content and executes an algorithm to extract specific scenes (address, backswing, top of swing, impact, finish).

[1033] 5. The server analyzes the images from each scene and evaluates them based on the swing diagnostic algorithm.

[1034] The system analyzes the swing form for each extracted scene image. The diagnostic algorithm evaluates the accuracy of posture and movement, identifying abnormalities and areas for improvement.

[1035] 6. The server generates diagnostic results and creates feedback such as "Your body is unbalanced at address" and "Your timing of impact is slow."

[1036] Based on the output of the diagnostic algorithm, specific feedback is generated to be provided to the user.

[1037] 7. The server sends the diagnostic results back to the smartphone.

[1038] The generated diagnostic results are sent back to the terminal as an HTTP response.

[1039] 8. The terminal displays the diagnostic results to the user and indicates areas for improvement.

[1040] The diagnostic results are displayed in the user interface, visually and textually presenting areas for improvement to the user. Annotations are added to the images in each scene to clearly indicate specific problems.

[1041] This system allows users to efficiently and objectively analyze their golf swing and visually understand areas for improvement without requiring expensive specialized equipment or expertise.

[1042] The following describes the processing flow.

[1043] Step 1: The user records a video of their golf swing using their smartphone.

[1044] Users record their golf swing as a video using their smartphone's camera function.

[1045] Step 2: The user opens the application and selects the recorded video.

[1046] Display the application's file selection screen and select the video file you recorded.

[1047] Step 3: The user uploads a video file.

[1048] Press the upload button in the application to send the selected video file to the server.

[1049] Step 4: The device sends the video file to the server.

[1050] The device uses an HTTP POST request to send the video file to the server.

[1051] Step 5: The server receives the video file.

[1052] The server retrieves and saves the video file from the received HTTP POST request.

[1053] Step 6: The server analyzes the video and extracts each scene.

[1054] The server analyzes the video frames to identify each scene—address, backswing, top of swing, impact, and finish—and extracts the corresponding frames.

[1055] Step 7: The server preprocesses the frames of each extracted scene.

[1056] The server standardizes the frames of each scene and converts them into a format that allows image analysis algorithms to function properly. For example, it performs image size unification and noise reduction.

[1057] Step 8: The server inputs the pre-processed images into the diagnostic algorithm.

[1058] The server inputs the pre-processed images of each scene into a swing diagnostic algorithm for evaluation.

[1059] Step 9: The server generates a diagnostic result based on the image analysis results.

[1060] Note) The server evaluates the analysis results and identifies anomalies in the swing form for each scene. It generates feedback on the identified anomalies and compiles them into a diagnostic report.

[1061] Step 10: The server sends the diagnostic results back to the terminal.

[1062] The server returns the diagnostic results to the terminal as an HTTP response.

[1063] Step 11: The terminal displays the diagnostic results to the user.

[1064] The terminal displays the received diagnostic results on the user interface, showing the user specific areas for improvement and evaluation of their swing form.

[1065] This processing step allows users to efficiently analyze their golf swing and receive specific feedback.

[1066] (Example 1)

[1067] Next, we will describe Example 1. In the following description, the data processing device 12 will be referred to as the "server," and the headset-type terminal 314 will be referred to as the "terminal."

[1068] Conventional technologies required expensive specialized equipment and the assistance of skilled technicians to analyze golf swings, making it difficult for ordinary users to easily evaluate and improve their swing form. Furthermore, there was a lack of methods to provide swing form diagnostic results in an intuitively understandable format. Therefore, there was a need for a technology that could analyze and evaluate golf swings easily and at low cost, allowing users to clearly identify areas for improvement themselves.

[1069] The identification process performed by the identification processing unit 290 of the data processing device 12 in Example 1 is realized by the following means.

[1070] In this invention, the server includes means for a user to upload a video or image of their swing, means for analyzing and extracting each scene (address, backswing, top, impact, finish) frame by frame from the uploaded video, means for executing a diagnostic algorithm to analyze the extracted scenes and evaluate the swing form, means for returning and displaying the generated diagnostic results to the user, means for generating specific feedback (e.g., issues with body balance or timing) based on the diagnostic results, and means for adding annotations to the images of each scene as feedback to indicate areas for improvement. This makes it possible for users to analyze their golf swing in detail and intuitively understand specific areas for improvement without using expensive specialized equipment.

[1071] A "user" refers to a person who uses the system to film their own golf swing and obtain the analysis results.

[1072] "Swing video or image" refers to a video file or still image file recorded by the user of their golf swing motion.

[1073] "Means for uploading" refers to the interface or function that allows users to provide videos or images they have taken to the system.

[1074] "Each scene" refers to the specific phases of a golf swing: address, backswing, top, impact, and finish.

[1075] "Means for analyzing and extracting frame by frame" refers to technologies and devices that divide a video into multiple still images and identify a specific scene from each frame.

[1076] A "diagnostic algorithm" refers to a calculation method or program that analyzes extracted scene images to evaluate golf swing form.

[1077] The means for returning and displaying the "generated diagnostic results" to the user refers to the communication method and display function for returning the analysis results from the diagnostic algorithm and making them viewable by the user.

[1078] "Specific feedback" refers to detailed observations and advice based on the diagnostic results, such as issues with body balance or timing.

[1079] "Annotation" refers to annotations, marks, and visual information added to images in each scene to indicate areas for improvement.

[1080] The present invention provides a system for users to record their golf swings as videos or images, and to analyze and evaluate those swings. This system consists of three main components: the user, the terminal, and the server.

[1081] User operations and terminal functions

[1082] Users record their golf swing as a video or image using a device such as a smartphone or tablet. They utilize the device's camera function for this purpose. For example, they can use the iPhone's camera app to shoot high-resolution video. Afterward, the user launches a dedicated application such as "Golf Swing Analyzer," selects the recorded video through the application's interface, and uploads it.

[1083] The terminal is responsible for sending the video file selected by the user to the server. Common communication technologies such as HTTP POST requests are used for this transmission. During transmission, the appropriate communication protocol is selected considering the file format and size. Furthermore, the terminal displays the diagnostic results received from the server on the user interface. Therefore, the terminal requires visual and textual display functions that allow the user to intuitively understand the diagnostic results.

[1084] Server analysis and diagnostic functions

[1085] The server plays a central role in receiving and analyzing videos or images sent from the terminal. After receiving the data, the server uses image analysis libraries such as OpenCV to divide the video into frames and extract each scene of the swing (address, backswing, top, impact, finish). A specific algorithm is used to extract each scene.

[1086] For each extracted scene, the server uses a generative AI model such as TensorFlow to evaluate the swing form. This diagnostic algorithm assesses the accuracy of posture and movement, and identifies abnormalities and areas for improvement. For example, it can evaluate things like "the body is unbalanced at address" or "the timing of impact is late."

[1087] Examples of specific cases and prompt statements

[1088] Further implementation examples include the following: A user uploads a video of their golf swing, the server analyzes it, and sends specific feedback back to the user's device. This feedback is displayed as annotations for each scene in the video, making it easy for the user to understand. This system allows users to obtain analysis and evaluation of their golf swing without needing expensive specialized equipment.

[1089] Examples of prompt messages are as follows:

[1090] Design a system that automatically analyzes the movements in each scene of a golf swing video recorded by a user on their iPhone, and diagnoses the swing form. Describe the specific analysis method and algorithm, and clearly specify the communication method between the server and client.

[1091] This invention allows users to efficiently and accurately analyze their golf swing and identify areas for improvement without requiring specialized knowledge or expensive equipment.

[1092] The flow of the specific processing in Example 1 will be explained using Figure 11.

[1093] Program processing flow

[1094] Step 1:

[1095] The user films their golf swing.

[1096] (Specific actions)

[1097] Users record their golf swings in video format using their smartphones or tablets. They use a camera app to capture the entire swing within the frame.

[1098] (input)

[1099] Videos filmed by users.

[1100] (output)

[1101] A video of a swing saved on a smartphone.

[1102] Step 2:

[1103] The user opens the application, selects the video they have recorded, and uploads it.

[1104] (Specific actions)

[1105] The user launches a dedicated app such as "Golf Swing Analyzer" and navigates to the "file selection screen" within the app's interface. The user selects the recorded video and clicks the "upload button".

[1106] (input)

[1107] A video of a swing saved on a smartphone.

[1108] (output)

[1109] The selected video will be displayed within the application, and an upload request will be generated.

[1110] Step 3:

[1111] The device sends the video file to the server.

[1112] (Specific actions)

[1113] The device sends the video file to the server using an HTTP POST request. During this process, the video file is appropriately split and transmitted according to the communication protocol.

[1114] (input)

[1115] The selected video file within the application, and the upload request.

[1116] (output)

[1117] Video file being sent to the server.

[1118] Step 4:

[1119] The server receives the video and analyzes it frame by frame.

[1120] (Specific actions)

[1121] The server saves the received video file and uses the OpenCV library to split the video into frames. Each frame is sequentially loaded into memory and used for analysis.

[1122] (input)

[1123] The video file sent to the server.

[1124] (output)

[1125] Video data divided into individual frames.

[1126] Step 5:

[1127] The server extracts each scene and applies a swing analysis algorithm.

[1128] (Specific actions)

[1129] The server analyzes each frame and detects specific scenes (address, backswing, top of trajectory, impact, finish). After this, a TensorFlow-based generative AI model is used to evaluate each scene.

[1130] (input)

[1131] Video data divided into individual frames.

[1132] (output)

[1133] Data for each detected scene and its diagnostic results.

[1134] Step 6:

[1135] The server generates diagnostic results and creates feedback for the user.

[1136] (Specific actions)

[1137] The server generates feedback based on the analysis results. For example, it creates diagnostic statements such as "body balance is off at address" or "impact timing is slow," and formats them in JSON format.

[1138] (input)

[1139] Data for each detected scene and its diagnostic results.

[1140] (output)

[1141] User feedback data.

[1142] Step 7:

[1143] The server sends the generated diagnostic results back to the user's terminal.

[1144] (Specific actions)

[1145] The server sends the generated diagnostic results back to the user's terminal as an HTTP response. The returned data is sent in an appropriate format, such as JSON or XML.

[1146] (input)

[1147] User feedback data.

[1148] (output)

[1149] Diagnostic results sent to the user's terminal.

[1150] Step 8:

[1151] The device displays the diagnostic results.

[1152] (Specific actions)

[1153] The terminal analyzes the received diagnostic results and displays them in the user interface. Because the user can visually confirm the diagnostic results, specific annotations are added for each scene.

[1154] (input)

[1155] Diagnostic results sent from the server.

[1156] (output)

[1157] Diagnostic results and annotations displayed in the user interface.

[1158] (Application Example 1)

[1159] Next, we will explain Application Example 1. In the following explanation, the data processing device 12 will be referred to as the "server," and the headset-type terminal 314 will be referred to as the "terminal."

[1160] Until now, there had been no system that could efficiently and quickly analyze golf swing form and provide users with useful feedback in real time. As a result, users needed expert advice or expensive equipment to improve their swing form. This meant that many amateur golfers did not have enough opportunities for self-improvement. Furthermore, because the feedback was not in real time, users could not immediately obtain information for improvement, which reduced the efficiency of their practice. In addition, the lack of a feedback system using wearable devices made it difficult for users to check swing improvement measures on the spot.

[1161] The specific processing performed by the specific processing unit 290 of the data processing device 12 in Application Example 1 is realized by the following means.

[1162] In this invention, the server includes means for a user to upload a video or image of their swing, means for extracting each scene (address, backswing, top, impact, finish) from the uploaded video, means for executing a diagnostic algorithm to analyze the extracted scenes and evaluate the swing form, and means for providing real-time feedback of the displayed diagnostic results to a wearable device worn by the user. This enables real-time feedback of golf swing form analysis and diagnostic results, allowing the user to immediately see areas for improvement and increase the efficiency of their practice.

[1163] A "user" is an individual or group that uses the system to upload videos or images of their golf swing and receive analysis and diagnostic results.

[1164] A "video" is a dynamic media file that displays a series of images taken by a user in sequence.

[1165] An "image" is a media file that includes still images taken by the user.

[1166] "Uploading" refers to the act of a user transferring video or image data from their device to a server.

[1167] A "scene" refers to a specific moment or stage in a swing, specifically the address, backswing, top of the swing, impact, and finish.

[1168] "Extraction" refers to taking a specific scene from an uploaded video.

[1169] "Analysis" refers to evaluating the extracted images and videos of scenes using a diagnostic algorithm.

[1170] A "diagnostic algorithm" is a specific set of calculation procedures or analytical models used to evaluate swing form and generate diagnostic results.

[1171] "Generation" refers to creating diagnostic results based on the analyzed data.

[1172] "Display" refers to outputting the generated diagnostic results to the terminal in a format that the user can review.

[1173] A "wearable device" refers to a device that a user can wear, typically in the form of glasses, a head-mounted display, or a wristband.

[1174] "Real-time" refers to processing and feedback being performed instantly without delay.

[1175] "Feedback" is the process of returning the generated diagnostic results and areas for improvement to the user.

[1176] This invention is a system that analyzes videos and images of golf swings to diagnose swing form, and is designed to be easily accessible to users. The main components are the user, a terminal, a server, and a wearable device.

[1177] User

[1178] Users record their golf swing using a device such as a smartphone or smart glasses, capturing it as a video or image. After recording, they upload the video or image using a dedicated application. This allows the system to input the user's swing form.

[1179] terminal

[1180] The device is used to send videos and images taken by the user to the server. HTTP POST requests are typically used for transmission. Furthermore, the device displays the diagnostic results received from the server on the user interface. Examples of devices include smartphones, smart glasses, and head-mounted displays.

[1181] server

[1182] The server plays a central role in analyzing videos and images sent from the terminal and evaluating the swing form. Specifically, it performs the following processes:

[1183] 1. When a video is sent, the server analyzes it frame by frame and extracts each scene of the swing (address, backswing, top, impact, finish).

[1184] 2. A diagnostic algorithm is run on the extracted scenes to evaluate the swing form of each scene. The software used includes OpenCV (image processing), TensorFlow and PyTorch (machine learning algorithms), and Flask (web framework).

[1185] 3. Generate diagnostic results and create feedback to provide to the user. This feedback will include specific areas for improvement and problems, and will be displayed in text and annotation format.

[1186] Wearable devices

[1187] Wearable devices are devices that users wear to receive real-time feedback. Examples include smart glasses and head-mounted displays. Diagnostic results and feedback transmitted from a server are displayed on these wearable devices in real time.

[1188] Specific example

[1189] As a concrete example, imagine a scenario where a user wears smart glasses at a golf driving range and performs a swing. The smart glasses film the user's swing in real time and send the video to a server. The server analyzes the video and diagnoses the swing form. As a result, feedback such as "Your body balance is off at address" or "Your top position is too high" is immediately displayed on the smart glasses' screen. The user can then correct their swing form on the spot, film again, and receive further evaluation.

[1190] Examples of prompts to input into a generative AI model

[1191] "Design an application that allows users to film their golf swing with smart glasses and receive real-time form analysis and diagnostic results. The system would involve filming the frame, uploading the footage to a server, and displaying the analyzed results on the smart glasses. Please provide a detailed explanation, including the necessary hardware and software, and the program's processing steps."

[1192] The flow of a specific process in Application Example 1 will be explained using Figure 12.

[1193] Step 1:

[1194] The user takes a video or image of their swing.

[1195] Input: A video or image of a swing taken using the camera of a smartphone or smart glasses.

[1196] Operation: The user performs a golf swing, and the camera records the motion. In the case of video, a series of movements are recorded continuously; in the case of still images, specific moments are saved as still images.

[1197] Output: The recorded video or image of the golf swing is saved on the device.

[1198] Step 2:

[1199] Users upload videos or images they have taken to the server via a dedicated application.

[1200] Input: Swing video or image taken after shooting.

[1201] Operation: The user opens the dedicated application, selects a video or image from the file selection screen, and presses the upload button. The file is sent to the server using a data transfer protocol such as an HTTP POST request.

[1202] Output: Video or image data is sent to the server.

[1203] Step 3:

[1204] The server analyzes the received video or images and extracts each scene of the swing.

[1205] Input: A swing video or image uploaded to the server.

[1206] Operation: The server divides the video frame by frame and detects important scenes in the swing (address, backswing, top, impact, finish). This analysis and extraction is performed using image processing libraries such as OpenCV.

[1207] Output: Frame images of each extracted scene.

[1208] Step 4:

[1209] The server analyzes each extracted scene and executes a diagnostic algorithm to evaluate the swing form.

[1210] Input: Frame images from each extracted scene.

[1211] Operation: The server uses machine learning models (e.g., TensorFlow, PyTorch) to evaluate the swing form. It analyzes the accuracy of posture and movement, and identifies abnormalities and areas for improvement.

[1212] Output: Swing form evaluation results for each scene.

[1213] Step 5:

[1214] The server generates diagnostic results based on the swing form evaluation and creates feedback to provide to the user.

[1215] Input: Swing form evaluation results.

[1216] Operation: The server generates specific feedback based on the evaluation results. For example, it might create text-based feedback such as "Your body balance is off at address" or "Your top position is too high."

[1217] Output: Generated diagnostic results and feedback.

[1218] Step 6:

[1219] The server sends diagnostic results to the terminal and wearable device, and displays feedback in real time.

[1220] Input: Generated diagnostic results and feedback.

[1221] Operation: The server sends the diagnostic results to the device (smartphone, smart glasses, etc.). The results are returned as an HTTP response, and in the case of wearable devices, the data is transferred via a dedicated API. The device or wearable device displays the received data.

[1222] Output: Diagnostic results and feedback displayed on the user's device or wearable device.

[1223] Step 7:

[1224] We will work on improving the swing based on the feedback displayed to the user.

[1225] Input: Diagnostic results and feedback displayed on a terminal or wearable device.

[1226] Operation: The user reviews the displayed feedback and implements measures to improve their swing form. If necessary, they film their swing again and use the system for further diagnosis.

[1227] Output: Improved swing form, revised diagnostic results.

[1228] Furthermore, an emotion engine that estimates the user's emotions may be incorporated. That is, the identification processing unit 290 may use the emotion identification model 59 to estimate the user's emotions and perform identification processing using the user's emotions.

[1229] The system of the present invention not only allows users to upload videos or images of their swings and analyzes and diagnoses their swing form, but also includes a function to recognize the user's emotions and adjust the diagnosis results based on those emotions. The operation of each element and the processing of the program are described below in natural language.

[1230] System Overview

[1231] This system consists of four main components: the user, the terminal, the server, and the emotion engine.

[1232] User

[1233] Users record videos or images of their golf swing using a device such as a smartphone or camera. After recording, they access a swing analysis application or web interface to select the file. In addition, the user's facial expressions are also recorded during the upload process.

[1234] terminal

[1235] The terminal's role is to send the video or image file selected by the user to the server. Communication methods such as HTTP POST requests are used for transmission. The terminal also displays the diagnostic results received from the server on the user interface.

[1236] server

[1237] The server plays a central role in receiving and analyzing videos or images sent from the terminal. If a video is sent, the server analyzes the frames and extracts each scene of the swing (address, backswing, top, impact, finish). It then runs a swing diagnostic algorithm on the extracted scenes and images to generate a diagnostic result. It also analyzes the user's emotional data and incorporates it into the diagnostic result.

[1238] Emotional Engine

[1239] The emotion engine analyzes the user's facial expressions to identify their emotional state. This allows the system to understand the user's emotions when they upload swing videos or images and adjust the swing analysis results accordingly.

[1240] Specific examples of program processing

[1241] 1. The user simultaneously records a video of their golf swing and their facial expressions using their smartphone.

[1242] The smartphone's camera function is used to simultaneously record the golf swing and facial expressions.

[1243] 2. The user opens the application, selects the recorded video, and uploads it.

[1244] Use the application's file selection screen to upload the recorded video file to the designated location.

[1245] 3. The device sends the selected video file to the server.

[1246] File transmission is performed using HTTP POST requests or other appropriate communication methods.

[1247] 4. The server receives the video file and extracts frames for each scene of the swing.

[1248] The system analyzes the video content and executes an algorithm to extract specific scenes (address, backswing, top of swing, impact, finish).

[1249] 5. The server preprocesses the frames of each extracted scene.

[1250] The frames of each scene are standardized and converted into a format that allows image analysis algorithms to function properly. For example, this may involve unifying image sizes and removing noise.

[1251] 6. The server inputs the pre-processed images into the diagnostic algorithm.

[1252] The images from each scene, after preprocessing is complete, are input into the swing analysis algorithm for evaluation.

[1253] 7. The emotion engine analyzes the user's facial expressions and identifies their emotional state.

[1254] The system analyzes captured facial expression data to classify emotions such as joy, surprise, sadness, and anger.

[1255] 8. The server adjusts the swing analysis results based on the output of the emotion engine.

[1256] The system takes emotional data into account and adds emotional feedback to the diagnostic results. For example, if the user has an anxious expression, it provides more positive feedback.

[1257] 9. The server generates the final diagnostic results and sends them back to the terminal.

[1258] The diagnostic results are sent back to the terminal as an HTTP response.

[1259] 10. The terminal displays the diagnostic results to the user.

[1260] The device displays the received diagnostic results in the user interface, showing the user specific areas for improvement in their swing form, evaluations, and emotion-based feedback.

[1261] This system allows users to efficiently analyze their golf swing and receive specific feedback, as well as flexible advice tailored to their emotions. This is expected to contribute to increased user motivation and reduced stress.

[1262] The following describes the processing flow.

[1263] Step 1: The user simultaneously films their golf swing and facial expression using their smartphone.

[1264] Users use their smartphone's camera function to record themselves and their facial expressions while performing a golf swing in video format.

[1265] Step 2: The user opens the application and selects the recorded video.

[1266] The user opens the swing analysis application and selects the recorded video file.

[1267] Step 3: The user uploads a video file.

[1268] The user presses the upload button within the application and sends the selected video file to the server.

[1269] Step 4: The device sends the video file to the server.

[1270] The device uses an HTTP POST request to send the video file to the server.

[1271] Step 5: The server receives the video file.

[1272] The server retrieves and saves the video file from the received HTTP POST request.

[1273] Step 6: The server analyzes the video and extracts each scene.

[1274] The server analyzes the video frames to identify each scene—address, backswing, top of swing, impact, and finish—and extracts the corresponding frames.

[1275] Step 7: The server preprocesses the frames of each extracted scene.

[1276] The server standardizes the frames of each scene and converts them into a format that allows image analysis algorithms to function properly. For example, it performs image size unification and noise reduction.

[1277] Step 8: The server inputs the pre-processed images into the diagnostic algorithm.

[1278] The server inputs the pre-processed images of each scene into a swing diagnostic algorithm for evaluation.

[1279] Step 9: The emotion engine analyzes the user's facial expressions.

[1280] An emotion engine installed on the server analyzes the user's facial expressions in the video and classifies them into emotions such as joy, surprise, sadness, and anger.

[1281] Step 10: The server adjusts the swing diagnosis results based on the output of the emotion engine.

[1282] The server considers emotional data obtained from facial expressions and includes emotion-based feedback in the generated swing diagnosis results. For example, if the user has an anxious expression, it will provide more positive feedback.

[1283] Step 11: The server generates the final diagnostic results and sends them back to the terminal.

[1284] The server returns the adjusted diagnostic results to the terminal as an HTTP response.

[1285] Step 12: The terminal displays the diagnostic results to the user.

[1286] The device displays the received diagnostic results in the user interface, visually showing specific areas for improvement and evaluation of the swing form, as well as feedback based on the user's emotions.

[1287] These detailed processing steps allow users to efficiently analyze their golf swing and receive specific, emotionally responsive feedback. This provides not only improvement to their swing form but also mental support.

[1288] (Example 2)

[1289] Next, we will describe Example 2. In the following description, the data processing device 12 will be referred to as the "server," and the headset-type terminal 314 will be referred to as the "terminal."

[1290] Conventional exercise form analysis systems fail to consider the user's emotional state during video analysis and evaluation, sometimes leading to user stress or decreased motivation. Furthermore, the lack of specific feedback for each scene of exercise form makes it difficult for users to identify areas for improvement. Therefore, there is a need to develop a system that recognizes the user's emotions and reflects them in the diagnostic results, thereby providing more effective feedback and improving user motivation.

[1291] The identification process performed by the identification processing unit 290 of the data processing device 12 in Example 2 is realized by the following means.

[1292] In this invention, the server includes means for a user to upload a video or image of exercise, means for extracting each scene (start, in motion, end) from the uploaded video, means for executing a diagnostic algorithm to analyze the extracted scenes and evaluate the exercise form, means for analyzing facial expressions to identify the emotional state, means for adjusting the diagnostic results considering the emotional state, and means for generating the diagnostic results and displaying them to the user. This enables more effective and motivating feedback for the user by analyzing the user's emotions and reflecting them in the diagnostic results of the exercise form.

[1293] A "user" refers to someone who uses the system to upload videos or images of their exercise and receives the analysis results.

[1294] "Movement" refers to specific actions such as a golf swing, and is the subject of analysis by the system.

[1295] "Video or image" refers to data in file format in which a user records their exercise movements.

[1296] "Means for uploading" refers to methods and devices for sending videos and images taken by users to a server.

[1297] "Means for extraction" refers to algorithms or methods used to select specific scenes from uploaded videos.

[1298] A "diagnostic algorithm" refers to an algorithm used to analyze extracted scenes and evaluate their movement form.

[1299] "Means for analyzing facial expressions" refers to methods or devices for analyzing a user's facial expressions to identify their emotional state.

[1300] "Means of adjusting diagnostic results by considering emotional state" refers to methods or algorithms that modify or correct diagnostic results based on the analyzed emotional state.

[1301] "Means for generating diagnostic results" refers to methods or devices for compiling and providing users with evaluation results of exercise form.

[1302] "Means for display" refers to methods or devices for visually presenting the generated diagnostic results to the user.

[1303] A "scene" refers to any frame in a video of an exercise in which a specific action takes place.

[1304] "Start," "In Progress," and "End" refer to the scenes in the video that indicate the first, middle, and final stages of the exercise, respectively.

[1305] The system of the present invention not only allows users to upload videos or images of their exercise and analyzes and diagnoses their exercise form, but also includes a function to recognize the user's emotions and adjust the diagnosis results based on those emotions. This system consists of four main components: the user, the terminal, the server, and the emotion engine.

[1306] User behavior

[1307] Users record videos or images of their exercise (e.g., a golf swing) using a device such as a smartphone or camera. After recording, they access an application or web interface for exercise analysis and select the files. In addition, the user's facial expressions are recorded simultaneously during the upload process.

[1308] Terminal operation

[1309] The terminal's role is to send the video or image file selected by the user to the server. Communication methods such as HTTP POST requests are used for transmission. The terminal also displays the diagnostic results received from the server on the user interface.

[1310] Server Role

[1311] The server plays a central role in receiving and analyzing videos or images sent from the terminal. Specifically, it performs the following steps:

[1312] 1. Once the video is sent, the server analyzes the frames and extracts each scene of the motion (e.g., start, in motion, end).

[1313] Example of a tool used: OpenCV

[1314] 2. The motion diagnosis algorithm is executed on the extracted scenes and images to generate the diagnosis results.

[1315] 3. Analyze user emotional data and incorporate it into the diagnostic results.

[1316] Example of tools used: Microsoft Azure Cognitive Services

[1317] Functions of the Emotion Engine

[1318] The emotion engine analyzes the user's facial expressions to identify their emotional state. This allows the system to understand the user's emotions when they upload exercise videos or images and adjust the exercise diagnosis results accordingly.

[1319] Specific examples of program processing

[1320] 1. The user simultaneously records a video of their exercise and their facial expressions using their smartphone.

[1321] The smartphone's camera function is used to record exercise and facial expressions simultaneously.

[1322] 2. The user opens the application, selects the recorded video, and uploads it.

[1323] Use the application's file selection screen to upload the recorded video file to the designated location.

[1324] 3. The device sends the selected video file to the server.

[1325] File transmission is performed using HTTP POST requests or other appropriate communication methods.

[1326] 4. The server receives the video file and extracts frames for each scene of the exercise.

[1327] The program analyzes the video content and executes an algorithm to extract specific scenes (start, during, and end). For example, it uses OpenCV.

[1328] 5. The server preprocesses the frames of each extracted scene.

[1329] The frames of each scene are standardized and converted into a format that allows image analysis algorithms to function properly. For example, this may involve unifying image sizes and removing noise.

[1330] 6. The server inputs the pre-processed images into the diagnostic algorithm.

[1331] The images from each scene, after preprocessing is complete, are input into a motion diagnosis algorithm for evaluation.

[1332] 7. The emotion engine analyzes the user's facial expressions and identifies their emotional state.

[1333] The system analyzes captured facial expression data to classify emotions such as joy, surprise, sadness, and anger. For example, it uses Microsoft Azure Cognitive Services.

[1334] 8. The server adjusts the motor diagnosis results based on the output of the emotion engine.

[1335] The system takes emotional data into account and adds emotional feedback to the diagnostic results. For example, if the user has an anxious expression, it provides more positive feedback.

[1336] 9. The server generates the final diagnostic results and sends them back to the terminal.

[1337] The diagnostic results are sent back to the terminal as an HTTP response.

[1338] 10. The terminal displays the diagnostic results to the user.

[1339] The device displays the received diagnostic results in the user interface, showing the user specific areas for improvement in their exercise form, evaluations, and emotion-based feedback.

[1340] Example of a prompt

[1341] "Please upload a video of your swing. We will analyze your swing form and your emotions, and provide feedback."

[1342] The flow of the specific processing in Example 2 will be explained using Figure 13.

[1343] Step 1:

[1344] The user simultaneously records a video of their exercise and their facial expressions using their smartphone.

[1345] Input: Smartphone camera function

[1346] Instructions: Mount your smartphone on a tripod and position it so that your whole body is visible. Press the record button to record your physical movements and facial expressions simultaneously.

[1347] Output: Exercise video file and facial expression video file

[1348] Step 2:

[1349] The user opens the application for swing analysis, selects the recorded video, and uploads it.

[1350] Input: Recorded exercise video file

[1351] How to do it: Tap the "Select Video" button in the application, choose the video file you've recorded, and then tap "Upload" to send it to the server.

[1352] Output: Exercise video file ready to upload

[1353] Step 3:

[1354] The device sends the selected video file to the server.

[1355] Input: Exercise video file selected by the user and uploaded by pressing the upload button.

[1356] Operation: The device sends a video file to the server using an HTTP POST request and displays the transmission progress.

[1357] Output: Exercise video file sent to the server

[1358] Step 4:

[1359] The server extracts each scene frame by frame from the received video file.

[1360] Input: Exercise video file sent to the server

[1361] Operation: The video is divided into frames using a video analysis library (e.g., OpenCV), and the start, middle, and end scenes are identified and extracted.

[1362] Output: Frame images divided by scene

[1363] Step 5:

[1364] The server preprocesses the extracted frame images.

[1365] Input: Frame images divided by scene

[1366] Operation: Unify image sizes and smooth images by applying a denoising filter. For example, unify the resolution to 640x480 pixels and run a denoising algorithm.

[1367] Output: Preprocessed frame image

[1368] Step 6:

[1369] The server inputs the pre-processed images into the motion diagnosis algorithm.

[1370] Input: Pre-processed frame image

[1371] Operation: Pre-processed frame images are passed to a motion diagnostic algorithm for evaluation scoring and analysis of movement form. For example, swing trajectory and body tilt are quantified.

[1372] Output: Evaluation results and scores of exercise form

[1373] Step 7:

[1374] The emotion engine analyzes the user's facial expressions to identify their emotional state.

[1375] Input: Facial expression video file being sent to the server

[1376] Operation: Uses an emotion analysis model (e.g., Microsoft Azure Cognitive Services) to classify emotions such as joy, surprise, sadness, and anger from facial expressions.

[1377] Output: Identified emotional state

[1378] Step 8:

[1379] The server adjusts the motor diagnosis results based on the output of the emotion engine.

[1380] Input: Evaluation results and scores of exercise form, identified emotional state

[1381] Operation: It takes emotional data into account and adjusts the diagnostic results to be positive or negative. For example, if the expression is anxious, it adds a positive message such as "Stay calm and continue practicing at this pace."

[1382] Output: Adjusted final diagnostic results

[1383] Step 9:

[1384] The server generates the final diagnostic results and sends them back to the terminal.

[1385] Input: Adjusted final diagnostic result

[1386] Function: Organizes the diagnostic results and sends them to the terminal in JSON format or another appropriate format.

[1387] Output: Final diagnostic results sent to the terminal

[1388] Step 10:

[1389] The terminal displays the diagnostic results to the user.

[1390] Input: Final diagnostic results received from the server

[1391] Operation: The user interface visually presents evaluation results of exercise form, areas for improvement, and emotion-based feedback. For example, it might display, "Your movement is generally good. Be careful as your body tends to move forward at the finish."

[1392] Output: Diagnostic results and feedback displayed to the user

[1393] Example of a prompt

[1394] "Please upload a video of your swing. We will analyze your swing form and your emotions, and provide feedback."

[1395] (Application Example 2)

[1396] Next, we will explain application example 2. In the following explanation, the data processing device 12 will be referred to as the "server," and the headset-type terminal 314 will be referred to as the "terminal."

[1397] Traditional swing analysis systems only evaluate the user's swing form and do not consider feedback based on the user's emotions, thus limiting their effectiveness in improving user motivation and reducing stress. Furthermore, they lacked the functionality to provide specific training methods, preventing them from directly contributing to user skill improvement.

[1398] The specific processing performed by the specific processing unit 290 of the data processing device 12 in Application Example 2 is realized by the following means.

[1399] In this invention, the server includes means for analyzing the user's facial expressions to identify their emotional state, means for adjusting the swing diagnosis results based on the emotional state, and means for providing specific training methods based on an evaluation of the swing form. This allows the user to receive flexible feedback tailored to their emotional state and specific methods for improvement.

[1400] A "user" is an individual or group that uses the system, uploads videos or images of their swing, and receives the analysis results.

[1401] A "video" is video data composed of a series of image frames, and is used to record the user's swing motion.

[1402] An "image" is visual data represented in a single still image format and is used to capture a specific scene of a swing.

[1403] A "scene" refers to a visual representation of each stage of the swing (address, backswing, top, impact, finish).

[1404] A "diagnostic algorithm" is a set of calculation procedures and rules used to analyze and evaluate swing form, and the system uses them to scrutinize each aspect of the swing.

[1405] "Facial expression" refers to the various muscle movements and arrangements that appear on a user's face, and serves as basic information for inferring their emotional state.

[1406] "Emotional state" refers to the internal psychological state analyzed from the user's facial expressions, and includes classifications such as joy, surprise, sadness, and anger.

[1407] "Training methods" refer to specific practice procedures and exercises aimed at improving swing form and technique.

[1408] In an embodiment for carrying out this invention, the system includes the following components.

[1409] User

[1410] A "user" is an individual or group that uses the system, and is the entity that takes videos or images of their swing using a device such as a smartphone and uploads them. Users can receive not only technical evaluations of their swing form, but also emotional feedback derived from their facial expressions.

[1411] terminal

[1412] A "device" is a device used by a user to upload data, and includes smartphones and tablets. A device utilizes the following hardware and software:

[1413] Hardware: Smartphones (iPhone, Android)

[1414] Software: Dedicated application (Golf Swing Master)

[1415] Its role is to handle everything from video recording to data transmission, and then receive diagnostic results from the server and display them to the user.

[1416] server

[1417] A "server" is a central processing unit that receives data transmitted from terminals and performs analysis and diagnosis. The server fulfills the following roles:

[1418] Swing Analysis: Frames are extracted from the video, and a diagnostic algorithm is executed to analyze each scene (address, backswing, top, impact, finish).

[1419] Emotion Recognition: Analyzes the user's facial expressions to identify their emotional state and reflect it in the swing analysis results.

[1420] Result Adjustment: The diagnostic results are adjusted based on emotional data to provide users with appropriate feedback.

[1421] Emotional Engine

[1422] The "emotion engine" is a software module that analyzes a user's facial expressions and classifies their emotional state. The emotion engine is used to understand the user's psychological state when uploading swing videos and images.

[1423] Training provided

[1424] The server provides specific training methods based on an evaluation of swing form. This allows users to learn concrete practical techniques that help improve their skills.

[1425] Examples

[1426] For example, a user films their swing at a golf driving range with their smartphone and uploads the video to a dedicated app. The device sends the video data to a server, which analyzes the video to evaluate the swing form and uses an emotion engine to analyze the user's facial expressions to identify their emotional state. As a result, in addition to the technical evaluation, the server provides feedback and training methods tailored to the user's emotional state. This feedback might include something like, "Your swing is good, but try to relax a little more. Refer to this video and practice swinging in a more relaxed state."

[1427] Examples of prompts for generative AI models

[1428] Create an application that analyzes a user's golf swing video and facial expressions to provide a technical evaluation of their swing form and emotionally responsive feedback.

[1429] 1. How to extract frames from a video.

[1430] 2. A method for identifying emotions by analyzing facial expressions.

[1431] 3. How to send swing analysis results and emotional feedback to the server.

[1432] 4. How to receive diagnostic results and how to display them in the user interface.

[1433] The flow of a specific process in Application Example 2 will be explained using Figure 14.

[1434] Step 1:

[1435] The user simultaneously records a video of their golf swing and their facial expressions using their smartphone.

[1436] Input: Video to be recorded and facial expressions

[1437] Output: Video files saved on the smartphone's storage.

[1438] Specific operation: The user uses the smartphone's camera app to record a video, and at the same time, their facial expressions are also recorded using the front camera.

[1439] Step 2:

[1440] The user opens the application, selects the video they have recorded, and uploads it.

[1441] Input: User-recorded video

[1442] Output: Video file uploaded to the application

[1443] Specific steps: Open the file selection screen within the application, select the recorded video file, and press the upload button.

[1444] Step 3:

[1445] The device sends the selected video file to the server.

[1446] Input: Uploaded video file

[1447] Output: Video data sent to the server

[1448] Specific action: Send the video file to the server using an HTTP POST request.

[1449] Step 4:

[1450] The server receives the video file and extracts frames for each scene of the swing.

[1451] Input: Video data sent from the device

[1452] Output: Frame images of each extracted scene

[1453] Specific operation: Use OpenCV to extract frames from the video and separate the address, backswing, top of swing, impact, and finish scenes.

[1454] Step 5:

[1455] The server preprocesses each frame of the extracted scene.

[1456] Input: Extracted frame image

[1457] Output: Preprocessed image data

[1458] Specific actions: Standardize frames, unify image sizes, and remove noise.

[1459] Step 6:

[1460] The server inputs the pre-processed images into the diagnostic algorithm.

[1461] Input: Preprocessed image data

[1462] Output: Swing form evaluation results

[1463] Specific operation: Pre-processed images are input into the diagnostic algorithm, and a technical evaluation of the swing form is performed.

[1464] Step 7:

[1465] The emotion engine analyzes the user's facial expressions to identify their emotional state.

[1466] Input: Facial expression data included in the video

[1467] Output: Classified emotional state

[1468] Specific operation: Facial expression data is input into an emotion recognition model to classify emotions such as joy, surprise, sadness, and anger.

[1469] Step 8:

[1470] The server adjusts the swing analysis results based on the output of the emotion engine.

[1471] Input: Swing form evaluation results and emotional state

[1472] Output: Final diagnostic results including emotional feedback

[1473] Specific actions: Consider emotional data and incorporate emotional feedback into the diagnostic results, generating feedback such as, "Your swing is good, but try to relax a little more."

[1474] Step 9:

[1475] The server generates the final diagnostic results and sends them back to the terminal.

[1476] Input: Final diagnostic results including emotional feedback

[1477] Output: Diagnostic results sent to the terminal

[1478] Specific operation: The final diagnostic results are sent back to the terminal using an HTTP response.

[1479] Step 10:

[1480] The terminal displays the diagnostic results to the user.

[1481] Input: Diagnostic results received from the server

[1482] Output: Diagnostic results displayed in the user interface

[1483] Specific action: Display the received feedback and swing evaluation in the application's user interface.

[1484] The specific processing unit 290 transmits the result of the specific processing to the headset terminal 314. In the headset terminal 314, the control unit 46A causes the speaker 240 and display 343 to output the result of the specific processing. The microphone 238 acquires audio indicating user input for the result of the specific processing. The control unit 46A transmits the audio data indicating user input acquired by the microphone 238 to the data processing unit 12. In the data processing unit 12, the specific processing unit 290 acquires the audio data.

[1485] Data generation model 58 is a type of so-called generative AI (Artificial Intelligence). One example of data generation model 58 is ChatGPT (Internet search<URL: https: / / openai.com / blog / chatgpt> ), Gemini (Internet search) <url: https: gemini.google.com ?hl="ja">Examples of generative AI include the following. The data generation model 58 is obtained by performing deep learning on a neural network. The data generation model 58 is input with prompts containing instructions, and with inference data such as audio data representing speech, text data representing text, and image data representing images. The data generation model 58 infers from the input inference data according to the instructions indicated by the prompts, and outputs the inference results in data formats such as audio data and text data. Here, inference refers to, for example, analysis, classification, prediction, and / or summarization.

[1486] In the above embodiment, an example was given in which specific processing is performed by the data processing device 12, but the technology of this disclosure is not limited thereto, and specific processing may also be performed by the headset terminal 314.

[1487] [Fourth Embodiment]

[1488] Figure 7 shows an example of the configuration of the data processing system 410 according to the fourth embodiment.

[1489] As shown in Figure 7, the data processing system 410 includes a data processing device 12 and a robot 414. An example of the data processing device 12 is a server.

[1490] The data processing device 12 comprises a computer 22, a database 24, and a communication interface 26. The computer 22 is an example of a "computer" related to the technology of this disclosure. The computer 22 comprises a processor 28, RAM 30, and storage 32. The processor 28, RAM 30, and storage 32 are connected to a bus 34. The database 24 and the communication interface 26 are also connected to the bus 34. The communication interface 26 is connected to a network 54. An example of the network 54 is a WAN (Wide Area Network) and / or a LAN (Local Area Network).

[1491] The robot 414 includes a computer 36, a microphone 238, a speaker 240, a camera 42, a communication interface 44, and a controlled object 443. The computer 36 includes a processor 46, RAM 48, and storage 50. The processor 46, RAM 48, and storage 50 are connected to a bus 52. The microphone 238, speaker 240, camera 42, and controlled object 443 are also connected to the bus 52.

[1492] The microphone 238 receives voice signals from the user 20 and receives instructions from the user 20. The microphone 238 captures the voice signals from the user 20, converts the captured voice into audio data, and outputs it to the processor 46. The speaker 240 outputs audio according to the instructions from the processor 46.

[1493] Camera 42 is a small digital camera equipped with an optical system including a lens, aperture, and shutter, and an image sensor such as a CMOS (Complementary Metal-Oxide-Semiconductor) image sensor or a CCD (Charge Coupled Device) image sensor, and captures images of the area around the user 20 (for example, an imaging range defined by a field of view equivalent to the width of a typical healthy person's field of vision).

[1494] Communication interface 44 is connected to network 54. Communication interfaces 44 and 26 are responsible for the exchange of various information between processor 46 and processor 28 via network 54. The exchange of various information between processor 46 and processor 28 using communication interfaces 44 and 26 is performed in a secure manner.

[1495] The controlled object 443 includes a display device, LEDs in the eyes, and motors that drive the arms, hands, and feet. The posture and gestures of the robot 414 are controlled by controlling the motors of the arms, hands, and feet. Some of the robot 414's emotions can be expressed by controlling these motors. Furthermore, the robot 414's facial expressions can also be expressed by controlling the illumination state of the LEDs in its eyes.

[1496] Figure 8 shows an example of the main functions of the data processing device 12 and the robot 414. As shown in Figure 8, the data processing device 12 performs specific processing using the processor 28. The storage 32 stores the specific processing program 56.

[1497] The specific processing program 56 is an example of a "program" relating to the technology of this disclosure. The processor 28 reads the specific processing program 56 from the storage 32 and executes the read specific processing program 56 on the RAM 30. The specific processing is realized by the processor 28 operating as a specific processing unit 290 in accordance with the specific processing program 56 executed on the RAM 30.

[1498] The storage 32 stores the data generation model 58 and the emotion identification model 59. The data generation model 58 and the emotion identification model 59 are used by the identification processing unit 290.

[1499] In robot 414, the processor 46 performs the reception output processing. The storage 50 stores the reception output program 60. The processor 46 reads the reception output program 60 from the storage 50 and executes the read reception output program 60 on the RAM 48. The reception output processing is realized by the processor 46 operating as a control unit 46A according to the reception output program 60 executed on the RAM 48.

[1500] Next, the specific processing performed by the specific processing unit 290 of the data processing device 12 will be described. In the following description, the data processing device 12 will be referred to as the "server" and the robot 414 as the "terminal".

[1501] The present invention provides a system for which a user uploads a video or image of their swing, and for which the swing form is analyzed and diagnosed. The operation of each element and the program's processing are described below in natural language.

[1502] System Overview

[1503] This system consists of three main components: users, terminals, and servers.

[1504] User

[1505] Users record videos or images of their golf swing using a device such as a smartphone or camera. After recording, they access a swing analysis application or web interface and select the file.

[1506] terminal

[1507] The terminal's role is to send the user-selected video or image file to the server. Communication methods such as HTTP POST requests are used for transmission. Furthermore, the terminal displays the diagnostic results received from the server on the user interface.

[1508] server

[1509] The server plays a central role in receiving and analyzing videos or images sent from the terminal. If a video is sent, the server analyzes the frames and extracts each scene of the swing (address, backswing, top, impact, finish). It then runs a swing diagnostic algorithm on the extracted scenes or images to generate a diagnostic result. The generated diagnostic result is then sent back to the terminal.

[1510] Specific examples of program processing

[1511] 1. The user films their golf swing with their smartphone.

[1512] Use your smartphone's camera function to record your swing motion in video format.

[1513] 2. The user opens the application, selects the recorded video, and uploads it.

[1514] Use the application's file selection screen to upload the recorded video file to the designated location.

[1515] 3. The device sends the selected video file to the server.

[1516] File transmission is performed using HTTP POST requests or other appropriate communication methods.

[1517] 4. The server receives the video and extracts frames for each scene of the swing.

[1518] The system analyzes the video content and executes an algorithm to extract specific scenes (address, backswing, top of swing, impact, finish).

[1519] 5. The server analyzes the images from each scene and evaluates them based on the swing diagnostic algorithm.

[1520] The system analyzes the swing form for each extracted scene image. The diagnostic algorithm evaluates the accuracy of posture and movement, identifying abnormalities and areas for improvement.

[1521] 6. The server generates diagnostic results and creates feedback such as "Your body is unbalanced at address" and "Your timing of impact is slow."

[1522] Based on the output of the diagnostic algorithm, specific feedback is generated to be provided to the user.

[1523] 7. The server sends the diagnostic results back to the smartphone.

[1524] The generated diagnostic results are sent back to the terminal as an HTTP response.

[1525] 8. The terminal displays the diagnostic results to the user and indicates areas for improvement.

[1526] The diagnostic results are displayed in the user interface, visually and textually presenting areas for improvement to the user. Annotations are added to the images in each scene to clearly indicate specific problems.

[1527] This system allows users to efficiently and objectively analyze their golf swing and visually understand areas for improvement without requiring expensive specialized equipment or expertise.

[1528] The following describes the processing flow.

[1529] Step 1: The user records a video of their golf swing using their smartphone.

[1530] Users record their golf swing as a video using their smartphone's camera function.

[1531] Step 2: The user opens the application and selects the recorded video.

[1532] Display the application's file selection screen and select the video file you recorded.

[1533] Step 3: The user uploads a video file.

[1534] Press the upload button in the application to send the selected video file to the server.

[1535] Step 4: The device sends the video file to the server.

[1536] The device uses an HTTP POST request to send the video file to the server.

[1537] Step 5: The server receives the video file.

[1538] The server retrieves and saves the video file from the received HTTP POST request.

[1539] Step 6: The server analyzes the video and extracts each scene.

[1540] The server analyzes the video frames to identify each scene—address, backswing, top of swing, impact, and finish—and extracts the corresponding frames.

[1541] Step 7: The server preprocesses the frames of each extracted scene.

[1542] The server standardizes the frames of each scene and converts them into a format that allows image analysis algorithms to function properly. For example, it performs image size unification and noise reduction.

[1543] Step 8: The server inputs the pre-processed images into the diagnostic algorithm.

[1544] The server inputs the pre-processed images of each scene into a swing diagnostic algorithm for evaluation.

[1545] Step 9: The server generates a diagnostic result based on the image analysis results.

[1546] Note) The server evaluates the analysis results and identifies anomalies in the swing form for each scene. It generates feedback on the identified anomalies and compiles them into a diagnostic report.

[1547] Step 10: The server sends the diagnostic results back to the terminal.

[1548] The server returns the diagnostic results to the terminal as an HTTP response.

[1549] Step 11: The terminal displays the diagnostic results to the user.

[1550] The terminal displays the received diagnostic results on the user interface, showing the user specific areas for improvement and evaluation of their swing form.

[1551] This processing step allows users to efficiently analyze their golf swing and receive specific feedback.

[1552] (Example 1)

[1553] Next, we will describe Example 1. In the following description, the data processing device 12 will be referred to as the "server" and the robot 414 as the "terminal".

[1554] Conventional technologies required expensive specialized equipment and the assistance of skilled technicians to analyze golf swings, making it difficult for ordinary users to easily evaluate and improve their swing form. Furthermore, there was a lack of methods to provide swing form diagnostic results in an intuitively understandable format. Therefore, there was a need for a technology that could analyze and evaluate golf swings easily and at low cost, allowing users to clearly identify areas for improvement themselves.

[1555] The identification process performed by the identification processing unit 290 of the data processing device 12 in Example 1 is realized by the following means.

[1556] In this invention, the server includes means for a user to upload a video or image of their swing, means for analyzing and extracting each scene (address, backswing, top, impact, finish) frame by frame from the uploaded video, means for executing a diagnostic algorithm to analyze the extracted scenes and evaluate the swing form, means for returning and displaying the generated diagnostic results to the user, means for generating specific feedback (e.g., issues with body balance or timing) based on the diagnostic results, and means for adding annotations to the images of each scene as feedback to indicate areas for improvement. This makes it possible for users to analyze their golf swing in detail and intuitively understand specific areas for improvement without using expensive specialized equipment.

[1557] A "user" refers to a person who uses the system to film their own golf swing and obtain the analysis results.

[1558] "Swing video or image" refers to a video file or still image file recorded by the user of their golf swing motion.

[1559] "Means for uploading" refers to the interface or function that allows users to provide videos or images they have taken to the system.

[1560] "Each scene" refers to the specific phases of a golf swing: address, backswing, top, impact, and finish.

[1561] "Means for analyzing and extracting frame by frame" refers to technologies and devices that divide a video into multiple still images and identify a specific scene from each frame.

[1562] A "diagnostic algorithm" refers to a calculation method or program that analyzes extracted scene images to evaluate golf swing form.

[1563] The means for returning and displaying the "generated diagnostic results" to the user refers to the communication method and display function for returning the analysis results from the diagnostic algorithm and making them viewable by the user.

[1564] "Specific feedback" refers to detailed observations and advice based on the diagnostic results, such as issues with body balance or timing.

[1565] "Annotation" refers to annotations, marks, and visual information added to images in each scene to indicate areas for improvement.

[1566] The present invention provides a system for users to record their golf swings as videos or images, and to analyze and evaluate those swings. This system consists of three main components: the user, the terminal, and the server.

[1567] User operations and terminal functions

[1568] Users record their golf swing as a video or image using a device such as a smartphone or tablet. They utilize the device's camera function for this purpose. For example, they can use the iPhone's camera app to shoot high-resolution video. Afterward, the user launches a dedicated application such as "Golf Swing Analyzer," selects the recorded video through the application's interface, and uploads it.

[1569] The terminal is responsible for sending the video file selected by the user to the server. Common communication technologies such as HTTP POST requests are used for this transmission. During transmission, the appropriate communication protocol is selected considering the file format and size. Furthermore, the terminal displays the diagnostic results received from the server on the user interface. Therefore, the terminal requires visual and textual display functions that allow the user to intuitively understand the diagnostic results.

[1570] Server analysis and diagnostic functions

[1571] The server plays a central role in receiving and analyzing videos or images sent from the terminal. After receiving the data, the server uses image analysis libraries such as OpenCV to divide the video into frames and extract each scene of the swing (address, backswing, top, impact, finish). A specific algorithm is used to extract each scene.

[1572] For each extracted scene, the server uses a generative AI model such as TensorFlow to evaluate the swing form. This diagnostic algorithm assesses the accuracy of posture and movement, and identifies abnormalities and areas for improvement. For example, it can evaluate things like "the body is unbalanced at address" or "the timing of impact is late."

[1573] Examples of specific cases and prompt statements

[1574] Further implementation examples include the following: A user uploads a video of their golf swing, the server analyzes it, and sends specific feedback back to the user's device. This feedback is displayed as annotations for each scene in the video, making it easy for the user to understand. This system allows users to obtain analysis and evaluation of their golf swing without needing expensive specialized equipment.

[1575] Examples of prompt messages are as follows:

[1576] Design a system that automatically analyzes the movements in each scene of a golf swing video recorded by a user on their iPhone, and diagnoses the swing form. Describe the specific analysis method and algorithm, and clearly specify the communication method between the server and client.

[1577] This invention allows users to efficiently and accurately analyze their golf swing and identify areas for improvement without requiring specialized knowledge or expensive equipment.

[1578] The flow of the specific processing in Example 1 will be explained using Figure 11.

[1579] Program processing flow

[1580] Step 1:

[1581] The user films their golf swing.

[1582] (Specific actions)

[1583] Users record their golf swings in video format using their smartphones or tablets. They use a camera app to capture the entire swing within the frame.

[1584] (input)

[1585] Videos filmed by users.

[1586] (output)

[1587] A video of a swing saved on a smartphone.

[1588] Step 2:

[1589] The user opens the application, selects the video they have recorded, and uploads it.

[1590] (Specific actions)

[1591] The user launches a dedicated app such as "Golf Swing Analyzer" and navigates to the "file selection screen" within the app's interface. The user selects the recorded video and clicks the "upload button".

[1592] (input)

[1593] A video of a swing saved on a smartphone.

[1594] (output)

[1595] The selected video will be displayed within the application, and an upload request will be generated.

[1596] Step 3:

[1597] The device sends the video file to the server.

[1598] (Specific actions)

[1599] The device sends the video file to the server using an HTTP POST request. During this process, the video file is appropriately split and transmitted according to the communication protocol.

[1600] (input)

[1601] The selected video file within the application, and the upload request.

[1602] (output)

[1603] Video file being sent to the server.

[1604] Step 4:

[1605] The server receives the video and analyzes it frame by frame.

[1606] (Specific actions)

[1607] The server saves the received video file and uses the OpenCV library to split the video into frames. Each frame is sequentially loaded into memory and used for analysis.

[1608] (input)

[1609] The video file sent to the server.

[1610] (output)

[1611] Video data divided into individual frames.

[1612] Step 5:

[1613] The server extracts each scene and applies a swing analysis algorithm.

[1614] (Specific actions)

[1615] The server analyzes each frame and detects specific scenes (address, backswing, top of trajectory, impact, finish). After this, a TensorFlow-based generative AI model is used to evaluate each scene.

[1616] (input)

[1617] Video data divided into individual frames.

[1618] (output)

[1619] Data for each detected scene and its diagnostic results.

[1620] Step 6:

[1621] The server generates diagnostic results and creates feedback for the user.

[1622] (Specific actions)

[1623] The server generates feedback based on the analysis results. For example, it creates diagnostic statements such as "body balance is off at address" or "impact timing is slow," and formats them in JSON format.

[1624] (input)

[1625] Data for each detected scene and its diagnostic results.

[1626] (output)

[1627] User feedback data.

[1628] Step 7:

[1629] The server sends the generated diagnostic results back to the user's terminal.

[1630] (Specific actions)

[1631] The server sends the generated diagnostic results back to the user's terminal as an HTTP response. The returned data is sent in an appropriate format, such as JSON or XML.

[1632] (input)

[1633] User feedback data.

[1634] (output)

[1635] Diagnostic results sent to the user's terminal.

[1636] Step 8:

[1637] The device displays the diagnostic results.

[1638] (Specific actions)

[1639] The terminal analyzes the received diagnostic results and displays them in the user interface. Because the user can visually confirm the diagnostic results, specific annotations are added for each scene.

[1640] (input)

[1641] Diagnostic results sent from the server.

[1642] (output)

[1643] Diagnostic results and annotations displayed in the user interface.

[1644] (Application Example 1)

[1645] Next, we will explain Application Example 1. In the following explanation, the data processing device 12 will be referred to as the "server" and the robot 414 as the "terminal".

[1646] Until now, there had been no system that could efficiently and quickly analyze golf swing form and provide users with useful feedback in real time. As a result, users needed expert advice or expensive equipment to improve their swing form. This meant that many amateur golfers did not have enough opportunities for self-improvement. Furthermore, because the feedback was not in real time, users could not immediately obtain information for improvement, which reduced the efficiency of their practice. In addition, the lack of a feedback system using wearable devices made it difficult for users to check swing improvement measures on the spot.

[1647] The specific processing performed by the specific processing unit 290 of the data processing device 12 in Application Example 1 is realized by the following means.

[1648] In this invention, the server includes means for a user to upload a video or image of their swing, means for extracting each scene (address, backswing, top, impact, finish) from the uploaded video, means for executing a diagnostic algorithm to analyze the extracted scenes and evaluate the swing form, and means for providing real-time feedback of the displayed diagnostic results to a wearable device worn by the user. This enables real-time feedback of golf swing form analysis and diagnostic results, allowing the user to immediately see areas for improvement and increase the efficiency of their practice.

[1649] A "user" is an individual or group that uses the system to upload videos or images of their golf swing and receive analysis and diagnostic results.

[1650] A "video" is a dynamic media file that displays a series of images taken by a user in sequence.

[1651] An "image" is a media file that includes still images taken by the user.

[1652] "Uploading" refers to the act of a user transferring video or image data from their device to a server.

[1653] A "scene" refers to a specific moment or stage in a swing, specifically the address, backswing, top of the swing, impact, and finish.

[1654] "Extraction" refers to taking a specific scene from an uploaded video.

[1655] "Analysis" refers to evaluating the extracted images and videos of scenes using a diagnostic algorithm.

[1656] A "diagnostic algorithm" is a specific set of calculation procedures or analytical models used to evaluate swing form and generate diagnostic results.

[1657] "Generation" refers to creating diagnostic results based on the analyzed data.

[1658] "Display" refers to outputting the generated diagnostic results to the terminal in a format that the user can review.

[1659] A "wearable device" refers to a device that a user can wear, typically in the form of glasses, a head-mounted display, or a wristband.

[1660] "Real-time" refers to processing and feedback being performed instantly without delay.

[1661] "Feedback" is the process of returning the generated diagnostic results and areas for improvement to the user.

[1662] This invention is a system that analyzes videos and images of golf swings to diagnose swing form, and is designed to be easily accessible to users. The main components are the user, a terminal, a server, and a wearable device.

[1663] User

[1664] Users record their golf swing using a device such as a smartphone or smart glasses, capturing it as a video or image. After recording, they upload the video or image using a dedicated application. This allows the system to input the user's swing form.

[1665] terminal

[1666] The device is used to send videos and images taken by the user to the server. HTTP POST requests are typically used for transmission. Furthermore, the device displays the diagnostic results received from the server on the user interface. Examples of devices include smartphones, smart glasses, and head-mounted displays.

[1667] server

[1668] The server plays a central role in analyzing videos and images sent from the terminal and evaluating the swing form. Specifically, it performs the following processes:

[1669] 1. When a video is sent, the server analyzes it frame by frame and extracts each scene of the swing (address, backswing, top, impact, finish).

[1670] 2. A diagnostic algorithm is run on the extracted scenes to evaluate the swing form of each scene. The software used includes OpenCV (image processing), TensorFlow and PyTorch (machine learning algorithms), and Flask (web framework).

[1671] 3. Generate diagnostic results and create feedback to provide to the user. This feedback will include specific areas for improvement and problems, and will be displayed in text and annotation format.

[1672] Wearable devices

[1673] Wearable devices are devices that users wear to receive real-time feedback. Examples include smart glasses and head-mounted displays. Diagnostic results and feedback transmitted from a server are displayed on these wearable devices in real time.

[1674] Specific example

[1675] As a concrete example, imagine a scenario where a user wears smart glasses at a golf driving range and performs a swing. The smart glasses film the user's swing in real time and send the video to a server. The server analyzes the video and diagnoses the swing form. As a result, feedback such as "Your body balance is off at address" or "Your top position is too high" is immediately displayed on the smart glasses' screen. The user can then correct their swing form on the spot, film again, and receive further evaluation.

[1676] Examples of prompts to input into a generative AI model

[1677] "Design an application that allows users to film their golf swing with smart glasses and receive real-time form analysis and diagnostic results. The system would involve filming the frame, uploading the footage to a server, and displaying the analyzed results on the smart glasses. Please provide a detailed explanation, including the necessary hardware and software, and the program's processing steps."

[1678] The flow of a specific process in Application Example 1 will be explained using Figure 12.

[1679] Step 1:

[1680] The user takes a video or image of their swing.

[1681] Input: A video or image of a swing taken using the camera of a smartphone or smart glasses.

[1682] Operation: The user performs a golf swing, and the camera records the motion. In the case of video, a series of movements are recorded continuously; in the case of still images, specific moments are saved as still images.

[1683] Output: The recorded video or image of the golf swing is saved on the device.

[1684] Step 2:

[1685] Users upload videos or images they have taken to the server via a dedicated application.

[1686] Input: Swing video or image taken after shooting.

[1687] Operation: The user opens the dedicated application, selects a video or image from the file selection screen, and presses the upload button. The file is sent to the server using a data transfer protocol such as an HTTP POST request.

[1688] Output: Video or image data is sent to the server.

[1689] Step 3:

[1690] The server analyzes the received video or images and extracts each scene of the swing.

[1691] Input: A swing video or image uploaded to the server.

[1692] Operation: The server divides the video frame by frame and detects important scenes in the swing (address, backswing, top, impact, finish). This analysis and extraction is performed using image processing libraries such as OpenCV.

[1693] Output: Frame images of each extracted scene.

[1694] Step 4:

[1695] The server analyzes each extracted scene and executes a diagnostic algorithm to evaluate the swing form.

[1696] Input: Frame images from each extracted scene.

[1697] Operation: The server uses machine learning models (e.g., TensorFlow, PyTorch) to evaluate the swing form. It analyzes the accuracy of posture and movement, and identifies abnormalities and areas for improvement.

[1698] Output: Swing form evaluation results for each scene.

[1699] Step 5:

[1700] The server generates diagnostic results based on the swing form evaluation and creates feedback to provide to the user.

[1701] Input: Swing form evaluation results.

[1702] Operation: The server generates specific feedback based on the evaluation results. For example, it might create text-based feedback such as "Your body balance is off at address" or "Your top position is too high."

[1703] Output: Generated diagnostic results and feedback.

[1704] Step 6:

[1705] The server sends diagnostic results to the terminal and wearable device, and displays feedback in real time.

[1706] Input: Generated diagnostic results and feedback.

[1707] Operation: The server sends the diagnostic results to the device (smartphone, smart glasses, etc.). The results are returned as an HTTP response, and in the case of wearable devices, the data is transferred via a dedicated API. The device or wearable device displays the received data.

[1708] Output: Diagnostic results and feedback displayed on the user's device or wearable device.

[1709] Step 7:

[1710] We will work on improving the swing based on the feedback displayed to the user.

[1711] Input: Diagnostic results and feedback displayed on a terminal or wearable device.

[1712] Operation: The user reviews the displayed feedback and implements measures to improve their swing form. If necessary, they film their swing again and use the system for further diagnosis.

[1713] Output: Improved swing form, revised diagnostic results.

[1714] Furthermore, an emotion engine that estimates the user's emotions may be incorporated. That is, the identification processing unit 290 may use the emotion identification model 59 to estimate the user's emotions and perform identification processing using the user's emotions.

[1715] The system of the present invention not only allows users to upload videos or images of their swings and analyzes and diagnoses their swing form, but also includes a function to recognize the user's emotions and adjust the diagnosis results based on those emotions. The operation of each element and the processing of the program are described below in natural language.

[1716] System Overview

[1717] This system consists of four main components: the user, the terminal, the server, and the emotion engine.

[1718] User

[1719] Users record videos or images of their golf swing using a device such as a smartphone or camera. After recording, they access a swing analysis application or web interface to select the file. In addition, the user's facial expressions are also recorded during the upload process.

[1720] terminal

[1721] The terminal's role is to send the video or image file selected by the user to the server. Communication methods such as HTTP POST requests are used for transmission. The terminal also displays the diagnostic results received from the server on the user interface.

[1722] server

[1723] The server plays a central role in receiving and analyzing videos or images sent from the terminal. If a video is sent, the server analyzes the frames and extracts each scene of the swing (address, backswing, top, impact, finish). It then runs a swing diagnostic algorithm on the extracted scenes and images to generate a diagnostic result. It also analyzes the user's emotional data and incorporates it into the diagnostic result.

[1724] Emotional Engine

[1725] The emotion engine analyzes the user's facial expressions to identify their emotional state. This allows the system to understand the user's emotions when they upload swing videos or images and adjust the swing analysis results accordingly.

[1726] Specific examples of program processing

[1727] 1. The user simultaneously records a video of their golf swing and their facial expressions using their smartphone.

[1728] The smartphone's camera function is used to simultaneously record the golf swing and facial expressions.

[1729] 2. The user opens the application, selects the recorded video, and uploads it.

[1730] Use the application's file selection screen to upload the recorded video file to the designated location.

[1731] 3. The device sends the selected video file to the server.

[1732] File transmission is performed using HTTP POST requests or other appropriate communication methods.

[1733] 4. The server receives the video file and extracts frames for each scene of the swing.

[1734] The system analyzes the video content and executes an algorithm to extract specific scenes (address, backswing, top of swing, impact, finish).

[1735] 5. The server preprocesses the frames of each extracted scene.

[1736] The frames of each scene are standardized and converted into a format that allows image analysis algorithms to function properly. For example, this may involve unifying image sizes and removing noise.

[1737] 6. The server inputs the pre-processed images into the diagnostic algorithm.

[1738] The images from each scene, after preprocessing is complete, are input into the swing analysis algorithm for evaluation.

[1739] 7. The emotion engine analyzes the user's facial expressions and identifies their emotional state.

[1740] The system analyzes captured facial expression data to classify emotions such as joy, surprise, sadness, and anger.

[1741] 8. The server adjusts the swing analysis results based on the output of the emotion engine.

[1742] The system takes emotional data into account and adds emotional feedback to the diagnostic results. For example, if the user has an anxious expression, it provides more positive feedback.

[1743] 9. The server generates the final diagnostic results and sends them back to the terminal.

[1744] The diagnostic results are sent back to the terminal as an HTTP response.

[1745] 10. The terminal displays the diagnostic results to the user.

[1746] The device displays the received diagnostic results in the user interface, showing the user specific areas for improvement in their swing form, evaluations, and emotion-based feedback.

[1747] This system allows users to efficiently analyze their golf swing and receive specific feedback, as well as flexible advice tailored to their emotions. This is expected to contribute to increased user motivation and reduced stress.

[1748] The following describes the processing flow.

[1749] Step 1: The user simultaneously films their golf swing and facial expression using their smartphone.

[1750] Users use their smartphone's camera function to record themselves and their facial expressions while performing a golf swing in video format.

[1751] Step 2: The user opens the application and selects the recorded video.

[1752] The user opens the swing analysis application and selects the recorded video file.

[1753] Step 3: The user uploads a video file.

[1754] The user presses the upload button within the application and sends the selected video file to the server.

[1755] Step 4: The device sends the video file to the server.

[1756] The device uses an HTTP POST request to send the video file to the server.

[1757] Step 5: The server receives the video file.

[1758] The server retrieves and saves the video file from the received HTTP POST request.

[1759] Step 6: The server analyzes the video and extracts each scene.

[1760] The server analyzes the video frames to identify each scene—address, backswing, top of swing, impact, and finish—and extracts the corresponding frames.

[1761] Step 7: The server preprocesses the frames of each extracted scene.

[1762] The server standardizes the frames of each scene and converts them into a format that allows image analysis algorithms to function properly. For example, it performs image size unification and noise reduction.

[1763] Step 8: The server inputs the pre-processed images into the diagnostic algorithm.

[1764] The server inputs the pre-processed images of each scene into a swing diagnostic algorithm for evaluation.

[1765] Step 9: The emotion engine analyzes the user's facial expressions.

[1766] An emotion engine installed on the server analyzes the user's facial expressions in the video and classifies them into emotions such as joy, surprise, sadness, and anger.

[1767] Step 10: The server adjusts the swing diagnosis results based on the output of the emotion engine.

[1768] The server considers emotional data obtained from facial expressions and includes emotion-based feedback in the generated swing diagnosis results. For example, if the user has an anxious expression, it will provide more positive feedback.

[1769] Step 11: The server generates the final diagnostic results and sends them back to the terminal.

[1770] The server returns the adjusted diagnostic results to the terminal as an HTTP response.

[1771] Step 12: The terminal displays the diagnostic results to the user.

[1772] The device displays the received diagnostic results in the user interface, visually showing specific areas for improvement and evaluation of the swing form, as well as feedback based on the user's emotions.

[1773] These detailed processing steps allow users to efficiently analyze their golf swing and receive specific, emotionally responsive feedback. This provides not only improvement to their swing form but also mental support.

[1774] (Example 2)

[1775] Next, we will describe Example 2. In the following description, the data processing device 12 will be referred to as the "server" and the robot 414 as the "terminal".

[1776] Conventional exercise form analysis systems fail to consider the user's emotional state during video analysis and evaluation, sometimes leading to user stress or decreased motivation. Furthermore, the lack of specific feedback for each scene of exercise form makes it difficult for users to identify areas for improvement. Therefore, there is a need to develop a system that recognizes the user's emotions and reflects them in the diagnostic results, thereby providing more effective feedback and improving user motivation.

[1777] The identification process performed by the identification processing unit 290 of the data processing device 12 in Example 2 is realized by the following means.

[1778] In this invention, the server includes means for a user to upload a video or image of exercise, means for extracting each scene (start, in motion, end) from the uploaded video, means for executing a diagnostic algorithm to analyze the extracted scenes and evaluate the exercise form, means for analyzing facial expressions to identify the emotional state, means for adjusting the diagnostic results considering the emotional state, and means for generating the diagnostic results and displaying them to the user. This enables more effective and motivating feedback for the user by analyzing the user's emotions and reflecting them in the diagnostic results of the exercise form.

[1779] A "user" refers to someone who uses the system to upload videos or images of their exercise and receives the analysis results.

[1780] "Movement" refers to specific actions such as a golf swing, and is the subject of analysis by the system.

[1781] "Video or image" refers to data in file format in which a user records their exercise movements.

[1782] "Means for uploading" refers to methods and devices for sending videos and images taken by users to a server.

[1783] "Means for extraction" refers to algorithms or methods used to select specific scenes from uploaded videos.

[1784] A "diagnostic algorithm" refers to an algorithm used to analyze extracted scenes and evaluate their movement form.

[1785] "Means for analyzing facial expressions" refers to methods or devices for analyzing a user's facial expressions to identify their emotional state.

[1786] "Means of adjusting diagnostic results by considering emotional state" refers to methods or algorithms that modify or correct diagnostic results based on the analyzed emotional state.

[1787] "Means for generating diagnostic results" refers to methods or devices for compiling and providing users with evaluation results of exercise form.

[1788] "Means for display" refers to methods or devices for visually presenting the generated diagnostic results to the user.

[1789] A "scene" refers to any frame in a video of an exercise in which a specific action takes place.

[1790] "Start," "In Progress," and "End" refer to the scenes in the video that indicate the first, middle, and final stages of the exercise, respectively.

[1791] The system of the present invention not only allows users to upload videos or images of their exercise and analyzes and diagnoses their exercise form, but also includes a function to recognize the user's emotions and adjust the diagnosis results based on those emotions. This system consists of four main components: the user, the terminal, the server, and the emotion engine.

[1792] User behavior

[1793] Users record videos or images of their exercise (e.g., a golf swing) using a device such as a smartphone or camera. After recording, they access an application or web interface for exercise analysis and select the files. In addition, the user's facial expressions are recorded simultaneously during the upload process.

[1794] Terminal operation

[1795] The terminal's role is to send the video or image file selected by the user to the server. Communication methods such as HTTP POST requests are used for transmission. The terminal also displays the diagnostic results received from the server on the user interface.

[1796] Server Role

[1797] The server plays a central role in receiving and analyzing videos or images sent from the terminal. Specifically, it performs the following steps:

[1798] 1. Once the video is sent, the server analyzes the frames and extracts each scene of the motion (e.g., start, in motion, end).

[1799] Example of a tool used: OpenCV

[1800] 2. The motion diagnosis algorithm is executed on the extracted scenes and images to generate the diagnosis results.

[1801] 3. Analyze user emotional data and incorporate it into the diagnostic results.

[1802] Example of tools used: Microsoft Azure Cognitive Services

[1803] Functions of the Emotion Engine

[1804] The emotion engine analyzes the user's facial expressions to identify their emotional state. This allows the system to understand the user's emotions when they upload exercise videos or images and adjust the exercise diagnosis results accordingly.

[1805] Specific examples of program processing

[1806] 1. The user simultaneously records a video of their exercise and their facial expressions using their smartphone.

[1807] The smartphone's camera function is used to record exercise and facial expressions simultaneously.

[1808] 2. The user opens the application, selects the recorded video, and uploads it.

[1809] Use the application's file selection screen to upload the recorded video file to the designated location.

[1810] 3. The device sends the selected video file to the server.

[1811] File transmission is performed using HTTP POST requests or other appropriate communication methods.

[1812] 4. The server receives the video file and extracts frames for each scene of the exercise.

[1813] The program analyzes the video content and executes an algorithm to extract specific scenes (start, during, and end). For example, it uses OpenCV.

[1814] 5. The server preprocesses the frames of each extracted scene.

[1815] The frames of each scene are standardized and converted into a format that allows image analysis algorithms to function properly. For example, this may involve unifying image sizes and removing noise.

[1816] 6. The server inputs the pre-processed images into the diagnostic algorithm.

[1817] The images from each scene, after preprocessing is complete, are input into a motion diagnosis algorithm for evaluation.

[1818] 7. The emotion engine analyzes the user's facial expressions and identifies their emotional state.

[1819] The system analyzes captured facial expression data to classify emotions such as joy, surprise, sadness, and anger. For example, it uses Microsoft Azure Cognitive Services.

[1820] 8. The server adjusts the motor diagnosis results based on the output of the emotion engine.

[1821] The system takes emotional data into account and adds emotional feedback to the diagnostic results. For example, if the user has an anxious expression, it provides more positive feedback.

[1822] 9. The server generates the final diagnostic results and sends them back to the terminal.

[1823] The diagnostic results are sent back to the terminal as an HTTP response.

[1824] 10. The terminal displays the diagnostic results to the user.

[1825] The device displays the received diagnostic results in the user interface, showing the user specific areas for improvement in their exercise form, evaluations, and emotion-based feedback.

[1826] Example of a prompt

[1827] "Please upload a video of your swing. We will analyze your swing form and your emotions, and provide feedback."

[1828] The flow of the specific processing in Example 2 will be explained using Figure 13.

[1829] Step 1:

[1830] The user simultaneously records a video of their exercise and their facial expressions using their smartphone.

[1831] Input: Smartphone camera function

[1832] Instructions: Mount your smartphone on a tripod and position it so that your whole body is visible. Press the record button to record your physical movements and facial expressions simultaneously.

[1833] Output: Exercise video file and facial expression video file

[1834] Step 2:

[1835] The user opens the application for swing analysis, selects the recorded video, and uploads it.

[1836] Input: Recorded exercise video file

[1837] How to do it: Tap the "Select Video" button in the application, choose the video file you've recorded, and then tap "Upload" to send it to the server.

[1838] Output: Exercise video file ready to upload

[1839] Step 3:

[1840] The device sends the selected video file to the server.

[1841] Input: Exercise video file selected by the user and uploaded by pressing the upload button.

[1842] Operation: The device sends a video file to the server using an HTTP POST request and displays the transmission progress.

[1843] Output: Exercise video file sent to the server

[1844] Step 4:

[1845] The server extracts each scene frame by frame from the received video file.

[1846] Input: Exercise video file sent to the server

[1847] Operation: The video is divided into frames using a video analysis library (e.g., OpenCV), and the start, middle, and end scenes are identified and extracted.

[1848] Output: Frame images divided by scene

[1849] Step 5:

[1850] The server preprocesses the extracted frame images.

[1851] Input: Frame images divided by scene

[1852] Operation: Unify image sizes and smooth images by applying a denoising filter. For example, unify the resolution to 640x480 pixels and run a denoising algorithm.

[1853] Output: Preprocessed frame image

[1854] Step 6:

[1855] The server inputs the pre-processed images into the motion diagnosis algorithm.

[1856] Input: Pre-processed frame image

[1857] Operation: Pre-processed frame images are passed to a motion diagnostic algorithm for evaluation scoring and analysis of movement form. For example, swing trajectory and body tilt are quantified.

[1858] Output: Evaluation results and scores of exercise form

[1859] Step 7:

[1860] The emotion engine analyzes the user's facial expressions to identify their emotional state.

[1861] Input: Facial expression video file being sent to the server

[1862] Operation: Uses an emotion analysis model (e.g., Microsoft Azure Cognitive Services) to classify emotions such as joy, surprise, sadness, and anger from facial expressions.

[1863] Output: Identified emotional state

[1864] Step 8:

[1865] The server adjusts the motor diagnosis results based on the output of the emotion engine.

[1866] Input: Evaluation results and scores of exercise form, identified emotional state

[1867] Operation: It takes emotional data into account and adjusts the diagnostic results to be positive or negative. For example, if the expression is anxious, it adds a positive message such as "Stay calm and continue practicing at this pace."

[1868] Output: Adjusted final diagnostic results

[1869] Step 9:

[1870] The server generates the final diagnostic results and sends them back to the terminal.

[1871] Input: Adjusted final diagnostic result

[1872] Function: Organizes the diagnostic results and sends them to the terminal in JSON format or another appropriate format.

[1873] Output: Final diagnostic results sent to the terminal

[1874] Step 10:

[1875] The terminal displays the diagnostic results to the user.

[1876] Input: Final diagnostic results received from the server

[1877] Operation: The user interface visually presents evaluation results of exercise form, areas for improvement, and emotion-based feedback. For example, it might display, "Your movement is generally good. Be careful as your body tends to move forward at the finish."

[1878] Output: Diagnostic results and feedback displayed to the user

[1879] Example of a prompt

[1880] "Please upload a video of your swing. We will analyze your swing form and your emotions, and provide feedback."

[1881] (Application Example 2)

[1882] Next, we will explain application example 2. In the following explanation, the data processing device 12 will be referred to as the "server" and the robot 414 as the "terminal".

[1883] Traditional swing analysis systems only evaluate the user's swing form and do not consider feedback based on the user's emotions, thus limiting their effectiveness in improving user motivation and reducing stress. Furthermore, they lacked the functionality to provide specific training methods, preventing them from directly contributing to user skill improvement.

[1884] The specific processing performed by the specific processing unit 290 of the data processing device 12 in Application Example 2 is realized by the following means.

[1885] In this invention, the server includes means for analyzing the user's facial expressions to identify their emotional state, means for adjusting the swing diagnosis results based on the emotional state, and means for providing specific training methods based on an evaluation of the swing form. This allows the user to receive flexible feedback tailored to their emotional state and specific methods for improvement.

[1886] A "user" is an individual or group that uses the system, uploads videos or images of their swing, and receives the analysis results.

[1887] A "video" is video data composed of a series of image frames, and is used to record the user's swing motion.

[1888] An "image" is visual data represented in a single still image format and is used to capture a specific scene of a swing.

[1889] A "scene" refers to a visual representation of each stage of the swing (address, backswing, top, impact, finish).

[1890] A "diagnostic algorithm" is a set of calculation procedures and rules used to analyze and evaluate swing form, and the system uses them to scrutinize each aspect of the swing.

[1891] "Facial expression" refers to the various muscle movements and arrangements that appear on a user's face, and serves as basic information for inferring their emotional state.

[1892] "Emotional state" refers to the internal psychological state analyzed from the user's facial expressions, and includes classifications such as joy, surprise, sadness, and anger.

[1893] "Training methods" refer to specific practice procedures and exercises aimed at improving swing form and technique.

[1894] In an embodiment for carrying out this invention, the system includes the following components.

[1895] User

[1896] A "user" is an individual or group that uses the system, and is the entity that takes videos or images of their swing using a device such as a smartphone and uploads them. Users can receive not only technical evaluations of their swing form, but also emotional feedback derived from their facial expressions.

[1897] terminal

[1898] A "device" is a device used by a user to upload data, and includes smartphones and tablets. A device utilizes the following hardware and software:

[1899] Hardware: Smartphones (iPhone, Android)

[1900] Software: Dedicated application (Golf Swing Master)

[1901] Its role is to handle everything from video recording to data transmission, and then receive diagnostic results from the server and display them to the user.

[1902] server

[1903] A "server" is a central processing unit that receives data transmitted from terminals and performs analysis and diagnosis. The server fulfills the following roles:

[1904] Swing Analysis: Frames are extracted from the video, and a diagnostic algorithm is executed to analyze each scene (address, backswing, top, impact, finish).

[1905] Emotion Recognition: Analyzes the user's facial expressions to identify their emotional state and reflect it in the swing analysis results.

[1906] Result Adjustment: The diagnostic results are adjusted based on emotional data to provide users with appropriate feedback.

[1907] Emotional Engine

[1908] The "emotion engine" is a software module that analyzes a user's facial expressions and classifies their emotional state. The emotion engine is used to understand the user's psychological state when uploading swing videos and images.

[1909] Training provided

[1910] The server provides specific training methods based on an evaluation of swing form. This allows users to learn concrete practical techniques that help improve their skills.

[1911] Examples

[1912] For example, a user films their swing at a golf driving range with their smartphone and uploads the video to a dedicated app. The device sends the video data to a server, which analyzes the video to evaluate the swing form and uses an emotion engine to analyze the user's facial expressions to identify their emotional state. As a result, in addition to the technical evaluation, the server provides feedback and training methods tailored to the user's emotional state. This feedback might include something like, "Your swing is good, but try to relax a little more. Refer to this video and practice swinging in a more relaxed state."

[1913] Examples of prompts for generative AI models

[1914] Create an application that analyzes a user's golf swing video and facial expressions to provide a technical evaluation of their swing form and emotionally responsive feedback.

[1915] 1. How to extract frames from a video.

[1916] 2. A method for identifying emotions by analyzing facial expressions.

[1917] 3. How to send swing analysis results and emotional feedback to the server.

[1918] 4. How to receive diagnostic results and how to display them in the user interface.

[1919] The flow of a specific process in Application Example 2 will be explained using Figure 14.

[1920] Step 1:

[1921] The user simultaneously records a video of their golf swing and their facial expressions using their smartphone.

[1922] Input: Video to be recorded and facial expressions

[1923] Output: Video files saved on the smartphone's storage.

[1924] Specific operation: The user uses the smartphone's camera app to record a video, and at the same time, their facial expressions are also recorded using the front camera.

[1925] Step 2:

[1926] The user opens the application, selects the video they have recorded, and uploads it.

[1927] Input: User-recorded video

[1928] Output: Video file uploaded to the application

[1929] Specific steps: Open the file selection screen within the application, select the recorded video file, and press the upload button.

[1930] Step 3:

[1931] The device sends the selected video file to the server.

[1932] Input: Uploaded video file

[1933] Output: Video data sent to the server

[1934] Specific action: Send the video file to the server using an HTTP POST request.

[1935] Step 4:

[1936] The server receives the video file and extracts frames for each scene of the swing.

[1937] Input: Video data sent from the device

[1938] Output: Frame images of each extracted scene

[1939] Specific operation: Use OpenCV to extract frames from the video and separate the address, backswing, top of swing, impact, and finish scenes.

[1940] Step 5:

[1941] The server preprocesses each frame of the extracted scene.

[1942] Input: Extracted frame image

[1943] Output: Preprocessed image data

[1944] Specific actions: Standardize frames, unify image sizes, and remove noise.

[1945] Step 6:

[1946] The server inputs the pre-processed images into the diagnostic algorithm.

[1947] Input: Preprocessed image data

[1948] Output: Swing form evaluation results

[1949] Specific operation: Pre-processed images are input into the diagnostic algorithm, and a technical evaluation of the swing form is performed.

[1950] Step 7:

[1951] The emotion engine analyzes the user's facial expressions to identify their emotional state.

[1952] Input: Facial expression data included in the video

[1953] Output: Classified emotional state

[1954] Specific operation: Facial expression data is input into an emotion recognition model to classify emotions such as joy, surprise, sadness, and anger.

[1955] Step 8:

[1956] The server adjusts the swing analysis results based on the output of the emotion engine.

[1957] Input: Swing form evaluation results and emotional state

[1958] Output: Final diagnostic results including emotional feedback

[1959] Specific actions: Consider emotional data and incorporate emotional feedback into the diagnostic results, generating feedback such as, "Your swing is good, but try to relax a little more."

[1960] Step 9:

[1961] The server generates the final diagnostic results and sends them back to the terminal.

[1962] Input: Final diagnostic results including emotional feedback

[1963] Output: Diagnostic results sent to the terminal

[1964] Specific operation: The final diagnostic results are sent back to the terminal using an HTTP response.

[1965] Step 10:

[1966] The terminal displays the diagnostic results to the user.

[1967] Input: Diagnostic results received from the server

[1968] Output: Diagnostic results displayed in the user interface

[1969] Specific action: Display the received feedback and swing evaluation in the application's user interface.

[1970] The specific processing unit 290 transmits the result of the specific processing to the robot 414. In the robot 414, the control unit 46A causes the speaker 240 and the controlled object 443 to output the result of the specific processing. The microphone 238 acquires audio indicating user input for the result of the specific processing. The control unit 46A transmits the audio data indicating user input acquired by the microphone 238 to the data processing unit 12. In the data processing unit 12, the specific processing unit 290 acquires the audio data.

[1971] Data generation model 58 is a type of so-called generative AI (Artificial Intelligence). One example of data generation model 58 is ChatGPT (Internet search<URL: https: / / openai.com / blog / chatgpt> ), Gemini (Internet search) <url: https: gemini.google.com ?hl="ja">Examples of generative AI include the following. The data generation model 58 is obtained by performing deep learning on a neural network. The data generation model 58 is input with prompts containing instructions, and with inference data such as audio data representing speech, text data representing text, and image data representing images. The data generation model 58 infers from the input inference data according to the instructions indicated by the prompts, and outputs the inference results in data formats such as audio data and text data. Here, inference refers to, for example, analysis, classification, prediction, and / or summarization.

[1972] In the above embodiment, an example was given in which specific processing is performed by the data processing device 12, but the technology of this disclosure is not limited thereto, and the specific processing may also be performed by the robot 414.

[1973] Furthermore, the emotion identification model 59, acting as an emotion engine, may determine the user's emotion according to a specific mapping. Specifically, the emotion identification model 59 may determine the user's emotion according to a specific mapping, which is an emotion map (see Figure 9). Similarly, the emotion identification model 59 may also determine the robot's emotion, and the identification processing unit 290 may perform identification processing using the robot's emotion.

[1974] Figure 9 shows an emotion map 400 in which multiple emotions are mapped. In the emotion map 400, emotions are arranged in concentric circles radiating from the center. The closer to the center of the concentric circles, the more primitive the emotions are located. Further out of the concentric circles, emotions representing states and actions arising from mental states are located. Emotion is a concept that includes feelings and mental states. On the left side of the concentric circles, emotions that are generally generated from reactions occurring in the brain are located. On the right side of the concentric circles, emotions that are generally induced by situational judgment are located. Above and below the concentric circles, emotions that are generally generated from reactions occurring in the brain and induced by situational judgment are located. In addition, the emotion of "pleasure" is located on the upper side of the concentric circles, and the emotion of "displeasure" is located on the lower side. Thus, in the emotion map 400, multiple emotions are mapped based on the structure in which emotions arise, and emotions that are likely to occur simultaneously are mapped close together.

[1975] These emotions are distributed at the 3 o'clock position on the Emotion Map 400, and usually fluctuate between feelings of security and anxiety. In the right half of the Emotion Map 400, situational awareness takes precedence over internal feelings, resulting in a calm impression.

[1976] The inside of the Emotion Map 400 represents inner thoughts, while the outside represents actions. Therefore, the further you go from the outside of the Emotion Map 400, the more visible (expressed in actions) your emotions become.

[1977] Here, human emotions are based on various balances, such as posture and blood sugar levels. When these balances deviate from the ideal, it results in discomfort, and when they approach the ideal, it results in pleasure. Similarly, in robots, cars, motorcycles, etc., emotions can be created based on various balances, such as posture and battery level. When these balances deviate from the ideal, it results in discomfort, and when they approach the ideal, it results in pleasure. The emotion map can be generated, for example, based on Dr. Mitsuyoshi's emotion map (Research on a system for analyzing brain physiological signals of speech emotion recognition and emotion, Tokushima University, doctoral dissertation: https: / / ci.nii.ac.jp / naid / 500000375379). The left half of the emotion map contains emotions belonging to a region called "response," where sensation is dominant. The right half of the emotion map contains emotions belonging to a region called "situation," where situational awareness is dominant.

[1978] The emotion map defines two emotions that promote learning. One is the emotion around the middle of the negative "repentance" and "reflection" on the situation side. In other words, it is when the robot experiences negative emotions such as "I never want to feel this way again" or "I don't want to be scolded again." The other is the emotion around the positive "desire" on the reaction side. In other words, it is when the robot has positive feelings such as "I want more" or "I want to know more."

[1979] The emotion identification model 59 inputs user input into a pre-trained neural network, obtains emotion values ​​representing each emotion shown in the emotion map 400, and determines the user's emotion. This neural network is pre-trained based on multiple training data sets, which are combinations of user input and emotion values ​​representing each emotion shown in the emotion map 400. Furthermore, this neural network is trained so that emotions located close together have similar values, as shown in the emotion map 900 in Figure 10. Figure 10 shows an example where multiple emotions such as "reassured," "calm," and "confident" have similar emotion values.

[1980] The above description primarily focuses on the functions of the data processing device 12 in relation to this disclosure. However, the system related to this disclosure is not necessarily implemented on a server. The system related to this disclosure may be implemented as a general information processing system. This disclosure may be implemented, for example, as a software program that runs on a personal computer or as an application that runs on a smartphone. The method related to this disclosure may be provided to users in SaaS (Software as a Service) format.

[1981] In the above embodiment, an example was given in which a specific process is performed by a single computer 22. However, the technology of this disclosure is not limited thereto, and a distributed processing of the specific process may be performed by multiple computers, including computer 22. For example, a data generation model 58 may be provided in an external device of the data processing device 12, and the external device may generate data according to the input data.

[1982] In the above embodiment, an example was given in which the specific processing program 56 is stored in the storage 32, but the technology of this disclosure is not limited thereto. For example, the specific processing program 56 may be stored in a portable, computer-readable, non-temporary storage medium such as a USB (Universal Serial Bus) memory. The specific processing program 56 stored in the non-temporary storage medium is installed in the computer 22 of the data processing device 12. The processor 28 executes specific processing according to the specific processing program 56.

[1983] Alternatively, the specific processing program 56 may be stored in a storage device such as a server connected to the data processing device 12 via the network 54, and the specific processing program 56 may be downloaded and installed on the computer 22 in response to a request from the data processing device 12.

[1984] Furthermore, it is not necessary to store the entirety of the specific processing program 56 in a storage device such as a server connected to the data processing device 12 via the network 54, or to store the entirety of the specific processing program 56 in the storage 32; it is acceptable to store only a portion of the specific processing program 56.

[1985] The following types of processors can be used as hardware resources to perform specific processing. Examples of processors include a CPU, a general-purpose processor that functions as a hardware resource to perform specific processing by executing software, i.e., a program. Other examples of processors include dedicated electrical circuits, such as FPGAs (Field-Programmable Gate Arrays), PLDs (Programmable Logic Devices), or ASICs (Application Specific Integrated Circuits), which have circuit configurations specifically designed to perform specific processing. All of these processors have built-in or connected memory, and all of them perform specific processing by using memory.

[1986] The hardware resource that performs a specific process may consist of one of these various processors, or it may consist of a combination of two or more processors of the same or different types (for example, a combination of multiple FPGAs, or a combination of a CPU and an FPGA). Alternatively, the hardware resource that performs a specific process may consist of a single processor.

[1987] Examples of configurations using a single processor include, firstly, a configuration in which one or more CPUs and software are combined to form a single processor, and this processor functions as a hardware resource that performs a specific process. Secondly, there is a configuration using a processor that realizes the functions of the entire system, including multiple hardware resources that perform a specific process, on a single IC chip, as exemplified by SoCs (System-on-a-chip). In this way, a specific process is realized using one or more of the above types of processors as hardware resources.

[1988] Furthermore, the hardware structure of these various processors can more specifically utilize electrical circuits that combine circuit elements such as semiconductor devices. Also, the specific processing described above is merely an example. Therefore, it goes without saying that unnecessary steps can be deleted, new steps added, or the processing order rearranged, as long as it does not deviate from the main purpose.

[1989] The descriptions and illustrations presented above are detailed explanations of the technical aspects of this disclosure and are merely examples of the technical aspects. For example, the above descriptions of the structure, function, operation, and effect are examples of the structure, function, operation, and effect of the technical aspects of this disclosure. Therefore, it goes without saying that you may delete unnecessary parts, add new elements, or replace elements in the descriptions and illustrations presented above, as long as you do not deviate from the essence of the technical aspects of this disclosure. Furthermore, in order to avoid confusion and facilitate understanding of the technical aspects of this disclosure, explanations of common technical knowledge and the like that do not require special explanation to enable the implementation of the technical aspects of this disclosure have been omitted from the descriptions and illustrations presented above.

[1990] All documents, patent applications, and technical standards described herein are incorporated by reference to the same extent as if each individual document, patent application, and technical standard were specifically and individually noted as being incorporated by reference.

[1991] The following is further disclosed regarding the embodiments described above.

[1992] (Claim 1)

[1993] A means for users to upload videos or images of their swings,

[1994] A method for extracting each scene (address, backswing, top, impact, finish) from uploaded videos,

[1995] A means for analyzing extracted scenes and executing a diagnostic algorithm to evaluate swing form,

[1996] A system that includes means for generating and displaying diagnostic results to the user.

[1997] (Claim 2)

[1998] The system according to claim 1, further comprising means for analyzing uploaded images and executing a diagnostic algorithm for evaluating swing form.

[1999] (Claim 3)

[2000] The system according to claim 1, further comprising means for adding annotations to images of each scene to indicate areas for improvement.

[2001] "Example 1"

[2002] (Claim 1)

[2003] A means for users to upload videos or images of their swings,

[2004] A method for analyzing and extracting each scene (address, backswing, top, impact, finish) frame by frame from uploaded videos,

[2005] A means for analyzing extracted scenes and executing a diagnostic algorithm to evaluate swing form,

[2006] A means for returning and displaying the generated diagnostic results to the user,

[2007] A means to generate specific feedback based on the diagnostic results (e.g., issues with body balance or timing),

[2008] A system that includes a means of adding annotations to images of each scene as feedback to indicate areas for improvement.

[2009] (Claim 2)

[2010] The system according to claim 1, further comprising means for analyzing uploaded images and executing a diagnostic algorithm for evaluating swing form.

[2011] (Claim 3)

[2012] The system according to claim 1, further comprising means for returning the generated diagnostic results to the user terminal and displaying them on the user interface.

[2013] "Application Example 1"

[2014] (Claim 1)

[2015] A means for users to upload videos or images of their swings,

[2016] A method for extracting each scene (address, backswing, top, impact, finish) from uploaded videos,

[2017] A means for analyzing extracted scenes and executing a diagnostic algorithm to evaluate swing form,

[2018] A means for generating and displaying diagnostic results to the user,

[2019] A system that includes a means of providing real-time feedback of the displayed diagnostic results to a wearable device worn by the user.

[2020] (Claim 2)

[2021] The system according to claim 1, further comprising means for analyzing uploaded images and executing a diagnostic algorithm for evaluating swing form, and means for displaying the feedback results on a wearable device.

[2022] (Claim 3)

[2023] The system according to claim 1, further comprising means for adding annotations to images of each scene to indicate areas for improvement, and means for displaying these annotations on a wearable device.

[2024] "Example 2 of combining an emotion engine"

[2025] (Claim 1)

[2026] A means for users to upload videos or images of their exercise,

[2027] A method for extracting each scene (start, in progress, end) from an uploaded video,

[2028] A means for analyzing extracted scenes and executing a diagnostic algorithm to evaluate movement form,

[2029] A method for analyzing facial expressions to identify emotional states,

[2030] A means of adjusting the diagnostic results to take into account the emotional state,

[2031] A system that includes means for generating and displaying diagnostic results to the user.

[2032] (Claim 2)

[2033] The system according to claim 1, further comprising means for analyzing uploaded images and executing a diagnostic algorithm for evaluating exercise form.

[2034] (Claim 3)

[2035] The system according to claim 1, further comprising means for adding annotations to images of each scene to indicate areas for improvement.

[2036] "Application example 2 when combining with an emotional engine"

[2037] (Claim 1)

[2038] A means for users to upload videos or images of their swings,

[2039] A method for extracting each scene (address, backswing, top, impact, finish) from uploaded videos,

[2040] A means for analyzing extracted scenes and executing a diagnostic algorithm to evaluate swing form,

[2041] A means of analyzing a user's facial expressions to identify their emotional state,

[2042] A means of adjusting swing analysis results based on emotional state,

[2043] A means for generating and displaying diagnostic results to the user,

[2044] A system that includes means for providing specific training methods based on an evaluation of swing form.

[2045] (Claim 2)

[2046] The system according to claim 1, further comprising means for analyzing uploaded images and executing a diagnostic algorithm for evaluating swing form.

[2047] (Claim 3)

[2048] The system according to claim 1, further comprising means for adding annotations to images of each scene to indicate areas for improvement, and means for providing feedback based on the user's emotional state. [Explanation of Symbols]

[2049] 10, 210, 310, 410 Data Processing Systems 12 Data Processing Devices 14 Smart Devices 214 Smart Glasses 314 Headset-type terminal 414 Robots< / url:> < / url:> < / url:> < / url:>

Claims

1. A means for users to upload videos or images of their swings, A method for extracting each scene from an uploaded video, A means for analyzing extracted scenes and executing a diagnostic algorithm to evaluate swing form, A system that includes means for generating and displaying diagnostic results to the user.

2. The system according to claim 1, further comprising means for analyzing uploaded images as they are and executing a diagnostic algorithm for evaluating swing form.

3. The system according to claim 1, further comprising means for adding annotations to images of each scene to indicate areas for improvement.

Citation Information

Patent Citations

  • Persona chatbot control method and system

    JP2022180282A