System
A system that analyzes user biometric and experience data to provide tailored, real-time feedback and ideal movement videos enhances batting and pitching practice by addressing the lack of individualized expert feedback in conventional methods.
Patent Information
- Application Number
- JP2024122692
- Authority / Receiving Office
- JP · JP
- Patent Type
- Applications
- Current Assignee / Owner
- Filing Date
- 2024-07-29
- Publication Date
- 2026-02-10
AI Technical Summary
Conventional batting and pitching practice methods lack specific, expert feedback tailored to individual characteristics, making it difficult to achieve efficient training, and users struggle to understand their ideal movements for improvement.
A system that inputs user biometric and experience information, records movements, analyzes the data, provides real-time feedback in various formats, allows selection of multiple advisors, and generates videos showing ideal movements, optimizing training based on individual needs.
Enables users to receive specialized and specific instruction, facilitating efficient training and improvement of batting and pitching techniques.
Smart Images

Figure 2026021010000001_ABST
Abstract
Description
[Technical Field]
[0001] The technology of the present disclosure relates to a system. [Background technology]
[0002] Patent document 1 discloses a persona chatbot control method performed by at least one processor, the method including the steps of receiving a user utterance, adding the user utterance to a prompt including an instruction sentence related to a description of the chatbot character, encoding the prompt, and inputting the encoded prompt into a language model to generate a chatbot utterance in response to the user utterance. [Prior art documents] [Patent documents]
[0003] [Patent Document 1] Japanese Patent Publication No. 2022-180282 Summary of the Invention [Problem to be solved by the invention]
[0004] Conventional batting center and pitching practice methods often rely on self-training or advice from friends, making it difficult to obtain specific, expert feedback. Furthermore, training methods tailored to individual characteristics such as age, gender, and baseball experience are not provided, making it difficult to achieve efficient practice. Furthermore, it is difficult for users to understand their ideal movements in a visual format and use them as guidelines for improvement. [Means for solving the problem]
[0005] The present invention solves the above-mentioned problems by providing a means for inputting a user's biometric information and experience information, a means for recording the user's movements, a means for analyzing the recorded movement data, a means for providing feedback on the analysis results in various formats, a means for allowing the user to select instruction from multiple advisors, and a means for generating a video showing the user's ideal movements. By analyzing the recorded movement data via a cloud server and providing the analysis results to the user in real time, optimal feedback tailored to individual characteristics is realized. This allows the user to receive specialized and specific instruction, enabling efficient training.
[0006] "Biometric information" refers to physical data about an individual, such as a user's age, gender, or medical limitations.
[0007] "Experience information" refers to historical data about the user's past practice and play, such as their baseball experience and athletic history.
[0008] "User's movements" refer to physical movements related to the sport the user plays, such as batting or pitching.
[0009] "Recording means" refers to a device or system for recording or saving user actions as sensor data.
[0010] "Analysis means" refers to algorithms or software used to analyze recorded performance data and extract technical improvements and features.
[0011] "Feedback means" refers to a method or system for notifying the user of the analysis results and providing them in the form of text, audio, images, video, etc.
[0012] "Advisor" refers to a virtual or real person who has a mentor profile that a user can select and who provides guidance.
[0013] "Instruction selection means" refers to an interface or system that allows a user to select from multiple advisors.
[0014] "Video Generation Means" refers to software and algorithms for generating visual content that shows the user's ideal movements.
[0015] "Cloud server" refers to the infrastructure and services provided for remotely processing and storing data. [Brief explanation of the drawings]
[0016] [Figure 1] 1 is a conceptual diagram showing an example of the configuration of a data processing system according to a first embodiment. [Figure 2] 1 is a conceptual diagram showing an example of main functions of a data processing device and a smart device according to a first embodiment. [Figure 3] FIG. 10 is a conceptual diagram showing an example of the configuration of a data processing system according to a second embodiment. [Figure 4] FIG. 10 is a conceptual diagram showing an example of main functions of a data processing device and smart glasses according to a second embodiment. [Figure 5] FIG. 10 is a conceptual diagram showing an example of the configuration of a data processing system according to a third embodiment. [Figure 6] FIG. 11 is a conceptual diagram showing an example of main functions of a data processing device and a headset-type terminal according to a third embodiment. [Figure 7] FIG. 10 is a conceptual diagram showing an example of the configuration of a data processing system according to a fourth embodiment. [Figure 8] FIG. 10 is a conceptual diagram showing an example of main functions of a data processing device and a robot according to a fourth embodiment. [Figure 9] 1 shows an emotion map onto which multiple emotions are mapped. [Figure 10] 1 shows an emotion map onto which multiple emotions are mapped. [Figure 11] FIG. 3 is a sequence diagram showing a processing flow of the data processing system according to the first embodiment. [Figure 12] FIG. 10 is a sequence diagram showing the flow of processing in the data processing system in Application Example 1. [Figure 13] FIG. 10 is a sequence diagram showing the flow of processing in the data processing system according to the second embodiment when an emotion engine is combined. [Figure 14] FIG. 10 is a sequence diagram showing the flow of processing in the data processing system in Application Example 2 when an emotion engine is combined. DETAILED DESCRIPTION OF THE INVENTION
[0017] An example of an embodiment of a system according to the technology of the present disclosure will be described below with reference to the accompanying drawings.
[0018] First, the terms used in the following description will be explained.
[0019] In the following embodiments, a coded processor (hereinafter simply referred to as a "processor") may be a single arithmetic device or a combination of multiple arithmetic devices. Furthermore, a processor may be a single type of arithmetic device or a combination of multiple types of arithmetic devices. Examples of arithmetic devices include a CPU (Central Processing Unit), a GPU (Graphics Processing Unit), a GPGPU (General-Purpose computing on Graphics Processing Units), and an APU (Accelerated Processing Unit).
[0020] In the following embodiments, a coded RAM (Random Access Memory) is a memory in which information is temporarily stored and is used as a working memory by a processor.
[0021] In the following embodiments, the coded storage is one or more non-volatile storage devices that store various programs, various parameters, etc. Examples of non-volatile storage devices include flash memory (SSD (Solid State Drive)), magnetic disks (e.g., hard disks), and magnetic tapes.
[0022] In the following embodiments, a communication I / F (Interface) with a symbol is an interface including a communication processor, an antenna, etc. The communication I / F controls communication between multiple computers. Examples of communication standards applied to the communication I / F include wireless communication standards including 5G (5th Generation Mobile Communication System), Wi-Fi (registered trademark), Bluetooth (registered trademark), etc.
[0023] In the following embodiments, "A and / or B" is synonymous with "at least one of A and B." In other words, "A and / or B" means that it may be only A, only B, or a combination of A and B. Furthermore, in this specification, the same concept as "A and / or B" is also applied when three or more things are expressed connected by "and / or."
[0024] [First embodiment]
[0025] FIG. 1 shows an example of the configuration of a data processing system 10 according to the first embodiment.
[0026] 1, a data processing system 10 includes a data processing device 12 and a smart device 14. An example of the data processing device 12 is a server.
[0027] The data processing device 12 includes a computer 22, a database 24, and a communication I / F 26. The computer 22 is an example of a "computer" according to the technology of the present disclosure. The computer 22 includes a processor 28, a RAM 30, and a storage 32. The processor 28, the RAM 30, and the storage 32 are connected to a bus 34. The database 24 and the communication I / F 26 are also connected to the bus 34. The communication I / F 26 is connected to a network 54. Examples of the network 54 include a WAN (Wide Area Network) and / or a LAN (Local Area Network).
[0028] The smart device 14 includes a computer 36, a reception device 38, an output device 40, a camera 42, and a communication I / F 44. The computer 36 includes a processor 46, a RAM 48, and a storage 50. The processor 46, the RAM 48, and the storage 50 are connected to a bus 52. The reception device 38, the output device 40, and the camera 42 are also connected to the bus 52.
[0029] The reception device 38 includes a touch panel 38A, a microphone 38B, and the like, and receives user input. The touch panel 38A detects contact with an indicator (for example, a pen or a finger) to receive user input by the touch of the indicator. The microphone 38B detects the user's voice to receive user input by voice. The control unit 46A transmits data indicating the user input received by the touch panel 38A and the microphone 38B to the data processing device 12. In the data processing device 12, the specific processing unit 290 acquires the data indicating the user input.
[0030] The output device 40 includes a display 40A and a speaker 40B, and presents data to the user 20 by outputting the data in a form of expression that the user 20 can perceive (for example, audio and / or text). The display 40A displays visible information such as text and images in accordance with instructions from the processor 46. The speaker 40B outputs audio in accordance with instructions from the processor 46. The camera 42 is a compact digital camera equipped with an optical system including a lens, aperture, and shutter, and an imaging element such as a CMOS (Complementary Metal-Oxide-Semiconductor) image sensor or a CCD (Charge Coupled Device) image sensor.
[0031] The communication I / F 44 is connected to a network 54. The communication I / Fs 44 and 26 control the exchange of various information between the processor 46 and the processor 28 via the network 54.
[0032] FIG. 2 shows an example of the main functions of the data processing device 12 and the smart device 14.
[0033] 2, in the data processing device 12, a specific process is performed by the processor 28. A specific processing program 56 is stored in the storage 32. The specific processing program 56 is an example of a "program" according to the technology of the present disclosure. The processor 28 reads the specific processing program 56 from the storage 32 and executes the read specific processing program 56 on the RAM 30. The specific process is realized by the processor 28 operating as a specific processing unit 290 in accordance with the specific processing program 56 executed on the RAM 30.
[0034] The storage 32 stores a data generation model 58 and an emotion identification model 59. The data generation model 58 and the emotion identification model 59 are used by the identification processing unit 290.
[0035] In the smart device 14, the processor 46 performs the reception output process. The storage 50 stores a reception output program 60. The reception output program 60 is used in conjunction with the specific processing program 56 by the data processing system 10. The processor 46 reads the reception output program 60 from the storage 50 and executes the read reception output program 60 on the RAM 48. The reception output process is realized by the processor 46 operating as the control unit 46A in accordance with the reception output program 60 executed on the RAM 48.
[0036] Next, a description will be given of the specific processing performed by the specific processing unit 290 of the data processing device 12. In the following description, the data processing device 12 will be referred to as a "server" and the smart device 14 will be referred to as a "terminal."
[0037] This invention is a system for efficiently improving batting and pitching form based on a user's biometric and experience information. The system records the user's movements, analyzes the data, and provides specific feedback in real time. It also provides a means for users to choose from multiple advisors and generates videos demonstrating ideal movements, promoting the user's technical improvement.
[0038] Overall system configuration
[0039] Users use a device (e.g., a smartphone) to access the system. A dedicated application is installed on the device, allowing users to input their biometric information and experience information. The input information is sent to the server and stored in a database.
[0040] The device is equipped with a camera and sensors to record the user's movements. The captured video data is sent to a cloud server and analyzed using Google's Gemini. The analysis results are provided to the user in real time as feedback in the form of text, audio, images, video, etc.
[0041] The system also has multiple advisor profiles, allowing users to receive guidance from their preferred advisor. Furthermore, it can generate videos demonstrating ideal movements, providing users with visual reference information.
[0042] Program processing
[0043] The program does the following:
[0044] 1. User information registration
[0045] The user uses the device to input their biometric and experience information. Once input is complete, the device sends the data to the server, which stores the received information in a database.
[0046] 2. Recording a video of your form
[0047] Users record their swings and throws in the form diagnostic booth at a batting center. The motions are recorded using the device's camera or sensor, and the data is sent to a server.
[0048] 3. Form video analysis
[0049] The server sends the received video data to Google Gemini and makes an analysis request. Google Gemini analyzes the motion and returns the results to the server. The analysis results are then processed into a user-friendly format.
[0050] 4. Providing real-time feedback
[0051] The processed analysis results are sent to the device and notified to the user in the form of text, audio, images, video, etc. The user can then try to improve the form based on this information.
[0052] 5. Select coaching mode
[0053] Users can select one of several advisors within the app, and the analysis results are customized based on the profile information of the selected advisor.
[0054] 6. Image training using videos
[0055] If the user wishes, an image video of the ideal swing or throwing technique will be generated and sent to the device, and the user can use this video as a reference for training.
[0056] Specific examples
[0057] For example, if a beginner user wants to improve their pitching technique, they first enter and save their basic information in the app. They then film a video of themselves pitching at a batting center and transfer it to the app. The server analyzes the video and provides specific advice in text and voice, such as "position your elbow a little higher." Additionally, if the user selects a "famous pitching coach," they will receive customized feedback from that coach. Furthermore, a video of the ideal pitching form is also provided, allowing the user to use it as a reference to improve their technique.
[0058] The system allows users to receive specific, expert feedback in real time, allowing them to efficiently improve their form.
[0059] The processing flow will be explained below.
[0060] Step 1:
[0061] User: Launch the app and enter basic information such as age, gender, baseball experience, and medical restrictions on the "New Registration" screen.
[0062] Step 2:
[0063] Terminal: When input is complete, the user presses the save button, and the input data is temporarily saved in local storage.
[0064] Step 3:
[0065] Terminal: Generates and sends a request to send the stored data to the server.
[0066] Step 4:
[0067] Server: Saves the received user data in a database and returns a confirmation response to the terminal that the data was saved successfully.
[0068] Step 5:
[0069] User: Enters the form diagnostic booth at the batting center, stands in the designated position, and begins swinging or throwing.
[0070] Step 6:
[0071] Device: When an action is initiated, the device uses cameras and sensors to record the user's actions in video format.
[0072] Step 7:
[0073] User: When the action is finished, press the stop recording button to end the recording.
[0074] Step 8:
[0075] Terminal: Saves the captured video data in local storage, generates a request to send the data to the server, and sends it.
[0076] Step 9:
[0077] Server: Temporarily stores the received video data and sends an analysis request to Google's Gemini API.
[0078] Step 10:
[0079] Server: Receives analysis results from Google's "Gemini".
[0080] Step 11:
[0081] Server: Processes the received analysis results into a user-friendly format (text, audio, image, video).
[0082] Step 12:
[0083] Server: Sends the processed analysis results to the terminal.
[0084] Step 13:
[0085] Device: Receives analysis results and displays / notifies them on the app.
[0086] Step 14:
[0087] User: Check the analysis results and use the improvements to improve their next swing or throw.
[0088] Step 15:
[0089] User: Select one of several advisors on the "Coaching Mode Selection" screen within the app.
[0090] Step 16:
[0091] Terminal: Based on the information of the selected advisor, the analysis results are customized and provided to the user.
[0092] Step 17:
[0093] Terminal: Sends a request to the server to generate an image video of the ideal swing or throwing based on the user's wishes.
[0094] Step 18:
[0095] Server: Generates an image video of the ideal form and sends it to the device.
[0096] Step 19:
[0097] Device: The received image video is played within the app and shown to the user.
[0098] Step 20:
[0099] User: Use the image video as a reference and train to improve their own form.
[0100] This allows users to receive specific, professional feedback in real time and efficiently improve their form.
[0101] Example 1
[0102] Next, a description will be given of Example 1. In the following description, the data processing device 12 will be referred to as a "server" and the smart device 14 will be referred to as a "terminal."
[0103] Conventional methods for improving batting and pitching form lack the means for users to accurately understand their own movements and receive effective feedback. Furthermore, opportunities to receive direct instruction from a coach are limited, making self-study to improve skills difficult. The objective of this invention is to provide a system that provides specific, specialized feedback in real time, allowing users to efficiently improve their movements.
[0104] The specific processing by the specific processing unit 290 of the data processing device 12 in the first embodiment is realized by the following means.
[0105] In this invention, the server includes means for inputting biometric information and experience information of the user, a terminal device for recording the user's movements, means for transmitting the recorded video data to the analytical model via the cloud server, means for processing the analysis results from the analytical model into a user-friendly format, means for feeding back the analysis results to the user in real time, means for enabling the user to select instruction from multiple instructors, and means for generating videos showing the user's ideal movements. This allows the user to receive specific and professional feedback in real time and efficiently improve their own movements.
[0106] "Biometric information" refers to information about the user's physical characteristics, such as the user's height, weight, athletic ability, and health condition.
[0107] "Experience information" is information about the user's experience, such as the user's sports history, skill level, and past training history.
[0108] A "terminal device" is a hardware device, such as a smartphone or tablet, that allows a user to input information and record actions.
[0109] A "cloud server" is a remote server connected via the Internet for storing, managing, and analyzing data.
[0110] An "analytics model" is a machine learning algorithm or data processing mechanism that analyzes a user's behavior data and provides form improvements and technical advice.
[0111] "Feedback" refers to providing users with information such as advice, comments, and analysis results generated based on the user's behavior data.
[0112] "Instructors" are coaches or advisors with specialized knowledge and experience that can be selected by the user.
[0113] "Videos showing ideal movements" are videos that show the ideal form and techniques that users should aim for, and are used to support the user's training.
[0114] The present invention provides a system for effectively improving a user's sports form based on biometric information and experience information. Specific embodiments for carrying out the present invention will be described in detail below.
[0115] Overall system configuration
[0116] Registering user information
[0117] The user starts a dedicated application using a device (e.g., a smartphone) and inputs biometric information (e.g., height, weight, health status) and experience information (e.g., sports history, skill level). This information is sent from the device to a server and stored in a database.
[0118] Recording actions
[0119] Users can record their batting and pitching movements at batting centers or practice fields using their smartphone cameras or dedicated sensors. The recorded video data is temporarily stored on the device and then uploaded to a cloud server.
[0120] Video Analysis
[0121] The server then sends the received video data to an analytical model (e.g., a machine learning algorithm) in the cloud and issues an analysis request. The analytical model analyzes the user's movements and extracts data on form flaws and areas for improvement. The extracted results are returned to the server, where they are further processed into a format that is easy for the user to understand (e.g., text, images, video).
[0122] Providing Feedback
[0123] The server then sends the processed analysis results to the device in real time. The device then displays the results and provides specific advice to the user via text or voice, such as "position your elbows a little higher." Real-time video feedback is also provided based on the analysis results.
[0124] Choosing a coaching mode
[0125] Users can browse the profiles of multiple instructors from the in-app menu and select their preferred instructor. The server generates customized feedback based on the profile information of the selected instructor and sends it to the device.
[0126] Image training using videos
[0127] It can generate images that allow users to reproduce their ideal movements. The server generates a video of the ideal form and sends it to the device. Users can use this video as a reference for training and improve their form.
[0128] Specific examples
[0129] For example, if a beginner user wants to improve their pitching form, they first enter and save their basic information into the app on their smartphone. Next, they use the smartphone camera to film their pitching at a batting center, and the video is sent from the app to a cloud server. The cloud server analyzes the video and sends specific advice, such as "position your elbow a little higher," to the user's smartphone in real time. Furthermore, if the user selects a well-known coach, that coach will provide them with customized feedback. Finally, a video of their ideal pitching form is provided to the user, who can use the video as a reference to improve their technique.
[0130] Generative AI model prompt example
[0131] "Please analyze the batting form. What is the best way to teach the player if their elbow is positioned too low?"
[0132] "Please tell me the ideal throwing motion for pitching form."
[0133] In this way, by inputting specific questions into the generative AI model, optimal feedback can be obtained.
[0134] The flow of the identification process in the first embodiment will be described with reference to FIG.
[0135] Program processing steps
[0136] Step 1: Register user information
[0137] Input: User biometric and experience information
[0138] Processing: The user launches a dedicated application using a device (e.g., a smartphone) and enters basic information such as height, weight, and sports history into a form. The device then sends the entered information to the server using the HTTPS protocol.
[0139] Output: The entered user information is saved on the cloud server.
[0140] Specific operation: When a user enters information on their smartphone and presses the "Send" button, the information is recorded in a database (e.g., MySQL) on the cloud server.
[0141] Step 2: Record a video of your form
[0142] Input: User action (e.g. swing or throw)
[0143] Processing: A user uses a smartphone camera app to record their movements at a batting cage or practice area. Once the recording is complete, it is temporarily saved as a video file on the device.
[0144] Output: The video data file will be saved on your device.
[0145] Specific operation: The user presses the record button on the camera, and when they finish recording their movements, they press the "Save" button to save the video data to their device.
[0146] Step 3: Submitting the form video
[0147] Input: Video data stored on the device
[0148] Processing: The device uploads the recorded video file to a cloud server using a secure file transfer protocol (e.g., HTTPS or SFTP).
[0149] Output: Video data is stored in cloud server storage (e.g. AWS S3).
[0150] Specific operation: The user presses the "upload" button within the app, and the video is sent to the cloud server.
[0151] Step 4: Analysis instructions for form videos
[0152] Input: Video data stored on a cloud server
[0153] Processing: The server sends the video data to the analytical model and makes an analysis request. The analytical model uses a motion analysis algorithm.
[0154] Output: The analysis request is sent to the analysis model.
[0155] Specific operation: The server automatically checks the stored video data and sends an API request to the analysis model.
[0156] Step 5: Analyzing your form video
[0157] Input: Video data and analysis request
[0158] Processing: The analytical model analyzes user behavior from video data and generates data on form parameters and areas for improvement.
[0159] Output: The analysis result data is sent back to the server.
[0160] Specific behavior: The analytical model performs behavior analysis, generates analysis results, and sends them back to the server.
[0161] Step 6: Processing the analysis results and generating feedback data
[0162] Input: Analysis result data
[0163] Processing: The server receives the analysis results data and processes it into a format that is easy for the user to understand (e.g., text, charts, videos).
[0164] Output: Processed feedback data is generated.
[0165] Specific operation: The server processes the analysis result data for feedback and converts it into the most suitable format for the user.
[0166] Step 7: Provide feedback
[0167] Input: Processed feedback data
[0168] Processing: The server sends the feedback data to the device using a real-time notification service (e.g., Firebase Cloud Messaging). The device displays the feedback data to the user.
[0169] Output: Feedback is displayed on the user's terminal.
[0170] Specific operation: A notification is sent to the user's smartphone, and when they open the app, advice such as "position your elbows a little higher" is displayed.
[0171] Step 8: Select a coaching mode
[0172] Input: User-selected mentor profile
[0173] What happens: Users select one of several mentors within the app, and feedback is customized based on their profile information.
[0174] Output: Customized feedback data is generated.
[0175] Specific operation: The server regenerates feedback using the profile information of the selected instructor and sends it to the terminal.
[0176] Step 9: Image training using videos
[0177] Input: User training request
[0178] Processing: The server generates a video demonstrating the ideal movement and sends it to the device. The user uses this video as a reference for training.
[0179] Output: A video showing the ideal behavior is displayed on the device.
[0180] Specific movements: Users play videos of ideal form in the app and train by imitating the movements.
[0181] Generative AI model prompt example
[0182] "Please analyze the batting form. What is the best way to teach the player if their elbow is positioned too low?"
[0183] "Please tell me the ideal throwing motion for pitching form."
[0184] (Application example 1)
[0185] Next, a description will be given of Application Example 1. In the following description, the data processing device 12 will be referred to as a "server" and the smart device 14 will be referred to as a "terminal."
[0186] In order to ensure the efficiency and safety of industrial robots, it is important to evaluate the accuracy and appropriateness of their movements in real time and provide feedback on areas for improvement. However, existing systems do not fully realize these functions, making it difficult to optimize robot movements.
[0187] The specific processing by the specific processing unit 290 of the data processing device 12 in the application example 1 is realized by the following means.
[0188] In this invention, the server includes a means for transmitting user information to the server and storing it in a database, a means for transmitting operation records to the server and analyzing them using a model, and a means for converting the analysis results into a user-friendly format in real time and providing the results, thereby enabling efficient and safe optimization of the operation of industrial robots.
[0189] "User's biometric information" refers to data related to the user's body, including information such as heart rate, body temperature, and muscle movements.
[0190] "Experience information" refers to the knowledge and experience that the user has accumulated up to now, and is data such as proficiency level and past performance records.
[0191] A "recording means" is a device that collects the user's actions using a video camera or a sensor.
[0192] "Means for analyzing" refers to algorithms or software for evaluating recorded motion data using analytical models.
[0193] "Feedback means" refers to means for providing analysis results to the user in the form of text, audio, images, video, etc.
[0194] The "means for enabling selection of instructor" refers to an interface that allows the user to select one instructor from multiple instructors and receive evaluation and advice from that instructor.
[0195] The "means for generating animation" refers to a system that creates animation to visually show the user's ideal movements.
[0196] A "server" is a computer system that processes, manages, and provides data over a network.
[0197] The "means for saving to a database" is a function for safely saving input user information to a database.
[0198] "Means of analysis using models" refers to means of analyzing data using technologies such as machine learning and artificial intelligence to generate insights and improvement suggestions.
[0199] "Means for converting the analysis results into a user-friendly format and providing them" refers to a method for converting the analysis results into an easy-to-understand and comprehend format and providing them to the user.
[0200] The "means for proposing improvements to operations" is a function for proposing specific improvements to the user based on the analysis results.
[0201] A specific system for implementing this invention includes the following components: First, a user uses a terminal to input their own biometric information and experience information. A dedicated application is installed on the terminal, which transmits the user's information to a server and stores it in a database. The system records the user's movements using a camera or sensor and evaluates the data using an analytical model. The analysis results are processed on a cloud server and converted into a user-friendly format. These results are then fed back to the user in real time.
[0202] The system also features an interface that allows users to select instruction from multiple instructors. Users can receive customized feedback from the instructor of their choice based on their own movement data. Furthermore, the system generates an ideal movement model based on the analysis results and makes suggestions for movement improvement based on that model.
[0203] Hardware and software used
[0204] Hardware
[0205] Camera (e.g. Logitech C920)
[0206] Computer system (e.g., CPU: Intel i7, RAM: 16GB, etc.)
[0207] Robots (e.g., industrial arm robots)
[0208] software
[0209] Python programming language
[0210] OpenCV library (motion capture and image processing)
[0211] Requests library (sending HTTP requests)
[0212] Server-side program (runs on a cloud server)
[0213] Analytical models (using machine learning and artificial intelligence techniques)
[0214] Processing example
[0215] For example, when improving the behavior of a welding robot in a factory to ensure accurate welding, a camera captures the robot's movements and sends the data to a cloud server for analysis. Specific feedback such as "The welding position needs to be moved 5 mm to the left" is provided as a result of the analysis. A video demonstrating the ideal welding behavior is also provided to the user.
[0216] Prompt Sentence Examples
[0217] "Please create a system that records the operations of welding robots in factories with a camera, analyzes them on a cloud server, and provides suggestions for improvement. Please provide specific feedback in real time based on the analysis results."
[0218] This system makes it possible to optimize the operation of industrial robots efficiently and safely.
[0219] The flow of the specific processing in the application example 1 will be described with reference to FIG.
[0220] Step 1:
[0221] Registering user information
[0222] The user inputs their biometric and experience information into the device. This input data includes physical information such as heart rate, body temperature, and muscle movement, as well as skill level and past performance records. The device sends this data to the server, which stores the received data in a database. The input of this process is the biometric and experience information entered by the user, and the output is the user information stored in the database.
[0223] Step 2:
[0224] Shooting a video of the operation
[0225] A camera records the user's actions. In this case, the object of the action is an industrial robot, for example, recording the robot's welding work. The camera (e.g., Logitech C920) captures the robot's actions and saves them as a video file. The input of this process is the robot's actions, and the output is the recorded video file.
[0226] Step 3:
[0227] Analysis of motion video
[0228] The device sends the recorded video file to a server. The server receives this video data and analyzes it using an analytical model (using machine learning and artificial intelligence techniques). The analytical model takes the video data as input and evaluates the accuracy and appropriateness of the movements, which then generates specific suggestions for improving the movements. The input to this process is the video file, and the output is the analysis results.
[0229] Step 4:
[0230] Providing real-time feedback
[0231] The server provides the analysis results to the user in real time by converting the analysis results into a user-friendly format and sending them to the device in the form of text, audio, images, video, etc. This allows the user to immediately understand specific areas for improvement. The input of this process is the analysis results, and the output is the feedback provided to the user.
[0232] Step 5:
[0233] Leader's Choice
[0234] The user can select one of several mentors within the app. Based on the profile information of the selected mentor, the server customizes the analysis results, allowing the user to receive advice and make this feedback more actionable. The input to this process is the user's mentor selection, and the output is customized feedback.
[0235] Step 6:
[0236] Generation of ideal behavior
[0237] The server generates an ideal motion model based on the user's motion data. Based on the generated model, a video demonstrating the ideal motion is created and sent to the device. The user can refer to this video to improve their own motion. The input to this process is the user's motion data, and the output is a video demonstrating the ideal motion.
[0238] Furthermore, an emotion engine that estimates the user's emotion may be combined. That is, the identification processing unit 290 may estimate the user's emotion using the emotion identification model 59 and perform identification processing using the user's emotion.
[0239] This invention is a system for efficiently improving batting and pitching form based on a user's biometric and experience information, and also combines it with an emotion engine that recognizes the user's emotions. The system records the user's movements, analyzes the data, and provides specific feedback in real time, adjusting the feedback content according to the user's emotional state. It also provides a means for users to choose instruction from multiple advisors and generates videos demonstrating ideal movements, promoting the user's improvement.
[0240] Overall system configuration
[0241] Users use a device (e.g., a smartphone) to access the system. A dedicated application is installed on the device, allowing users to input their biometric information and experience information. The input information is sent to the server and stored in a database.
[0242] The device is equipped with a camera and sensors to record the user's movements. The captured video data is sent to a cloud server and analyzed using a motion analysis system. The analysis results are provided to the user in real time as feedback in the form of text, audio, images, video, etc.
[0243] The system also has multiple advisor profiles, allowing users to receive guidance from their preferred advisor. Furthermore, it can generate videos demonstrating ideal movements, providing users with visual reference information.
[0244] In addition, the emotion engine analyzes the user's emotional state in real time and adjusts the feedback content based on the results, allowing users to receive feedback that is appropriate for their own psychological state, resulting in even more effective training.
[0245] Program processing
[0246] The program does the following:
[0247] 1. User information registration
[0248] The user uses the device to input their biometric and experience information. Once input is complete, the device sends the data to the server, which stores the received information in a database.
[0249] 2. Recording a video of your form
[0250] Users record their swings and throws in the form diagnostic booth at a batting center. The motions are recorded using the device's camera or sensor, and the data is sent to a server.
[0251] 3. Form video analysis
[0252] The server sends the received video data to the motion analysis system and issues an analysis request. The system analyzes the motion and returns the results to the server, which then processes the analysis results into a user-friendly format.
[0253] 4. Emotion Analysis
[0254] The emotion engine analyzes emotional data in real time from the user's movements, facial expressions, voice, etc., and sends the results to the server.
[0255] 5. Providing Feedback
[0256] The processed analysis results are combined with the emotion analysis results and sent to the device, which then notifies the user in the form of text, audio, images, video, etc. The user can then try to improve the form based on this information.
[0257] 6. Select coaching mode
[0258] Users select one of several advisors within the app, and the analysis results and emotional data are customized based on the profile information of the selected advisor.
[0259] 7. Image training using videos
[0260] If the user wishes, an image video of the ideal swing or throwing technique will be generated and sent to the device, and the user can use this video as a reference for training.
[0261] Specific examples
[0262] For example, if a beginner user wants to improve their pitching technique, they first enter and save their basic information into the app. They then take a video of themselves pitching at a batting center and transfer it to the app. The server analyzes the video and provides specific advice, such as "position your elbow a little higher," via text and voice.
[0263] At the same time, if the emotion engine determines that the user's emotional state is "down," more encouraging words and positive feedback will be provided. Additionally, if the user selects a "famous pitching coach," customized feedback from that coach will be provided. Furthermore, videos of ideal pitching form are also provided, so users can use them as reference to improve their technique.
[0264] This system allows users to receive real-time, specific, and psychologically relevant feedback, allowing them to efficiently improve their form.
[0265] The processing flow will be explained below.
[0266] Step 1:
[0267] User: Launch the app and enter basic information such as age, gender, baseball experience, and medical restrictions on the "New Registration" screen.
[0268] Step 2:
[0269] Terminal: When input is complete, the user presses the save button, and the input data is temporarily saved in local storage.
[0270] Step 3:
[0271] Terminal: Generates and sends a request to send the stored data to the server.
[0272] Step 4:
[0273] Server: Saves the received user data in a database and returns a confirmation response to the terminal that the data was saved successfully.
[0274] Step 5:
[0275] User: Enters the form diagnostic booth at the batting center, stands in the designated position, and begins swinging or throwing.
[0276] Step 6:
[0277] Device: When an action is initiated, the device uses cameras and sensors to record the user's actions in video format.
[0278] Step 7:
[0279] User: When the action is finished, press the stop recording button to end the recording.
[0280] Step 8:
[0281] Terminal: Saves the captured video data in local storage, generates a request to send the data to the server, and sends it.
[0282] Step 9:
[0283] Server: Temporarily stores the received video data and sends an analysis request to the motion analysis system.
[0284] Step 10:
[0285] Server: Receives analysis results from the motion analysis system.
[0286] Step 11:
[0287] Server: Processes the received analysis results into a user-friendly format (text, audio, image, video).
[0288] Step 12:
[0289] Terminal: Activates an emotion engine that recognizes emotions from the user's movements, facial expressions, and voice, and acquires emotional data.
[0290] Step 13:
[0291] Server: Analyzes the emotion data sent from the emotion engine and uses it to adjust the feedback content.
[0292] Step 14:
[0293] Server: Combines the processed analysis results with the emotion analysis results and sends them to the device.
[0294] Step 15:
[0295] Terminal: Feedback based on analysis results and emotional data is displayed and notified to the user in the form of text, audio, images, video, etc.
[0296] Step 16:
[0297] User: Check the feedback and incorporate it into their next swing or throw.
[0298] Step 17:
[0299] User: Select one of several advisors on the "Coaching Mode Selection" screen within the app.
[0300] Step 18:
[0301] Terminal: Based on the information of the selected advisor, the analysis results and emotional data are customized and provided to the user.
[0302] Step 19:
[0303] Terminal: Sends a request to the server to generate an image video of the ideal swing or throwing based on the user's wishes.
[0304] Step 20:
[0305] Server: Generates an image video of the ideal form and sends it to the device.
[0306] Step 21:
[0307] Device: The received image video is played within the app and shown to the user.
[0308] Step 22:
[0309] User: Use the image video as a reference and train to improve their own form.
[0310] This allows users to receive real-time, specific, and psychologically relevant feedback, allowing them to efficiently improve their form.
[0311] Example 2
[0312] Next, a description will be given of Example 2. In the following description, the data processing device 12 will be referred to as a "server" and the smart device 14 will be referred to as a "terminal."
[0313] Conventional form improvement systems have limited dimensionality in the user's motion analysis and feedback, and in many cases, the feedback does not take into account the user's psychological state, resulting in reduced training efficiency. Furthermore, specialized equipment is often used for motion analysis and evaluation, which places constraints on cost and usage environment. Furthermore, it is difficult to select detailed guidance from multiple advisors.
[0314] The specific processing by the specific processing unit 290 of the data processing device 12 in the second embodiment is realized by the following means.
[0315] In this invention, the server includes means for inputting biometric information and experience information of the user, means for recording the user's movements, means for analyzing the recorded movement data, means for adjusting the feedback content based on the analysis results and the user's emotional state and providing the feedback in various formats, means for allowing the user to select advice from multiple advisors, and means for generating videos showing the user's ideal movements. This allows the user to receive individual and optimal feedback in real time based on the results of their own movement analysis and psychological state, enabling them to improve their form efficiently and effectively.
[0316] "Biometric information" is information that indicates the physical condition of the user, and examples include height, weight, and heart rate.
[0317] "Experience information" is information about the user's skills and career history, and examples include sports history and practice frequency.
[0318] "Motion data" is data that records the user's physical movements, and includes video and sensor data acquired using cameras and sensors.
[0319] "Analysis means" refers to a method or system for analyzing recorded motion data and providing the results.
[0320] "Feedback means" refers to a means of providing information to users based on the analysis results, and includes systems that display information in a variety of formats, such as text, audio, images, and video.
[0321] "Emotional state" refers to a user's psychological state and includes methods and systems for analyzing that state.
[0322] "Advisor" refers to an expert who provides technical guidance and advice to users, and includes their profile information.
[0323] "Video generation means" refers to a system or method for generating videos that show ideal movements to a user.
[0324] This invention is a system for efficiently improving batting and pitching form based on a user's biometric information and experience information, and also combines it with an emotion engine that recognizes the user's emotions. The specific system configuration and operation are described in detail below.
[0325] The system mainly includes the following elements:
[0326] 1. Means for inputting user biometric and experience information
[0327] 2. A means of recording user actions
[0328] 3. Means of analyzing recorded motion data
[0329] 4. A method for adjusting feedback content based on analysis results and the user's emotional state and providing feedback in various formats
[0330] 5. A means to allow students to choose guidance from multiple advisors
[0331] 6. A method for generating videos showing ideal user behavior
[0332] Entering biometric and experience information
[0333] Users use a dedicated application on a device (e.g., a smartphone) to input their own biometric information (e.g., height, weight, heart rate) and experience information (e.g., sports history, practice frequency). This information is sent from the device to a server and stored in a database.
[0334] Recording actions
[0335] For example, users can use the device's camera and sensors to record their swings and pitching movements in a form diagnostic booth installed at a batting center. The recorded video data is then sent to a cloud server.
[0336] Analysis of behavioral data
[0337] The server receives the video data sent to the cloud server and sends an analysis request to the motion analysis system. This motion analysis system analyzes the user's movements using specific algorithms and machine learning models. The analysis results are then processed into a format that is easy for the user to understand.
[0338] Emotion Analysis
[0339] The emotion engine analyzes the user's emotions in real time based on their facial expressions and voice data. The analyzed emotional state data is sent to the server and integrated with the movement analysis results.
[0340] Providing Feedback
[0341] The server generates feedback content based on the results of motion analysis and emotion analysis and sends it to the device. The device receives this data and notifies the user of the feedback in the form of text, audio, images, video, etc. This allows the user to receive specific feedback in real time that is appropriate to their psychological state.
[0342] Coaching mode and video image training
[0343] Within the app, users can select their preferred instructor from multiple advisors. Feedback is customized based on the profile of the selected advisor. Furthermore, if the user wishes, a video showing ideal swing and pitching techniques is generated and sent to the device. Users can use this video as a reference for their training.
[0344] Specific examples
[0345] For example, if a beginner user wants to improve their pitching technique, they first enter their basic information into the app. They then film a video of themselves pitching at a batting center and transfer it to the app. The server analyzes the video and provides specific advice in text and voice, such as "position your elbow a little higher." At the same time, if the emotion engine determines that the user's emotional state is "depressed," more encouraging words and positive feedback are provided.
[0346] Users can also select a "famous pitching coach" to receive customized feedback from that coach, and videos of ideal pitching form are also provided, allowing users to use these to improve their technique.
[0347] Prompt Sentence Examples
[0348] Examples of prompts for a generative AI model include:
[0349] "Please analyze my pitching video and let me know how I can improve. Also, please tailor your feedback to take into account my current emotional state."
[0350] By combining the above-mentioned methods, users can receive specific feedback in real time that is appropriate to their psychological state, allowing them to efficiently improve their form.
[0351] The flow of the identification process in the second embodiment will be described with reference to FIG.
[0352] Step 1:
[0353] Registering user information
[0354] Input: The user uses the device to input their biometric information (e.g., height, weight, heart rate) and experience information (e.g., sports history, practice frequency) into a dedicated app.
[0355] Action: The user enters the required information into the input fields and clicks the "Submit" button.
[0356] Data processing and calculation: The terminal formats the input information and generates data for transmission.
[0357] Output: The terminal sends the generated transmission data to the server.
[0358] Step 2:
[0359] Retention of Information
[0360] Input: User information data sent from the device.
[0361] How it works: The server stores the received data in a database.
[0362] Data processing and calculation: Analyzes received data and converts it into the appropriate database format.
[0363] Output: Successfully saved data is confirmed and a save confirmation notification is generated.
[0364] Step 3:
[0365] Form video recording
[0366] Input: The user has a terminal.
[0367] Actions: The user takes a video of their swing and pitching using the device's camera and sensors in the batting center's diagnostic booth. They press the record button to record their movements. When they're finished recording, they press the "stop" button.
[0368] Data processing and calculation: The device temporarily stores the recorded video data and encodes it for transfer.
[0369] Output: Send the encoded video data to the server.
[0370] Step 4:
[0371] Video data storage and analysis requests
[0372] Input: Video data sent from the device.
[0373] Operation: The server receives the video data and temporarily stores it.
[0374] Data processing and calculation: Generates requests to the motion analysis system to analyze the received video data.
[0375] Output: A request is sent to the behavior analysis system.
[0376] Step 5:
[0377] Motion analysis
[0378] Input: Video data analysis request sent from the server.
[0379] Motion: The motion analysis system analyzes the received video data using machine learning algorithms.
[0380] Data processing and calculation: The motion analysis system extracts motion elements from the video and generates analysis results.
[0381] Output: The generated behavior analysis results are sent back to the server.
[0382] Step 6:
[0383] Emotion analysis
[0384] Input: User facial and voice data.
[0385] How it works: The device uses a camera and microphone to record the user's facial expressions and voice data in real time.
[0386] Data processing and calculation: The device formats the recorded emotion data for transmission and sends it to the server. The server then sends an analysis request to the emotion analysis system.
[0387] Output: The emotion analysis system sends the analysis results to the server, and the final emotion data is generated.
[0388] Step 7:
[0389] Generating and Providing Feedback
[0390] Input: Motion analysis results and emotion analysis results.
[0391] Behavior: The server integrates the results of behavior analysis and emotion analysis and generates feedback for the user.
[0392] Data processing and calculation: Processing the feedback content into text, audio, image, or video format.
[0393] Output: The generated feedback content is sent to the device, which displays or plays it.
[0394] Step 8:
[0395] Choosing a coaching mode
[0396] Input: Profile information for multiple advisors.
[0397] How it works: A user selects their preferred advisor within the app.
[0398] Data processing and calculation: The server customizes the feedback content based on the information of the selected advisor.
[0399] Output: The customized feedback is sent to the user's device and displayed.
[0400] Step 9:
[0401] Creating image training videos
[0402] Input: The user's desired action information.
[0403] Action: The user selects the "Image Training" option.
[0404] Data processing and calculation: The server requests a video showing the ideal behavior from the generation engine and receives the generated video data.
[0405] Output: Send the generated image training video to the device.
[0406] (Application example 2)
[0407] Next, a description will be given of Application Example 2. In the following description, the data processing device 12 will be referred to as a "server" and the smart device 14 will be referred to as a "terminal."
[0408] Improving work efficiency and ensuring safety in factories are key challenges. In particular, there is a need to analyze the movements of robots and workers in real time and provide accurate feedback based on that analysis. It is also necessary to provide feedback that takes into account the mental state of workers in order to create a more effective work environment.
[0409] The specification processing by the specification processing unit 290 of the data processing device 12 in Application Example 2 is realized by the following means. In this invention, the server includes means for inputting biometric information and experience information of the user, means for recording the user's movements, means for analyzing the recorded movement data, means for feeding back the analysis results in various formats, means for allowing the user to select advice from multiple advisors, means for generating a video showing the user's ideal movements, means for analyzing the user's emotional state in real time, and means for adjusting the feedback content based on the user's emotional state. This enables effective feedback that takes into account the movement analysis results and the user's emotional state.
[0410] "User's biological information" is data relating to the user's physical condition, such as heart rate, respiratory rate, and body temperature.
[0411] "User experience information" is information such as the work content and proficiency level of the user that the user has previously experienced.
[0412] "User actions" refer to specific movements or tasks performed by the user.
[0413] "Motion data" refers to video recordings and measurement data of the user's movements.
[0414] "Emotional state" refers to the user's emotional or psychological state.
[0415] "Analysis" is the process of analyzing collected data and extracting meaningful information.
[0416] "Feedback" refers to the act of returning analysis results, advice, etc. to the user.
[0417] An "advisor" is an expert with specific knowledge and experience who provides guidance and advice to users.
[0418] "Ideal movement" refers to an efficient and safe movement pattern that optimizes work performance.
[0419] An "emotion engine" is a program or algorithm that analyzes a user's emotional state in real time.
[0420] A "cloud server" is an external server that analyzes and stores data via the Internet.
[0421] "Real-time" refers to processing occurring immediately or with a very short time delay.
[0422] System Overview
[0423] This invention is a system for improving work efficiency and safety in factories. The system uses a computer terminal (e.g., smart glasses or a head-mounted display) to record the user's movements and analyzes the collected data on a cloud server. The analysis results are fed back to the user in real time, and the feedback content is adjusted according to the user's emotional state.
[0424] Hardware and Software
[0425] The system includes the following major hardware and software:
[0426] Smart glasses or head-mounted displays: equipped with cameras and sensors to record user movements.
[0427] Cloud Server: Computing resources for analyzing recorded data and generating feedback.
[0428] Motion analysis model: A machine learning model trained using TensorFlow.
[0429] Emotion analysis model: Face detection using Dlib and emotion recognition model using TensorFlow.
[0430] Processing flow
[0431] 1. Data Entry:
[0432] The user uses the device to input biometric data (e.g., heart rate, body temperature) and past experience information. The data is sent from the device to a cloud server, where it is stored in a database.
[0433] 2. Operation Record:
[0434] The device records the user's actions in real time, and the recorded data is sent to a cloud server.
[0435] 3. Data Analysis:
[0436] The cloud server analyzes the received motion data using a motion analysis model, and simultaneously analyzes the user's emotional state using face detection and emotion recognition.
[0437] 4. Providing Feedback:
[0438] Feedback is generated in the form of text, audio, images, and video and is provided to the user in real time, with the feedback content adjusted according to the user's emotional state.
[0439] Specific examples
[0440] For example, the system records the movements of a robot packing items along a conveyor belt in a factory. The movement data is analyzed on a cloud server, and if there is any waste in the movement, the system provides specific advice such as "Please reduce your arm movements to reduce waste in movement." Furthermore, if signs of fatigue are detected from the worker's facial expressions or movements, it can generate emotion-based feedback such as "Please take a break."
[0441] Example prompts for generative AI models
[0442] "Please explain how the system analyzes the operational data of robots working in a factory and provides advice on improving efficiency. In addition, please incorporate a system that analyzes the emotional state of workers and provides appropriate feedback. This system will operate using smart glasses or a head-mounted display."
[0443] The flow of the specific processing in the application example 2 will be described with reference to FIG.
[0444] Step 1:
[0445] The user wears smart glasses or a head-mounted display and inputs biometric information (heart rate, body temperature, etc.) and experience information. The device sends this data to a cloud server, which then stores the received data in a database. At this stage, the input is the user's biometric information and experience information, and the output is data stored in the cloud server.
[0446] Step 2:
[0447] The device worn by the user uses built-in cameras and sensors to record the user's movements in real time. The recorded video data is converted into an appropriate format and sent to a cloud server. At this stage, the input is the user's movement data, and the output is data sent to the cloud server.
[0448] Step 3:
[0449] The cloud server inputs the transmitted motion data into a motion analysis model and performs motion analysis. The motion analysis model extracts motion characteristics (e.g., elbow angle, hand movement, etc.) and calculates indicators related to efficiency and safety. The input at this stage is the motion data, and the output is the motion analysis results.
[0450] Step 4:
[0451] In parallel with the motion analysis, the cloud server uses an emotion engine to analyze the user's emotional state from facial expression and voice data. The emotion analysis model analyzes facial features (e.g., facial muscle movements) and determines the user's emotion (e.g., fatigue, stress). The input at this stage is facial expression data and voice data, and the output is the emotion analysis results.
[0452] Step 5:
[0453] The cloud server integrates the results of the motion analysis and emotion analysis to generate appropriate feedback. The generated feedback is sent to the user's device in the form of text, audio, image, or video. For example, feedback such as "Please reduce your arm movements" or "Please take a break" is provided. The input at this stage is the results of the motion analysis and emotion analysis, and the output is the feedback content.
[0454] Step 6:
[0455] The user receives feedback through the device and modifies their behavior to improve efficiency and safety. If the user's emotional state is determined to be "fatigued," they receive feedback that reflects their emotions, such as encouragement or a prompt to take a break. The input at this stage is the feedback content, and the output is a change in the user's behavior.
[0456] The specific processing unit 290 transmits the result of the specific processing to the smart device 14. In the smart device 14, the control unit 46A causes the output device 40 to output the result of the specific processing. The microphone 38B acquires audio indicating a user input regarding the result of the specific processing. The control unit 46A transmits audio data indicating the user input acquired by the microphone 38B to the data processing device 12. In the data processing device 12, the specific processing unit 290 acquires the audio data.
[0457] The data generation model 58 is a so-called generative AI (Artificial Intelligence). An example of the data generation model 58 is ChatGPT (Internet Search<URL: https: / / openai.com / blog / chatgpt> ), Gemini (Internet search <url: https: gemini.google.com ?hl="ja">) and other generation AIs. The data generation model 58 is obtained by performing deep learning on a neural network. A prompt including an instruction is input to the data generation model 58, and inference data such as voice data indicating voice, text data indicating text, and image data indicating an image is also input. The data generation model 58 performs inference on the input inference data in accordance with the instruction indicated by the prompt, and outputs the inference result in a data format such as voice data and text data. Here, inference refers to, for example, analysis, classification, prediction, and / or summarization.
[0458] In the above embodiment, an example in which the specific process is performed by the data processing device 12 has been given, but the technology of the present disclosure is not limited to this, and the specific process may be performed by the smart device 14.
[0459] [Second embodiment]
[0460] FIG. 3 shows an example of the configuration of a data processing system 210 according to the second embodiment.
[0461] 3, the data processing system 210 includes the data processing device 12 and smart glasses 214. An example of the data processing device 12 is a server.
[0462] The data processing device 12 includes a computer 22, a database 24, and a communication I / F 26. The computer 22 is an example of a "computer" according to the technology of the present disclosure. The computer 22 includes a processor 28, a RAM 30, and a storage 32. The processor 28, the RAM 30, and the storage 32 are connected to a bus 34. The database 24 and the communication I / F 26 are also connected to the bus 34. The communication I / F 26 is connected to a network 54. Examples of the network 54 include a WAN (Wide Area Network) and / or a LAN (Local Area Network).
[0463] The smart glasses 214 include a computer 36, a microphone 238, a speaker 240, a camera 42, and a communication I / F 44. The computer 36 includes a processor 46, a RAM 48, and a storage 50. The processor 46, the RAM 48, and the storage 50 are connected to a bus 52. The microphone 238, the speaker 240, and the camera 42 are also connected to the bus 52.
[0464] The microphone 238 receives instructions and the like from the user 20 by receiving voice uttered by the user 20. The microphone 238 captures the voice uttered by the user 20, converts the captured voice into audio data, and outputs it to the processor 46. The speaker 240 outputs audio in accordance with instructions from the processor 46.
[0465] Camera 42 is a small digital camera equipped with an optical system including a lens, aperture, and shutter, and an imaging element such as a CMOS (Complementary Metal-Oxide-Semiconductor) image sensor or a CCD (Charge Coupled Device) image sensor, and captures images of the surroundings of user 20 (for example, an imaging range defined by an angle of view equivalent to the field of vision of a typical healthy person).
[0466] The communication I / F 44 is connected to a network 54. The communication I / Fs 44 and 26 are responsible for the exchange of various information between the processor 46 and the processor 28 via the network 54. The exchange of various information between the processor 46 and the processor 28 using the communication I / Fs 44 and 26 is carried out in a secure state.
[0467] Fig. 4 shows an example of the main functions of the data processing device 12 and the smart glasses 214. As shown in Fig. 4, in the data processing device 12, a specific process is performed by the processor 28. A specific process program 56 is stored in the storage 32.
[0468] The specific processing program 56 is an example of a "program" according to the technology of the present disclosure. The processor 28 reads the specific processing program 56 from the storage 32 and executes the read specific processing program 56 on the RAM 30. The specific processing is realized by the processor 28 operating as a specific processing unit 290 in accordance with the specific processing program 56 executed on the RAM 30.
[0469] The storage 32 stores a data generation model 58 and an emotion identification model 59. The data generation model 58 and the emotion identification model 59 are used by the identification processing unit 290.
[0470] In the smart glasses 214, the reception output process is performed by the processor 46. A reception output program 60 is stored in the storage 50. The processor 46 reads the reception output program 60 from the storage 50 and executes the read reception output program 60 on the RAM 48. The reception output process is realized by the processor 46 operating as the control unit 46A in accordance with the reception output program 60 executed on the RAM 48.
[0471] Next, a description will be given of the identification process performed by the identification processing unit 290 of the data processing device 12. In the following description, the data processing device 12 will be referred to as the "server" and the smart glasses 214 will be referred to as the "terminal."
[0472] This invention is a system for efficiently improving batting and pitching form based on a user's biometric and experience information. The system records the user's movements, analyzes the data, and provides specific feedback in real time. It also provides a means for users to choose from multiple advisors and generates videos demonstrating ideal movements, promoting the user's technical improvement.
[0473] Overall system configuration
[0474] Users use a device (e.g., a smartphone) to access the system. A dedicated application is installed on the device, allowing users to input their biometric information and experience information. The input information is sent to the server and stored in a database.
[0475] The device is equipped with a camera and sensors to record the user's movements. The captured video data is sent to a cloud server and analyzed using Google's Gemini. The analysis results are provided to the user in real time as feedback in the form of text, audio, images, video, etc.
[0476] The system also has multiple advisor profiles, allowing users to receive guidance from their preferred advisor. Furthermore, it can generate videos demonstrating ideal movements, providing users with visual reference information.
[0477] Program processing
[0478] The program does the following:
[0479] 1. User information registration
[0480] The user uses the device to input their biometric and experience information. Once input is complete, the device sends the data to the server, which stores the received information in a database.
[0481] 2. Recording a video of your form
[0482] Users record their swings and throws in the form diagnostic booth at a batting center. The motions are recorded using the device's camera or sensor, and the data is sent to a server.
[0483] 3. Form video analysis
[0484] The server sends the received video data to Google Gemini and makes an analysis request. Google Gemini analyzes the motion and returns the results to the server. The analysis results are then processed into a user-friendly format.
[0485] 4. Providing real-time feedback
[0486] The processed analysis results are sent to the device and notified to the user in the form of text, audio, images, video, etc. The user can then try to improve the form based on this information.
[0487] 5. Select coaching mode
[0488] Users can select one of several advisors within the app, and the analysis results are customized based on the profile information of the selected advisor.
[0489] 6. Image training using videos
[0490] If the user wishes, an image video of the ideal swing or throwing technique will be generated and sent to the device, and the user can use this video as a reference for training.
[0491] Specific examples
[0492] For example, if a beginner user wants to improve their pitching technique, they first enter and save their basic information in the app. They then film a video of themselves pitching at a batting center and transfer it to the app. The server analyzes the video and provides specific advice in text and voice, such as "position your elbow a little higher." Additionally, if the user selects a "famous pitching coach," they will receive customized feedback from that coach. Furthermore, a video of the ideal pitching form is also provided, allowing the user to use it as a reference to improve their technique.
[0493] The system allows users to receive specific, expert feedback in real time, allowing them to efficiently improve their form.
[0494] The processing flow will be explained below.
[0495] Step 1:
[0496] User: Launch the app and enter basic information such as age, gender, baseball experience, and medical restrictions on the "New Registration" screen.
[0497] Step 2:
[0498] Terminal: When input is complete, the user presses the save button, and the input data is temporarily saved in local storage.
[0499] Step 3:
[0500] Terminal: Generates and sends a request to send the stored data to the server.
[0501] Step 4:
[0502] Server: Saves the received user data in a database and returns a confirmation response to the terminal that the data was saved successfully.
[0503] Step 5:
[0504] User: Enters the form diagnostic booth at the batting center, stands in the designated position, and begins swinging or throwing.
[0505] Step 6:
[0506] Device: When an action is initiated, the device uses cameras and sensors to record the user's actions in video format.
[0507] Step 7:
[0508] User: When the action is finished, press the stop recording button to end the recording.
[0509] Step 8:
[0510] Terminal: Saves the captured video data in local storage, generates a request to send the data to the server, and sends it.
[0511] Step 9:
[0512] Server: Temporarily stores the received video data and sends an analysis request to Google's Gemini API.
[0513] Step 10:
[0514] Server: Receives analysis results from Google's "Gemini".
[0515] Step 11:
[0516] Server: Processes the received analysis results into a user-friendly format (text, audio, image, video).
[0517] Step 12:
[0518] Server: Sends the processed analysis results to the terminal.
[0519] Step 13:
[0520] Device: Receives analysis results and displays / notifies them on the app.
[0521] Step 14:
[0522] User: Check the analysis results and use the improvements to improve their next swing or throw.
[0523] Step 15:
[0524] User: Select one of several advisors on the "Coaching Mode Selection" screen within the app.
[0525] Step 16:
[0526] Terminal: Based on the information of the selected advisor, the analysis results are customized and provided to the user.
[0527] Step 17:
[0528] Terminal: Sends a request to the server to generate an image video of the ideal swing or throwing based on the user's wishes.
[0529] Step 18:
[0530] Server: Generates an image video of the ideal form and sends it to the device.
[0531] Step 19:
[0532] Device: The received image video is played within the app and shown to the user.
[0533] Step 20:
[0534] User: Use the image video as a reference and train to improve their own form.
[0535] This allows users to receive specific, professional feedback in real time and efficiently improve their form.
[0536] Example 1
[0537] Next, a description will be given of Example 1. In the following description, the data processing device 12 will be referred to as a "server" and the smart glasses 214 will be referred to as a "terminal."
[0538] Conventional methods for improving batting and pitching form lack the means for users to accurately understand their own movements and receive effective feedback. Furthermore, opportunities to receive direct instruction from a coach are limited, making self-study to improve skills difficult. The objective of this invention is to provide a system that provides specific, specialized feedback in real time, allowing users to efficiently improve their movements.
[0539] The specific processing by the specific processing unit 290 of the data processing device 12 in the first embodiment is realized by the following means.
[0540] In this invention, the server includes means for inputting biometric information and experience information of the user, a terminal device for recording the user's movements, means for transmitting the recorded video data to the analytical model via the cloud server, means for processing the analysis results from the analytical model into a user-friendly format, means for feeding back the analysis results to the user in real time, means for enabling the user to select instruction from multiple instructors, and means for generating videos showing the user's ideal movements. This allows the user to receive specific and professional feedback in real time and efficiently improve their own movements.
[0541] "Biometric information" refers to information about the user's physical characteristics, such as the user's height, weight, athletic ability, and health condition.
[0542] "Experience information" is information about the user's experience, such as the user's sports history, skill level, and past training history.
[0543] A "terminal device" is a hardware device, such as a smartphone or tablet, that allows a user to input information and record actions.
[0544] A "cloud server" is a remote server connected via the Internet for storing, managing, and analyzing data.
[0545] An "analytics model" is a machine learning algorithm or data processing mechanism that analyzes a user's behavior data and provides form improvements and technical advice.
[0546] "Feedback" refers to providing users with information such as advice, comments, and analysis results generated based on the user's behavior data.
[0547] "Instructors" are coaches or advisors with specialized knowledge and experience that can be selected by the user.
[0548] "Videos showing ideal movements" are videos that show the ideal form and techniques that users should aim for, and are used to support the user's training.
[0549] The present invention provides a system for effectively improving a user's sports form based on biometric information and experience information. Specific embodiments for carrying out the present invention will be described in detail below.
[0550] Overall system configuration
[0551] Registering user information
[0552] The user starts a dedicated application using a device (e.g., a smartphone) and inputs biometric information (e.g., height, weight, health status) and experience information (e.g., sports history, skill level). This information is sent from the device to a server and stored in a database.
[0553] Recording actions
[0554] Users can record their batting and pitching movements at batting centers or practice fields using their smartphone cameras or dedicated sensors. The recorded video data is temporarily stored on the device and then uploaded to a cloud server.
[0555] Video Analysis
[0556] The server then sends the received video data to an analytical model (e.g., a machine learning algorithm) in the cloud and issues an analysis request. The analytical model analyzes the user's movements and extracts data on form flaws and areas for improvement. The extracted results are returned to the server, where they are further processed into a format that is easy for the user to understand (e.g., text, images, video).
[0557] Providing Feedback
[0558] The server then sends the processed analysis results to the device in real time. The device then displays the results and provides specific advice to the user via text or voice, such as "position your elbows a little higher." Real-time video feedback is also provided based on the analysis results.
[0559] Choosing a coaching mode
[0560] Users can browse the profiles of multiple instructors from the in-app menu and select their preferred instructor. The server generates customized feedback based on the profile information of the selected instructor and sends it to the device.
[0561] Image training using videos
[0562] It can generate images that allow users to reproduce their ideal movements. The server generates a video of the ideal form and sends it to the device. Users can use this video as a reference for training and improve their form.
[0563] Specific examples
[0564] For example, if a beginner user wants to improve their pitching form, they first enter and save their basic information into the app on their smartphone. Next, they use the smartphone camera to film their pitching at a batting center, and the video is sent from the app to a cloud server. The cloud server analyzes the video and sends specific advice, such as "position your elbow a little higher," to the user's smartphone in real time. Furthermore, if the user selects a well-known coach, that coach will provide them with customized feedback. Finally, a video of their ideal pitching form is provided to the user, who can use the video as a reference to improve their technique.
[0565] Generative AI model prompt example
[0566] "Please analyze the batting form. What is the best way to teach the player if their elbow is positioned too low?"
[0567] "Please tell me the ideal throwing motion for pitching form."
[0568] In this way, by inputting specific questions into the generative AI model, optimal feedback can be obtained.
[0569] The flow of the identification process in the first embodiment will be described with reference to FIG.
[0570] Program processing steps
[0571] Step 1: Register user information
[0572] Input: User biometric and experience information
[0573] Processing: The user launches a dedicated application using a device (e.g., a smartphone) and enters basic information such as height, weight, and sports history into a form. The device then sends the entered information to the server using the HTTPS protocol.
[0574] Output: The entered user information is saved on the cloud server.
[0575] Specific operation: When a user enters information on their smartphone and presses the "Send" button, the information is recorded in a database (e.g., MySQL) on the cloud server.
[0576] Step 2: Record a video of your form
[0577] Input: User action (e.g. swing or throw)
[0578] Processing: A user uses a smartphone camera app to record their movements at a batting cage or practice area. Once the recording is complete, it is temporarily saved as a video file on the device.
[0579] Output: The video data file will be saved on your device.
[0580] Specific operation: The user presses the record button on the camera, and when they finish recording their movements, they press the "Save" button to save the video data to their device.
[0581] Step 3: Submitting the form video
[0582] Input: Video data stored on the device
[0583] Processing: The device uploads the recorded video file to a cloud server using a secure file transfer protocol (e.g., HTTPS or SFTP).
[0584] Output: Video data is stored in cloud server storage (e.g. AWS S3).
[0585] Specific operation: The user presses the "upload" button within the app, and the video is sent to the cloud server.
[0586] Step 4: Analysis instructions for form videos
[0587] Input: Video data stored on a cloud server
[0588] Processing: The server sends the video data to the analytical model and makes an analysis request. The analytical model uses a motion analysis algorithm.
[0589] Output: The analysis request is sent to the analysis model.
[0590] Specific operation: The server automatically checks the stored video data and sends an API request to the analysis model.
[0591] Step 5: Analyzing your form video
[0592] Input: Video data and analysis request
[0593] Processing: The analytical model analyzes user behavior from video data and generates data on form parameters and areas for improvement.
[0594] Output: The analysis result data is sent back to the server.
[0595] Specific behavior: The analytical model performs behavior analysis, generates analysis results, and sends them back to the server.
[0596] Step 6: Processing the analysis results and generating feedback data
[0597] Input: Analysis result data
[0598] Processing: The server receives the analysis results data and processes it into a format that is easy for the user to understand (e.g., text, charts, videos).
[0599] Output: Processed feedback data is generated.
[0600] Specific operation: The server processes the analysis result data for feedback and converts it into the most suitable format for the user.
[0601] Step 7: Provide feedback
[0602] Input: Processed feedback data
[0603] Processing: The server sends the feedback data to the device using a real-time notification service (e.g., Firebase Cloud Messaging). The device displays the feedback data to the user.
[0604] Output: Feedback is displayed on the user's terminal.
[0605] Specific operation: A notification is sent to the user's smartphone, and when they open the app, advice such as "position your elbows a little higher" is displayed.
[0606] Step 8: Select a coaching mode
[0607] Input: User-selected mentor profile
[0608] What happens: Users select one of several mentors within the app, and feedback is customized based on their profile information.
[0609] Output: Customized feedback data is generated.
[0610] Specific operation: The server regenerates feedback using the profile information of the selected instructor and sends it to the terminal.
[0611] Step 9: Image training using videos
[0612] Input: User training request
[0613] Processing: The server generates a video demonstrating the ideal movement and sends it to the device. The user uses this video as a reference for training.
[0614] Output: A video showing the ideal behavior is displayed on the device.
[0615] Specific movements: Users play videos of ideal form in the app and train by imitating the movements.
[0616] Generative AI model prompt example
[0617] "Please analyze the batting form. What is the best way to teach the player if their elbow is positioned too low?"
[0618] "Please tell me the ideal throwing motion for pitching form."
[0619] (Application example 1)
[0620] Next, a description will be given of Application Example 1. In the following description, the data processing device 12 will be referred to as a "server" and the smart glasses 214 will be referred to as a "terminal."
[0621] In order to ensure the efficiency and safety of industrial robots, it is important to evaluate the accuracy and appropriateness of their movements in real time and provide feedback on areas for improvement. However, existing systems do not fully realize these functions, making it difficult to optimize robot movements.
[0622] The specific processing by the specific processing unit 290 of the data processing device 12 in the application example 1 is realized by the following means.
[0623] In this invention, the server includes a means for transmitting user information to the server and storing it in a database, a means for transmitting operation records to the server and analyzing them using a model, and a means for converting the analysis results into a user-friendly format in real time and providing the results, thereby enabling efficient and safe optimization of the operation of industrial robots.
[0624] "User's biometric information" refers to data related to the user's body, including information such as heart rate, body temperature, and muscle movements.
[0625] "Experience information" refers to the knowledge and experience that the user has accumulated up to now, and is data such as proficiency level and past performance records.
[0626] A "recording means" is a device that collects the user's actions using a video camera or a sensor.
[0627] "Means for analyzing" refers to algorithms or software for evaluating recorded motion data using analytical models.
[0628] "Feedback means" refers to means for providing analysis results to the user in the form of text, audio, images, video, etc.
[0629] The "means for enabling selection of instructor" refers to an interface that allows the user to select one instructor from multiple instructors and receive evaluation and advice from that instructor.
[0630] The "means for generating animation" refers to a system that creates animation to visually show the user's ideal movements.
[0631] A "server" is a computer system that processes, manages, and provides data over a network.
[0632] The "means for saving to a database" is a function for safely saving input user information to a database.
[0633] "Means of analysis using models" refers to means of analyzing data using technologies such as machine learning and artificial intelligence to generate insights and improvement suggestions.
[0634] "Means for converting the analysis results into a user-friendly format and providing them" refers to a method for converting the analysis results into an easy-to-understand and comprehend format and providing them to the user.
[0635] The "means for proposing improvements to operations" is a function for proposing specific improvements to the user based on the analysis results.
[0636] A specific system for implementing this invention includes the following components: First, a user uses a terminal to input their own biometric information and experience information. A dedicated application is installed on the terminal, which transmits the user's information to a server and stores it in a database. The system records the user's movements using a camera or sensor and evaluates the data using an analytical model. The analysis results are processed on a cloud server and converted into a user-friendly format. These results are then fed back to the user in real time.
[0637] The system also features an interface that allows users to select instruction from multiple instructors. Users can receive customized feedback from the instructor of their choice based on their own movement data. Furthermore, the system generates an ideal movement model based on the analysis results and makes suggestions for movement improvement based on that model.
[0638] Hardware and software used
[0639] Hardware
[0640] Camera (e.g. Logitech C920)
[0641] Computer system (e.g., CPU: Intel i7, RAM: 16GB, etc.)
[0642] Robots (e.g., industrial arm robots)
[0643] software
[0644] Python programming language
[0645] OpenCV library (motion capture and image processing)
[0646] Requests library (sending HTTP requests)
[0647] Server-side program (runs on a cloud server)
[0648] Analytical models (using machine learning and artificial intelligence techniques)
[0649] Processing example
[0650] For example, when improving the behavior of a welding robot in a factory to ensure accurate welding, a camera captures the robot's movements and sends the data to a cloud server for analysis. Specific feedback such as "The welding position needs to be moved 5 mm to the left" is provided as a result of the analysis. A video demonstrating the ideal welding behavior is also provided to the user.
[0651] Prompt Sentence Examples
[0652] "Please create a system that records the operations of welding robots in factories with a camera, analyzes them on a cloud server, and provides suggestions for improvement. Please provide specific feedback in real time based on the analysis results."
[0653] This system makes it possible to optimize the operation of industrial robots efficiently and safely.
[0654] The flow of the specific processing in the application example 1 will be described with reference to FIG.
[0655] Step 1:
[0656] Registering user information
[0657] The user inputs their biometric and experience information into the device. This input data includes physical information such as heart rate, body temperature, and muscle movement, as well as skill level and past performance records. The device sends this data to the server, which stores the received data in a database. The input of this process is the biometric and experience information entered by the user, and the output is the user information stored in the database.
[0658] Step 2:
[0659] Shooting a video of the operation
[0660] A camera records the user's actions. In this case, the object of the action is an industrial robot, for example, recording the robot's welding work. The camera (e.g., Logitech C920) captures the robot's actions and saves them as a video file. The input of this process is the robot's actions, and the output is the recorded video file.
[0661] Step 3:
[0662] Analysis of motion video
[0663] The device sends the recorded video file to a server. The server receives this video data and analyzes it using an analytical model (using machine learning and artificial intelligence techniques). The analytical model takes the video data as input and evaluates the accuracy and appropriateness of the movements, which then generates specific suggestions for improving the movements. The input to this process is the video file, and the output is the analysis results.
[0664] Step 4:
[0665] Providing real-time feedback
[0666] The server provides the analysis results to the user in real time by converting the analysis results into a user-friendly format and sending them to the device in the form of text, audio, images, video, etc. This allows the user to immediately understand specific areas for improvement. The input of this process is the analysis results, and the output is the feedback provided to the user.
[0667] Step 5:
[0668] Leader's Choice
[0669] The user can select one of several mentors within the app. Based on the profile information of the selected mentor, the server customizes the analysis results, allowing the user to receive advice and make this feedback more actionable. The input to this process is the user's mentor selection, and the output is customized feedback.
[0670] Step 6:
[0671] Generation of ideal behavior
[0672] The server generates an ideal motion model based on the user's motion data. Based on the generated model, a video demonstrating the ideal motion is created and sent to the device. The user can refer to this video to improve their own motion. The input to this process is the user's motion data, and the output is a video demonstrating the ideal motion.
[0673] Furthermore, an emotion engine that estimates the user's emotion may be further combined. That is, the identification processing unit 290 may estimate the user's emotion using the emotion identification model 59, and perform identification processing using the user's emotion.
[0674] This invention is a system for efficiently improving batting and pitching form based on a user's biometric and experience information, and also combines it with an emotion engine that recognizes the user's emotions. The system records the user's movements, analyzes the data, and provides specific feedback in real time, adjusting the feedback content according to the user's emotional state. It also provides a means for users to choose instruction from multiple advisors and generates videos demonstrating ideal movements, promoting the user's improvement.
[0675] Overall system configuration
[0676] Users use a device (e.g., a smartphone) to access the system. A dedicated application is installed on the device, allowing users to input their biometric information and experience information. The input information is sent to the server and stored in a database.
[0677] The device is equipped with a camera and sensors to record the user's movements. The captured video data is sent to a cloud server and analyzed using a motion analysis system. The analysis results are provided to the user in real time as feedback in the form of text, audio, images, video, etc.
[0678] The system also has multiple advisor profiles, allowing users to receive guidance from their preferred advisor. Furthermore, it can generate videos demonstrating ideal movements, providing users with visual reference information.
[0679] In addition, the emotion engine analyzes the user's emotional state in real time and adjusts the feedback content based on the results, allowing users to receive feedback that is appropriate for their own psychological state, resulting in even more effective training.
[0680] Program processing
[0681] The program does the following:
[0682] 1. User information registration
[0683] The user uses the device to input their biometric and experience information. Once input is complete, the device sends the data to the server, which stores the received information in a database.
[0684] 2. Recording a video of your form
[0685] Users record their swings and throws in the form diagnostic booth at a batting center. The motions are recorded using the device's camera or sensor, and the data is sent to a server.
[0686] 3. Form video analysis
[0687] The server sends the received video data to the motion analysis system and issues an analysis request. The system analyzes the motion and returns the results to the server, which then processes the analysis results into a user-friendly format.
[0688] 4. Emotion Analysis
[0689] The emotion engine analyzes emotional data in real time from the user's movements, facial expressions, voice, etc., and sends the results to the server.
[0690] 5. Providing Feedback
[0691] The processed analysis results are combined with the emotion analysis results and sent to the device, which then notifies the user in the form of text, audio, images, video, etc. The user can then try to improve the form based on this information.
[0692] 6. Select coaching mode
[0693] Users select one of several advisors within the app, and the analysis results and emotional data are customized based on the profile information of the selected advisor.
[0694] 7. Image training using videos
[0695] If the user wishes, an image video of the ideal swing or throwing technique will be generated and sent to the device, and the user can use this video as a reference for training.
[0696] Specific examples
[0697] For example, if a beginner user wants to improve their pitching technique, they first enter and save their basic information into the app. They then take a video of themselves pitching at a batting center and transfer it to the app. The server analyzes the video and provides specific advice, such as "position your elbow a little higher," via text and voice.
[0698] At the same time, if the emotion engine determines that the user's emotional state is "down," more encouraging words and positive feedback will be provided. Additionally, if the user selects a "famous pitching coach," customized feedback from that coach will be provided. Furthermore, videos of ideal pitching form are also provided, so users can use them as reference to improve their technique.
[0699] This system allows users to receive real-time, specific, and psychologically relevant feedback, allowing them to efficiently improve their form.
[0700] The processing flow will be explained below.
[0701] Step 1:
[0702] User: Launch the app and enter basic information such as age, gender, baseball experience, and medical restrictions on the "New Registration" screen.
[0703] Step 2:
[0704] Terminal: When input is complete, the user presses the save button, and the input data is temporarily saved in local storage.
[0705] Step 3:
[0706] Terminal: Generates and sends a request to send the stored data to the server.
[0707] Step 4:
[0708] Server: Saves the received user data in a database and returns a confirmation response to the terminal that the data was saved successfully.
[0709] Step 5:
[0710] User: Enters the form diagnostic booth at the batting center, stands in the designated position, and begins swinging or throwing.
[0711] Step 6:
[0712] Device: When an action is initiated, the device uses cameras and sensors to record the user's actions in video format.
[0713] Step 7:
[0714] User: When the action is finished, press the stop recording button to end the recording.
[0715] Step 8:
[0716] Terminal: Saves the captured video data in local storage, generates a request to send the data to the server, and sends it.
[0717] Step 9:
[0718] Server: Temporarily stores the received video data and sends an analysis request to the motion analysis system.
[0719] Step 10:
[0720] Server: Receives analysis results from the motion analysis system.
[0721] Step 11:
[0722] Server: Processes the received analysis results into a user-friendly format (text, audio, image, video).
[0723] Step 12:
[0724] Terminal: Activates an emotion engine that recognizes emotions from the user's movements, facial expressions, and voice, and acquires emotional data.
[0725] Step 13:
[0726] Server: Analyzes the emotion data sent from the emotion engine and uses it to adjust the feedback content.
[0727] Step 14:
[0728] Server: Combines the processed analysis results with the emotion analysis results and sends them to the device.
[0729] Step 15:
[0730] Terminal: Feedback based on analysis results and emotional data is displayed and notified to the user in the form of text, audio, images, video, etc.
[0731] Step 16:
[0732] User: Check the feedback and incorporate it into their next swing or throw.
[0733] Step 17:
[0734] User: Select one of several advisors on the "Coaching Mode Selection" screen within the app.
[0735] Step 18:
[0736] Terminal: Based on the information of the selected advisor, the analysis results and emotional data are customized and provided to the user.
[0737] Step 19:
[0738] Terminal: Sends a request to the server to generate an image video of the ideal swing or throwing based on the user's wishes.
[0739] Step 20:
[0740] Server: Generates an image video of the ideal form and sends it to the device.
[0741] Step 21:
[0742] Device: The received image video is played within the app and shown to the user.
[0743] Step 22:
[0744] User: Use the image video as a reference and train to improve their own form.
[0745] This allows users to receive real-time, specific, and psychologically relevant feedback, allowing them to efficiently improve their form.
[0746] Example 2
[0747] Next, a description will be given of Example 2. In the following description, the data processing device 12 will be referred to as a "server" and the smart glasses 214 will be referred to as a "terminal."
[0748] Conventional form improvement systems have limited dimensionality in the user's motion analysis and feedback, and in many cases, the feedback does not take into account the user's psychological state, resulting in reduced training efficiency. Furthermore, specialized equipment is often used for motion analysis and evaluation, which places constraints on cost and usage environment. Furthermore, it is difficult to select detailed guidance from multiple advisors.
[0749] The specific processing by the specific processing unit 290 of the data processing device 12 in the second embodiment is realized by the following means.
[0750] In this invention, the server includes means for inputting biometric information and experience information of the user, means for recording the user's movements, means for analyzing the recorded movement data, means for adjusting the feedback content based on the analysis results and the user's emotional state and providing the feedback in various formats, means for allowing the user to select advice from multiple advisors, and means for generating videos showing the user's ideal movements. This allows the user to receive individual and optimal feedback in real time based on the results of their own movement analysis and psychological state, enabling them to improve their form efficiently and effectively.
[0751] "Biometric information" is information that indicates the physical condition of the user, and examples include height, weight, and heart rate.
[0752] "Experience information" is information about the user's skills and career history, and examples include sports history and practice frequency.
[0753] "Motion data" is data that records the user's physical movements, and includes video and sensor data acquired using cameras and sensors.
[0754] "Analysis means" refers to a method or system for analyzing recorded motion data and providing the results.
[0755] "Feedback means" refers to a means of providing information to users based on the analysis results, and includes systems that display information in a variety of formats, such as text, audio, images, and video.
[0756] "Emotional state" refers to a user's psychological state and includes methods and systems for analyzing that state.
[0757] "Advisor" refers to an expert who provides technical guidance and advice to users, and includes their profile information.
[0758] "Video generation means" refers to a system or method for generating videos that show ideal movements to a user.
[0759] This invention is a system for efficiently improving batting and pitching form based on a user's biometric information and experience information, and also combines it with an emotion engine that recognizes the user's emotions. The specific system configuration and operation are described in detail below.
[0760] The system mainly includes the following elements:
[0761] 1. Means for inputting user biometric and experience information
[0762] 2. A means of recording user actions
[0763] 3. Means of analyzing recorded motion data
[0764] 4. A method for adjusting feedback content based on analysis results and the user's emotional state and providing feedback in various formats
[0765] 5. A means to allow students to choose guidance from multiple advisors
[0766] 6. A method for generating videos showing ideal user behavior
[0767] Entering biometric and experience information
[0768] Users use a dedicated application on a device (e.g., a smartphone) to input their own biometric information (e.g., height, weight, heart rate) and experience information (e.g., sports history, practice frequency). This information is sent from the device to a server and stored in a database.
[0769] Recording actions
[0770] For example, users can use the device's camera and sensors to record their swings and pitching movements in a form diagnostic booth installed at a batting center. The recorded video data is then sent to a cloud server.
[0771] Analysis of behavioral data
[0772] The server receives the video data sent to the cloud server and sends an analysis request to the motion analysis system. This motion analysis system analyzes the user's movements using specific algorithms and machine learning models. The analysis results are then processed into a format that is easy for the user to understand.
[0773] Emotion Analysis
[0774] The emotion engine analyzes the user's emotions in real time based on their facial expressions and voice data. The analyzed emotional state data is sent to the server and integrated with the movement analysis results.
[0775] Providing Feedback
[0776] The server generates feedback content based on the results of motion analysis and emotion analysis and sends it to the device. The device receives this data and notifies the user of the feedback in the form of text, audio, images, video, etc. This allows the user to receive specific feedback in real time that is appropriate to their psychological state.
[0777] Coaching mode and video image training
[0778] Within the app, users can select their preferred instructor from multiple advisors. Feedback is customized based on the profile of the selected advisor. Furthermore, if the user wishes, a video showing ideal swing and pitching techniques is generated and sent to the device. Users can use this video as a reference for their training.
[0779] Specific examples
[0780] For example, if a beginner user wants to improve their pitching technique, they first enter their basic information into the app. They then film a video of themselves pitching at a batting center and transfer it to the app. The server analyzes the video and provides specific advice in text and voice, such as "position your elbow a little higher." At the same time, if the emotion engine determines that the user's emotional state is "depressed," more encouraging words and positive feedback are provided.
[0781] Users can also select a "famous pitching coach" to receive customized feedback from that coach, and videos of ideal pitching form are also provided, allowing users to use these to improve their technique.
[0782] Prompt Sentence Examples
[0783] Examples of prompts for a generative AI model include:
[0784] "Please analyze my pitching video and let me know how I can improve. Also, please tailor your feedback to take into account my current emotional state."
[0785] By combining the above-mentioned methods, users can receive specific feedback in real time that is appropriate to their psychological state, allowing them to efficiently improve their form.
[0786] The flow of the identification process in the second embodiment will be described with reference to FIG.
[0787] Step 1:
[0788] Registering user information
[0789] Input: The user uses the device to input their biometric information (e.g., height, weight, heart rate) and experience information (e.g., sports history, practice frequency) into a dedicated app.
[0790] Action: The user enters the required information into the input fields and clicks the "Submit" button.
[0791] Data processing and calculation: The terminal formats the input information and generates data for transmission.
[0792] Output: The terminal sends the generated transmission data to the server.
[0793] Step 2:
[0794] Retention of Information
[0795] Input: User information data sent from the device.
[0796] How it works: The server stores the received data in a database.
[0797] Data processing and calculation: Analyzes received data and converts it into the appropriate database format.
[0798] Output: Successfully saved data is confirmed and a save confirmation notification is generated.
[0799] Step 3:
[0800] Form video recording
[0801] Input: The user has a terminal.
[0802] Actions: The user takes a video of their swing and pitching using the device's camera and sensors in the batting center's diagnostic booth. They press the record button to record their movements. When they're finished recording, they press the "stop" button.
[0803] Data processing and calculation: The device temporarily stores the recorded video data and encodes it for transfer.
[0804] Output: Send the encoded video data to the server.
[0805] Step 4:
[0806] Video data storage and analysis requests
[0807] Input: Video data sent from the device.
[0808] Operation: The server receives the video data and temporarily stores it.
[0809] Data processing and calculation: Generates requests to the motion analysis system to analyze the received video data.
[0810] Output: A request is sent to the behavior analysis system.
[0811] Step 5:
[0812] Motion analysis
[0813] Input: Video data analysis request sent from the server.
[0814] Motion: The motion analysis system analyzes the received video data using machine learning algorithms.
[0815] Data processing and calculation: The motion analysis system extracts motion elements from the video and generates analysis results.
[0816] Output: The generated behavior analysis results are sent back to the server.
[0817] Step 6:
[0818] Emotion analysis
[0819] Input: User facial and voice data.
[0820] How it works: The device uses a camera and microphone to record the user's facial expressions and voice data in real time.
[0821] Data processing and calculation: The device formats the recorded emotion data for transmission and sends it to the server. The server then sends an analysis request to the emotion analysis system.
[0822] Output: The emotion analysis system sends the analysis results to the server, and the final emotion data is generated.
[0823] Step 7:
[0824] Generating and Providing Feedback
[0825] Input: Motion analysis results and emotion analysis results.
[0826] Behavior: The server integrates the results of behavior analysis and emotion analysis and generates feedback for the user.
[0827] Data processing and calculation: Processing the feedback content into text, audio, image, or video format.
[0828] Output: The generated feedback content is sent to the device, which displays or plays it.
[0829] Step 8:
[0830] Choosing a coaching mode
[0831] Input: Profile information for multiple advisors.
[0832] How it works: A user selects their preferred advisor within the app.
[0833] Data processing and calculation: The server customizes the feedback content based on the information of the selected advisor.
[0834] Output: The customized feedback is sent to the user's device and displayed.
[0835] Step 9:
[0836] Creating image training videos
[0837] Input: The user's desired action information.
[0838] Action: The user selects the "Image Training" option.
[0839] Data processing and calculation: The server requests a video showing the ideal behavior from the generation engine and receives the generated video data.
[0840] Output: Send the generated image training video to the device.
[0841] (Application example 2)
[0842] Next, a description will be given of Application Example 2. In the following description, the data processing device 12 will be referred to as a "server" and the smart glasses 214 will be referred to as a "terminal."
[0843] Improving work efficiency and ensuring safety in factories are key challenges. In particular, there is a need to analyze the movements of robots and workers in real time and provide accurate feedback based on that analysis. It is also necessary to provide feedback that takes into account the mental state of workers in order to create a more effective work environment.
[0844] The specification processing by the specification processing unit 290 of the data processing device 12 in Application Example 2 is realized by the following means. In this invention, the server includes means for inputting biometric information and experience information of the user, means for recording the user's movements, means for analyzing the recorded movement data, means for feeding back the analysis results in various formats, means for allowing the user to select advice from multiple advisors, means for generating a video showing the user's ideal movements, means for analyzing the user's emotional state in real time, and means for adjusting the feedback content based on the user's emotional state. This enables effective feedback that takes into account the movement analysis results and the user's emotional state.
[0845] "User's biological information" is data relating to the user's physical condition, such as heart rate, respiratory rate, and body temperature.
[0846] "User experience information" is information such as the work content and proficiency level of the user that the user has previously experienced.
[0847] "User actions" refer to specific movements or tasks performed by the user.
[0848] "Motion data" refers to video recordings and measurement data of the user's movements.
[0849] "Emotional state" refers to the user's emotional or psychological state.
[0850] "Analysis" is the process of analyzing collected data and extracting meaningful information.
[0851] "Feedback" refers to the act of returning analysis results, advice, etc. to the user.
[0852] An "advisor" is an expert with specific knowledge and experience who provides guidance and advice to users.
[0853] "Ideal movement" refers to an efficient and safe movement pattern that optimizes work performance.
[0854] An "emotion engine" is a program or algorithm that analyzes a user's emotional state in real time.
[0855] A "cloud server" is an external server that analyzes and stores data via the Internet.
[0856] "Real-time" refers to processing occurring immediately or with a very short time delay.
[0857] System Overview
[0858] This invention is a system for improving work efficiency and safety in factories. The system uses a computer terminal (e.g., smart glasses or a head-mounted display) to record the user's movements and analyzes the collected data on a cloud server. The analysis results are fed back to the user in real time, and the feedback content is adjusted according to the user's emotional state.
[0859] Hardware and Software
[0860] The system includes the following major hardware and software:
[0861] Smart glasses or head-mounted displays: equipped with cameras and sensors to record user movements.
[0862] Cloud Server: Computing resources for analyzing recorded data and generating feedback.
[0863] Motion analysis model: A machine learning model trained using TensorFlow.
[0864] Emotion analysis model: Face detection using Dlib and emotion recognition model using TensorFlow.
[0865] Processing flow
[0866] 1. Data Entry:
[0867] The user uses the device to input biometric data (e.g., heart rate, body temperature) and past experience information. The data is sent from the device to a cloud server, where it is stored in a database.
[0868] 2. Operation Record:
[0869] The device records the user's actions in real time, and the recorded data is sent to a cloud server.
[0870] 3. Data Analysis:
[0871] The cloud server analyzes the received motion data using a motion analysis model, and simultaneously analyzes the user's emotional state using face detection and emotion recognition.
[0872] 4. Providing Feedback:
[0873] Feedback is generated in the form of text, audio, images, and video and is provided to the user in real time, with the feedback content adjusted according to the user's emotional state.
[0874] Specific examples
[0875] For example, the system records the movements of a robot packing items along a conveyor belt in a factory. The movement data is analyzed on a cloud server, and if there is any waste in the movement, the system provides specific advice such as "Please reduce your arm movements to reduce waste in movement." Furthermore, if signs of fatigue are detected from the worker's facial expressions or movements, it can generate emotion-based feedback such as "Please take a break."
[0876] Example prompts for generative AI models
[0877] "Please explain how the system analyzes the operational data of robots working in a factory and provides advice on improving efficiency. In addition, please incorporate a system that analyzes the emotional state of workers and provides appropriate feedback. This system will operate using smart glasses or a head-mounted display."
[0878] The flow of the specific processing in the application example 2 will be described with reference to FIG.
[0879] Step 1:
[0880] The user wears smart glasses or a head-mounted display and inputs biometric information (heart rate, body temperature, etc.) and experience information. The device sends this data to a cloud server, which then stores the received data in a database. At this stage, the input is the user's biometric information and experience information, and the output is data stored in the cloud server.
[0881] Step 2:
[0882] The device worn by the user uses built-in cameras and sensors to record the user's movements in real time. The recorded video data is converted into an appropriate format and sent to a cloud server. At this stage, the input is the user's movement data, and the output is data sent to the cloud server.
[0883] Step 3:
[0884] The cloud server inputs the transmitted motion data into a motion analysis model and performs motion analysis. The motion analysis model extracts motion characteristics (e.g., elbow angle, hand movement, etc.) and calculates indicators related to efficiency and safety. The input at this stage is the motion data, and the output is the motion analysis results.
[0885] Step 4:
[0886] In parallel with the motion analysis, the cloud server uses an emotion engine to analyze the user's emotional state from facial expression and voice data. The emotion analysis model analyzes facial features (e.g., facial muscle movements) and determines the user's emotion (e.g., fatigue, stress). The input at this stage is facial expression data and voice data, and the output is the emotion analysis results.
[0887] Step 5:
[0888] The cloud server integrates the results of the motion analysis and emotion analysis to generate appropriate feedback. The generated feedback is sent to the user's device in the form of text, audio, image, or video. For example, feedback such as "Please reduce your arm movements" or "Please take a break" is provided. The input at this stage is the results of the motion analysis and emotion analysis, and the output is the feedback content.
[0889] Step 6:
[0890] The user receives feedback through the device and modifies their behavior to improve efficiency and safety. If the user's emotional state is determined to be "fatigued," they receive feedback that reflects their emotions, such as encouragement or a prompt to take a break. The input at this stage is the feedback content, and the output is a change in the user's behavior.
[0891] The specific processing unit 290 transmits the result of the specific processing to the smart glasses 214. In the smart glasses 214, the control unit 46A causes the speaker 240 to output the result of the specific processing. The microphone 238 acquires audio indicating a user input regarding the result of the specific processing. The control unit 46A transmits audio data indicating the user input acquired by the microphone 238 to the data processing device 12. In the data processing device 12, the specific processing unit 290 acquires the audio data.
[0892] The data generation model 58 is a so-called generative AI (Artificial Intelligence). An example of the data generation model 58 is ChatGPT (Internet Search<URL: https: / / openai.com / blog / chatgpt> ), Gemini (Internet search <url: https: gemini.google.com ?hl="ja">) and other generation AIs. The data generation model 58 is obtained by performing deep learning on a neural network. A prompt including an instruction is input to the data generation model 58, and inference data such as voice data indicating voice, text data indicating text, and image data indicating an image is also input. The data generation model 58 performs inference on the input inference data in accordance with the instruction indicated by the prompt, and outputs the inference result in a data format such as voice data and text data. Here, inference refers to, for example, analysis, classification, prediction, and / or summarization.
[0893] In the above embodiment, an example in which the specific processing is performed by the data processing device 12 has been given, but the technology of the present disclosure is not limited to this, and the specific processing may be performed by the smart glasses 214.
[0894] [Third embodiment]
[0895] FIG. 5 shows an example of the configuration of a data processing system 310 according to the third embodiment.
[0896] 5, the data processing system 310 includes the data processing device 12 and a headset type terminal 314. An example of the data processing device 12 is a server.
[0897] The data processing device 12 includes a computer 22, a database 24, and a communication I / F 26. The computer 22 is an example of a "computer" according to the technology of the present disclosure. The computer 22 includes a processor 28, a RAM 30, and a storage 32. The processor 28, the RAM 30, and the storage 32 are connected to a bus 34. The database 24 and the communication I / F 26 are also connected to the bus 34. The communication I / F 26 is connected to a network 54. Examples of the network 54 include a WAN (Wide Area Network) and / or a LAN (Local Area Network).
[0898] The headset type terminal 314 includes a computer 36, a microphone 238, a speaker 240, a camera 42, a communication I / F 44, and a display 343. The computer 36 includes a processor 46, a RAM 48, and a storage 50. The processor 46, the RAM 48, and the storage 50 are connected to a bus 52. The microphone 238, the speaker 240, the camera 42, and the display 343 are also connected to the bus 52.
[0899] The microphone 238 receives instructions and the like from the user 20 by receiving voice uttered by the user 20. The microphone 238 captures the voice uttered by the user 20, converts the captured voice into audio data, and outputs it to the processor 46. The speaker 240 outputs audio in accordance with instructions from the processor 46.
[0900] Camera 42 is a small digital camera equipped with an optical system including a lens, aperture, and shutter, and an imaging element such as a CMOS (Complementary Metal-Oxide-Semiconductor) image sensor or a CCD (Charge Coupled Device) image sensor, and captures images of the surroundings of user 20 (for example, an imaging range defined by an angle of view equivalent to the field of vision of a typical healthy person).
[0901] The communication I / F 44 is connected to a network 54. The communication I / Fs 44 and 26 are responsible for the exchange of various information between the processor 46 and the processor 28 via the network 54. The exchange of various information between the processor 46 and the processor 28 using the communication I / Fs 44 and 26 is carried out in a secure state.
[0902] Fig. 6 shows an example of the main functions of the data processing device 12 and the headset type terminal 314. As shown in Fig. 6, in the data processing device 12, a specific process is performed by the processor 28. A specific process program 56 is stored in the storage 32.
[0903] The specific processing program 56 is an example of a "program" according to the technology of the present disclosure. The processor 28 reads the specific processing program 56 from the storage 32 and executes the read specific processing program 56 on the RAM 30. The specific processing is realized by the processor 28 operating as a specific processing unit 290 in accordance with the specific processing program 56 executed on the RAM 30.
[0904] The storage 32 stores a data generation model 58 and an emotion identification model 59. The data generation model 58 and the emotion identification model 59 are used by the identification processing unit 290.
[0905] In the headset type terminal 314, a reception output process is performed by the processor 46. A reception output program 60 is stored in the storage 50. The processor 46 reads the reception output program 60 from the storage 50 and executes the read reception output program 60 on the RAM 48. The reception output process is realized by the processor 46 operating as the control unit 46A in accordance with the reception output program 60 executed on the RAM 48.
[0906] Next, a description will be given of the identification process performed by the identification processing unit 290 of the data processing device 12. In the following description, the data processing device 12 will be referred to as the "server" and the headset type terminal 314 will be referred to as the "terminal."
[0907] This invention is a system for efficiently improving batting and pitching form based on a user's biometric and experience information. The system records the user's movements, analyzes the data, and provides specific feedback in real time. It also provides a means for users to choose from multiple advisors and generates videos demonstrating ideal movements, promoting the user's technical improvement.
[0908] Overall system configuration
[0909] Users use a device (e.g., a smartphone) to access the system. A dedicated application is installed on the device, allowing users to input their biometric information and experience information. The input information is sent to the server and stored in a database.
[0910] The device is equipped with a camera and sensors to record the user's movements. The captured video data is sent to a cloud server and analyzed using Google's Gemini. The analysis results are provided to the user in real time as feedback in the form of text, audio, images, video, etc.
[0911] The system also has multiple advisor profiles, allowing users to receive guidance from their preferred advisor. Furthermore, it can generate videos demonstrating ideal movements, providing users with visual reference information.
[0912] Program processing
[0913] The program does the following:
[0914] 1. User information registration
[0915] The user uses the device to input their biometric and experience information. Once input is complete, the device sends the data to the server, which stores the received information in a database.
[0916] 2. Recording a video of your form
[0917] Users record their swings and throws in the form diagnostic booth at a batting center. The motions are recorded using the device's camera or sensor, and the data is sent to a server.
[0918] 3. Form video analysis
[0919] The server sends the received video data to Google Gemini and makes an analysis request. Google Gemini analyzes the motion and returns the results to the server. The analysis results are then processed into a user-friendly format.
[0920] 4. Providing real-time feedback
[0921] The processed analysis results are sent to the device and notified to the user in the form of text, audio, images, video, etc. The user can then try to improve the form based on this information.
[0922] 5. Select coaching mode
[0923] Users can select one of several advisors within the app, and the analysis results are customized based on the profile information of the selected advisor.
[0924] 6. Image training using videos
[0925] If the user wishes, an image video of the ideal swing or throwing technique will be generated and sent to the device, and the user can use this video as a reference for training.
[0926] Specific examples
[0927] For example, if a beginner user wants to improve their pitching technique, they first enter and save their basic information in the app. They then film a video of themselves pitching at a batting center and transfer it to the app. The server analyzes the video and provides specific advice in text and voice, such as "position your elbow a little higher." Additionally, if the user selects a "famous pitching coach," they will receive customized feedback from that coach. Furthermore, a video of the ideal pitching form is also provided, allowing the user to use it as a reference to improve their technique.
[0928] The system allows users to receive specific, expert feedback in real time, allowing them to efficiently improve their form.
[0929] The processing flow will be explained below.
[0930] Step 1:
[0931] User: Launch the app and enter basic information such as age, gender, baseball experience, and medical restrictions on the "New Registration" screen.
[0932] Step 2:
[0933] Terminal: When input is complete, the user presses the save button, and the input data is temporarily saved in local storage.
[0934] Step 3:
[0935] Terminal: Generates and sends a request to send the stored data to the server.
[0936] Step 4:
[0937] Server: Saves the received user data in a database and returns a confirmation response to the terminal that the data was saved successfully.
[0938] Step 5:
[0939] User: Enters the form diagnostic booth at the batting center, stands in the designated position, and begins swinging or throwing.
[0940] Step 6:
[0941] Device: When an action is initiated, the device uses cameras and sensors to record the user's actions in video format.
[0942] Step 7:
[0943] User: When the action is finished, press the stop recording button to end the recording.
[0944] Step 8:
[0945] Terminal: Saves the captured video data in local storage, generates a request to send the data to the server, and sends it.
[0946] Step 9:
[0947] Server: Temporarily stores the received video data and sends an analysis request to Google's Gemini API.
[0948] Step 10:
[0949] Server: Receives analysis results from Google's "Gemini".
[0950] Step 11:
[0951] Server: Processes the received analysis results into a user-friendly format (text, audio, image, video).
[0952] Step 12:
[0953] Server: Sends the processed analysis results to the terminal.
[0954] Step 13:
[0955] Device: Receives analysis results and displays / notifies them on the app.
[0956] Step 14:
[0957] User: Check the analysis results and use the improvements to improve their next swing or throw.
[0958] Step 15:
[0959] User: Select one of several advisors on the "Coaching Mode Selection" screen within the app.
[0960] Step 16:
[0961] Terminal: Based on the information of the selected advisor, the analysis results are customized and provided to the user.
[0962] Step 17:
[0963] Terminal: Sends a request to the server to generate an image video of the ideal swing or throwing based on the user's wishes.
[0964] Step 18:
[0965] Server: Generates an image video of the ideal form and sends it to the device.
[0966] Step 19:
[0967] Device: The received image video is played within the app and shown to the user.
[0968] Step 20:
[0969] User: Use the image video as a reference and train to improve their own form.
[0970] This allows users to receive specific, professional feedback in real time and efficiently improve their form.
[0971] Example 1
[0972] Next, a description will be given of Example 1. In the following description, the data processing device 12 will be referred to as a "server" and the headset type terminal 314 will be referred to as a "terminal."
[0973] Conventional methods for improving batting and pitching form lack the means for users to accurately understand their own movements and receive effective feedback. Furthermore, opportunities to receive direct instruction from a coach are limited, making self-study to improve skills difficult. The objective of this invention is to provide a system that provides specific, specialized feedback in real time, allowing users to efficiently improve their movements.
[0974] The specific processing by the specific processing unit 290 of the data processing device 12 in the first embodiment is realized by the following means.
[0975] In this invention, the server includes means for inputting biometric information and experience information of the user, a terminal device for recording the user's movements, means for transmitting the recorded video data to the analytical model via the cloud server, means for processing the analysis results from the analytical model into a user-friendly format, means for feeding back the analysis results to the user in real time, means for enabling the user to select instruction from multiple instructors, and means for generating videos showing the user's ideal movements. This allows the user to receive specific and professional feedback in real time and efficiently improve their own movements.
[0976] "Biometric information" refers to information about the user's physical characteristics, such as the user's height, weight, athletic ability, and health condition.
[0977] "Experience information" is information about the user's experience, such as the user's sports history, skill level, and past training history.
[0978] A "terminal device" is a hardware device, such as a smartphone or tablet, that allows a user to input information and record actions.
[0979] A "cloud server" is a remote server connected via the Internet for storing, managing, and analyzing data.
[0980] An "analytics model" is a machine learning algorithm or data processing mechanism that analyzes a user's behavior data and provides form improvements and technical advice.
[0981] "Feedback" refers to providing users with information such as advice, comments, and analysis results generated based on the user's behavior data.
[0982] "Instructors" are coaches or advisors with specialized knowledge and experience that can be selected by the user.
[0983] "Videos showing ideal movements" are videos that show the ideal form and techniques that users should aim for, and are used to support the user's training.
[0984] The present invention provides a system for effectively improving a user's sports form based on biometric information and experience information. Specific embodiments for carrying out the present invention will be described in detail below.
[0985] Overall system configuration
[0986] Registering user information
[0987] The user starts a dedicated application using a device (e.g., a smartphone) and inputs biometric information (e.g., height, weight, health status) and experience information (e.g., sports history, skill level). This information is sent from the device to a server and stored in a database.
[0988] Recording actions
[0989] Users can record their batting and pitching movements at batting centers or practice fields using their smartphone cameras or dedicated sensors. The recorded video data is temporarily stored on the device and then uploaded to a cloud server.
[0990] Video Analysis
[0991] The server then sends the received video data to an analytical model (e.g., a machine learning algorithm) in the cloud and issues an analysis request. The analytical model analyzes the user's movements and extracts data on form flaws and areas for improvement. The extracted results are returned to the server, where they are further processed into a format that is easy for the user to understand (e.g., text, images, video).
[0992] Providing Feedback
[0993] The server then sends the processed analysis results to the device in real time. The device then displays the results and provides specific advice to the user via text or voice, such as "position your elbows a little higher." Real-time video feedback is also provided based on the analysis results.
[0994] Choosing a coaching mode
[0995] Users can browse the profiles of multiple instructors from the in-app menu and select their preferred instructor. The server generates customized feedback based on the profile information of the selected instructor and sends it to the device.
[0996] Image training using videos
[0997] It can generate images that allow users to reproduce their ideal movements. The server generates a video of the ideal form and sends it to the device. Users can use this video as a reference for training and improve their form.
[0998] Specific examples
[0999] For example, if a beginner user wants to improve their pitching form, they first enter and save their basic information into the app on their smartphone. Next, they use the smartphone camera to film their pitching at a batting center, and the video is sent from the app to a cloud server. The cloud server analyzes the video and sends specific advice, such as "position your elbow a little higher," to the user's smartphone in real time. Furthermore, if the user selects a well-known coach, that coach will provide them with customized feedback. Finally, a video of their ideal pitching form is provided to the user, who can use the video as a reference to improve their technique.
[1000] Generative AI model prompt example
[1001] "Please analyze the batting form. What is the best way to teach the player if their elbow is positioned too low?"
[1002] "Please tell me the ideal throwing motion for pitching form."
[1003] In this way, by inputting specific questions into the generative AI model, optimal feedback can be obtained.
[1004] The flow of the identification process in the first embodiment will be described with reference to FIG.
[1005] Program processing steps
[1006] Step 1: Register user information
[1007] Input: User biometric and experience information
[1008] Processing: The user launches a dedicated application using a device (e.g., a smartphone) and enters basic information such as height, weight, and sports history into a form. The device then sends the entered information to the server using the HTTPS protocol.
[1009] Output: The entered user information is saved on the cloud server.
[1010] Specific operation: When a user enters information on their smartphone and presses the "Send" button, the information is recorded in a database (e.g., MySQL) on the cloud server.
[1011] Step 2: Record a video of your form
[1012] Input: User action (e.g. swing or throw)
[1013] Processing: A user uses a smartphone camera app to record their movements at a batting cage or practice area. Once the recording is complete, it is temporarily saved as a video file on the device.
[1014] Output: The video data file will be saved on your device.
[1015] Specific operation: The user presses the record button on the camera, and when they finish recording their movements, they press the "Save" button to save the video data to their device.
[1016] Step 3: Submitting the form video
[1017] Input: Video data stored on the device
[1018] Processing: The device uploads the recorded video file to a cloud server using a secure file transfer protocol (e.g., HTTPS or SFTP).
[1019] Output: Video data is stored in cloud server storage (e.g. AWS S3).
[1020] Specific operation: The user presses the "upload" button within the app, and the video is sent to the cloud server.
[1021] Step 4: Analysis instructions for form videos
[1022] Input: Video data stored on a cloud server
[1023] Processing: The server sends the video data to the analytical model and makes an analysis request. The analytical model uses a motion analysis algorithm.
[1024] Output: The analysis request is sent to the analysis model.
[1025] Specific operation: The server automatically checks the stored video data and sends an API request to the analysis model.
[1026] Step 5: Analyzing your form video
[1027] Input: Video data and analysis request
[1028] Processing: The analytical model analyzes user behavior from video data and generates data on form parameters and areas for improvement.
[1029] Output: The analysis result data is sent back to the server.
[1030] Specific behavior: The analytical model performs behavior analysis, generates analysis results, and sends them back to the server.
[1031] Step 6: Processing the analysis results and generating feedback data
[1032] Input: Analysis result data
[1033] Processing: The server receives the analysis results data and processes it into a format that is easy for the user to understand (e.g., text, charts, videos).
[1034] Output: Processed feedback data is generated.
[1035] Specific operation: The server processes the analysis result data for feedback and converts it into the most suitable format for the user.
[1036] Step 7: Provide feedback
[1037] Input: Processed feedback data
[1038] Processing: The server sends the feedback data to the device using a real-time notification service (e.g., Firebase Cloud Messaging). The device displays the feedback data to the user.
[1039] Output: Feedback is displayed on the user's terminal.
[1040] Specific operation: A notification is sent to the user's smartphone, and when they open the app, advice such as "position your elbows a little higher" is displayed.
[1041] Step 8: Select a coaching mode
[1042] Input: User-selected mentor profile
[1043] What happens: Users select one of several mentors within the app, and feedback is customized based on their profile information.
[1044] Output: Customized feedback data is generated.
[1045] Specific operation: The server regenerates feedback using the profile information of the selected instructor and sends it to the terminal.
[1046] Step 9: Image training using videos
[1047] Input: User training request
[1048] Processing: The server generates a video demonstrating the ideal movement and sends it to the device. The user uses this video as a reference for training.
[1049] Output: A video showing the ideal behavior is displayed on the device.
[1050] Specific movements: Users play videos of ideal form in the app and train by imitating the movements.
[1051] Generative AI model prompt example
[1052] "Please analyze the batting form. What is the best way to teach the player if their elbow is positioned too low?"
[1053] "Please tell me the ideal throwing motion for pitching form."
[1054] (Application example 1)
[1055] Next, a description will be given of Application Example 1. In the following description, the data processing device 12 will be referred to as a "server" and the headset type terminal 314 will be referred to as a "terminal."
[1056] In order to ensure the efficiency and safety of industrial robots, it is important to evaluate the accuracy and appropriateness of their movements in real time and provide feedback on areas for improvement. However, existing systems do not fully realize these functions, making it difficult to optimize robot movements.
[1057] The specific processing by the specific processing unit 290 of the data processing device 12 in the application example 1 is realized by the following means.
[1058] In this invention, the server includes a means for transmitting user information to the server and storing it in a database, a means for transmitting operation records to the server and analyzing them using a model, and a means for converting the analysis results into a user-friendly format in real time and providing the results, thereby enabling efficient and safe optimization of the operation of industrial robots.
[1059] "User's biometric information" refers to data related to the user's body, including information such as heart rate, body temperature, and muscle movements.
[1060] "Experience information" refers to the knowledge and experience that the user has accumulated up to now, and is data such as proficiency level and past performance records.
[1061] A "recording means" is a device that collects the user's actions using a video camera or a sensor.
[1062] "Means for analyzing" refers to algorithms or software for evaluating recorded motion data using analytical models.
[1063] "Feedback means" refers to means for providing analysis results to the user in the form of text, audio, images, video, etc.
[1064] The "means for enabling selection of instructor" refers to an interface that allows the user to select one instructor from multiple instructors and receive evaluation and advice from that instructor.
[1065] The "means for generating animation" refers to a system that creates animation to visually show the user's ideal movements.
[1066] A "server" is a computer system that processes, manages, and provides data over a network.
[1067] The "means for saving to a database" is a function for safely saving input user information to a database.
[1068] "Means of analysis using models" refers to means of analyzing data using technologies such as machine learning and artificial intelligence to generate insights and improvement suggestions.
[1069] "Means for converting the analysis results into a user-friendly format and providing them" refers to a method for converting the analysis results into an easy-to-understand and comprehend format and providing them to the user.
[1070] The "means for proposing improvements to operations" is a function for proposing specific improvements to the user based on the analysis results.
[1071] A specific system for implementing this invention includes the following components: First, a user uses a terminal to input their own biometric information and experience information. A dedicated application is installed on the terminal, which transmits the user's information to a server and stores it in a database. The system records the user's movements using a camera or sensor and evaluates the data using an analytical model. The analysis results are processed on a cloud server and converted into a user-friendly format. These results are then fed back to the user in real time.
[1072] The system also features an interface that allows users to select instruction from multiple instructors. Users can receive customized feedback from the instructor of their choice based on their own movement data. Furthermore, the system generates an ideal movement model based on the analysis results and makes suggestions for movement improvement based on that model.
[1073] Hardware and software used
[1074] Hardware
[1075] Camera (e.g. Logitech C920)
[1076] Computer system (e.g., CPU: Intel i7, RAM: 16GB, etc.)
[1077] Robots (e.g., industrial arm robots)
[1078] software
[1079] Python programming language
[1080] OpenCV library (motion capture and image processing)
[1081] Requests library (sending HTTP requests)
[1082] Server-side program (runs on a cloud server)
[1083] Analytical models (using machine learning and artificial intelligence techniques)
[1084] Processing example
[1085] For example, when improving the behavior of a welding robot in a factory to ensure accurate welding, a camera captures the robot's movements and sends the data to a cloud server for analysis. Specific feedback such as "The welding position needs to be moved 5 mm to the left" is provided as a result of the analysis. A video demonstrating the ideal welding behavior is also provided to the user.
[1086] Prompt Sentence Examples
[1087] "Please create a system that records the operations of welding robots in factories with a camera, analyzes them on a cloud server, and provides suggestions for improvement. Please provide specific feedback in real time based on the analysis results."
[1088] This system makes it possible to optimize the operation of industrial robots efficiently and safely.
[1089] The flow of the specific processing in the application example 1 will be described with reference to FIG.
[1090] Step 1:
[1091] Registering user information
[1092] The user inputs their biometric and experience information into the device. This input data includes physical information such as heart rate, body temperature, and muscle movement, as well as skill level and past performance records. The device sends this data to the server, which stores the received data in a database. The input of this process is the biometric and experience information entered by the user, and the output is the user information stored in the database.
[1093] Step 2:
[1094] Shooting a video of the operation
[1095] A camera records the user's actions. In this case, the object of the action is an industrial robot, for example, recording the robot's welding work. The camera (e.g., Logitech C920) captures the robot's actions and saves them as a video file. The input of this process is the robot's actions, and the output is the recorded video file.
[1096] Step 3:
[1097] Analysis of motion video
[1098] The device sends the recorded video file to a server. The server receives this video data and analyzes it using an analytical model (using machine learning and artificial intelligence techniques). The analytical model takes the video data as input and evaluates the accuracy and appropriateness of the movements, which then generates specific suggestions for improving the movements. The input to this process is the video file, and the output is the analysis results.
[1099] Step 4:
[1100] Providing real-time feedback
[1101] The server provides the analysis results to the user in real time by converting the analysis results into a user-friendly format and sending them to the device in the form of text, audio, images, video, etc. This allows the user to immediately understand specific areas for improvement. The input of this process is the analysis results, and the output is the feedback provided to the user.
[1102] Step 5:
[1103] Leader's Choice
[1104] The user can select one of several mentors within the app. Based on the profile information of the selected mentor, the server customizes the analysis results, allowing the user to receive advice and make this feedback more actionable. The input to this process is the user's mentor selection, and the output is customized feedback.
[1105] Step 6:
[1106] Generation of ideal behavior
[1107] The server generates an ideal motion model based on the user's motion data. Based on the generated model, a video demonstrating the ideal motion is created and sent to the device. The user can refer to this video to improve their own motion. The input to this process is the user's motion data, and the output is a video demonstrating the ideal motion.
[1108] Furthermore, an emotion engine that estimates the user's emotion may be further combined. That is, the identification processing unit 290 may estimate the user's emotion using the emotion identification model 59, and perform identification processing using the user's emotion.
[1109] This invention is a system for efficiently improving batting and pitching form based on a user's biometric and experience information, and also combines it with an emotion engine that recognizes the user's emotions. The system records the user's movements, analyzes the data, and provides specific feedback in real time, adjusting the feedback content according to the user's emotional state. It also provides a means for users to choose instruction from multiple advisors and generates videos demonstrating ideal movements, promoting the user's improvement.
[1110] Overall system configuration
[1111] Users use a device (e.g., a smartphone) to access the system. A dedicated application is installed on the device, allowing users to input their biometric information and experience information. The input information is sent to the server and stored in a database.
[1112] The device is equipped with a camera and sensors to record the user's movements. The captured video data is sent to a cloud server and analyzed using a motion analysis system. The analysis results are provided to the user in real time as feedback in the form of text, audio, images, video, etc.
[1113] The system also has multiple advisor profiles, allowing users to receive guidance from their preferred advisor. Furthermore, it can generate videos demonstrating ideal movements, providing users with visual reference information.
[1114] In addition, the emotion engine analyzes the user's emotional state in real time and adjusts the feedback content based on the results, allowing users to receive feedback that is appropriate for their own psychological state, resulting in even more effective training.
[1115] Program processing
[1116] The program does the following:
[1117] 1. User information registration
[1118] The user uses the device to input their biometric and experience information. Once input is complete, the device sends the data to the server, which stores the received information in a database.
[1119] 2. Recording a video of your form
[1120] Users record their swings and throws in the form diagnostic booth at a batting center. The motions are recorded using the device's camera or sensor, and the data is sent to a server.
[1121] 3. Form video analysis
[1122] The server sends the received video data to the motion analysis system and issues an analysis request. The system analyzes the motion and returns the results to the server, which then processes the analysis results into a user-friendly format.
[1123] 4. Emotion Analysis
[1124] The emotion engine analyzes emotional data in real time from the user's movements, facial expressions, voice, etc., and sends the results to the server.
[1125] 5. Providing Feedback
[1126] The processed analysis results are combined with the emotion analysis results and sent to the device, which then notifies the user in the form of text, audio, images, video, etc. The user can then try to improve the form based on this information.
[1127] 6. Select coaching mode
[1128] Users select one of several advisors within the app, and the analysis results and emotional data are customized based on the profile information of the selected advisor.
[1129] 7. Image training using videos
[1130] If the user wishes, an image video of the ideal swing or throwing technique will be generated and sent to the device, and the user can use this video as a reference for training.
[1131] Specific examples
[1132] For example, if a beginner user wants to improve their pitching technique, they first enter and save their basic information into the app. They then take a video of themselves pitching at a batting center and transfer it to the app. The server analyzes the video and provides specific advice, such as "position your elbow a little higher," via text and voice.
[1133] At the same time, if the emotion engine determines that the user's emotional state is "down," more encouraging words and positive feedback will be provided. Additionally, if the user selects a "famous pitching coach," customized feedback from that coach will be provided. Furthermore, videos of ideal pitching form are also provided, so users can use them as reference to improve their technique.
[1134] This system allows users to receive real-time, specific, and psychologically relevant feedback, allowing them to efficiently improve their form.
[1135] The processing flow will be explained below.
[1136] Step 1:
[1137] User: Launch the app and enter basic information such as age, gender, baseball experience, and medical restrictions on the "New Registration" screen.
[1138] Step 2:
[1139] Terminal: When input is complete, the user presses the save button, and the input data is temporarily saved in local storage.
[1140] Step 3:
[1141] Terminal: Generates and sends a request to send the stored data to the server.
[1142] Step 4:
[1143] Server: Saves the received user data in a database and returns a confirmation response to the terminal that the data was saved successfully.
[1144] Step 5:
[1145] User: Enters the form diagnostic booth at the batting center, stands in the designated position, and begins swinging or throwing.
[1146] Step 6:
[1147] Device: When an action is initiated, the device uses cameras and sensors to record the user's actions in video format.
[1148] Step 7:
[1149] User: When the action is finished, press the stop recording button to end the recording.
[1150] Step 8:
[1151] Terminal: Saves the captured video data in local storage, generates a request to send the data to the server, and sends it.
[1152] Step 9:
[1153] Server: Temporarily stores the received video data and sends an analysis request to the motion analysis system.
[1154] Step 10:
[1155] Server: Receives analysis results from the motion analysis system.
[1156] Step 11:
[1157] Server: Processes the received analysis results into a user-friendly format (text, audio, image, video).
[1158] Step 12:
[1159] Terminal: Activates an emotion engine that recognizes emotions from the user's movements, facial expressions, and voice, and acquires emotional data.
[1160] Step 13:
[1161] Server: Analyzes the emotion data sent from the emotion engine and uses it to adjust the feedback content.
[1162] Step 14:
[1163] Server: Combines the processed analysis results with the emotion analysis results and sends them to the device.
[1164] Step 15:
[1165] Terminal: Feedback based on analysis results and emotional data is displayed and notified to the user in the form of text, audio, images, video, etc.
[1166] Step 16:
[1167] User: Check the feedback and incorporate it into their next swing or throw.
[1168] Step 17:
[1169] User: Select one of several advisors on the "Coaching Mode Selection" screen within the app.
[1170] Step 18:
[1171] Terminal: Based on the information of the selected advisor, the analysis results and emotional data are customized and provided to the user.
[1172] Step 19:
[1173] Terminal: Sends a request to the server to generate an image video of the ideal swing or throwing based on the user's wishes.
[1174] Step 20:
[1175] Server: Generates an image video of the ideal form and sends it to the device.
[1176] Step 21:
[1177] Device: The received image video is played within the app and shown to the user.
[1178] Step 22:
[1179] User: Use the image video as a reference and train to improve their own form.
[1180] This allows users to receive real-time, specific, and psychologically relevant feedback, allowing them to efficiently improve their form.
[1181] Example 2
[1182] Next, a description will be given of Example 2. In the following description, the data processing device 12 will be referred to as a "server" and the headset type terminal 314 will be referred to as a "terminal."
[1183] Conventional form improvement systems have limited dimensionality in the user's motion analysis and feedback, and in many cases, the feedback does not take into account the user's psychological state, resulting in reduced training efficiency. Furthermore, specialized equipment is often used for motion analysis and evaluation, which places constraints on cost and usage environment. Furthermore, it is difficult to select detailed guidance from multiple advisors.
[1184] The specific processing by the specific processing unit 290 of the data processing device 12 in the second embodiment is realized by the following means.
[1185] In this invention, the server includes means for inputting biometric information and experience information of the user, means for recording the user's movements, means for analyzing the recorded movement data, means for adjusting the feedback content based on the analysis results and the user's emotional state and providing the feedback in various formats, means for allowing the user to select advice from multiple advisors, and means for generating videos showing the user's ideal movements. This allows the user to receive individual and optimal feedback in real time based on the results of their own movement analysis and psychological state, enabling them to improve their form efficiently and effectively.
[1186] "Biometric information" is information that indicates the physical condition of the user, and examples include height, weight, and heart rate.
[1187] "Experience information" is information about the user's skills and career history, and examples include sports history and practice frequency.
[1188] "Motion data" is data that records the user's physical movements, and includes video and sensor data acquired using cameras and sensors.
[1189] "Analysis means" refers to a method or system for analyzing recorded motion data and providing the results.
[1190] "Feedback means" refers to a means of providing information to users based on the analysis results, and includes systems that display information in a variety of formats, such as text, audio, images, and video.
[1191] "Emotional state" refers to a user's psychological state and includes methods and systems for analyzing that state.
[1192] "Advisor" refers to an expert who provides technical guidance and advice to users, and includes their profile information.
[1193] "Video generation means" refers to a system or method for generating videos that show ideal movements to a user.
[1194] This invention is a system for efficiently improving batting and pitching form based on a user's biometric information and experience information, and also combines it with an emotion engine that recognizes the user's emotions. The specific system configuration and operation are described in detail below.
[1195] The system mainly includes the following elements:
[1196] 1. Means for inputting user biometric and experience information
[1197] 2. A means of recording user actions
[1198] 3. Means of analyzing recorded motion data
[1199] 4. A method for adjusting feedback content based on analysis results and the user's emotional state and providing feedback in various formats
[1200] 5. A means to allow students to choose guidance from multiple advisors
[1201] 6. A method for generating videos showing ideal user behavior
[1202] Entering biometric and experience information
[1203] Users use a dedicated application on a device (e.g., a smartphone) to input their own biometric information (e.g., height, weight, heart rate) and experience information (e.g., sports history, practice frequency). This information is sent from the device to a server and stored in a database.
[1204] Recording actions
[1205] For example, users can use the device's camera and sensors to record their swings and pitching movements in a form diagnostic booth installed at a batting center. The recorded video data is then sent to a cloud server.
[1206] Analysis of behavioral data
[1207] The server receives the video data sent to the cloud server and sends an analysis request to the motion analysis system. This motion analysis system analyzes the user's movements using specific algorithms and machine learning models. The analysis results are then processed into a format that is easy for the user to understand.
[1208] Emotion Analysis
[1209] The emotion engine analyzes the user's emotions in real time based on their facial expressions and voice data. The analyzed emotional state data is sent to the server and integrated with the movement analysis results.
[1210] Providing Feedback
[1211] The server generates feedback content based on the results of motion analysis and emotion analysis and sends it to the device. The device receives this data and notifies the user of the feedback in the form of text, audio, images, video, etc. This allows the user to receive specific feedback in real time that is appropriate to their psychological state.
[1212] Coaching mode and video image training
[1213] Within the app, users can select their preferred instructor from multiple advisors. Feedback is customized based on the profile of the selected advisor. Furthermore, if the user wishes, a video showing ideal swing and pitching techniques is generated and sent to the device. Users can use this video as a reference for their training.
[1214] Specific examples
[1215] For example, if a beginner user wants to improve their pitching technique, they first enter their basic information into the app. They then film a video of themselves pitching at a batting center and transfer it to the app. The server analyzes the video and provides specific advice in text and voice, such as "position your elbow a little higher." At the same time, if the emotion engine determines that the user's emotional state is "depressed," more encouraging words and positive feedback are provided.
[1216] Users can also select a "famous pitching coach" to receive customized feedback from that coach, and videos of ideal pitching form are also provided, allowing users to use these to improve their technique.
[1217] Prompt Sentence Examples
[1218] Examples of prompts for a generative AI model include:
[1219] "Please analyze my pitching video and let me know how I can improve. Also, please tailor your feedback to take into account my current emotional state."
[1220] By combining the above-mentioned methods, users can receive specific feedback in real time that is appropriate to their psychological state, allowing them to efficiently improve their form.
[1221] The flow of the identification process in the second embodiment will be described with reference to FIG.
[1222] Step 1:
[1223] Registering user information
[1224] Input: The user uses the device to input their biometric information (e.g., height, weight, heart rate) and experience information (e.g., sports history, practice frequency) into a dedicated app.
[1225] Action: The user enters the required information into the input fields and clicks the "Submit" button.
[1226] Data processing and calculation: The terminal formats the input information and generates data for transmission.
[1227] Output: The terminal sends the generated transmission data to the server.
[1228] Step 2:
[1229] Retention of Information
[1230] Input: User information data sent from the device.
[1231] How it works: The server stores the received data in a database.
[1232] Data processing and calculation: Analyzes received data and converts it into the appropriate database format.
[1233] Output: Successfully saved data is confirmed and a save confirmation notification is generated.
[1234] Step 3:
[1235] Form video recording
[1236] Input: The user has a terminal.
[1237] Actions: The user takes a video of their swing and pitching using the device's camera and sensors in the batting center's diagnostic booth. They press the record button to record their movements. When they're finished recording, they press the "stop" button.
[1238] Data processing and calculation: The device temporarily stores the recorded video data and encodes it for transfer.
[1239] Output: Send the encoded video data to the server.
[1240] Step 4:
[1241] Video data storage and analysis requests
[1242] Input: Video data sent from the device.
[1243] Operation: The server receives the video data and temporarily stores it.
[1244] Data processing and calculation: Generates requests to the motion analysis system to analyze the received video data.
[1245] Output: A request is sent to the behavior analysis system.
[1246] Step 5:
[1247] Motion analysis
[1248] Input: Video data analysis request sent from the server.
[1249] Motion: The motion analysis system analyzes the received video data using machine learning algorithms.
[1250] Data processing and calculation: The motion analysis system extracts motion elements from the video and generates analysis results.
[1251] Output: The generated behavior analysis results are sent back to the server.
[1252] Step 6:
[1253] Emotion analysis
[1254] Input: User facial and voice data.
[1255] How it works: The device uses a camera and microphone to record the user's facial expressions and voice data in real time.
[1256] Data processing and calculation: The device formats the recorded emotion data for transmission and sends it to the server. The server then sends an analysis request to the emotion analysis system.
[1257] Output: The emotion analysis system sends the analysis results to the server, and the final emotion data is generated.
[1258] Step 7:
[1259] Generating and Providing Feedback
[1260] Input: Motion analysis results and emotion analysis results.
[1261] Behavior: The server integrates the results of behavior analysis and emotion analysis and generates feedback for the user.
[1262] Data processing and calculation: Processing the feedback content into text, audio, image, or video format.
[1263] Output: The generated feedback content is sent to the device, which displays or plays it.
[1264] Step 8:
[1265] Choosing a coaching mode
[1266] Input: Profile information for multiple advisors.
[1267] How it works: A user selects their preferred advisor within the app.
[1268] Data processing and calculation: The server customizes the feedback content based on the information of the selected advisor.
[1269] Output: The customized feedback is sent to the user's device and displayed.
[1270] Step 9:
[1271] Creating image training videos
[1272] Input: The user's desired action information.
[1273] Action: The user selects the "Image Training" option.
[1274] Data processing and calculation: The server requests a video showing the ideal behavior from the generation engine and receives the generated video data.
[1275] Output: Send the generated image training video to the device.
[1276] (Application example 2)
[1277] Next, a description will be given of Application Example 2. In the following description, the data processing device 12 will be referred to as a "server" and the headset type terminal 314 will be referred to as a "terminal."
[1278] Improving work efficiency and ensuring safety in factories are key challenges. In particular, there is a need to analyze the movements of robots and workers in real time and provide accurate feedback based on that analysis. It is also necessary to provide feedback that takes into account the mental state of workers in order to create a more effective work environment.
[1279] The specification processing by the specification processing unit 290 of the data processing device 12 in Application Example 2 is realized by the following means. In this invention, the server includes means for inputting biometric information and experience information of the user, means for recording the user's movements, means for analyzing the recorded movement data, means for feeding back the analysis results in various formats, means for allowing the user to select advice from multiple advisors, means for generating a video showing the user's ideal movements, means for analyzing the user's emotional state in real time, and means for adjusting the feedback content based on the user's emotional state. This enables effective feedback that takes into account the movement analysis results and the user's emotional state.
[1280] "User's biological information" is data relating to the user's physical condition, such as heart rate, respiratory rate, and body temperature.
[1281] "User experience information" is information such as the work content and proficiency level of the user that the user has previously experienced.
[1282] "User actions" refer to specific movements or tasks performed by the user.
[1283] "Motion data" refers to video recordings and measurement data of the user's movements.
[1284] "Emotional state" refers to the user's emotional or psychological state.
[1285] "Analysis" is the process of analyzing collected data and extracting meaningful information.
[1286] "Feedback" refers to the act of returning analysis results, advice, etc. to the user.
[1287] An "advisor" is an expert with specific knowledge and experience who provides guidance and advice to users.
[1288] "Ideal movement" refers to an efficient and safe movement pattern that optimizes work performance.
[1289] An "emotion engine" is a program or algorithm that analyzes a user's emotional state in real time.
[1290] A "cloud server" is an external server that analyzes and stores data via the Internet.
[1291] "Real-time" refers to processing occurring immediately or with a very short time delay.
[1292] System Overview
[1293] This invention is a system for improving work efficiency and safety in factories. The system uses a computer terminal (e.g., smart glasses or a head-mounted display) to record the user's movements and analyzes the collected data on a cloud server. The analysis results are fed back to the user in real time, and the feedback content is adjusted according to the user's emotional state.
[1294] Hardware and Software
[1295] The system includes the following major hardware and software:
[1296] Smart glasses or head-mounted displays: equipped with cameras and sensors to record user movements.
[1297] Cloud Server: Computing resources for analyzing recorded data and generating feedback.
[1298] Motion analysis model: A machine learning model trained using TensorFlow.
[1299] Emotion analysis model: Face detection using Dlib and emotion recognition model using TensorFlow.
[1300] Processing flow
[1301] 1. Data Entry:
[1302] The user uses the device to input biometric data (e.g., heart rate, body temperature) and past experience information. The data is sent from the device to a cloud server, where it is stored in a database.
[1303] 2. Operation Record:
[1304] The device records the user's actions in real time, and the recorded data is sent to a cloud server.
[1305] 3. Data Analysis:
[1306] The cloud server analyzes the received motion data using a motion analysis model, and simultaneously analyzes the user's emotional state using face detection and emotion recognition.
[1307] 4. Providing Feedback:
[1308] Feedback is generated in the form of text, audio, images, and video and is provided to the user in real time, with the feedback content adjusted according to the user's emotional state.
[1309] Specific examples
[1310] For example, the system records the movements of a robot packing items along a conveyor belt in a factory. The movement data is analyzed on a cloud server, and if there is any waste in the movement, the system provides specific advice such as "Please reduce your arm movements to reduce waste in movement." Furthermore, if signs of fatigue are detected from the worker's facial expressions or movements, it can generate emotion-based feedback such as "Please take a break."
[1311] Example prompts for generative AI models
[1312] "Please explain how the system analyzes the operational data of robots working in a factory and provides advice on improving efficiency. In addition, please incorporate a system that analyzes the emotional state of workers and provides appropriate feedback. This system will operate using smart glasses or a head-mounted display."
[1313] The flow of the specific processing in the application example 2 will be described with reference to FIG.
[1314] Step 1:
[1315] The user wears smart glasses or a head-mounted display and inputs biometric information (heart rate, body temperature, etc.) and experience information. The device sends this data to a cloud server, which then stores the received data in a database. At this stage, the input is the user's biometric information and experience information, and the output is data stored in the cloud server.
[1316] Step 2:
[1317] The device worn by the user uses built-in cameras and sensors to record the user's movements in real time. The recorded video data is converted into an appropriate format and sent to a cloud server. At this stage, the input is the user's movement data, and the output is data sent to the cloud server.
[1318] Step 3:
[1319] The cloud server inputs the transmitted motion data into a motion analysis model and performs motion analysis. The motion analysis model extracts motion characteristics (e.g., elbow angle, hand movement, etc.) and calculates indicators related to efficiency and safety. The input at this stage is the motion data, and the output is the motion analysis results.
[1320] Step 4:
[1321] In parallel with the motion analysis, the cloud server uses an emotion engine to analyze the user's emotional state from facial expression and voice data. The emotion analysis model analyzes facial features (e.g., facial muscle movements) and determines the user's emotion (e.g., fatigue, stress). The input at this stage is facial expression data and voice data, and the output is the emotion analysis results.
[1322] Step 5:
[1323] The cloud server integrates the results of the motion analysis and emotion analysis to generate appropriate feedback. The generated feedback is sent to the user's device in the form of text, audio, image, or video. For example, feedback such as "Please reduce your arm movements" or "Please take a break" is provided. The input at this stage is the results of the motion analysis and emotion analysis, and the output is the feedback content.
[1324] Step 6:
[1325] The user receives feedback through the device and modifies their behavior to improve efficiency and safety. If the user's emotional state is determined to be "fatigued," they receive feedback that reflects their emotions, such as encouragement or a prompt to take a break. The input at this stage is the feedback content, and the output is a change in the user's behavior.
[1326] The specific processing unit 290 transmits the result of the specific processing to the headset type terminal 314. In the headset type terminal 314, the control unit 46A causes the speaker 240 and the display 343 to output the result of the specific processing. The microphone 238 acquires audio indicating a user input regarding the result of the specific processing. The control unit 46A transmits audio data indicating the user input acquired by the microphone 238 to the data processing device 12. In the data processing device 12, the specific processing unit 290 acquires the audio data.
[1327] The data generation model 58 is a so-called generative AI (Artificial Intelligence). An example of the data generation model 58 is ChatGPT (Internet Search<URL: https: / / openai.com / blog / chatgpt> ), Gemini (Internet search <url: https: gemini.google.com ?hl="ja">) and other generation AIs. The data generation model 58 is obtained by performing deep learning on a neural network. A prompt including an instruction is input to the data generation model 58, and inference data such as voice data indicating voice, text data indicating text, and image data indicating an image is also input. The data generation model 58 performs inference on the input inference data in accordance with the instruction indicated by the prompt, and outputs the inference result in a data format such as voice data and text data. Here, inference refers to, for example, analysis, classification, prediction, and / or summarization.
[1328] In the above embodiment, an example was given in which the specific processing is performed by the data processing device 12, but the technology of the present disclosure is not limited to this, and the specific processing may be performed by the headset type terminal 314.
[1329] [Fourth embodiment]
[1330] FIG. 7 shows an example of the configuration of a data processing system 410 according to the fourth embodiment.
[1331] 7, a data processing system 410 includes a data processing device 12 and a robot 414. An example of the data processing device 12 is a server.
[1332] The data processing device 12 includes a computer 22, a database 24, and a communication I / F 26. The computer 22 is an example of a "computer" according to the technology of the present disclosure. The computer 22 includes a processor 28, a RAM 30, and a storage 32. The processor 28, the RAM 30, and the storage 32 are connected to a bus 34. The database 24 and the communication I / F 26 are also connected to the bus 34. The communication I / F 26 is connected to a network 54. Examples of the network 54 include a WAN (Wide Area Network) and / or a LAN (Local Area Network).
[1333] The robot 414 includes a computer 36, a microphone 238, a speaker 240, a camera 42, a communication I / F 44, and a control target 443. The computer 36 includes a processor 46, a RAM 48, and a storage 50. The processor 46, the RAM 48, and the storage 50 are connected to a bus 52. The microphone 238, the speaker 240, the camera 42, and the control target 443 are also connected to the bus 52.
[1334] The microphone 238 receives instructions and the like from the user 20 by receiving voice uttered by the user 20. The microphone 238 captures the voice uttered by the user 20, converts the captured voice into audio data, and outputs it to the processor 46. The speaker 240 outputs audio in accordance with instructions from the processor 46.
[1335] Camera 42 is a small digital camera equipped with an optical system including a lens, aperture, and shutter, and an imaging element such as a CMOS (Complementary Metal-Oxide-Semiconductor) image sensor or a CCD (Charge Coupled Device) image sensor, and captures images of the surroundings of user 20 (for example, an imaging range defined by an angle of view equivalent to the field of vision of a typical healthy person).
[1336] The communication I / F 44 is connected to a network 54. The communication I / Fs 44 and 26 are responsible for the exchange of various information between the processor 46 and the processor 28 via the network 54. The exchange of various information between the processor 46 and the processor 28 using the communication I / Fs 44 and 26 is carried out in a secure state.
[1337] The control object 443 includes a display device, LEDs in the eyes, and motors for driving the arms, hands, and feet. The posture and gestures of the robot 414 are controlled by controlling the motors of the arms, hands, and feet. Some of the emotions of the robot 414 can be expressed by controlling these motors. In addition, the facial expressions of the robot 414 can also be expressed by controlling the light emission state of the LEDs in the eyes of the robot 414.
[1338] Fig. 8 shows an example of the main functions of the data processing device 12 and the robot 414. As shown in Fig. 8, in the data processing device 12, a specific process is performed by the processor 28. A specific process program 56 is stored in the storage 32.
[1339] The specific processing program 56 is an example of a "program" according to the technology of the present disclosure. The processor 28 reads the specific processing program 56 from the storage 32 and executes the read specific processing program 56 on the RAM 30. The specific processing is realized by the processor 28 operating as a specific processing unit 290 in accordance with the specific processing program 56 executed on the RAM 30.
[1340] The storage 32 stores a data generation model 58 and an emotion identification model 59. The data generation model 58 and the emotion identification model 59 are used by the identification processing unit 290.
[1341] In the robot 414, the processor 46 performs the reception output process. A reception output program 60 is stored in the storage 50. The processor 46 reads the reception output program 60 from the storage 50 and executes the read reception output program 60 on the RAM 48. The reception output process is realized by the processor 46 operating as the control unit 46A in accordance with the reception output program 60 executed on the RAM 48.
[1342] Next, a description will be given of the specific processing performed by the specific processing unit 290 of the data processing device 12. In the following description, the data processing device 12 will be referred to as a "server" and the robot 414 will be referred to as a "terminal."
[1343] This invention is a system for efficiently improving batting and pitching form based on a user's biometric and experience information. The system records the user's movements, analyzes the data, and provides specific feedback in real time. It also provides a means for users to choose from multiple advisors and generates videos demonstrating ideal movements, promoting the user's technical improvement.
[1344] Overall system configuration
[1345] Users use a device (e.g., a smartphone) to access the system. A dedicated application is installed on the device, allowing users to input their biometric information and experience information. The input information is sent to the server and stored in a database.
[1346] The device is equipped with a camera and sensors to record the user's movements. The captured video data is sent to a cloud server and analyzed using Google's Gemini. The analysis results are provided to the user in real time as feedback in the form of text, audio, images, video, etc.
[1347] The system also has multiple advisor profiles, allowing users to receive guidance from their preferred advisor. Furthermore, it can generate videos demonstrating ideal movements, providing users with visual reference information.
[1348] Program processing
[1349] The program does the following:
[1350] 1. User information registration
[1351] The user uses the device to input their biometric and experience information. Once input is complete, the device sends the data to the server, which stores the received information in a database.
[1352] 2. Recording a video of your form
[1353] Users record their swings and throws in the form diagnostic booth at a batting center. The motions are recorded using the device's camera or sensor, and the data is sent to a server.
[1354] 3. Form video analysis
[1355] The server sends the received video data to Google Gemini and makes an analysis request. Google Gemini analyzes the motion and returns the results to the server. The analysis results are then processed into a user-friendly format.
[1356] 4. Providing real-time feedback
[1357] The processed analysis results are sent to the device and notified to the user in the form of text, audio, images, video, etc. The user can then try to improve the form based on this information.
[1358] 5. Select coaching mode
[1359] Users can select one of several advisors within the app, and the analysis results are customized based on the profile information of the selected advisor.
[1360] 6. Image training using videos
[1361] If the user wishes, an image video of the ideal swing or throwing technique will be generated and sent to the device, and the user can use this video as a reference for training.
[1362] Specific examples
[1363] For example, if a beginner user wants to improve their pitching technique, they first enter and save their basic information in the app. They then film a video of themselves pitching at a batting center and transfer it to the app. The server analyzes the video and provides specific advice in text and voice, such as "position your elbow a little higher." Additionally, if the user selects a "famous pitching coach," they will receive customized feedback from that coach. Furthermore, a video of the ideal pitching form is also provided, allowing the user to use it as a reference to improve their technique.
[1364] The system allows users to receive specific, expert feedback in real time, allowing them to efficiently improve their form.
[1365] The processing flow will be explained below.
[1366] Step 1:
[1367] User: Launch the app and enter basic information such as age, gender, baseball experience, and medical restrictions on the "New Registration" screen.
[1368] Step 2:
[1369] Terminal: When input is complete, the user presses the save button, and the input data is temporarily saved in local storage.
[1370] Step 3:
[1371] Terminal: Generates and sends a request to send the stored data to the server.
[1372] Step 4:
[1373] Server: Saves the received user data in a database and returns a confirmation response to the terminal that the data was saved successfully.
[1374] Step 5:
[1375] User: Enters the form diagnostic booth at the batting center, stands in the designated position, and begins swinging or throwing.
[1376] Step 6:
[1377] Device: When an action is initiated, the device uses cameras and sensors to record the user's actions in video format.
[1378] Step 7:
[1379] User: When the action is finished, press the stop recording button to end the recording.
[1380] Step 8:
[1381] Terminal: Saves the captured video data in local storage, generates a request to send the data to the server, and sends it.
[1382] Step 9:
[1383] Server: Temporarily stores the received video data and sends an analysis request to Google's Gemini API.
[1384] Step 10:
[1385] Server: Receives analysis results from Google's "Gemini".
[1386] Step 11:
[1387] Server: Processes the received analysis results into a user-friendly format (text, audio, image, video).
[1388] Step 12:
[1389] Server: Sends the processed analysis results to the terminal.
[1390] Step 13:
[1391] Device: Receives analysis results and displays / notifies them on the app.
[1392] Step 14:
[1393] User: Check the analysis results and use the improvements to improve their next swing or throw.
[1394] Step 15:
[1395] User: Select one of several advisors on the "Coaching Mode Selection" screen within the app.
[1396] Step 16:
[1397] Terminal: Based on the information of the selected advisor, the analysis results are customized and provided to the user.
[1398] Step 17:
[1399] Terminal: Sends a request to the server to generate an image video of the ideal swing or throwing based on the user's wishes.
[1400] Step 18:
[1401] Server: Generates an image video of the ideal form and sends it to the device.
[1402] Step 19:
[1403] Device: The received image video is played within the app and shown to the user.
[1404] Step 20:
[1405] User: Use the image video as a reference and train to improve their own form.
[1406] This allows users to receive specific, professional feedback in real time and efficiently improve their form.
[1407] Example 1
[1408] Next, a description will be given of Example 1. In the following description, the data processing device 12 will be referred to as a "server" and the robot 414 will be referred to as a "terminal."
[1409] Conventional methods for improving batting and pitching form lack the means for users to accurately understand their own movements and receive effective feedback. Furthermore, opportunities to receive direct instruction from a coach are limited, making self-study to improve skills difficult. The objective of this invention is to provide a system that provides specific, specialized feedback in real time, allowing users to efficiently improve their movements.
[1410] The specific processing by the specific processing unit 290 of the data processing device 12 in the first embodiment is realized by the following means.
[1411] In this invention, the server includes means for inputting biometric information and experience information of the user, a terminal device for recording the user's movements, means for transmitting the recorded video data to the analytical model via the cloud server, means for processing the analysis results from the analytical model into a user-friendly format, means for feeding back the analysis results to the user in real time, means for enabling the user to select instruction from multiple instructors, and means for generating videos showing the user's ideal movements. This allows the user to receive specific and professional feedback in real time and efficiently improve their own movements.
[1412] "Biometric information" refers to information about the user's physical characteristics, such as the user's height, weight, athletic ability, and health condition.
[1413] "Experience information" is information about the user's experience, such as the user's sports history, skill level, and past training history.
[1414] A "terminal device" is a hardware device, such as a smartphone or tablet, that allows a user to input information and record actions.
[1415] A "cloud server" is a remote server connected via the Internet for storing, managing, and analyzing data.
[1416] An "analytics model" is a machine learning algorithm or data processing mechanism that analyzes a user's behavior data and provides form improvements and technical advice.
[1417] "Feedback" refers to providing users with information such as advice, comments, and analysis results generated based on the user's behavior data.
[1418] "Instructors" are coaches or advisors with specialized knowledge and experience that can be selected by the user.
[1419] "Videos showing ideal movements" are videos that show the ideal form and techniques that users should aim for, and are used to support the user's training.
[1420] The present invention provides a system for effectively improving a user's sports form based on biometric information and experience information. Specific embodiments for carrying out the present invention will be described in detail below.
[1421] Overall system configuration
[1422] Registering user information
[1423] The user starts a dedicated application using a device (e.g., a smartphone) and inputs biometric information (e.g., height, weight, health status) and experience information (e.g., sports history, skill level). This information is sent from the device to a server and stored in a database.
[1424] Recording actions
[1425] Users can record their batting and pitching movements at batting centers or practice fields using their smartphone cameras or dedicated sensors. The recorded video data is temporarily stored on the device and then uploaded to a cloud server.
[1426] Video Analysis
[1427] The server then sends the received video data to an analytical model (e.g., a machine learning algorithm) in the cloud and issues an analysis request. The analytical model analyzes the user's movements and extracts data on form flaws and areas for improvement. The extracted results are returned to the server, where they are further processed into a format that is easy for the user to understand (e.g., text, images, video).
[1428] Providing Feedback
[1429] The server then sends the processed analysis results to the device in real time. The device then displays the results and provides specific advice to the user via text or voice, such as "position your elbows a little higher." Real-time video feedback is also provided based on the analysis results.
[1430] Choosing a coaching mode
[1431] Users can browse the profiles of multiple instructors from the in-app menu and select their preferred instructor. The server generates customized feedback based on the profile information of the selected instructor and sends it to the device.
[1432] Image training using videos
[1433] It can generate images that allow users to reproduce their ideal movements. The server generates a video of the ideal form and sends it to the device. Users can use this video as a reference for training and improve their form.
[1434] Specific examples
[1435] For example, if a beginner user wants to improve their pitching form, they first enter and save their basic information into the app on their smartphone. Next, they use the smartphone camera to film their pitching at a batting center, and the video is sent from the app to a cloud server. The cloud server analyzes the video and sends specific advice, such as "position your elbow a little higher," to the user's smartphone in real time. Furthermore, if the user selects a well-known coach, that coach will provide them with customized feedback. Finally, a video of their ideal pitching form is provided to the user, who can use the video as a reference to improve their technique.
[1436] Generative AI model prompt example
[1437] "Please analyze the batting form. What is the best way to teach the player if their elbow is positioned too low?"
[1438] "Please tell me the ideal throwing motion for pitching form."
[1439] In this way, by inputting specific questions into the generative AI model, optimal feedback can be obtained.
[1440] The flow of the identification process in the first embodiment will be described with reference to FIG.
[1441] Program processing steps
[1442] Step 1: Register user information
[1443] Input: User biometric and experience information
[1444] Processing: The user launches a dedicated application using a device (e.g., a smartphone) and enters basic information such as height, weight, and sports history into a form. The device then sends the entered information to the server using the HTTPS protocol.
[1445] Output: The entered user information is saved on the cloud server.
[1446] Specific operation: When a user enters information on their smartphone and presses the "Send" button, the information is recorded in a database (e.g., MySQL) on the cloud server.
[1447] Step 2: Record a video of your form
[1448] Input: User action (e.g. swing or throw)
[1449] Processing: A user uses a smartphone camera app to record their movements at a batting cage or practice area. Once the recording is complete, it is temporarily saved as a video file on the device.
[1450] Output: The video data file will be saved on your device.
[1451] Specific operation: The user presses the record button on the camera, and when they finish recording their movements, they press the "Save" button to save the video data to their device.
[1452] Step 3: Submitting the form video
[1453] Input: Video data stored on the device
[1454] Processing: The device uploads the recorded video file to a cloud server using a secure file transfer protocol (e.g., HTTPS or SFTP).
[1455] Output: Video data is stored in cloud server storage (e.g. AWS S3).
[1456] Specific operation: The user presses the "upload" button within the app, and the video is sent to the cloud server.
[1457] Step 4: Analysis instructions for form videos
[1458] Input: Video data stored on a cloud server
[1459] Processing: The server sends the video data to the analytical model and makes an analysis request. The analytical model uses a motion analysis algorithm.
[1460] Output: The analysis request is sent to the analysis model.
[1461] Specific operation: The server automatically checks the stored video data and sends an API request to the analysis model.
[1462] Step 5: Analyzing your form video
[1463] Input: Video data and analysis request
[1464] Processing: The analytical model analyzes user behavior from video data and generates data on form parameters and areas for improvement.
[1465] Output: The analysis result data is sent back to the server.
[1466] Specific behavior: The analytical model performs behavior analysis, generates analysis results, and sends them back to the server.
[1467] Step 6: Processing the analysis results and generating feedback data
[1468] Input: Analysis result data
[1469] Processing: The server receives the analysis results data and processes it into a format that is easy for the user to understand (e.g., text, charts, videos).
[1470] Output: Processed feedback data is generated.
[1471] Specific operation: The server processes the analysis result data for feedback and converts it into the most suitable format for the user.
[1472] Step 7: Provide feedback
[1473] Input: Processed feedback data
[1474] Processing: The server sends the feedback data to the device using a real-time notification service (e.g., Firebase Cloud Messaging). The device displays the feedback data to the user.
[1475] Output: Feedback is displayed on the user's terminal.
[1476] Specific operation: A notification is sent to the user's smartphone, and when they open the app, advice such as "position your elbows a little higher" is displayed.
[1477] Step 8: Select a coaching mode
[1478] Input: User-selected mentor profile
[1479] What happens: Users select one of several mentors within the app, and feedback is customized based on their profile information.
[1480] Output: Customized feedback data is generated.
[1481] Specific operation: The server regenerates feedback using the profile information of the selected instructor and sends it to the terminal.
[1482] Step 9: Image training using videos
[1483] Input: User training request
[1484] Processing: The server generates a video demonstrating the ideal movement and sends it to the device. The user uses this video as a reference for training.
[1485] Output: A video showing the ideal behavior is displayed on the device.
[1486] Specific movements: Users play videos of ideal form in the app and train by imitating the movements.
[1487] Generative AI model prompt example
[1488] "Please analyze the batting form. What is the best way to teach the player if their elbow is positioned too low?"
[1489] "Please tell me the ideal throwing motion for pitching form."
[1490] (Application example 1)
[1491] Next, a description will be given of Application Example 1. In the following description, the data processing device 12 will be referred to as a "server" and the robot 414 will be referred to as a "terminal."
[1492] In order to ensure the efficiency and safety of industrial robots, it is important to evaluate the accuracy and appropriateness of their movements in real time and provide feedback on areas for improvement. However, existing systems do not fully realize these functions, making it difficult to optimize robot movements.
[1493] The specific processing by the specific processing unit 290 of the data processing device 12 in the application example 1 is realized by the following means.
[1494] In this invention, the server includes a means for transmitting user information to the server and storing it in a database, a means for transmitting operation records to the server and analyzing them using a model, and a means for converting the analysis results into a user-friendly format in real time and providing the results, thereby enabling efficient and safe optimization of the operation of industrial robots.
[1495] "User's biometric information" refers to data related to the user's body, including information such as heart rate, body temperature, and muscle movements.
[1496] "Experience information" refers to the knowledge and experience that the user has accumulated up to now, and is data such as proficiency level and past performance records.
[1497] A "recording means" is a device that collects the user's actions using a video camera or a sensor.
[1498] "Means for analyzing" refers to algorithms or software for evaluating recorded motion data using analytical models.
[1499] "Feedback means" refers to means for providing analysis results to the user in the form of text, audio, images, video, etc.
[1500] The "means for enabling selection of instructor" refers to an interface that allows the user to select one instructor from multiple instructors and receive evaluation and advice from that instructor.
[1501] The "means for generating animation" refers to a system that creates animation to visually show the user's ideal movements.
[1502] A "server" is a computer system that processes, manages, and provides data over a network.
[1503] The "means for saving to a database" is a function for safely saving input user information to a database.
[1504] "Means of analysis using models" refers to means of analyzing data using technologies such as machine learning and artificial intelligence to generate insights and improvement suggestions.
[1505] "Means for converting the analysis results into a user-friendly format and providing them" refers to a method for converting the analysis results into an easy-to-understand and comprehend format and providing them to the user.
[1506] The "means for proposing improvements to operations" is a function for proposing specific improvements to the user based on the analysis results.
[1507] A specific system for implementing this invention includes the following components: First, a user uses a terminal to input their own biometric information and experience information. A dedicated application is installed on the terminal, which transmits the user's information to a server and stores it in a database. The system records the user's movements using a camera or sensor and evaluates the data using an analytical model. The analysis results are processed on a cloud server and converted into a user-friendly format. These results are then fed back to the user in real time.
[1508] The system also features an interface that allows users to select instruction from multiple instructors. Users can receive customized feedback from the instructor of their choice based on their own movement data. Furthermore, the system generates an ideal movement model based on the analysis results and makes suggestions for movement improvement based on that model.
[1509] Hardware and software used
[1510] Hardware
[1511] Camera (e.g. Logitech C920)
[1512] Computer system (e.g., CPU: Intel i7, RAM: 16GB, etc.)
[1513] Robots (e.g., industrial arm robots)
[1514] software
[1515] Python programming language
[1516] OpenCV library (motion capture and image processing)
[1517] Requests library (sending HTTP requests)
[1518] Server-side program (runs on a cloud server)
[1519] Analytical models (using machine learning and artificial intelligence techniques)
[1520] Processing example
[1521] For example, when improving the behavior of a welding robot in a factory to ensure accurate welding, a camera captures the robot's movements and sends the data to a cloud server for analysis. Specific feedback such as "The welding position needs to be moved 5 mm to the left" is provided as a result of the analysis. A video demonstrating the ideal welding behavior is also provided to the user.
[1522] Prompt Sentence Examples
[1523] "Please create a system that records the operations of welding robots in factories with a camera, analyzes them on a cloud server, and provides suggestions for improvement. Please provide specific feedback in real time based on the analysis results."
[1524] This system makes it possible to optimize the operation of industrial robots efficiently and safely.
[1525] The flow of the specific processing in the application example 1 will be described with reference to FIG.
[1526] Step 1:
[1527] Registering user information
[1528] The user inputs their biometric and experience information into the device. This input data includes physical information such as heart rate, body temperature, and muscle movement, as well as skill level and past performance records. The device sends this data to the server, which stores the received data in a database. The input of this process is the biometric and experience information entered by the user, and the output is the user information stored in the database.
[1529] Step 2:
[1530] Shooting a video of the operation
[1531] A camera records the user's actions. In this case, the object of the action is an industrial robot, for example, recording the robot's welding work. The camera (e.g., Logitech C920) captures the robot's actions and saves them as a video file. The input of this process is the robot's actions, and the output is the recorded video file.
[1532] Step 3:
[1533] Analysis of motion video
[1534] The device sends the recorded video file to a server. The server receives this video data and analyzes it using an analytical model (using machine learning and artificial intelligence techniques). The analytical model takes the video data as input and evaluates the accuracy and appropriateness of the movements, which then generates specific suggestions for improving the movements. The input to this process is the video file, and the output is the analysis results.
[1535] Step 4:
[1536] Providing real-time feedback
[1537] The server provides the analysis results to the user in real time by converting the analysis results into a user-friendly format and sending them to the device in the form of text, audio, images, video, etc. This allows the user to immediately understand specific areas for improvement. The input of this process is the analysis results, and the output is the feedback provided to the user.
[1538] Step 5:
[1539] Leader's Choice
[1540] The user can select one of several mentors within the app. Based on the profile information of the selected mentor, the server customizes the analysis results, allowing the user to receive advice and make this feedback more actionable. The input to this process is the user's mentor selection, and the output is customized feedback.
[1541] Step 6:
[1542] Generation of ideal behavior
[1543] The server generates an ideal motion model based on the user's motion data. Based on the generated model, a video demonstrating the ideal motion is created and sent to the device. The user can refer to this video to improve their own motion. The input to this process is the user's motion data, and the output is a video demonstrating the ideal motion.
[1544] Furthermore, an emotion engine that estimates the user's emotion may be further combined. That is, the identification processing unit 290 may estimate the user's emotion using the emotion identification model 59, and perform identification processing using the user's emotion.
[1545] This invention is a system for efficiently improving batting and pitching form based on a user's biometric and experience information, and also combines it with an emotion engine that recognizes the user's emotions. The system records the user's movements, analyzes the data, and provides specific feedback in real time, adjusting the feedback content according to the user's emotional state. It also provides a means for users to choose instruction from multiple advisors and generates videos demonstrating ideal movements, promoting the user's improvement.
[1546] Overall system configuration
[1547] Users use a device (e.g., a smartphone) to access the system. A dedicated application is installed on the device, allowing users to input their biometric information and experience information. The input information is sent to the server and stored in a database.
[1548] The device is equipped with a camera and sensors to record the user's movements. The captured video data is sent to a cloud server and analyzed using a motion analysis system. The analysis results are provided to the user in real time as feedback in the form of text, audio, images, video, etc.
[1549] The system also has multiple advisor profiles, allowing users to receive guidance from their preferred advisor. Furthermore, it can generate videos demonstrating ideal movements, providing users with visual reference information.
[1550] In addition, the emotion engine analyzes the user's emotional state in real time and adjusts the feedback content based on the results, allowing users to receive feedback that is appropriate for their own psychological state, resulting in even more effective training.
[1551] Program processing
[1552] The program does the following:
[1553] 1. User information registration
[1554] The user uses the device to input their biometric and experience information. Once input is complete, the device sends the data to the server, which stores the received information in a database.
[1555] 2. Recording a video of your form
[1556] Users record their swings and throws in the form diagnostic booth at a batting center. The motions are recorded using the device's camera or sensor, and the data is sent to a server.
[1557] 3. Form video analysis
[1558] The server sends the received video data to the motion analysis system and issues an analysis request. The system analyzes the motion and returns the results to the server, which then processes the analysis results into a user-friendly format.
[1559] 4. Emotion Analysis
[1560] The emotion engine analyzes emotional data in real time from the user's movements, facial expressions, voice, etc., and sends the results to the server.
[1561] 5. Providing Feedback
[1562] The processed analysis results are combined with the emotion analysis results and sent to the device, which then notifies the user in the form of text, audio, images, video, etc. The user can then try to improve the form based on this information.
[1563] 6. Select coaching mode
[1564] Users select one of several advisors within the app, and the analysis results and emotional data are customized based on the profile information of the selected advisor.
[1565] 7. Image training using videos
[1566] If the user wishes, an image video of the ideal swing or throwing technique will be generated and sent to the device, and the user can use this video as a reference for training.
[1567] Specific examples
[1568] For example, if a beginner user wants to improve their pitching technique, they first enter and save their basic information into the app. They then take a video of themselves pitching at a batting center and transfer it to the app. The server analyzes the video and provides specific advice, such as "position your elbow a little higher," via text and voice.
[1569] At the same time, if the emotion engine determines that the user's emotional state is "down," more encouraging words and positive feedback will be provided. Additionally, if the user selects a "famous pitching coach," customized feedback from that coach will be provided. Furthermore, videos of ideal pitching form are also provided, so users can use them as reference to improve their technique.
[1570] This system allows users to receive real-time, specific, and psychologically relevant feedback, allowing them to efficiently improve their form.
[1571] The processing flow will be explained below.
[1572] Step 1:
[1573] User: Launch the app and enter basic information such as age, gender, baseball experience, and medical restrictions on the "New Registration" screen.
[1574] Step 2:
[1575] Terminal: When input is complete, the user presses the save button, and the input data is temporarily saved in local storage.
[1576] Step 3:
[1577] Terminal: Generates and sends a request to send the stored data to the server.
[1578] Step 4:
[1579] Server: Saves the received user data in a database and returns a confirmation response to the terminal that the data was saved successfully.
[1580] Step 5:
[1581] User: Enters the form diagnostic booth at the batting center, stands in the designated position, and begins swinging or throwing.
[1582] Step 6:
[1583] Device: When an action is initiated, the device uses cameras and sensors to record the user's actions in video format.
[1584] Step 7:
[1585] User: When the action is finished, press the stop recording button to end the recording.
[1586] Step 8:
[1587] Terminal: Saves the captured video data in local storage, generates a request to send the data to the server, and sends it.
[1588] Step 9:
[1589] Server: Temporarily stores the received video data and sends an analysis request to the motion analysis system.
[1590] Step 10:
[1591] Server: Receives analysis results from the motion analysis system.
[1592] Step 11:
[1593] Server: Processes the received analysis results into a user-friendly format (text, audio, image, video).
[1594] Step 12:
[1595] Terminal: Activates an emotion engine that recognizes emotions from the user's movements, facial expressions, and voice, and acquires emotional data.
[1596] Step 13:
[1597] Server: Analyzes the emotion data sent from the emotion engine and uses it to adjust the feedback content.
[1598] Step 14:
[1599] Server: Combines the processed analysis results with the emotion analysis results and sends them to the device.
[1600] Step 15:
[1601] Terminal: Feedback based on analysis results and emotional data is displayed and notified to the user in the form of text, audio, images, video, etc.
[1602] Step 16:
[1603] User: Check the feedback and incorporate it into their next swing or throw.
[1604] Step 17:
[1605] User: Select one of several advisors on the "Coaching Mode Selection" screen within the app.
[1606] Step 18:
[1607] Terminal: Based on the information of the selected advisor, the analysis results and emotional data are customized and provided to the user.
[1608] Step 19:
[1609] Terminal: Sends a request to the server to generate an image video of the ideal swing or throwing based on the user's wishes.
[1610] Step 20:
[1611] Server: Generates an image video of the ideal form and sends it to the device.
[1612] Step 21:
[1613] Device: The received image video is played within the app and shown to the user.
[1614] Step 22:
[1615] User: Use the image video as a reference and train to improve their own form.
[1616] This allows users to receive real-time, specific, and psychologically relevant feedback, allowing them to efficiently improve their form.
[1617] Example 2
[1618] Next, a description will be given of Example 2. In the following description, the data processing device 12 will be referred to as a "server" and the robot 414 will be referred to as a "terminal."
[1619] Conventional form improvement systems have limited dimensionality in the user's motion analysis and feedback, and in many cases, the feedback does not take into account the user's psychological state, resulting in reduced training efficiency. Furthermore, specialized equipment is often used for motion analysis and evaluation, which places constraints on cost and usage environment. Furthermore, it is difficult to select detailed guidance from multiple advisors.
[1620] The specific processing by the specific processing unit 290 of the data processing device 12 in the second embodiment is realized by the following means.
[1621] In this invention, the server includes means for inputting biometric information and experience information of the user, means for recording the user's movements, means for analyzing the recorded movement data, means for adjusting the feedback content based on the analysis results and the user's emotional state and providing the feedback in various formats, means for allowing the user to select advice from multiple advisors, and means for generating videos showing the user's ideal movements. This allows the user to receive individual and optimal feedback in real time based on the results of their own movement analysis and psychological state, enabling them to improve their form efficiently and effectively.
[1622] "Biometric information" is information that indicates the physical condition of the user, and examples include height, weight, and heart rate.
[1623] "Experience information" is information about the user's skills and career history, and examples include sports history and practice frequency.
[1624] "Motion data" is data that records the user's physical movements, and includes video and sensor data acquired using cameras and sensors.
[1625] "Analysis means" refers to a method or system for analyzing recorded motion data and providing the results.
[1626] "Feedback means" refers to a means of providing information to users based on the analysis results, and includes systems that display information in a variety of formats, such as text, audio, images, and video.
[1627] "Emotional state" refers to a user's psychological state and includes methods and systems for analyzing that state.
[1628] "Advisor" refers to an expert who provides technical guidance and advice to users, and includes their profile information.
[1629] "Video generation means" refers to a system or method for generating videos that show ideal movements to a user.
[1630] This invention is a system for efficiently improving batting and pitching form based on a user's biometric information and experience information, and also combines it with an emotion engine that recognizes the user's emotions. The specific system configuration and operation are described in detail below.
[1631] The system mainly includes the following elements:
[1632] 1. Means for inputting user biometric and experience information
[1633] 2. A means of recording user actions
[1634] 3. Means of analyzing recorded motion data
[1635] 4. A method for adjusting feedback content based on analysis results and the user's emotional state and providing feedback in various formats
[1636] 5. A means to allow students to choose guidance from multiple advisors
[1637] 6. A method for generating videos showing ideal user behavior
[1638] Entering biometric and experience information
[1639] Users use a dedicated application on a device (e.g., a smartphone) to input their own biometric information (e.g., height, weight, heart rate) and experience information (e.g., sports history, practice frequency). This information is sent from the device to a server and stored in a database.
[1640] Recording actions
[1641] For example, users can use the device's camera and sensors to record their swings and pitching movements in a form diagnostic booth installed at a batting center. The recorded video data is then sent to a cloud server.
[1642] Analysis of behavioral data
[1643] The server receives the video data sent to the cloud server and sends an analysis request to the motion analysis system. This motion analysis system analyzes the user's movements using specific algorithms and machine learning models. The analysis results are then processed into a format that is easy for the user to understand.
[1644] Emotion Analysis
[1645] The emotion engine analyzes the user's emotions in real time based on their facial expressions and voice data. The analyzed emotional state data is sent to the server and integrated with the movement analysis results.
[1646] Providing Feedback
[1647] The server generates feedback content based on the results of motion analysis and emotion analysis and sends it to the device. The device receives this data and notifies the user of the feedback in the form of text, audio, images, video, etc. This allows the user to receive specific feedback in real time that is appropriate to their psychological state.
[1648] Coaching mode and video image training
[1649] Within the app, users can select their preferred instructor from multiple advisors. Feedback is customized based on the profile of the selected advisor. Furthermore, if the user wishes, a video showing ideal swing and pitching techniques is generated and sent to the device. Users can use this video as a reference for their training.
[1650] Specific examples
[1651] For example, if a beginner user wants to improve their pitching technique, they first enter their basic information into the app. They then film a video of themselves pitching at a batting center and transfer it to the app. The server analyzes the video and provides specific advice in text and voice, such as "position your elbow a little higher." At the same time, if the emotion engine determines that the user's emotional state is "depressed," more encouraging words and positive feedback are provided.
[1652] Users can also select a "famous pitching coach" to receive customized feedback from that coach, and videos of ideal pitching form are also provided, allowing users to use these to improve their technique.
[1653] Prompt Sentence Examples
[1654] Examples of prompts for a generative AI model include:
[1655] "Please analyze my pitching video and let me know how I can improve. Also, please tailor your feedback to take into account my current emotional state."
[1656] By combining the above-mentioned methods, users can receive specific feedback in real time that is appropriate to their psychological state, allowing them to efficiently improve their form.
[1657] The flow of the identification process in the second embodiment will be described with reference to FIG.
[1658] Step 1:
[1659] Registering user information
[1660] Input: The user uses the device to input their biometric information (e.g., height, weight, heart rate) and experience information (e.g., sports history, practice frequency) into a dedicated app.
[1661] Action: The user enters the required information into the input fields and clicks the "Submit" button.
[1662] Data processing and calculation: The terminal formats the input information and generates data for transmission.
[1663] Output: The terminal sends the generated transmission data to the server.
[1664] Step 2:
[1665] Retention of Information
[1666] Input: User information data sent from the device.
[1667] How it works: The server stores the received data in a database.
[1668] Data processing and calculation: Analyzes received data and converts it into the appropriate database format.
[1669] Output: Successfully saved data is confirmed and a save confirmation notification is generated.
[1670] Step 3:
[1671] Form video recording
[1672] Input: The user has a terminal.
[1673] Actions: The user takes a video of their swing and pitching using the device's camera and sensors in the batting center's diagnostic booth. They press the record button to record their movements. When they're finished recording, they press the "stop" button.
[1674] Data processing and calculation: The device temporarily stores the recorded video data and encodes it for transfer.
[1675] Output: Send the encoded video data to the server.
[1676] Step 4:
[1677] Video data storage and analysis requests
[1678] Input: Video data sent from the device.
[1679] Operation: The server receives the video data and temporarily stores it.
[1680] Data processing and calculation: Generates requests to the motion analysis system to analyze the received video data.
[1681] Output: A request is sent to the behavior analysis system.
[1682] Step 5:
[1683] Motion analysis
[1684] Input: Video data analysis request sent from the server.
[1685] Motion: The motion analysis system analyzes the received video data using machine learning algorithms.
[1686] Data processing and calculation: The motion analysis system extracts motion elements from the video and generates analysis results.
[1687] Output: The generated behavior analysis results are sent back to the server.
[1688] Step 6:
[1689] Emotion analysis
[1690] Input: User facial and voice data.
[1691] How it works: The device uses a camera and microphone to record the user's facial expressions and voice data in real time.
[1692] Data processing and calculation: The device formats the recorded emotion data for transmission and sends it to the server. The server then sends an analysis request to the emotion analysis system.
[1693] Output: The emotion analysis system sends the analysis results to the server, and the final emotion data is generated.
[1694] Step 7:
[1695] Generating and Providing Feedback
[1696] Input: Motion analysis results and emotion analysis results.
[1697] Behavior: The server integrates the results of behavior analysis and emotion analysis and generates feedback for the user.
[1698] Data processing and calculation: Processing the feedback content into text, audio, image, or video format.
[1699] Output: The generated feedback content is sent to the device, which displays or plays it.
[1700] Step 8:
[1701] Choosing a coaching mode
[1702] Input: Profile information for multiple advisors.
[1703] How it works: A user selects their preferred advisor within the app.
[1704] Data processing and calculation: The server customizes the feedback content based on the information of the selected advisor.
[1705] Output: The customized feedback is sent to the user's device and displayed.
[1706] Step 9:
[1707] Creating image training videos
[1708] Input: The user's desired action information.
[1709] Action: The user selects the "Image Training" option.
[1710] Data processing and calculation: The server requests a video showing the ideal behavior from the generation engine and receives the generated video data.
[1711] Output: Send the generated image training video to the device.
[1712] (Application example 2)
[1713] Next, a description will be given of Application Example 2. In the following description, the data processing device 12 will be referred to as a "server" and the robot 414 will be referred to as a "terminal."
[1714] Improving work efficiency and ensuring safety in factories are key challenges. In particular, there is a need to analyze the movements of robots and workers in real time and provide accurate feedback based on that analysis. It is also necessary to provide feedback that takes into account the mental state of workers in order to create a more effective work environment.
[1715] The specification processing by the specification processing unit 290 of the data processing device 12 in Application Example 2 is realized by the following means. In this invention, the server includes means for inputting biometric information and experience information of the user, means for recording the user's movements, means for analyzing the recorded movement data, means for feeding back the analysis results in various formats, means for allowing the user to select advice from multiple advisors, means for generating a video showing the user's ideal movements, means for analyzing the user's emotional state in real time, and means for adjusting the feedback content based on the user's emotional state. This enables effective feedback that takes into account the movement analysis results and the user's emotional state.
[1716] "User's biological information" is data relating to the user's physical condition, such as heart rate, respiratory rate, and body temperature.
[1717] "User experience information" is information such as the work content and proficiency level of the user that the user has previously experienced.
[1718] "User actions" refer to specific movements or tasks performed by the user.
[1719] "Motion data" refers to video recordings and measurement data of the user's movements.
[1720] "Emotional state" refers to the user's emotional or psychological state.
[1721] "Analysis" is the process of analyzing collected data and extracting meaningful information.
[1722] "Feedback" refers to the act of returning analysis results, advice, etc. to the user.
[1723] An "advisor" is an expert with specific knowledge and experience who provides guidance and advice to users.
[1724] "Ideal movement" refers to an efficient and safe movement pattern that optimizes work performance.
[1725] An "emotion engine" is a program or algorithm that analyzes a user's emotional state in real time.
[1726] A "cloud server" is an external server that analyzes and stores data via the Internet.
[1727] "Real-time" refers to processing occurring immediately or with a very short time delay.
[1728] System Overview
[1729] This invention is a system for improving work efficiency and safety in factories. The system uses a computer terminal (e.g., smart glasses or a head-mounted display) to record the user's movements and analyzes the collected data on a cloud server. The analysis results are fed back to the user in real time, and the feedback content is adjusted according to the user's emotional state.
[1730] Hardware and Software
[1731] The system includes the following major hardware and software:
[1732] Smart glasses or head-mounted displays: equipped with cameras and sensors to record user movements.
[1733] Cloud Server: Computing resources for analyzing recorded data and generating feedback.
[1734] Motion analysis model: A machine learning model trained using TensorFlow.
[1735] Emotion analysis model: Face detection using Dlib and emotion recognition model using TensorFlow.
[1736] Processing flow
[1737] 1. Data Entry:
[1738] The user uses the device to input biometric data (e.g., heart rate, body temperature) and past experience information. The data is sent from the device to a cloud server, where it is stored in a database.
[1739] 2. Operation Record:
[1740] The device records the user's actions in real time, and the recorded data is sent to a cloud server.
[1741] 3. Data Analysis:
[1742] The cloud server analyzes the received motion data using a motion analysis model, and simultaneously analyzes the user's emotional state using face detection and emotion recognition.
[1743] 4. Providing Feedback:
[1744] Feedback is generated in the form of text, audio, images, and video and is provided to the user in real time, with the feedback content adjusted according to the user's emotional state.
[1745] Specific examples
[1746] For example, the system records the movements of a robot packing items along a conveyor belt in a factory. The movement data is analyzed on a cloud server, and if there is any waste in the movement, the system provides specific advice such as "Please reduce your arm movements to reduce waste in movement." Furthermore, if signs of fatigue are detected from the worker's facial expressions or movements, it can generate emotion-based feedback such as "Please take a break."
[1747] Example prompts for generative AI models
[1748] "Please explain how the system analyzes the operational data of robots working in a factory and provides advice on improving efficiency. In addition, please incorporate a system that analyzes the emotional state of workers and provides appropriate feedback. This system will operate using smart glasses or a head-mounted display."
[1749] The flow of the specific processing in the application example 2 will be described with reference to FIG.
[1750] Step 1:
[1751] The user wears smart glasses or a head-mounted display and inputs biometric information (heart rate, body temperature, etc.) and experience information. The device sends this data to a cloud server, which then stores the received data in a database. At this stage, the input is the user's biometric information and experience information, and the output is data stored in the cloud server.
[1752] Step 2:
[1753] The device worn by the user uses built-in cameras and sensors to record the user's movements in real time. The recorded video data is converted into an appropriate format and sent to a cloud server. At this stage, the input is the user's movement data, and the output is data sent to the cloud server.
[1754] Step 3:
[1755] The cloud server inputs the transmitted motion data into a motion analysis model and performs motion analysis. The motion analysis model extracts motion characteristics (e.g., elbow angle, hand movement, etc.) and calculates indicators related to efficiency and safety. The input at this stage is the motion data, and the output is the motion analysis results.
[1756] Step 4:
[1757] In parallel with the motion analysis, the cloud server uses an emotion engine to analyze the user's emotional state from facial expression and voice data. The emotion analysis model analyzes facial features (e.g., facial muscle movements) and determines the user's emotion (e.g., fatigue, stress). The input at this stage is facial expression data and voice data, and the output is the emotion analysis results.
[1758] Step 5:
[1759] The cloud server integrates the results of the motion analysis and emotion analysis to generate appropriate feedback. The generated feedback is sent to the user's device in the form of text, audio, image, or video. For example, feedback such as "Please reduce your arm movements" or "Please take a break" is provided. The input at this stage is the results of the motion analysis and emotion analysis, and the output is the feedback content.
[1760] Step 6:
[1761] The user receives feedback through the device and modifies their behavior to improve efficiency and safety. If the user's emotional state is determined to be "fatigued," they receive feedback that reflects their emotions, such as encouragement or a prompt to take a break. The input at this stage is the feedback content, and the output is a change in the user's behavior.
[1762] The specific processing unit 290 transmits the result of the specific processing to the robot 414. In the robot 414, the control unit 46A causes the speaker 240 and the control target 443 to output the result of the specific processing. The microphone 238 acquires voice indicating a user input regarding the result of the specific processing. The control unit 46A transmits voice data indicating the user input acquired by the microphone 238 to the data processing device 12. In the data processing device 12, the specific processing unit 290 acquires the voice data.
[1763] The data generation model 58 is a so-called generative AI (Artificial Intelligence). An example of the data generation model 58 is ChatGPT (Internet Search<URL: https: / / openai.com / blog / chatgpt> ), Gemini (Internet search <url: https: gemini.google.com ?hl="ja">) and other generation AIs. The data generation model 58 is obtained by performing deep learning on a neural network. A prompt including an instruction is input to the data generation model 58, and inference data such as voice data indicating voice, text data indicating text, and image data indicating an image is also input. The data generation model 58 performs inference on the input inference data in accordance with the instruction indicated by the prompt, and outputs the inference result in a data format such as voice data and text data. Here, inference refers to, for example, analysis, classification, prediction, and / or summarization.
[1764] In the above embodiment, an example was given in which the specific processing is performed by the data processing device 12, but the technology of the present disclosure is not limited to this, and the specific processing may be performed by the robot 414.
[1765] The emotion identification model 59 as an emotion engine may determine the user's emotion according to a specific mapping. Specifically, the emotion identification model 59 may determine the user's emotion according to an emotion map (see FIG. 9), which is a specific mapping. Similarly, the emotion identification model 59 may determine the robot's emotion, and the identification processing unit 290 may perform identification processing using the robot's emotion.
[1766] FIG. 9 is a diagram illustrating an emotion map 400 on which multiple emotions are mapped. In the emotion map 400, emotions are arranged in concentric circles radiating from the center. Emotions closer to the center of the concentric circles are more primitive. Emotions representing states and actions arising from a state of mind are arranged on the outer edges of the concentric circles. The concept of emotion includes both affect and mental states. Emotions generally generated from reactions occurring in the brain are arranged on the left side of the concentric circles. Emotions generally induced by situational judgment are arranged on the right side of the concentric circles. Emotions generally generated from reactions occurring in the brain and induced by situational judgment are arranged on the upper and lower sides of the concentric circles. Furthermore, the emotion of "pleasure" is arranged on the upper side of the concentric circles, and the emotion of "discomfort" is arranged on the lower side. In this way, in the emotion map 400, multiple emotions are mapped based on the structure by which emotions are generated, and emotions that tend to occur simultaneously are mapped close to each other.
[1767] These emotions are distributed in the 3 o'clock direction on emotion map 400, and typically fluctuate between relief and anxiety. In the right half of emotion map 400, situational awareness dominates over internal sensations, resulting in a sense of calm.
[1768] The inside of emotion map 400 represents what is going on in the mind, and the outside of emotion map 400 represents behavior, so the further you go outside emotion map 400, the more visible the emotions become (the more they are expressed in behavior).
[1769] Human emotions are based on various balances, such as posture and blood sugar levels. When these balances deviate from the ideal, a state of discomfort is indicated, and when they approach the ideal, a state of pleasure is indicated. Emotions can also be created for robots, automobiles, and motorcycles, based on various balances, such as posture and remaining battery life. When these balances deviate from the ideal, a state of discomfort is indicated, and when they approach the ideal, a state of pleasure is indicated. An emotion map can be generated, for example, based on Dr. Mitsuyoshi's emotion map (Research on Voice Emotion Recognition and Emotional Brain Physiological Signal Analysis Systems, Tokushima University, Doctoral Dissertation: https: / / ci.nii.ac.jp / naid / 500000375379). The left half of the emotion map lists emotions belonging to the "reaction" domain, where sensation is dominant. The right half of the emotion map lists emotions belonging to the "situation" domain, where situational awareness is dominant.
[1770] The emotion map defines two emotions that promote learning. One is a negative emotion on the situation side, around the middle of "repentance" or "reflection." In other words, this occurs when the robot experiences negative emotions such as "I never want to feel this way again" or "I don't want to be scolded again." The other is a positive emotion on the response side, around "desire." In other words, this occurs when the robot experiences positive feelings such as "I want more" or "I want to know more."
[1771] The emotion identification model 59 inputs user input into a pre-trained neural network, obtains emotion values indicating each emotion shown in the emotion map 400, and determines the user's emotion. This neural network is pre-trained based on multiple pieces of training data that are combinations of user input and emotion values indicating each emotion shown in the emotion map 400. Furthermore, this neural network is trained so that emotions that are located close to each other have similar values, as in the emotion map 900 shown in FIG. 10. FIG. 10 shows an example in which multiple emotions, "relieved," "calm," and "reassuring," have similar emotion values.
[1772] The system according to the present disclosure has been described above mainly with respect to the functions of the data processing device 12, but the system according to the present disclosure is not necessarily implemented on a server. The system according to the present disclosure may be implemented as a general information processing system. The present disclosure may be implemented, for example, as a software program running on a personal computer or an application running on a smartphone, etc. The method according to the present disclosure may be provided to users in the form of SaaS (Software as a Service).
[1773] In the above embodiment, an example was given in which the specific processing is performed by one computer 22, but the technology of the present disclosure is not limited to this, and the specific processing may be distributed and performed by a plurality of computers including the computer 22. For example, the data generation model 58 may be provided in an external device of the data processing device 12, and data may be generated in the external device in accordance with input data.
[1774] In the above embodiment, an example in which the specific processing program 56 is stored in the storage 32 has been described, but the technology of the present disclosure is not limited to this. For example, the specific processing program 56 may be stored in a portable, computer-readable, non-transitory storage medium such as a USB (Universal Serial Bus) memory. The specific processing program 56 stored in the non-transitory storage medium is installed in the computer 22 of the data processing device 12. The processor 28 executes the specific processing in accordance with the specific processing program 56.
[1775] Alternatively, the specific processing program 56 may be stored in a storage device such as a server connected to the data processing device 12 via the network 54, and the specific processing program 56 may be downloaded and installed on the computer 22 in response to a request from the data processing device 12.
[1776] It is not necessary to store all of the specific processing program 56 in a storage device such as a server connected to the data processing device 12 via the network 54, or to store all of the specific processing program 56 in the storage 32; only a portion of the specific processing program 56 may be stored.
[1777] The hardware resource for executing a specific process can be any of the following processors: An example of a processor is a CPU, which is a general-purpose processor that functions as a hardware resource for executing a specific process by executing software, i.e., a program. Another example of a processor is a dedicated electrical circuit, such as an FPGA (Field-Programmable Gate Array), a PLD (Programmable Logic Device), or an ASIC (Application Specific Integrated Circuit), which is a processor with a circuit configuration designed specifically for executing a specific process. Each processor has built-in or connected memory, and each processor uses the memory to execute the specific process.
[1778] The hardware resource that executes the specific processing may be configured with one of these various processors, or may be configured with a combination of two or more processors of the same or different types (for example, a combination of multiple FPGAs, or a combination of a CPU and an FPGA). Also, the hardware resource that executes the specific processing may be a single processor.
[1779] As an example of a system configured with a single processor, first, one processor is configured by combining one or more CPUs and software, and this processor functions as a hardware resource that executes a specific process. Second, there is a system that uses a processor that realizes the functions of an entire system including multiple hardware resources that execute a specific process on a single IC chip, as typified by SoC (System-on-a-chip). In this way, a specific process is realized using one or more of the above-mentioned various processors as hardware resources.
[1780] Furthermore, the hardware structure of these various processors can be, more specifically, an electric circuit that combines circuit elements such as semiconductor devices. The specific processing described above is merely an example. Therefore, it goes without saying that unnecessary steps may be deleted, new steps may be added, or the processing order may be rearranged, without departing from the spirit of the invention.
[1781] The above-described description and illustrations are a detailed explanation of the parts related to the technology of the present disclosure and are merely an example of the technology of the present disclosure. For example, the above description of the configuration, functions, actions, and effects is an explanation of an example of the configuration, functions, actions, and effects of the parts related to the technology of the present disclosure. Therefore, it goes without saying that unnecessary parts may be deleted, new elements may be added, or replacements may be made to the above-described description and illustrations within the scope of the gist of the technology of the present disclosure. Furthermore, to avoid confusion and facilitate understanding of the parts related to the technology of the present disclosure, the above-described description and illustrations omit explanations of common technical knowledge that do not require particular explanation to enable the implementation of the technology of the present disclosure.
[1782] All publications, patent applications, and technical standards mentioned in this specification are herein incorporated by reference to the same extent as if each individual publication, patent application, or technical standard was specifically and individually indicated to be incorporated by reference.
[1783] The following is further disclosed regarding the above embodiment.
[1784] (Claim 1)
[1785] means for inputting biometric information and experience information of a user;
[1786] means for recording user actions;
[1787] means for analyzing the recorded motion data;
[1788] A means of providing feedback on analysis results in a variety of formats,
[1789] a means for allowing the selection of guidance from multiple advisors;
[1790] means for generating a video showing an ideal user movement;
[1791] A system including:
[1792] (Claim 2)
[1793] 10. The system of claim 1, wherein the recorded motion data is analyzed via a cloud server.
[1794] (Claim 3)
[1795] 10. The system of claim 1, wherein the analysis results are provided to the user in real time.
[1796] "Example 1"
[1797] (Claim 1)
[1798] means for inputting biometric information and experience information of a user;
[1799] a terminal device that records user actions;
[1800] A means for transmitting the recorded video data to an analysis model via a cloud server;
[1801] A means of processing the analytical results from the analytical model into a user-friendly format; and
[1802] A means of providing real-time feedback of analysis results to users;
[1803] a means for enabling a selection of instruction from a plurality of instructors;
[1804] means for generating a video showing an ideal user movement;
[1805] A system including:
[1806] (Claim 2)
[1807] 10. The system of claim 1, wherein a machine learning model is used to analyze the recorded video data.
[1808] (Claim 3)
[1809] The system according to claim 1, wherein the analysis results are provided to the user in a variety of formats, such as text, audio, images, and videos.
[1810] "Application Example 1"
[1811] (Claim 1)
[1812] means for inputting biometric information and experience information of a user;
[1813] means for recording user actions;
[1814] means for analyzing the recorded motion data;
[1815] A means of providing feedback on analysis results in a variety of formats,
[1816] a means for enabling a selection of instruction from a plurality of instructors;
[1817] means for generating a video showing an ideal user movement;
[1818] means for transmitting user information to a server and storing the information in a database;
[1819] a means for transmitting the operation record to a server and analyzing the operation record using a model;
[1820] A means to process and provide the analysis results in a user-friendly format in real time,
[1821] A means for proposing improvements to the behavior based on an ideal behavior model based on the analysis results;
[1822] A system including:
[1823] (Claim 2)
[1824] 10. The system of claim 1, wherein the recorded motion data is analyzed via a network.
[1825] (Claim 3)
[1826] 10. The system of claim 1, wherein the analysis results are provided to the user in real time.
[1827] "Example 2: Combining Emotion Engines"
[1828] (Claim 1)
[1829] means for inputting biometric information and experience information of a user;
[1830] means for recording user actions;
[1831] means for analyzing the recorded motion data;
[1832] a means for adjusting the feedback content based on the analysis result and the user's emotional state and providing the feedback in various formats;
[1833] a means for allowing the selection of guidance from multiple advisors;
[1834] means for generating a video showing an ideal user movement;
[1835] A system including:
[1836] (Claim 2)
[1837] 10. The system of claim 1, wherein the recorded motion data is analyzed via a cloud server.
[1838] (Claim 3)
[1839] 10. The system of claim 1, wherein the analysis results are provided to the user in real time.
[1840] "Application example 2 when combining emotion engines"
[1841] (Claim 1)
[1842] means for inputting biometric information and experience information of a user;
[1843] means for recording user actions;
[1844] means for analyzing the recorded motion data;
[1845] A means of providing feedback on analysis results in a variety of formats,
[1846] a means for allowing the selection of guidance from multiple advisors;
[1847] means for generating a video showing an ideal user movement;
[1848] means for analyzing the user's emotional state in real time;
[1849] means for adjusting the feedback content based on the user's emotional state;
[1850] A system including:
[1851] (Claim 2)
[1852] 10. The system of claim 1, wherein the recorded motion data is analyzed via a cloud server.
[1853] (Claim 3)
[1854] 10. The system of claim 1, wherein the analysis results are provided to the user in real time. [Explanation of symbols]
[1855] 10, 210, 310, 410 Data Processing Systems 12 Data Processing Device 14 Smart Devices 214 Smart Glasses 314 Headset-type terminal 414 Robot< / url:> < / url:> < / url:> < / url:>
Claims
1. means for inputting biometric information and experience information of a user; means for recording user actions; means for analyzing the recorded motion data; A means of providing feedback on analysis results in a variety of formats, a means for allowing the selection of guidance from multiple advisors; means for generating a video showing an ideal user movement; A system including:
2. The system of claim 1 , wherein the recorded motion data is analyzed via a cloud server.
3. The system of claim 1, wherein the analysis results are provided to the user in real time.
Citation Information
Patent Citations
Persona chatbot control method and system
JP2022180282A